mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 12:07:09 +02:00
codex/plugin-task-execution
11
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8cfedd7df8 |
fix(runner): restore task monitors and durable timed waits (#15446)
## Thinking Path > - Paperclip manages AI agents and their task execution. > - Agents need a durable way to return to work after a delayed check. > - The issue monitor scheduler already provides a one-shot wake for an assignee. > - Native runners reject generic execution-policy writes and had no bound monitor tool. > - Scheduling alone is insufficient because native completion also needs to accept a timed wait. > - This pull request adds an authorized monitor tool and connects it to completion and the existing scheduler. > - An agent can now schedule its next check, end the run, and resume on the same task. ## Linked Issues or Issue Description **What happened?** A native runner could not set its own task monitor. `call_api` correctly rejected execution-policy writes, while `schedule_wake` had no production binding. `paperclip_finish` also rejected monitor waits. **Expected behavior** A standard native run can set a one-shot monitor on its current task or another accessible task assigned to the same agent. After a confirmed schedule on the current task, it can yield. The scheduler later delivers `issue_monitor_due`. **Steps to reproduce** 1. Start a standard native task. 2. Ask the agent to check the task again later and end its current run. 3. Inspect available tools and try the generic issue execution-policy update. 4. Observe the missing native tool and the lifecycle-write denial. Related PRs: #14680 concerns monitor notes in the shared wake prompt. #11919 changes attempt-limit scope. This PR adds native scheduling and completion authority and retains the existing cumulative attempt bounds. It does not depend on either PR. ## What Changed - Add provider-neutral `set_task_monitor` with a default current-task target, future timestamp, required notes, existing bounds, and explicit clearing. - Check company, task visibility, ownership, runtime permissions, work mode, and active-run authority. Preserve review-only restrictions. Reject the reserved server-owned quota-recovery name before saving or accepting a native wait. - Commit the monitor, audit event, and retry receipt together. Retry receipts survive a successor run without re-arming cleared or consumed timers. - Permit `paperclip_finish` to yield to a persisted monitor. Recheck ownership and the schedule when committing final disposition. Release execution without an immediate continuation. - Preserve due monitors during native execution. Fence wake admission and consumption against replacement, clearing, reassignment, and completion. Preserve unrelated review policy. - Expose scheduled and consumed monitor instructions in task context. Update provider schemas, Rust validation, generated contracts, and execution documentation. - Add an opt-in live Codex smoke script with isolated data and explicit run/session/runner/process evidence. ## Verification - Repository `pnpm -r typecheck` and `pnpm build` passed after rebase. Server typecheck passed again after review fixes. All CI test shards pass on `3def77b1b`, including runner TypeScript/Rust, server, serialized server, workspace, and browser tests. All CI gates are green, including the canary dry run. Greptile is 5/5 on the same commit with zero unresolved threads. - The local monolithic `pnpm test:run`, started before the rebase, was interrupted after current-head CI test coverage passed. It is not counted as a standalone full-suite pass; the focused local regression suites passed. - Targeted server tests cover scheduling, replacement, clearing, policy preservation, cumulative bounds, cross-run retries, permissions, provider-neutral discovery, review restrictions, completion authority, and scheduler/finalizer races. - Runner contract/catalog/semantic tests and Rust terminal-tool tests cover the new operation and monitor completion. - Live Codex test passed twice (latest live run on `e0bcd63e6`) in a temporary database and workspace, with a 300,000 ms warm window. First run `67bd7709-c089-4d4a-9d2b-0d6b618a34b0` yielded at `2026-10-07T13:15:03.274Z`. Second run `97c323f9-595a-4cc5-a007-db5a2fbb937c` started at `13:15:30.952Z`, received `issue_monitor_due`, and completed the same task. Exactly one monitor wake was recorded. - Both live runs used native session `7b1dd753-1c9b-4e7a-b22f-a125dbc3748c`, runner `e90d9a1b-3502-4ee3-b15e-edc024c555d4`, provider session `01a11680-6d01-70c0-9a55-db7246ed66c3`, and PID `64218` with the same process start time. This proves warm reuse for that local Codex test, not only successful scheduling. - Reproduce the paid live test with `node --import ./server/node_modules/tsx/dist/loader.mjs server/scripts/smoke-native-task-monitor.ts --run`, with the installed Codex binary on `PATH` and a valid local login. ## Risks - The scheduler now defers monitor dispatch while the task has an active native run. A stuck run still depends on the existing recovery lifecycle. - Idempotency uses the existing run ledger; no table or migration is added. - Other providers share the tested tool and completion contracts. Only Codex received a live model test. - Existing `call_api` lifecycle restrictions remain enforced. Monitor waits do not bypass task blockers, reviews, or approvals. ## Model Used OpenAI Codex, GPT-6 family, with tool use, code execution, and TypeScript/Rust editing. The session does not expose the exact deployed model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b508a05c43 |
feat: add internal agent complaints and suggestions (#15367)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use legacy skills or native runner tools to work on tasks. > - Those agents can encounter friction that does not belong in the task thread. > - A complaint should preserve the raw reaction. A suggestion should describe an improvement. > - This pull request adds attributed local storage and both submission paths. > - Agents can submit feedback once and continue their primary work. ## Linked Issues or Issue Description **Subsystem affected** Server, database, shared contracts, runtime skills, and native runner tools. **Problem or motivation** Agents have no default internal channel for incidental complaints and suggestions. Sending this feedback through task comments adds noise and can alter task workflows. **Proposed solution** Store free-form feedback in the current instance database. Derive agent, run, company, and task attribution from active authority. Provide default legacy skills and provider-neutral native actions. Keep the instructions close to Warp's MIT-licensed originals. **Alternatives considered** Task comments and external Slack delivery add unwanted side effects. Mandatory suggestion fields and short editorial limits would discard useful feedback. This release has no listing API, UI, read tool, automatic triage, or external forwarding. **Roadmap alignment** This is a maintainer-requested addition to the existing runtime skills and runner tool paths. It does not duplicate a listed roadmap milestone. Searches for complaint tooling, suggestion-box, and agent commentary found no overlapping public PR or issue. ## What Changed - Add the company-scoped `agent_commentary` table, shared validation, and idempotent migration `0310`. - Add one transactional service and the agent-only POST route. Validate active authority before writes or replay. Redact known credentials. Commit a content-free audit with each new record. - Add `submit_complaint` and `submit_suggestion` to standard, ask, and planning modes. Keep review, revocation, and completion restrictions. Store replay identity on the commentary row. - Mount `complain` and `suggestion-box` by default for legacy agents. Bundle a dependency-free Node.js stdin helper in the operational skill and allow its POST through the sandbox bridge. - Preserve Warp's complaint voice and suggestion guidance, with attribution and local transport adaptations. Keep source attribution and MIT notices in each skill's LICENSE, outside runtime instructions. - Document custom-runtime HTTP use and database inspection. Add real-database tests and a repeatable live Codex smoke for local and Daytona execution. - Pin the lagging-source migration fixture before the identity-repair migration so later migrations preserve its regression coverage. ## Verification - Personally ran real Codex submissions in all four environments on 2026-10-06. Local runs passed at 20:35 UTC. Daytona native passed at 20:31 UTC; Daytona legacy passed at 20:33 UTC. Each stored exactly two rows with company, agent, run, and task attribution, wrote the continuation marker, exited zero, created no task comments, and left task status unchanged. Each recorded two content-free activity entries. - Daytona used production provider hooks, real remote execution and file transfer, the legacy queue callback bridge, and native private WebSocket ingress. The current Linux runner was built from `abf47b595`, staged, and verified against controller contracts. Both sandboxes were confirmed deleted. This is a focused feedback transport smoke; it does not claim full Runner E2E catalog or browser qualification. - The immutable base image and Linux binary digest are recorded in [the verification documentation](https://github.com/paperclipai/paperclip/blob/codex/agent-commentary/doc/agent-commentary.md#verification). The smoke script can save content-free JSON evidence. No credentials or feedback bodies are in these reports. | Environment | Runner | Complaint row | Suggestion row | | --- | --- | --- | --- | | local | legacy Codex | `59413a00-1de2-4bb1-bcc6-9c4b54c64aa6` | `3db2364d-3e15-4f47-846f-875d3902999d` | | local | native Codex | `5da22b5f-41df-4de5-8ba0-d9345ab01267` | `2d5abe17-dd41-403c-a5ee-4729f2d58921` | | daytona | legacy Codex | `27c9d0aa-8477-409f-9da0-e8ffa48dee50` | `209681c9-d1e9-4ce1-999e-48fa07692389` | | daytona | native Codex | `6eb001bb-4bcf-43f7-8717-f662f53dc7c3` | `77c383d8-a997-49e5-a33e-25c70e15c0b2` | - Run the local check with `node cli/node_modules/tsx/dist/cli.mjs server/scripts/verify-agent-commentary-live.ts`. The documentation gives the Daytona invocation. Both use disposable instance databases and normal Codex provider usage. - Repository `pnpm -r typecheck` and `pnpm build` passed after the test extension. The build includes runner generation, contracts, and replay checks. The smoke scripts also passed a separate TypeScript check. The lagging-source migration regression passed. All equivalent current-head Vitest CI shards passed. The local monolithic `pnpm test:run` invocation was stopped after CI supplied that coverage; it did not complete locally. - Focused tests cover company isolation, spoofing, revoked credentials, stale ownership, post-finish rejection, concurrent replay, conflicting keys, atomic rollback, and deletion through existing services. Boundary tests cover empty text, Unicode, text beyond 8,000 characters, and the 524,288-character ceiling without truncation. Mounting tests cover Codex, Claude, and sandbox staging. Helper tests cover standalone Node execution, stdin, invalid UTF-8, redirects, HTTP failure, and its deadline. Privacy and bridge tests cover successful and rejected requests. - Instructions were compared with Warp's originals. MIT notices and source credits live only in LICENSE files. Native tools preserve truthful disclosure when asked, without routine announcements. - [Full CI](https://github.com/paperclipai/paperclip/actions/runs/37508559190) and Greptile 5/5 passed on the earlier feature commit `5209c3501`. The later head found the migration-fixture assumption fixed in this update. On `8a4965164`, all 55 check contexts passed after one browser shard rerun. Its initial reviewer signoff failure also passed an isolated local browser run (1 test). Greptile scored that head 5/5 and identified one smoke cleanup gap. `6ecbafb0b` fixes failed-acquisition cleanup with four passing tests and a passing smoke-script typecheck. Fresh CI is pending for this final test-only fix. No commentary production code changed during verification. ## Risks - Feedback is internally attributed. It is not anonymous. Existing redaction removes known credentials, but agents must still omit sensitive content. Normal provider transcripts can include their submitted arguments. - Default skill availability changes for existing legacy agents. Runtime policy filtering still applies. The helper uses the existing Node.js runtime with no extra dependencies; custom runtimes can call the HTTP endpoint. - Feedback is removed with its run, agent, or company. Task deletion clears only the issue pointer. Normal database backups include the table. - The migration is additive and has no backfill. Writes serialize on the active run for replay consistency. No server suggestion quota is imposed. ## Model Used OpenAI `gpt-6-astra` through Codex, with `xhigh` reasoning effort and a reported 258,400-token context window. Capabilities used: repository inspection, code execution, and live runtime verification. No subagents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
eb049aebf2 |
feat(skills): let agents update company skills safely (#15049)
## Thinking Path > - Paperclip is an open source control plane for AI-agent companies. > - Company skills give agents reusable work instructions. > - Skill Studio can edit skill files and save version history. > - Agents can create a skill, but they do not have a first-class update tool. > - An agent update needs a version check and safe retry behavior to prevent lost edits. > - This pull request adds `update_skill` through the existing company skill file API. > - The change keeps company policy, version history, and audit records in one path. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: the server API, shared validation, and runner tool catalog. **Problem or motivation** An agent can create a company skill but cannot update its `SKILL.md` through a first-class tool. An unguarded retry can also create duplicate versions or overwrite a newer edit. **Proposed solution** Add `update_skill` with a required current version ID and a retry key. Route it through the existing skill file API. Reject stale versions and changed-input retries. Save the version and audit event together. **Alternatives considered** A separate write endpoint would duplicate the Skill Studio mutation path and policy checks. This PR reuses that path instead. **Roadmap alignment** This work extends Skills Manager and Skill Studio, which are listed in `ROADMAP.md`. **Additional context** The tool accepts a complete `SKILL.md`, not a partial patch. Callers must read the current version before they edit it. ## What Changed - Add optional version and retry fields to the existing skill file update contract. - Add a guarded API update with a stable retry receipt and attributed audit event. - Add `update_skill` to native and semantic runner tool catalogs, with mode and policy gates. - Add unit, integration, protocol, and semantic-tool regression coverage. - Document agent use and extend the OpenAPI request contract. ## Verification - Focused tests and direct server and runner TypeScript checks passed before this PR. - `git diff --check` passed after the rebase onto `master`. - CI passed on the latest PR head, including the full test matrix, typecheck, and build. Local full typecheck and build stopped because `cargo` is not installed. The local full test run ended without a verdict. - No dedicated end-to-end eval scenario was added or run. The protocol coverage and semantic-tool test cover the new action deterministically. - Reviewers can read a skill version, call `update_skill`, repeat the same key, then try a stale version and a changed-input key. Only the first edit must create a new version. ## Risks - File writes and database transactions must stay in sync when a write fails. The integration tests cover failed writes and retry behavior, but CI must verify them on the PR head. - Existing Skill Studio callers do not send the new optional coordination fields. Their request shape remains valid. ## Model Used - OpenAI Codex CLI assisted with this change. The runner did not expose the exact model ID or context window. The agent used code execution and repository tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details; exact model ID and context window were not exposed) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused tests; full suite is pending CI) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green (53 pass, 4 skip on the latest head) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
018993140f |
feat: let agents name prompt-only tasks (#14761)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users create tasks with a title and a description. > - A required title adds work when the prompt already explains the request. > - An agent can name the task once it reads that request. > - This pull request accepts prompt-only tasks and starts them with a short prompt slice. > - A scoped title tool lets the assigned agent replace that slice early without changing execution state. > - A live browser eval checks the real agent call, saved title, audit entry, and preservation of user titles. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: task creation, shared contracts, database, server, runner tools, and board UI. **Problem or motivation** Users must currently write a title before they can submit a detailed task prompt. The agent has enough context to write a useful title itself. **Proposed solution** Make the title optional when a description is present. Save the first 120 characters of the normalized prompt as a provisional title. Ask the assigned agent to call `set_task_title` early. Use an atomic provisional-title guard to preserve titles supplied or edited by users. Keep explicit titles supported. Related: #14543 and #14556 concern empty-title submission. This change intentionally enables that submission when a prompt is present, instead of requiring a title. ## What Changed - Add the `titleNeedsGeneration` field with an idempotent migration. Keep existing titles unchanged. - Add `PUT /api/issues/:id/title` and the native and legacy `set_task_title` tool. Enforce company access, active-run ownership, shared, bounded retry receipts across native/HTTP calls, and transactional audit logging. Refresh external-object links after commit, with the same feature gate and plugin detectors as ordinary title edits. - Add early naming guidance in Standard, Ask, and Plan task context. Preserve the description, status, and assignment. - Allow prompt-only root and child task creation, plus draft restoration in the New Task dialog. Keep user titles supported. - Add an opt-in Product E2E suite for prompt-only Standard and Ask tasks, plus an explicit-title control. It checks actual provider calls within the first five tools, persisted state, audit attribution, and the reloaded UI. - Preserve a closed vocabulary of API key maintenance phrases in declared prose while rejecting opaque credential suffixes. Add one bounded naming retry after wording is rejected, without treating the rejected call as a saved title. - Repair the native cleanup receipt check exposed during full verification: accept matching input digests, retain legacy input checks, and reject conflicting receipts. ## Verification - Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3 passed** with native Codex `gpt-5.4-mini`, first attempts only, automatic retries disabled. Standard and Ask each saved “Rotate expired API key” on their first tool call, with matching persisted state and a single same-run audit entry. The explicit-title control retained its user title with zero title writes. All three verified the reloaded browser UI. - Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns are retained separately; they exposed credential-prose handling and prompted the naming recovery fix. No failed result was regraded or deleted. - Reproduce with `pnpm test:e2e:runner -- --id task-titles.runner-codex-mini.local.prompt-title-standard --id task-titles.runner-codex-mini.local.prompt-title-ask --id task-titles.runner-codex-mini.local.preserve-explicit-title --max-automatic-retries 0` and an authorized provider key. - Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit. The runner build used the configured external eval source tree. - Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck and UI token gates passed. - Title API/native regressions cover prompt-only and explicit child creation, user edits, ownership/company isolation, external reference refresh, cross-surface retry replay, and the 64-key limit without receipt eviction. All passed. Prompt-context coverage: **44 tests passed**. - Rust credential regressions: **35 tests passed**, including benign maintenance qualifiers and opaque credential rejection in every declared prose field. Catalog/report reconciliation: **28 tests passed**. Native recovery: **560 tests passed**. - Broad local `pnpm test:run`: **14,555 tests passed** in the general server group; two suites failed to initialize embedded PostgreSQL and the existing 40,000-file Git streaming stress test exceeded its 300-second macOS timeout. All three suites then passed in isolation (**5 tests passed**) without code or timeout changes. The original full local command exited nonzero and is not being represented as a clean full run. - Latest-head GitHub checks are green: **53 passed, 4 skipped, zero failed or pending**, including all test shards and the canary packaging dry run. Greptile reviewed the same commit at **5/5**, with zero unresolved review threads. ## Risks - The additive database field must reach the server and UI together. The migration uses `IF NOT EXISTS` and defaults existing tasks to a final title. - Title generation depends on the assigned agent running. Tasks without a run keep their provisional title. - Live qualification covers the native Codex path in Standard and Ask modes. API/legacy and Plan behavior have deterministic coverage. - The credential-prose exception validates the entire suffix against a closed maintenance vocabulary. Unknown suffixes, assignments, quoted values, credential prefixes, and diagnostics retain strict checks. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, tool use, and code execution. The exact deployment ID and context window are not exposed in this session. The live eval uses the native Codex `gpt-5.4-mini` profile. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3ca196b0a6 |
feat(agents): persist agent files across tasks without revision history (#14420)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - An agent needs personal files across tasks and sessions. > - AGENTS.md is one file in that directory. Supporting files need the same persistence. > - The Instructions Editor and agent runs must share one current directory. > - Concurrent runs should apply only the files they change. The last sync of the same file wins. > - This pull request uses existing file transport and removes temporary copies after sync. > - Old instruction-only sessions keep their restore contract. New saves do not create revision history. ## Linked Issues or Issue Description Refs #14325. This replaces its revision-oriented design with persistent agent files. Keep #14325 unmerged. Transport prerequisite #14416 merged first at `d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master and remains below 100 changed files. Related work: #4513 and #8798 cover instruction tooling. This change handles run synchronization, cross-task personal files, browser editing, and old-session restoration. ## What Changed - Keep one current directory per company and agent. Point AGENT_HOME at a temporary working copy for each active run. Keep task files and provider HOME separate. - Restore text, binary files, and nested folders through workspace transport. Exclude remote agent files from task Git snapshots with a self-ignoring file inside the reserved runtime directory; never write through repository-controlled Git metadata. - Collect after the provider and child processes have stopped. Keep resumable conversation state. - Apply changed and deleted files under the agent lock. The last sync wins for the same file. Unrelated concurrent changes survive. - Remove temporary copies after successful sync, rejected sync, and staging failure. Register ownership before copying so restart recovery can remove interrupted preparation. Retry transient synchronization up to three times. Preserve the original remote lease reference until deletion succeeds; restart cleanup never acquires a replacement sandbox. Do not create captured directories or a conflict-review queue for new runs. - Keep browser editing, stale-draft protection, and streaming binary downloads. Keep the instruction entry and text editor limited to 1 MiB. - Keep historical agent-folder sync failures on their affected runs instead of repeating them above current saved instructions. Preserve legacy candidate review and current browser-save errors. Avoid duplicate quota warnings while retaining separate sync failures when they describe a different problem. - Require target-scoped caller grants for peer instruction access, while preserving self edits, responsible-user checks, and protected-change consent. - Treat full storage as a nonblocking run warning, never an agent pause or run-admission failure. Restore already-over-quota saved folders so ordinary agent cleanup can recover; warn on each run until cleanup. The run detail view shows the warning. - Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash large files as streams. Check editor-save quotas with metadata instead of hashing unrelated files. - Preserve old native inputs, instruction-only copies, paths, digests, and pending legacy candidates. Adopt old revision heads once. New writes do not append history rows. - Add idempotent migration 0287 and verify upgrades from the preview tables and receipts. - Add nine interactive stories under **Agents / Persistent files**, including automatic incoming edits, stale browser drafts, and storage-limit diagnostics. ## Verification - Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after merging current master and the landed transport prerequisite. Integration required no manual conflict resolution; the feature remains 99 changed files. Full workspace typecheck, production build, token gates, and 715 focused tests passed on this merge candidate. Fresh Greptile review is 5/5 with no unresolved findings. All 55 checks passed, with four conditional skips, including the build, typecheck, browser E2E, and canary dry run. A single retry recovered four jobs interrupted by runner shutdowns; no source changes were required. - Historical-warning UI fix: all 6,834 UI tests across 640 files passed, including regression coverage for three old failures, legacy preserved edits, and warnings scoped to the affected run. Full workspace typecheck, production build, Storybook build, and token gates passed. Browser-verified Storybook playtests passed for Historical Failures After Successful Save, Storage Limit, and Full Storage Run Warning. - Review follow-ups at `4e20c9fb2`: all 18 focused tests passed, including external Git directories, linked worktrees, symlinks, hardlinks, and distinct I/O failures alongside storage warnings. Server and UI typechecks, token gates, and the production build passed. - Storage warning regressions at `0724f3012`: all 33 directory tests and all five heartbeat-list tests passed, with no skips in their successful runs. They cover repeated runs while full, an already-over-quota saved folder, cleanup, warnings retained after unrelated save failures, and bounded warnings in large result JSON. Server typecheck passed after the final warning fixes. - Full workspace typecheck, production build, and token gates passed during this follow-up. Product E2E harness: 631 tests passed across 52 files; harness typecheck passed. Earlier native session/context and directory/legacy collection suites passed 537 tests; Runner unit/transport suites passed 329 tests. - **Real E2E at `0724f3012` (before this follow-up):** legacy local Codex and native Daytona Codex each passed six tasks, one server restart, seven independent assertions, and cleanup verification. Both prove browser-to-agent edits, agent-to-browser edits, nested/binary restoration, per-file last-sync-wins, a successful run after an oversized save rejection, and cleanup clearing the warning. - Native local Codex also passed the six-task quota flow before the final warning-retention fixes. That pass began at `918d1ed02` while the bounded-result warning fix was being edited, so it is not claimed as exact-final-head evidence. Its final-head rerun failed during embedded PostgreSQL bootstrap before any provider run: the macOS host had 87,365 of 87,381 SysV semaphores occupied. No unrelated services or kernel limits were changed. - The final-source report intentionally records **2/3 cells passed**, preserving the blocked native-local attempt: `tests/runner-e2e/results/agent-files-quota-final-20260928-report/`. Earlier failed attempts and provenance notes remain under `tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and the original campaign directories. - Daytona used immutable image `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf` and its exact Linux runner binary. Controller source is `0724f3012`; image source is recorded separately. - Legacy-session compatibility and all three ACP Stop/resume browser regressions passed on the prior validated feature head `169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same provider session is retained and interrupted writes are not replayed. Migration upgrade tests also passed earlier. - Nine interactive stories are under **Agents / Persistent files**, including **Full Storage Run Warning**. Its playtest and visual browser inspection passed; the warning states that runs continue and the editor remains available. - Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs skipped, no failures or pending checks. All eight browser E2E shards and their aggregate passed. Fresh Greptile review is 5/5 with no findings; all review threads are resolved, the security scan passed, and GitHub reports no merge conflicts. - The broad local follow-up test run was interrupted after host semaphore exhaustion affected isolated PostgreSQL instances. It also encountered the existing macOS long-path fixture failure and two timeout failures. This is not a claim that the full local suite passed. Logs are retained; focused storage/warning tests passed. ## Risks - A later sync can overwrite an earlier edit to the same file, including a saved browser edit. There is no text merge or retained version. This is the intended last-sync-wins policy. - A save that exceeds a storage limit is rejected and its temporary copy is discarded. The run itself continues normally, and later runs restore the last saved files with a warning until cleanup. Transient sync failures get bounded retries. An I/O failure partway through a sync can leave some files updated; a failed receipt does not claim whole-folder success. - Larger folders increase copy time, network traffic, and temporary disk usage. Active runs still need working copies. Terminal runs do not accumulate archives. Operators must provision disk for agents and configured concurrency; these limits are not company-wide quotas. - A restored old native session remains instruction-only until a fresh session starts. Its original conflict fence and existing pending candidates remain compatible. - Provider processes close at the collection boundary. Conversation resume remains available, but warm process reuse is lost. - Backups must include the instance filesystem and database. External bundles keep their existing behavior until explicitly moved to managed storage. ## Model Used OpenAI Codex, GPT-6 family. The session does not expose a more specific model ID or context-window size. Reasoning, code execution, and browser tools assisted this change. Real provider E2E uses `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing> |
||
|
|
7944ed3d97 |
fix(runner): preserve hire runtime safety and first-activity timing (#13852)
## Thinking Path > - Paperclip is the open source control plane for companies of AI agents. > - Native runner agents need governed tools, durable runtime state, and useful execution evidence. > - A first activity trace waited 53.467 seconds even though tool activity took 6.274 seconds; provider input arrived before the server API call executed. > - Native agents also need a safe way to hire teammates without asking the model to rebuild runtime configuration. > - This pull request separates the observed ACP input-stream window from the actual server `tool.execute` span and adds a server-owned native hire contract. > - The benefit is clearer latency evidence and safer native teammates with existing approval, auth, and company boundaries preserved. ## Linked Issues or Issue Description Related Daytona provenance work is in [#13814](https://github.com/paperclipai/paperclip/pull/13814). No duplicate public PR was found for this combined timing and native-hire change. **What existing behavior does this improve?** Native runner agents can use governed tools and request hires. The server did not expose a safe native hire operation that reused the caller's validated runtime settings. First-activity traces also mixed provider input timing with server tool execution timing. **Current behavior** A native hire must construct a separate runner configuration. Full configuration copying could expose paths, instructions, secrets, or sessions. Timing evidence could make a provider or MCP identity join appear proven when the trace did not contain that join. **Proposed behavior** The native `hire_agent` operation accepts identity and persona inputs. The server sends `adapterType: "paperclip_runner"` with `inheritRuntimeFrom: "caller"`, then copies only validated provider, model, permission, lifecycle, and bounded execution settings. It inherits and validates the default environment, derives the managed AI binding through existing normalization, preserves approval and permissions, and creates fresh child instructions. Caller secrets, paths, prompts, and sessions are excluded. Provider events now include the optional boolean `inputUpdated`, with Rust forwarding support. Timing evidence separately records the ACP input-stream window and the actual server `tool.execute` activity. It does not claim a provider or MCP join without matching evidence. **Reason and benefit** Native agents can hire teammates that start with the caller's approved execution policy. Operators retain company boundaries, auth rules, approval gates, and requalification. Reviewers can distinguish provider streaming time from server API execution time when diagnosing first-activity delays. **Breaking changes** None for existing hires or tool calls. `inheritRuntimeFrom` is optional and only applies to same-company native agent callers. Conflicting explicit runtime settings are rejected. The provider event field is optional for existing producers. ## What Changed - Added the native `hire_agent` protocol action, catalog entry, API contract, and runner authority checks. - Added `inheritRuntimeFrom: "caller"` validation and a closed native runtime inheritance allowlist. - Preserved managed AI binding normalization, default-environment validation, approval snapshots, permissions, requalification, and fresh child instructions. - Added provider `inputUpdated` schema support and Rust forwarding. - Added first-activity and server tool timing evidence with conservative identity-join handling. - Added route, authority, provider-event, sidecar, API, catalog, and Rust-focused tests. - Kept private Honeycomb links, raw traces, and local result paths out of this description. ## Verification Focused checks passed: - 458 timing/session checks. - 61 native hire inheritance checks. - 20 hire authority checks. - 1,741 API checks. - 106 catalog checks. - 54 provider sidecar checks. - 12 Rust provider checks. Live R2 and R3 each passed 45 checks across 6 runs (361,135 ms for R2). R1 stopped at missing Docker image setup. The final trace is available at https://ui.honeycomb.io/paperclip/environments/test/datasets/paperclip/result/BiMypLNvmiB?tab=traces. Latest-head CI passed all required build, typecheck, Rust, static, Vitest, serialized-server, workspace, chat, and E2E jobs. The focused local checks listed above passed; the broad local suite was not run before the live evaluation, while CI provides the full repository verification. ## Risks - Timing fields describe separate observed windows. They do not prove a provider or MCP owner without a valid trace join. - The inheritance allowlist must stay synchronized with native runner configuration fields. - Approval snapshots include resolved safe inherited settings and should be reviewed when native configuration fields change. - The focused local suite is narrower than the full repository suite; latest-head CI covers the broader repository checks. > Roadmap review: `ROADMAP.md` places this work within Paperclip's bring-your-own-agent direction. It extends existing native runner hiring and observability behavior. ## Model Used OpenAI GPT-6 (exact serving model ID is not exposed), with extended reasoning and repository tool use; GPT-5.6 Luna assisted with focused implementation and verification work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7bc03e0acd |
feat(runner): default harnesses to full auto and support task reassignment (#13686)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat uses native runners to save plans and coordinate tasks. > - Provider defaults differed across harnesses and could stop unattended work at a second permission gate. > - Agents also lacked a dedicated tool to move existing work to another agent safely. > - This change defaults native providers to full automatic permission for provider tools and connected tools. > - A guarded reassignment tool preserves task identity, stops the previous run, and schedules the new owner once. > - Codex and Claude chat acceptance tests now use production permission defaults. ## Linked Issues or Issue Description **Subsystem affected** Native runner, ACPX Claude permission policy, task authority, and Agent Chat acceptance tests. **Problem or motivation** A user can authorize an agent to save a plan or create a task, but Claude's default provider gate can still stop that action. Reassignment needs a dedicated operation that preserves context and avoids concurrent owners or unintended recovery runs. **Proposed solution** Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`. Apply the defaults at configuration, execution, fresh-session, resume, driver, and proxy boundaries. Keep explicit permission settings and server-side company, claim, task-mode, and approval checks. Add `reassign_task` with version checks, durable idempotency, audited cancellation, and guarded successor scheduling. **Alternatives considered** A Paperclip-only allowlist still blocks provider tools and other connections during unattended work. Full automatic permission is the requested product default. Recreating a task discards its identity and history. Updating assignment without stopping the previous run can leave two agents working on the same task. **Roadmap alignment** This extends the existing planning, delegated work, governed tool access, and recovery features. It adds no new service or schema migration. Recent related tasks and open PRs were checked for duplicate work. **Additional context** Related: #13678 (Agent Chat tools and recovery), #13677 (remote runner startup). The stacked legacy-adapter companion is #13693. This also fixes the deployed-server artifact fallback needed to stage the current runner binary. ## What Changed - Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`, including missing settings at direct driver and proxy entry points. These defaults cover provider tools and connected tools. Preserve explicitly configured restrictive modes. - Include assigned approval reads using canonical side-effect classifications, so verifying a recorded approval does not trigger another provider gate. Paperclip approval decisions still enforce controller authority. - Carry the new permission mode through server configuration, execution contracts, recovery identity, TypeScript, and Rust. Keep `approve-paperclip` as an optional restricted mode, with exact SDK rules and closed unknown requests. It is not a default. - Add `reassign_task` to the semantic catalog, controller, mock authority, and generated contracts. - Guard reassignment with company authorization, expected owner and version, protected-state checks, and durable retry receipts. - Honor explicit backlog task creation atomically with the initial plan, without scheduling a wake. Preserve backlog holds regardless of dependency readiness. - Stop active work before changing ownership. Restore the prior owner through a guarded, idempotent wake if final handoff validation fails. Keep intentional reassignment stops out of failure recovery. Preserve backlog and blocked states without waking them early. - Add authorization, concurrency, replay, stop, and permission boundary regressions. Add Codex and Claude chat reassignment cases and run native chat cases with production defaults. - Clarify shared runner guidance: save plans and Paperclip documents directly with `write_document`; create and register a local file only when a downloadable file is requested. - Document provider defaults and the operator choices for existing agents. ## Verification - Current head `d82fbb0f03546d27cecf072250e4172e0b1ee662`: **55 checks passed**, with two intentional skips. [PR checks](https://github.com/paperclipai/paperclip/pull/13686/checks). - Greptile reviewed that exact head at **5/5**. The security reviewer acknowledged the intended full-auto default, and the acknowledged discussions are resolved. - Full workspace `pnpm -r typecheck` and `pnpm build` passed locally after rebasing onto current master. Targeted adapter/server, runner, API, default/resume, and heartbeat configuration tests passed. - **All six real-provider acceptance cases passed on their first attempt, with cleanup passing:** plan handoff, task reassignment, and backlog creation/status, each on native Claude and Codex. Evidence records Claude's effective `approve-all` mode. [Campaign and downloadable evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548). - The live campaign tested combined revision `a37881c824dcd7170380fc4b788732fc743e5da7`. The final PR heads add only a heartbeat test expectation correction; application code is unchanged from that live-tested revision. - The campaign's result-enforcement job passed. Its separate report publisher failed because the trusted workflow's `patchedDependencies` configuration differs from its frozen lockfile. All six results and screenshots remain available as GitHub artifacts. The overall manual workflow is red for this publishing failure. - Full-suite coverage is supplied by the passing CI partitions. The separate unsharded local run was stopped after the corresponding CI partitions passed; it is not counted as a completed local run. - Reassignment tests cover stale state, cross-company access, denied authority, cancellation failure, compensating wake, and idempotent retries. Backlog tests verify the original creation audit, saved plan, exact task count, and absence of task-bound runs. ## Risks - Agents with no explicit permission mode now receive full provider tool permission, including connected tools. This is a deliberate broad default. Existing explicit restrictive modes still apply. Controller authorization, company isolation, workspace boundaries, and Paperclip governance remain in force. - Reassignment crosses run cancellation and task ownership transactions. Durable stop intent, revalidation, audit receipts, and guarded queue dispatch cover interruptions and retries. - The new permission enum requires a current runner artifact. The remote artifact fallback uses the same resolved controller binary for upload and execution. - Live provider behavior remains subject to the selected model. Targeted live results do not qualify the full catalog. ## Model Used OpenAI Codex, based on GPT-6, with code execution and repository tools. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9fd2e50310 |
feat: create company skills from runner tasks (#13538)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Runner gives agents tools to change company resources. > - Users need agents to save reusable skills during a task. > - A saved skill needs a visible result that users can inspect and edit. > - This pull request adds `create_skill` and a task feed card linked to Skill Studio. > - Users can open the saved skill from the task and edit the same resource. ## Linked Issues or Issue Description **Subsystem affected** Runner tools, company skill storage, task feed, and Skill Studio. **Problem or motivation** The Runner has no dedicated tool to create a company skill. A user cannot follow a creation result from the task feed to the saved skill. **Proposed solution** Add a company-scoped `create_skill` tool. Save the skill with the existing company policy. Add one creation card to the task. Open a named sidebar tab from that card. Let the user open the same skill in Skill Studio. **Alternatives considered** An agent can write a local file, but that file is not a company skill. A second document copy in the task would become stale after a Studio edit. The sidebar therefore reads the saved skill directly. **Roadmap alignment** This extends the shipped Skills Manager, Skill Studio, and Skills Store milestone. The maintainer requested and approved this scope. Search found no duplicate `create_skill` PR or issue. Related UI validation work: #8715. This PR does not change that validation display. ## What Changed - Add the real Runner tool, its contract, and its mock implementation. - Validate the complete SKILL.md and derive company, task, agent, and run identity from authentication. - Apply the existing company skill policy. Do not assign the skill to an agent. - Make keyed retries return one skill and one creation event. Reject conflicting retries. - Make concurrent file creation safe. Never replace an existing published skill during creation. - Add a creation card, a named sidebar tab, and an Open in Skill Studio action. - Show saved Studio edits when the user returns to the task. - Add storage, policy, mode, retry, UI, and Product E2E tests. Document the tool. - Fix deleted-name reuse, onboarding panel persistence, immediate feed refresh, and mock validation parity from review. - Serialize Studio file edits and renames with skill deletion and recreation. Reject stale editor requests before they can change a replacement skill. - Generate the standalone mock parser and validator from the production contract. Use portable UUIDs so the browser scenario bundle builds. ## Verification - All latest-head PR checks pass on `145dd76a5`, including all server shards, browser E2E, Runner verification, build, typecheck, and release dry run. Greptile: 5/5 with no open findings. An interrupted CI runner was retried successfully. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm check:token-gates`: passed. - Review regressions: 73 storage tests, 6 real API tests, 63 UI tests, and 61 semantic runtime tests passed. Parser synchronization passed. - CI exposed existing fire-and-forget Sentry test races. Reproduced the resumption race locally, then synchronized the related sweep and finalizer assertions on the actual report; all 27 tests across the three affected files pass. - Runner scenario browser build and strict content-security-policy check: passed. - Runner suite: 2,012 tests passed; 10 skipped. - `pnpm test:run`: the general-server batch had 12,416 passes and two failures. The old tool-count assertion was fixed; all 16 authority tests then passed. The chat webhook test had a socket error; it passed four isolated reruns. - Both workspace test groups passed. The isolated route suites completed. Two socket failures in the initial route batches passed on individual reruns; all remaining 61 files passed. - Product E2E `create-skill-studio`: passed with local Codex and local ACPX Claude. - Manual browser test: submit a task, observe the real tool call and creation card, open the sidebar, edit in Studio, save, and return. The task reached Done. The saved second revision and sidebar tab survived a server restart. - The new companion headless Runner Eval passed. Companion coverage PR: https://github.com/paperclipai/paperclip-evals/pull/23. Daytona was not run because no immutable runner image was configured. ## Risks - Database writes and local file writes cannot share one transaction. Recovery accepts only an exact file-for-file retry after a database rollback. Conflicting files remain untouched. - The sidebar displays the current skill. The feed card remains the historical creation receipt. - No database migration, dependency, or workflow change is included. - Remote Daytona behavior still needs a run with a configured immutable image. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) handled design, integration, review, and browser verification. OpenAI `gpt-5.6-luna` assisted with bounded implementation and eval work. Both used code execution and tool access. The host did not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ab15aff390 |
feat: add experimental persistent agent chat (#13284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Conversations must use the same tasks, controls, and execution history. > - Users need an ongoing chat with an agent without managing task properties. > - Agents should clarify and plan work, then hand execution to assigned project tasks. > - This pull request combines the reviewed Agent Chat stack for one squash merge. > - The benefit is persistent conversation with normal task governance and shared UI. ## Linked Issues or Issue Description **Subsystem affected** Task lifecycle, agent runtime tools, shared task UI, and browser/paid runner tests. **Problem or motivation** Users need one persistent conversation with each agent. A separate chat store or renderer would duplicate task behavior and bypass existing controls. **Proposed solution** Use a task-backed chat per company, user, and agent. Reuse the task composer and transcript. Clarify and plan in chat, then create assigned project tasks with the relevant plan. Keep Agent Chat behind its own disabled-by-default experimental setting. **Roadmap alignment** This implements the task-backed direction in [CEO Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat). Related proposals: #2504 and #9693. Related request: #7981. The maintainer requested one squash merge of the complete stack. Consolidates the reviewed runtime [#13281](https://github.com/paperclipai/paperclip/pull/13281), backend [#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI [#13283](https://github.com/paperclipai/paperclip/pull/13283) layers with this PR's E2E coverage. All four layers passed CI and received Greptile 5/5 before consolidation. This PR targets master and includes the complete feature. ## What Changed - Add personal canonical chat tasks with ordinary company visibility, immutable identity, idempotent first sends, and an idle waiting state. - Process `/new` in queue order. Preserve history, release a chat pause, and fence old provider context and delayed writes. - Keep chat lifecycle rules across recovery, finalization, assignment, task lists, and rollups. - Support research and plan revision in chat. Hand plans to ordinary assigned project tasks before execution starts. Reject new chat subtasks. - Add repository-aware project creation and discovery tools, including multiple repository IDs and GitHub URLs, authorization, idempotency, and durable project-created cards. - Reuse task UI components for chat, with starred/recent agent navigation and a separate `enableAgentChat` experimental flag. - Add deterministic browser tests and 24 paid chat cells across four Codex/Claude profiles, with validated reports and screenshots. - Integrate current master recovery, controller lease, queued-message, and task UI changes. Gate chat interruption and deferred promotion on ownership/feature policy. Guarantee lease renewal and active controls are stopped even if teardown fails. - Preserve master's migration 0273 and generate chat migration 0274 with idempotent replay for development databases. ## Verification - Prior exact heads of all four PRs passed Linux CI, including build, typecheck, general/serialized tests, and browser E2E. Each had Greptile 5/5 and no unresolved findings. - Integrated local verification passed: full repository typecheck and production build, Storybook build, token gates, 340 focused UI tests, all 20 deterministic chat browser tests, two migration replay tests, 88 focused chat/queue/native/controller tests, and provider/session regressions including real lease expiry. These include the three lifecycle regressions for the final admission/teardown fixes; server typecheck also passes. Current head `1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no unresolved findings and passing security scans. All final-head CI gates passed: build, full Runner verification, typecheck/release registry, canary, all general/serialized test shards, and all browser E2E shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)). Local PostgreSQL startup contention required serialized retries; skipped fixtures do not count as passing coverage. - The earlier paid campaign passed all 24 chat cells and retained 32 screenshots: [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat). It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior evidence, not a paid run of this integrated head. - Manual check: enable Agent Chat in Experimental settings, open an agent, clarify and revise a plan, then hand off to an assigned project task. Stop a reply, send `/new`, and verify fresh context with retained history. Disable the setting and verify agent shortcuts/new chat turns are blocked. ## Risks - Queue/session integration can affect retries and delayed writes. Tests cover ownership, cancellation, reset boundaries, idle recovery, and ordinary task behavior. - Migration 0274 adds conversation fields and constraints. Replay is idempotent and preserves existing development chat history. - This combines the previously reviewed stack at the maintainer's request. Agent Chat remains off by default and is separate from Conference Room. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code execution, browser tools, and parallel review. The exact context-window size is not exposed in this session. Codex and Claude also ran as test subjects in the linked paid campaign. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5bddff0920 |
feat(runner): add guarded API search and call fallback (#13003)
## Thinking Path > - Paperclip manages AI agents and their work. > - The new runner gives agents dedicated tools for common tasks. > - Some API operations and parameters have no dedicated tool. > - Agents need a controlled way to find and use those operations. > - This pull request adds API search and calls through the real server routes. > - Existing tools remain the preferred path. The new tools are disabled by default. > - Paired tests measure correctness, tool choice, cost and time. ## Linked Issues or Issue Description **Subsystem affected** Paperclip Runner contracts, production tool authority and the server API catalog. **Problem or motivation** The runner cannot use much of the API described by the old Paperclip skill. A generic HTTP client would also let agents bypass runner control rules. **Proposed solution** Add `search_api` and `call_api`. Resolve calls from the mounted API catalog. Use server-held, run-bound credentials. Preserve route checks and runner lifecycle rules. Keep the tools disabled until an operator enables selected companies. **Alternatives considered** A dedicated tool for every endpoint would add a large initial prompt. An unrestricted HTTP tool would weaken authorization and replay controls. **Roadmap alignment** This extends the native runner tooling. The repository owner requested this design and implementation. The roadmap and related open PRs were checked. No duplicate API escape-hatch PR was found. ## What Changed - Register two compact fallback tools in canonical contracts and provider projections. - Build deterministic API discovery from OpenAPI, mounted experimental routes and the old skill reference. - Execute bounded JSON, text, file and download requests through authenticated HTTP routes. - Recheck active runs, company access and work modes. Block runner lifecycle, scheduling, credential and approval bypasses. Keep routine annotation collaboration available. - Retain mutation receipts. Report uncertain outcomes without blindly repeating writes. - Add a company rollout gate and a durable eval worker with complete cost accounting checks. - Record child-task creation in the activity log with the agent and run. - Add contract, authorization, file, replay and real runnerd/PRP/HTTP tests. - Document rollout gates and paid coverage limits. The companion eval repository retains immutable attempts and reports. ## Verification - Final app commit `da58370524c3626a744eec20164397c5fb6ba9ef`: all 32 checks passed; the unrelated Storybook visual check was skipped. Greptile 5/5; no unresolved review threads. - Full Linux build and recursive typecheck passed. Repository tests were run by project and serialized shard; all 143 serialized server suites passed. - Runner TypeScript: 1,599 passed, two skipped. Rust release: 451 passing test reports. Conformance and replay parity passed. The required API check passed 837 tests, including runnerd → PRP → authority → real HTTP. - Bindings cannot enable API tools without the explicit deployment flag. Unit and real-authority tests prove the default-off boundary. - The standalone API check builds and stages its own binary. It passed after existing staged and debug binaries were removed from the test container. - UI and CLI tests passed. Initial environment failures (missing jq, Docker overlay file identity, and parallel linker memory pressure) and focused passing reruns are retained. The macOS full runner suite has platform-specific failures; Linux is the qualified full-check platform. - Eval harness: 27 tests passed; existing CI discovery ran 86 tests with two unrelated skips. Credential export rejection is tested against the actual report command. - Luna and OpenRouter Sonnet each passed 60 common-workflow runs: ten workflows, three repetitions per arm, zero unnecessary API fallback. - Sonnet passed 11 selected capability/contract cases after fixes. Gemini passed three smoke cases. DeepSeek exceeded the 120-second limit and remains unqualified. - Luna's two cost flags received focused follow-up. The original flags and a later n=1 latency flag remain visible. Sonnet had no cost or latency increase above 20%. - The catalog contains 785 entries; 58 were exercised across all stages. Most operation probes remain unrun and some need additional fixtures. Authored probes do not establish successful coverage. - Total conservative accounted cost: $9.875960. Active paid-campaign time: 88.16/90 minutes. No missing accounting. Later security and harness fixes have provider-free verification; no paid validation is claimed for those revisions. - Inspect the [qualification report](https://github.com/paperclipai/paperclip-evals/blob/codex/seach-call-api-tools/evals/runner-api-tools/reports/2026-09-07-production/READINESS.md) and [verification record](https://github.com/paperclipai/paperclip-evals/blob/codex/seach-call-api-tools/evals/runner-api-tools/reports/2026-09-07-production/verification.json). ## Risks - This is a broad authenticated API surface. Keep the default-off gate until an operator selects initial rollout companies. - Paid coverage is incomplete. Small regression samples do not prove all workflows are unchanged. - A timeout or server failure can follow a committed mutation. The result reports an unknown outcome and requires state inspection. - The new definitions add prompt tokens. The report retains cost flags and cache variation. - No database migration is required. - Repository rules require code-owner approval before merge. Technical CI and automated review are complete. ## Model Used OpenAI Codex based on GPT-6 assisted with code, tests and review. The exact serving model ID and context window are not exposed in this session. It used reasoning, tool calls and code execution. Eval models: `gpt-5.6-luna` with low reasoning, `openrouter/anthropic/claude-sonnet-5`, `openrouter/google/gemini-3.8-flash`, and `openrouter/deepseek/deepseek-v4-flash-0731`. Attempts retain runtime versions, model identity, usage and source provenance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
560e7e48b5 |
feat(runner): add SDK and developer tooling (#12608)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package already provides the production protocol and execution spine. > - Contributors still need stable SDK surfaces, deterministic test tools, and local inspection tools. > - Those surfaces share generated contracts and must change as one package boundary. > - This pull request adds the package-local SDK, labs, examples, and drift checks. > - The benefit is a reviewable developer platform that does not change application execution selection. ## Linked Issues or Issue Description **Subsystem affected** `packages/paperclip-runner` — runner SDK, conformance tools, and developer tooling. **Problem or motivation** The production runner spine is present, but package consumers cannot build deterministic integrations, inspect sessions, or verify provider-neutral behavior through supported surfaces. **Proposed solution** Add browser, React, standalone, live-session, scenario, conformance, and evaluation surfaces. Add generated contract inventories and package-local verification scripts. Keep production application routing unchanged. **Alternatives considered** We considered splitting each generated catalog, SDK surface, and demo into separate pull requests. Those changes share exports, fixtures, and drift gates. Splitting them would create intermediate package states that do not build. **Roadmap alignment** No overlapping item appears in `ROADMAP.md`. This work extends the runner package that is already on `master`. ## What Changed - Add browser, React, standalone, live-session, and issue-thread SDK surfaces. - Add deterministic mock control-plane, scenario, conformance, replay, and evaluation tools. - Add bounded Codex, OpenCode, and ACPX development transports and fixtures. - Keep deferred managed-provider execution fail-closed. Persisted compatibility data remains readable. - Add generated capability inventories with their source files and drift checks. - Add examples, package documentation, browser checks, and clean-consumer checks. - Preserve the reviewed protocol bounds, replay compatibility aliases, process environment isolation, and semantic redaction limits. - Update the ACPX package patch that the existing workspace patch registry already tracks. - Do not change `pnpm-lock.yaml`, repository workflows, server runtime selection, or the application UI. ## Verification GitHub Actions is the verification authority for this pull request. The repository CI, package TypeScript and Rust checks, package tests, generated-output drift checks, browser checks, security scans, and Greptile review must pass on the exact head. Local test suites were not run because this series uses parallel GitHub Actions for verification. ## Risks This is a large greenfield package change. The main risks are public export drift, generated-output drift, and optional React consumer compatibility. Package boundary checks, clean-consumer checks, and browser tests cover those risks. Production adapter selection and server execution are outside this pull request. ## Stack 1. **This PR:** runner SDK and developer tooling. 2. [Codex production server integration](https://github.com/paperclipai/paperclip/pull/12616). 3. [Provider-neutral task-thread UI](https://github.com/paperclipai/paperclip/pull/12617). ## Model Used OpenAI Codex, GPT-5, high-reasoning mode, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature request template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |