mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 11:13:44 +02:00
e095b84dabb34c790faeba36a0b8030cb4e17f44
1001
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e095b84dab |
feat(connections): connect services from native task feeds (#13058)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents use connections to reach external services. > - A fresh native task can have no service tools installed. > - The agent needs a way to discover services and ask the responsible person for access. > - This pull request brings the existing connection-intent flow into native task execution. > - The person can connect from the task, and the agent can continue with updated tools. ## Linked Issues or Issue Description **Subsystem affected** Native runner tool authority, connection intents, task interactions, and shared connection setup. **Problem or motivation** A task that needs an unconnected service cannot finish its work. Leaving the task to configure access also loses context. A resolved request must survive a restart and resume the correct agent once. **Proposed solution** Expose connection discovery and access requests as server-owned native tools. Render a durable task card and use the shared setup dialog. Persist outcome delivery and start a fresh provider session after access is ready. **Alternatives considered** Sending the person to the Connections page adds navigation and does not solve continuation. Polling for authorization consumes runs and can create duplicate requests. **Roadmap alignment** This extends the existing connection-intent runtime and setup experience. It reuses the shared access model and the native runner. Related: #12345, #12347. The service-slug fix in #12906 is related but separate. Companion evaluation PR: https://github.com/paperclipai/paperclip-evals/pull/21. ## What Changed - Expose `connections_search` and `connection_request` with server-bound company, task, agent, and responsible user. Preserve the legacy entry points. - Discover catalog services and authorized custom connections. Check installation, identity, health, and executable permissions before reporting ready. - Keep pending cards through ordinary messages. Reuse requests and retire stale ownership. Put Connect at the right of Not now. - Reuse the shared setup flow in a task dialog. Keep access additive and default to the requesting agent. Recover from cancelled or blocked OAuth windows with a new-tab fallback. - Persist outcome delivery with an idempotent wake key. Resume in a fresh session and recheck ownership before dispatch. - Add native browser fixtures, offline Storybook states, server contracts, and evaluation fixtures. Update guidance and documentation. ## Verification - `pnpm build`: passed after replaying the change on current master. - `pnpm -r typecheck`: passed. - `pnpm check:token-gates`: passed. - `pnpm --filter @paperclipai/ui build-storybook`: passed. - New continuation-policy regression cases: 16 passed. - Docker-backed PostgreSQL regressions passed for requester-only OAuth access, assignment-only expiry, terminal expiry, and credential-free setup metadata. - Shared setup and task-card UI tests: 121 passed, including configured MCP reconnect URL recovery and preserving user edits across refetch. - Storybook browser checks: all 119 passed on the latest reconnect fix. - `pnpm test:run`: 4,734 tests passed in the first server group, but embedded PostgreSQL startup failures and resulting cleanup errors prevented a complete local pass. All Linux CI lanes passed on the latest reviewed commit. One external-object route test returned an unexplained 500 on the first run; it passed twice locally and the failed shard passed on retry without code changes. - Earlier feature-checkout evidence: three deterministic native browser journeys passed, including restart delivery and an actual fixture tool result. Legacy scripted coverage also passed. All 59 added stories were inspected in light and dark themes. - Live Notion testing recorded successful provider reads. The manual test used a local-trusted instance. It does not prove authenticated/cloud deployment or every provider journey. - Native browser rerun reached the embedded PostgreSQL startup limit before bootstrap, so the latest checkout’s full native browser journey remains unverified. Both OAuth page/task regression cases passed against isolated Docker-backed PostgreSQL 17. They verify no premature task access, requester-only completion, additive retries, and reconnect preservation. - Applied both new migrations twice to isolated PostgreSQL 17. Foreign keys remained intact, duplicate active delivery keys were rejected, and failed delivery records did not block retries. Reviewer path: start a fresh test drive, enable the native runner, use an agent that can perform work directly, and ask it to summarize a Notion page. Connect from the card, then verify the resumed provider call and source-linked answer. The default test-drive CEO is instructed to delegate, so it can introduce an unrelated hiring step. ## Risks - Two additive migrations create durable deliveries and a partial unique wake index. They are idempotent. The wake index can require a maintenance window on large tables because migrations run in a transaction. - OAuth and continuation cross asynchronous boundaries. Tests cover ownership changes, retries, additive access, and restart delivery; live provider behavior still varies. - The latest requester-scope fix has not yet been exercised through live OAuth. GitHub, API-key, authenticated-user, and all recovery journeys are not claimed as verified. ## Model Used OpenAI GPT-6-based Codex assisted with implementation, tests, and review using tools and code execution. The runtime does not expose the exact model version, context window, or reasoning setting. Live evaluation used `gpt-5.6-luna`; manual native testing used `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used and disclosed unavailable runtime details - [x] I have checked ROADMAP.md and confirmed this extends existing connection work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [ ] I have run all required tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation - [x] I have considered and documented risks - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ebaeba40ee |
feat: simplify agent onboarding and configuration (#13011)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators create agents and configure their runtimes in the board UI. > - The old creation flow presents several choices and a large form before an agent can start. > - The existing onboarding controls already provide clear provider connection steps. > - This pull request uses those controls in a new-agent wizard and organizes the full configuration pages. > - Operators can connect, test, save, and assign a first task while keeping the existing configuration tools. ## Linked Issues or Issue Description Related: #10974. That earlier open PR also reorganizes agent configuration. This PR follows the reviewed Storybook designs for agent creation and the current configuration tabs. **What existing behavior does this improve?** Agent creation, provider connection, runtime tests, and full agent configuration. **Current behavior** The creation dialog leads to a large manual configuration form. Provider login controls differ from onboarding. Environment variables and secret access appear in separate places. **Proposed behavior** Choose a name and adapter. Connect Claude or Codex through the existing onboarding controls. Configure and test the runtime, save the agent, and open a task dialog with that agent assigned. Use the same design on the existing configuration tabs. **Reason and benefit** The first setup asks for fewer decisions. The full editor keeps instructions, skills, runtime controls, secret access, permissions, keys, and revisions available in clear sections. **Breaking changes** The board creation and configuration layouts change. The test-environment API adds an optional, allowlisted `testCredentials` field for one-shot probes. Database contracts stay the same. Native ACPX tests now reject unsupported local platforms before a CLI login can mask the runtime restriction. ## What Changed - Added a new-agent wizard with numbered steps, adapter branding, provider connections, editable model choices, runtime tests, and confirmation. - Added Codex app-server, Claude ACPX, and OpenCode runner choices. - Stored API credentials through existing secret APIs and persisted references in agent configuration. New setup keys are isolated from credentials used by existing agents. - Preserved external-agent invitations beside the wizard, including optional messages, one-time prompts, and clipboard fallback. - Added OpenRouter provider and secret bindings for Pi and OpenCode. - Added adapter-specific prerequisite fields for Cursor, Gemini, Kimi, and Hermes. Cursor Cloud keys are saved as new organization secrets. - Fixed Cursor Cloud repository field mapping, omitted empty remote environment values, and added useful model and repository error messages. - Preserved complete MCP assignments when multiple valid profiles contain more than 250 tools in total. Generated profiles retain exact tool selectors. - Added service branding and deployment-aware adapter choices. Cloud setup offers Claude, Codex, and OpenCode; local native runners require the experimental setting. - Made the agent list responsive at intermediate widths. - Applied the reviewed design to the real agent configuration pages. Kept the instruction editor, skills, and existing mutations. - Combined secret access and environment variables under one Save and Discard action. - Added interactive Storybook screens for setup, configuration, confirmation, authentication, and test results. - Fixed Pi provider-error parsing and thinking-effort persistence. Native ACPX validates Linux x64 on the actual local, SSH, or sandbox target. - Redacted the complete transient probe-credential field from HTTP error logs, including rejected provider names. ## Verification - Current head `df0292fe6` has a fresh Greptile 5/5 review with no unresolved findings. All 31 executed CI checks passed, including the aggregate verification gate and all browser E2E shards. Storybook visual regression is skipped by its workflow; the local Storybook build passed. - Browser tests completed real assigned tasks with direct Codex, Claude, OpenCode, Pi, and native Codex. - Verified external-agent invitation generation and automatic prompt copying in the live browser. - Pi and OpenCode used an existing OpenRouter secret. Browser checks covered save and reload, instruction edits, skill selection, environment-variable Save and Discard, and assigned task creation. - Invalid Claude API credentials remained on the connection step with an error. A live Pi/OpenRouter invalid-key probe returned a provider failure and left the user-secret inventory unchanged (zero entries before and after). - Full workspace typecheck and build passed after rebasing onto current master. After review fixes, server and UI typechecks, token gates, and the full build passed again. Storybook built successfully. - All 5,542 local UI tests passed. The Cursor Cloud and Pi adapter regressions passed all 24 tests. Review regressions passed 69 server tests and all 18 agent-list tests. - The local full test command ran 6,971 general server tests successfully. Editing review fixes during that long run caused nine tests to use stale modules; fresh isolated runs passed. An unrelated embedded-Postgres fixture hit the host shared-memory limit; its 15 affected tests passed when the fixture groups ran separately. - Local workspace groups passed after rerunning 18 CLI tests sequentially to avoid host database limits and parallel-load timeouts. The local full command stopped at the general server phase, so serialized server verification comes from the five passing CI shards. - Browser testing at 390px confirmed that the agent action menu opens and the page has no horizontal overflow. CI browser E2E shards passed. - Review the `Onboarding / New agent` and `Agents / Configuration refresh` Storybook groups. In the real app, create an agent, run its connection test, save it, assign a task, and reload its configuration. ## Risks - This changes the main agent setup and configuration UI. Regression tests cover routing, persistence, secret bindings, and form actions. - Native Claude ACPX requires Linux x64. Direct Claude works on macOS. Remote checks execute a bounded platform probe and reject unsupported or unverified targets. - A native OpenCode task reached the provider context limit because of its tool payload. Its provider connection test passed. Direct OpenCode completed a task. This existing native execution limit is not fixed here. - Claude and Codex connection keys use the existing user-secret store. Other runtime setup keys use distinct organization secrets. Existing credentials are never rotated. Probes do not store entered keys. Failed agent creation removes newly staged credentials. - Cursor Cloud has not completed a live task. Its authenticated account still needs GitHub repository access. The live run passed MCP provisioning, remote environment validation, and explicit Auto model selection before the repository prerequisite blocked execution. - Generated runtime MCP profiles can exceed the public profile-edit request limit. They still contain exact catalog selectors and preserve permission boundaries. - No database migration, dependency, lockfile, or workflow changes are included. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository tools, shell execution, and browser automation. The runtime did not expose the exact model ID or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f65991a5f1 |
refactor(server): move scheduled-retry and queued-run dispatch into a run-dispatch module (#12920)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server heartbeat service dispatches scheduled retries and queued runs. > - The service kept policy decisions and database writes in one large file. > - This layout made policy branches harder to test and transaction boundaries harder to inspect. > - This pull request moves the policy rules and database transactions into a run-dispatch module. > - The benefit is a smaller service, pure policy tests, and clear transaction ownership. ## Linked Issues or Issue Description **What existing behavior does this improve?** The server heartbeat service promotes scheduled retries and cancels stale queued runs. **Subsystem affected** server/ — REST API and orchestration services. **Current behavior** The heartbeat service contains the policy rules and the database writes for these dispatch paths. **Proposed behavior** A run-dispatch module owns pure policy functions and semantic database transactions. The public service contracts stay unchanged. **Reason and benefit** The new layout separates branch rules from database effects. It makes each policy branch easier to test and keeps each operation’s row writes in one transaction. **Breaking changes** None. The public service contracts stay unchanged. ## What Changed - Move scheduled-retry promotion and queued-run staleness rules into pure functions. - Add table-driven unit tests for each policy branch. - Move promotion and cancellation writes into semantic transactions. - Keep row locking, company isolation, and post-commit effects unchanged. ## Verification - `node scripts/check-module-boundaries.mjs` passes. - `tsc --noEmit` from `server/` reports no errors. - The focused server test command passes 228 tests in six files. - Full pull request CI passes. - Greptile reports 5/5, and all review threads are resolved. ## Risks The main risk concerns changed transaction boundaries in scheduled-retry promotion and queued-run cancellation. The focused tests retain coverage for locking, transactionality, company isolation, and transport contracts. The public service contracts do not change. ## Model Used Codex, GPT-5, with code execution and tool use. The context window is not provided by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I have addressed all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5752d6bd93 |
fix(heartbeat): block runs on a stuck sandbox plugin and re-enable errored bundled plugins at boot (#12957)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents run inside environments. A sandbox environment gets its
sandbox from a provider plugin (for example the bundled
`paperclip.kubernetes-sandbox-provider`), and every run starts by
acquiring a lease through that plugin.
> - When a plugin activation fails once (on a hosted deployment: one
`RPC call "initialize" timed out after 15000ms`), the loader calls
`markError`. That persists `status = error` on the plugin row and
switches off worker auto-restart. Boot activation (`loadAll`), the
bundled-plugin bootstrap and the lazy worker recovery all consider only
`ready` plugins, so the plugin stays in `error` across restarts until an
operator enables it by hand.
> - Every run that needs the provider then fails before dispatch with
`Sandbox provider "kubernetes" is installed via plugin "...", but that
plugin is currently error.` That message matches neither the retryable
classifier (`... but its worker is not running`) nor any configuration
classifier, so the run is recorded as a plain `setup_failed`, the issue
is released, and the scheduler dispatches the same failing run again on
the next tick. On the hosted deployment one company produced about
11,300 identical failed runs, one every 30 seconds, for a week (#12953
is a customer's report of the same condition).
> - Two gaps cause this: the heartbeat treats a condition that only an
operator can change as a transient setup failure, and the bundled-plugin
bootstrap never gives a plugin in `error` another chance even though the
bundle ships with the release image.
> - This pull request classifies the "installed but not ready" lease
failure as `configuration_incomplete`, so the existing recovery path
moves the issue to `blocked` with one recovery action and an actionable
notice; and it re-enables a bundled plugin found in `error` once per
boot, so the next server restart heals the plugin.
> - The benefit is that a stuck provider plugin surfaces as one blocked
issue per task with clear next steps, instead of an endless stream of
identical failed runs, and a restart repairs the plugin without an
operator having to know the plugin API.
## Linked Issues or Issue Description
- Refs #12953 — hosted report: "that plugin is currently error" on every
run for six days, including runs that were retried by hand. This PR
stops the retry loop (issue goes to `blocked`) and makes a server
restart re-activate the bundled plugin. It does not change how a managed
Kubernetes environment is provisioned for a company, which the same
report also mentions.
- Related PR: #9760 pauses the agent for the permanent `Adapter "..." is
not in the configured adapter registry` setup failure. This PR handles a
different permanent condition (plugin not `ready`) and routes it through
the existing `configuration_incomplete` recovery path (issue-level block
with a recovery action) rather than an agent-level pause, because the
gap is on the plugin, not on the agent. The two do not overlap in code
paths.
- No existing issue covers the bundled-plugin re-enable. Bug
description:
**What happened**
A bundled sandbox provider plugin went to `status = error` after one
failed activation. It stayed in `error` across every later server
restart. Every run for every agent on that provider failed lease
acquisition in under a second with `... but that plugin is currently
error.` (`setup_failed`), and the heartbeat kept dispatching new runs
that failed the same way.
**Expected behavior**
A run that fails because its provider plugin is not `ready` is recorded
as a configuration gap and the issue is moved to `blocked` with a notice
that names the plugin and its status, so no further runs are dispatched
until an operator acts. A bundled plugin left in `error` gets a fresh
activation attempt on the next boot.
**Steps to reproduce**
1. Install a sandbox provider plugin and create a sandbox environment
that uses it; make it an agent's default environment.
2. Set the plugin row's status to `error` (or make its worker fail
`initialize` once so the loader does it).
3. Assign an issue to the agent and let the heartbeat run it.
4. Observe: the run fails with `... but that plugin is currently error.`
as `setup_failed`, the issue is released, and the next tick dispatches
another run that fails the same way. Restart the server: the plugin is
still `error`.
**Paperclip version**
master at
|
||
|
|
c723bb4dfc |
fix(skills): reuse validated runtime revisions during preparation (#13042)
## Thinking Path > - Paperclip manages AI agents and prepares their runtime inputs before each turn. > - Shared company skills are part of those inputs for native and legacy adapters. > - Runtime materialization refreshed the full inventory again for every declared file. > - Remote skill directories were also downloaded and rebuilt on every turn. > - Measured preparation took 42–73 seconds while runner execution took 7–9 seconds. > - This change reads the inventory once and reuses validated installed revisions. > - Agents retain their selected skills while repeated preparation avoids upstream work. ## Linked Issues or Issue Description **What happened?** One 114-skill preparation performed 407 inventory refreshes, 48 directory rebuilds, and 388 GitHub file fetches. Reusing existing local copies took 151 ms. **Expected behavior** Each listing refreshes inventory once. Unchanged installed remote revisions reuse complete, validated local copies. Local edits remain visible. Explicit updates select new revisions. **Steps to reproduce** 1. Import GitHub skills with supporting files. 2. Run an agent turn, then run another with the same installed revisions. 3. Observe repeated inventory scans, downloads, and runtime directory replacement before execution. Related prior attempts: #2330 and #9268 (still open; #9268 last updated July 9). Those use a marker compared with `updatedAt`. This patch follows the required content validation, immutable revision, company isolation, atomic publication, and read-only semantics, and removes refresh-per-file multiplication. ## What Changed - Split public file reading from reading an already loaded skill. Runtime listing refreshes inventory once. - Add a company-scoped revision cache with file manifests outside the delivered skill directory. Fingerprints omit cosmetic metadata. - Validate exact file inventory, sizes, and hashes before warm reuse. Reject traversal and symlinks. Stage complete builds and serialize atomic publication across processes. - Preserve local/catalog direct sources, stored Markdown fallback, explicit version snapshots, and legacy mutable-ref compatibility. Report missing supporting files and keep older valid revisions readable. - Clean both runtime layouts on rename/removal and record `skills.prepare` under preparation timing. - Add service/cache regressions and an isolated 114-skill benchmark, including a new-process warm run. ## Verification - Final targeted skill-service/cache/trace validation: 86 tests pass (61 embedded-PostgreSQL service tests, 19 cache tests, 6 trace tests). Database tests executed rather than skipped. Focused skill routes, adapter selection, and native runtime context also pass. - `pnpm -r typecheck` and `pnpm build` pass locally at `22caa1fe4`. - The full `pnpm test:run` matrix passes on supported Linux CI at the final head: [CI run](https://github.com/paperclipai/paperclip/actions/runs/34236097762). Local full-suite execution encountered PostgreSQL startup contention, a random allocated-port boundary, and a socket hang-up; every affected suite passed on an isolated rerun. The interrupted local serialized run is not claimed as a complete local pass. - Repeatable benchmark: `pnpm --filter @paperclipai/server exec tsx ../scripts/benchmark-skill-preparation.ts`. Mixed 114-skill inventory with 429 remote files on Linux: cold 286 ms, warm median 96 ms / maximum 153 ms including a new process. Every warm sample performs one refresh, zero upstream fetches/rebuilds, and reports no missing entries; content assertions pass. - Controlled deployment against the previously deployed revision completed with zero lost runs. Real inventory: 114 skills, 670 declared files; 402 cached files match the prior installed copies byte-for-byte. Ten post-deployment warm preparations: median 129 ms / maximum 208 ms; new-process warm 194 ms, zero downloads/rebuilds/missing entries. - Five sequential real browser questions persisted in 10.6–20.7 s (median 12.2 s), versus 50–83 s before. Skill preparation median 240 ms, with one 2.37 s outlier. Total preparation median 3.337 s / maximum 8.728 s **does not fully meet** the <3 s / <5 s target. The excluded historical-run redaction query takes about 1.36 s per scan at two preparation call sites; wider application latency coincided with the outlier, without a cache rebuild. These residuals are reported rather than discarded. - Disposable skill reimport verified through actual selected-skill runs: the next run read the changed code. Fixture removed and agent configuration verified unchanged. - Greptile 5/5, zero unresolved review threads, all final-head CI checks green. ## Risks - Cold preparation still requires upstream availability for supporting files. An unavailable revision is reported missing and never falls back to an older revision. - Valid older revisions and quarantined invalid entries consume additive disk space until skill cleanup. An abruptly killed publisher can leave a lock that requires operator cleanup after confirming its PID is dead. - Warm validation reads all cached file bytes. Very large inventories still have proportional local I/O cost. - No HTTP API, schema, agent configuration, or first-party Telemetry changes. OpenTelemetry retains its operator endpoint gate. ## Model Used OpenAI GPT-6 in Codex, with reasoning, repository inspection, code editing, and test execution. The exact serving snapshot and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted and isolated reruns; full Linux CI matrix passes, local full-run caveats above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b97101893f |
feat(projects): select multiple GitHub source repositories (#13010)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Projects give tasks a common source repository and execution context. > - The current project form asks for a raw URL and unrelated metadata. > - Teams need to select several repos from GitHub connections they can use. > - This pull request implements the reviewed project form and repository editor. > - The server checks credential ownership and shared audiences before discovery. > - Existing workspace URLs and runtime identity rules remain compatible. ## Linked Issues or Issue Description **Problem or motivation** Project creation accepts one raw repository URL. It does not help users select repos from their usable GitHub connections or attach several repos together. **Proposed solution** Add a shared GitHub repository picker to project creation and Configuration. Support multiple selections, transactional persistence, and the existing GitHub setup flow. Simplify the project form and Configuration tab as reviewed. **Alternatives considered** Keep a raw URL field or add a separate repository table. The existing workspace collection already supports several repositories and keeps legacy URLs compatible. **Roadmap alignment** This builds on the shipped MCP Tool Gateway and Apps capability. It does not change runtime credential delegation. Related work: #11662 addresses the existing dialog's viewport limits. #4552 addresses generic Git URLs; this change preserves those URLs in existing workspaces. ## What Changed - Add company-scoped repository discovery from usable personal and shared GitHub grants, with provider-ID deduplication, PAT pagination, and partial failure handling. - Document the repository endpoints and board access requirements in OpenAPI. - Validate new selections and save projects with multiple repository workspaces in one transaction. Preserve legacy URLs and existing selections whose access was lost. - Implement the reviewed Create project dialog, shared repository editor, scrolling, and mobile layout. - Move repositories above environment variables, remove Status and Goals controls and env help paragraphs, move Created to the bottom, and redirect Overview to Configuration. - Reuse GitHub setup in dialogs, preserve project drafts, and verify popup completion through the API. - Replace the configuration story's DOM adapter with explicit production composition. Keep the reviewed mobile and short-viewport stories. ## Verification - Passed: `pnpm build`, `pnpm -r typecheck`, `pnpm build-storybook`, and `pnpm check:token-gates`. - Passed: focused repository access, database persistence, configuration, and connection setup tests. - Passed: `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/project-repositories.spec.ts`. - The browser tests use a real temporary server/database. They cover create, forty persisted repos, mobile scrolling, save/reload, legacy URL editing, and rejection without a partial project. - GitHub responses and popup completion use deterministic fixtures. No real GitHub account was authorized by the test suite. - All CI general, serialized server, and browser test shards pass on the final commit. - The local full-suite run overlapped review edits and was stopped; fresh repository, OpenAPI, UI/CLI, and connection tests pass. Unrelated local worker, built-in-agent, and routine timing/socket failures passed isolated reruns. - Final commit `1b3308dca`: all CI gates pass, including build, runner verification, typecheck, canary dry run, and security checks. Greptile is 5/5 with no unresolved review threads. - Storybook visual regression is opt-in and was skipped by CI; the Storybook build passed locally. ## Risks - Repository discovery depends on provider availability. Failed connections are reported while successful results stay usable. - Selections identify source workspaces; they do not grant agents new credentials. The existing primary-workspace and responsible-user identity rules still apply. - No database migration is needed. Existing API status, goals, dates, and manual workspace URLs remain supported. ## Model Used OpenAI Codex, based on GPT-6, with repository inspection, code execution, and browser tools. The runtime does not expose a more specific model deployment ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
297d8741f5 |
fix: resolve duplicate connections to the same GitHub account (#13022)
## Thinking Path > - Paperclip lets people share agents while keeping GitHub access personal. > - Each managed Git or GitHub operation selects an eligible connection grant. > - Connecting the same GitHub account twice creates two grants. > - The old resolver counted grants and rejected them as competing identities. > - Managed commands then ran anonymously and reported a misleading login failure. > - This change compares GitHub account IDs and selects one eligible grant for the same account. > - The benefit is reliable access after reconnecting, with clear diagnostics for real failures. ## Linked Issues or Issue Description Refs #13005. **What happened?** Two active connections owned by one Paperclip user pointed to the same GitHub account. Managed Git refused both as ambiguous. The agent could not push, although the account was connected and had repository access. **Expected behavior** Multiple grants for the same GitHub account resolve to one eligible authorization. Different accounts remain ambiguous. Unavailable access explains its cause without blocking unrelated work. **Steps to reproduce** 1. Connect the same GitHub account twice for one Paperclip user and allow the shared agent through both connection audiences. 2. Start an instruction as that user. 3. Run managed gh or git push. Before this fix, no credential is provided. ## What Changed - Compare stable GitHub account IDs when more than one eligible grant exists. Never deduplicate by login alone. - Prefer an available grant, then the newest authorization with a stable ID tie-breaker. Refresh and webhook timestamps do not change the selection. - Keep the selected credential and connection policy together. Do not combine permissions or fall back from a dedicated account to a personal account. - Print the redacted unavailable reason in managed command output. Unrelated local operations still work anonymously. - Add database and executable launcher regressions, and document selection behavior. ## Verification - Final `pnpm -r typecheck` and `pnpm build` passed. - Fourteen operation credential integration tests passed, covering duplicate personal/dedicated grants, incomplete credentials, distinct accounts with the same login, missing identity metadata, revocation, membership, connection audiences, and A → B → A steering. Existing Git credential and gateway suites and both executable launcher tests also passed. - The local broad test run encountered three embedded-Postgres lifecycle timeouts and stale modules from edits made during that run. A fresh process rerun of all four affected suites passed all 35 tests. The full Node 24 CI test matrix passed on the final commit. - CI passed all 31 checks on `797973b30beb16ba5fa69ed281835e1ab812b449` (Storybook visual regression was correctly skipped). An unrelated Company Settings UI test failed once; the focused local reproduction and rerun of its CI shard both passed without code changes. - Fresh Greptile review of the final commit: 5/5, with no open findings. Security checks passed. - Live acceptance passed with both duplicate connections enabled: managed `gh api user` returned the expected account, managed `git push` succeeded, and the agent created #13023 and pushed its review fixes. No host login or credential changes were used. - Applied the final source/compiled patch to the affected instance with backups, after confirming no runs were active. Restarted service health and the final resolver selection were verified. The patch is an overlay on the existing deployment; this PR supplies the upstream fix. ## Risks The resolver selects one authorization for an already permitted GitHub account. It does not combine repository permissions across connections. If the selected authorization has narrower access, that operation can still be denied by GitHub. Different provider account IDs and unknown duplicate identities continue to fail closed. No schema, host credential, or connection permission changes are included. ## Model Used OpenAI GPT-6 through Codex assisted implementation and verification with shell, database, and browser tools. The exact model variant and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1cc45086d3 |
feat: use the responsible person's GitHub for shared agent operations (#13005)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Several people can send instructions to the same agent and task. > - A fixed GitHub token in the provider process can keep the first person's access after another person's message is accepted. > - Task ownership cannot select credentials for each accepted instruction or preserve the identity of an operation already in progress. > - This pull request records ordered execution identity contexts and resolves credentials when managed Git, gh, or GitHub tools start. > - The benefit is automatic personal GitHub access for shared agents, with durable continuation rules and no teammate credential fallback. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: orchestration, connection grants, database, runtime adapters, native runners, and run details. **Problem or motivation** A shared agent must use the person whose instructions it has accepted. A queued message must retain its author. A retry or approval without new instructions must retain the originating identity. GitHub must remain optional for ordinary work. **Proposed solution** Persist execution identity separately from task ownership. Give new processes a run-scoped broker capability and token-free managed launchers. Capture identity at operation start. Keep an explicit dedicated-agent grant as an override. Show redacted diagnostics in run details. **Alternatives considered** Per-task ownership, fixed provider tokens, and mutable repository author configuration do not handle accepted steering or concurrent operations. A manual account-selection action would add unnecessary setup to each turn. **Roadmap alignment** This completes the existing Multiple Human Users, MCP Tool Gateway & Apps, Secrets Manager, and Self-healing Runs capabilities. The implementation follows the maintainer-approved plan. Related work: Refs #12843, Refs #12907. Existing proposals #4618 and #8945 cover per-agent or per-worktree author configuration. This change instead follows the accepted human instruction across runtime types. Refs #11831 for governed personal connection delegation; this change preserves connection audience checks and does not use standing delegation as a personal credential fallback. ## What Changed - Add durable, ordered identity contexts and active run references. Preserve message authors through consolidation, steering, retries, delegation, approvals, routines, and restart. - Add an authenticated operation-time GitHub credential broker and local/remote managed git and gh launchers. Keep personal tokens out of the long-lived provider process. - Resolve GitHub gateway and server-side Git operations through the same responsible-person or dedicated-grant selection rules. - Make absent and unavailable GitHub credentials non-blocking at generic startup. Clear host and prior-person credentials. Keep anonymous Git access where supported. - Add run-detail identity history and the dedicated-account warning. Keep task ownership and queue-versus-steer decisions unchanged. - Preserve personal OAuth declarations through connection edits. Retain exact selected grants in the gateway. - Fix continuation races found during real acceptance: verify a warm owner before credential rotation, and wait for bounded durable runner suspension before the next run starts. - Make migrations replay-safe. Retain identity through agent/run deletion, remove it with its company, and clean terminal launcher directories before releasing execution environments. Document coordinated release and rollback. ## Verification - Full workspace typecheck, build, and token gates passed. The complete local suite passed in its normal test groups: 17,120 passing tests, including all 143 serialized server suites. After integrating the newly merged runner API work, full local typecheck and build passed again, along with 890 focused integration tests. All 31 checks on the integrated revision passed, including build, browser E2E, release registry, canary dry run, typecheck, security and all test suites. Greptile is 5/5 with all review threads resolved. - Current focused checks passed: 142 native executor tests, 67 runtime lifecycle tests, 9 durable identity tests, 75 credential/routine tests, 19 low-trust/resumption tests, and the executable migration replay test. - Authenticated browser acceptance with two Paperclip users and two GitHub accounts on one shared native agent passed. Real commits and pushes followed A → B accepted steering → queued A continuation in the same saved conversation. GitHub commit author and committer identities matched all three operations. Both runs succeeded and task ownership stayed unchanged. - Real GitHub MCP calls switched from A to B after accepted steering. A delegated subtask retained its originating identity across a server restart. - Disabling B's GitHub connection left ordinary work successful. Managed gh was unauthenticated and the provider had no inherited GH_TOKEN or GITHUB_TOKEN. - The browser displayed run-detail diagnostics and the exact dedicated-account warning. A final controller-restart check followed by another-person continuation retained the conversation, selected the correct GitHub login and Git author, and removed each terminal launcher directory. - Company-lifetime migration and all five previously failing CI suites passed locally (167 tests). Same-token gateway A → B → A and six broker/launcher boundary tests passed. - Remote callback, launcher, sandbox, and runtime contract tests passed. Both native and legacy Codex completed actual Daytona executions on the integrated revision ([campaign results](https://github.com/paperclipai/paperclip/actions/runs/34155056509)). The remote package-manager shim staging regression also passed locally. ## Risks - Deploy the migrations, server broker, launchers, and runner artifacts together. Existing processes finish with their original contract. New managed processes need the broker endpoint for GitHub operations. - Finish or stop new managed executions before rolling application code back. Keep the additive schema and identity history during rollback. - Scripts that require a persistent raw GH_TOKEN must use managed git, gh, or GitHub gateway tools. Run capabilities authorize code executing within that run to acquire its current identity; this is not hostile-code isolation within one execution principal. Managed commands prevent automatic credential carryover; arbitrary code deliberately copying a credential is outside that boundary. - Uncertain steering acknowledgement deliberately holds new credential acquisition until reconciliation. Already-started operations retain their captured identity. - GitHub private access and provider outages can still fail the specific operation that needs them. Dedicated grant failure does not fall back to personal access. ## Model Used OpenAI GPT-6 through Codex assisted implementation, review, shell execution, and browser acceptance. The exact model variant and context-window size are not exposed in this session. Tool use included TypeScript and Rust tests, database integration tests, GitHub CLI, and authenticated browser control. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5bddff0920 |
feat(runner): add guarded API search and call fallback (#13003)
## Thinking Path > - Paperclip manages AI agents and their work. > - The new runner gives agents dedicated tools for common tasks. > - Some API operations and parameters have no dedicated tool. > - Agents need a controlled way to find and use those operations. > - This pull request adds API search and calls through the real server routes. > - Existing tools remain the preferred path. The new tools are disabled by default. > - Paired tests measure correctness, tool choice, cost and time. ## Linked Issues or Issue Description **Subsystem affected** Paperclip Runner contracts, production tool authority and the server API catalog. **Problem or motivation** The runner cannot use much of the API described by the old Paperclip skill. A generic HTTP client would also let agents bypass runner control rules. **Proposed solution** Add `search_api` and `call_api`. Resolve calls from the mounted API catalog. Use server-held, run-bound credentials. Preserve route checks and runner lifecycle rules. Keep the tools disabled until an operator enables selected companies. **Alternatives considered** A dedicated tool for every endpoint would add a large initial prompt. An unrestricted HTTP tool would weaken authorization and replay controls. **Roadmap alignment** This extends the native runner tooling. The repository owner requested this design and implementation. The roadmap and related open PRs were checked. No duplicate API escape-hatch PR was found. ## What Changed - Register two compact fallback tools in canonical contracts and provider projections. - Build deterministic API discovery from OpenAPI, mounted experimental routes and the old skill reference. - Execute bounded JSON, text, file and download requests through authenticated HTTP routes. - Recheck active runs, company access and work modes. Block runner lifecycle, scheduling, credential and approval bypasses. Keep routine annotation collaboration available. - Retain mutation receipts. Report uncertain outcomes without blindly repeating writes. - Add a company rollout gate and a durable eval worker with complete cost accounting checks. - Record child-task creation in the activity log with the agent and run. - Add contract, authorization, file, replay and real runnerd/PRP/HTTP tests. - Document rollout gates and paid coverage limits. The companion eval repository retains immutable attempts and reports. ## Verification - Final app commit `da58370524c3626a744eec20164397c5fb6ba9ef`: all 32 checks passed; the unrelated Storybook visual check was skipped. Greptile 5/5; no unresolved review threads. - Full Linux build and recursive typecheck passed. Repository tests were run by project and serialized shard; all 143 serialized server suites passed. - Runner TypeScript: 1,599 passed, two skipped. Rust release: 451 passing test reports. Conformance and replay parity passed. The required API check passed 837 tests, including runnerd → PRP → authority → real HTTP. - Bindings cannot enable API tools without the explicit deployment flag. Unit and real-authority tests prove the default-off boundary. - The standalone API check builds and stages its own binary. It passed after existing staged and debug binaries were removed from the test container. - UI and CLI tests passed. Initial environment failures (missing jq, Docker overlay file identity, and parallel linker memory pressure) and focused passing reruns are retained. The macOS full runner suite has platform-specific failures; Linux is the qualified full-check platform. - Eval harness: 27 tests passed; existing CI discovery ran 86 tests with two unrelated skips. Credential export rejection is tested against the actual report command. - Luna and OpenRouter Sonnet each passed 60 common-workflow runs: ten workflows, three repetitions per arm, zero unnecessary API fallback. - Sonnet passed 11 selected capability/contract cases after fixes. Gemini passed three smoke cases. DeepSeek exceeded the 120-second limit and remains unqualified. - Luna's two cost flags received focused follow-up. The original flags and a later n=1 latency flag remain visible. Sonnet had no cost or latency increase above 20%. - The catalog contains 785 entries; 58 were exercised across all stages. Most operation probes remain unrun and some need additional fixtures. Authored probes do not establish successful coverage. - Total conservative accounted cost: $9.875960. Active paid-campaign time: 88.16/90 minutes. No missing accounting. Later security and harness fixes have provider-free verification; no paid validation is claimed for those revisions. - Inspect the [qualification report](https://github.com/paperclipai/paperclip-evals/blob/codex/seach-call-api-tools/evals/runner-api-tools/reports/2026-09-07-production/READINESS.md) and [verification record](https://github.com/paperclipai/paperclip-evals/blob/codex/seach-call-api-tools/evals/runner-api-tools/reports/2026-09-07-production/verification.json). ## Risks - This is a broad authenticated API surface. Keep the default-off gate until an operator selects initial rollout companies. - Paid coverage is incomplete. Small regression samples do not prove all workflows are unchanged. - A timeout or server failure can follow a committed mutation. The result reports an unknown outcome and requires state inspection. - The new definitions add prompt tokens. The report retains cost flags and cache variation. - No database migration is required. - Repository rules require code-owner approval before merge. Technical CI and automated review are complete. ## Model Used OpenAI Codex based on GPT-6 assisted with code, tests and review. The exact serving model ID and context window are not exposed in this session. It used reasoning, tool calls and code execution. Eval models: `gpt-5.6-luna` with low reasoning, `openrouter/anthropic/claude-sonnet-5`, `openrouter/google/gemini-3.8-flash`, and `openrouter/deepseek/deepseek-v4-flash-0731`. Attempts retain runtime versions, model identity, usage and source provenance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f6a211479f |
fix: share current CLI runtimes across sandbox adapters (#12994)
## Thinking Path - Paperclip Runner needs its runtime preinstalled for fast sandbox startup. - Native and local adapters should launch one current CLI installation per provider. - An older global copy can shadow that installation, and exact native compatibility pins must match it. - Update the qualified releases and binary digests, expose shared CLI entrypoints from the provider pack, and prefer the image-owned bin directory. - Keep dependency installation in the image build; task startup only discovers, links, and verifies artifacts. ## Linked Issues or Issue Description **What happened?** Remote native startup rejected a stale global Codex, while CLI-only images lacked runnerd entirely. **Expected behavior** An image-baked runtime starts without uploading binaries or installing packages. All adapters share the same current provider CLI. **Steps to reproduce** Start a native remote task with the old global Codex and the updated runtime available only under `/opt/paperclip-runner/bin`. **Paperclip version or commit** Discovery behavior at `54a99d884`. **Deployment mode** Docker with a remote sandbox. ## What Changed - Prefer `/opt/paperclip-runner/bin`, then the user's local bin directory, then PATH. Existing metadata and version validation remains in force. - Qualify Codex 0.153.4, OpenCode 1.18.29, and Claude SDK 0.3.263 / CLI 2.1.263. Update binary digests, TypeScript/Rust checks, registry defaults, and the displayed OpenCode version together. - Share Codex and Claude's native executable with the ACP bridges through exact dependency overrides. Preserve the separately qualified ACP bridge implementations and their security patches. - Expose shared provider-pack CLI launchers; fail the pack build if Codex ACP resolves a separate Codex installation. Update the eval image's other agent CLIs to current stable releases and remove duplicate global provider installs. - Document the single-current-CLI policy in source comments and development guidance. Latest stable releases are resolved at review/build preparation and pinned; task startup never auto-updates. ## Verification - Native-session and adapter-registry suites: 158 tests passed. - Provider suites: 88 tests passed, 7 Linux-only checks skipped on macOS. One existing macOS temporary-path alias assertion passed when rerun with canonical `TMPDIR=/private/tmp`. - Package-contract and OpenCode materialization tests: 11 passed. - Full typecheck, build, and token gates passed. Rust native-provider/recovery tests: 19 passed. - Broad local suite: 5,974 passed, 23 failed, 41 skipped. Failures are in unchanged macOS workspace/path/port and connection suites; focused runtime tests pass. All latest-head Linux PR checks passed, including the full test shards, typecheck, build, runner verification, browser suites, and canary dry run. - The standalone fleet image built with one current provider CLI each and passed native Codex/Claude binary-integrity checks. A disposable Daytona sandbox reported ready in 798 ms; its baked runner completed an API-key `gpt-5.6-luna` turn in 2,430 ms and returned the expected marker with a usage receipt. No runtime artifacts were uploaded or installed. - The normal shared `codex exec` entrypoint also completed an API-key `gpt-5.6-luna` turn in 2,321 ms. - Both image builds verify the complete generated lockfile against a reviewed SHA-256 before package installation or lifecycle execution. Root lockfile changes remain CI-owned. Merge and rollout remain on hold for operator review. ## Risks - Updating provider CLIs changes their behavior for all adapters; version probes and live native smoke testing are required before image promotion. - The image-owned directory takes precedence. Its entries must launch the same shared CLI as the global PATH, not a private older/newer copy. - Application qualification pins and the deployed image must move together. No startup fallback installation is added. - No schema or authentication-policy changes. ## Model Used OpenAI GPT-6 (Codex). The session does not expose a more specific model ID or context-window size. Used reasoning, repository inspection, code execution, and browser verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
932ddb7b37 |
feat: browse GitHub repository access across organizations (#12998)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - GitHub connections give agents access to approved repositories. > - One GitHub identity can use installations across several organizations. > - The permissions page linked to one installation and showed an unfiltered list. > - Users could not easily find another organization or inspect a large selection. > - This change adds account filtering, search, and access configuration links. > - Users can inspect repository access in one compact view. ## Linked Issues or Issue Description Refs #12993. Related repository-catalog work in #11228 and #11234 was checked. This change only improves the existing GitHub connection permissions page. **What existing behavior does this improve?** The GitHub connection permissions page and its repository display metadata. **Current behavior** The page links directly to an existing installation. The repository list has no account filter, search, height limit, or private-repository marker. Refresh access occupies a separate section. **Proposed behavior** Show all authorized repositories by default. Filter by account or organization and search by name. Open GitHub's account chooser to configure access across organizations. Show GitHub icons and private-repository locks. Keep refresh beside configuration and limit the visible list to about ten rows. **Reason and benefit** Users can find repositories across organizations and configure missing access without creating another GitHub identity. Large repository lists no longer fill the page. **Breaking changes** None. Repository display metadata gains an optional private flag. Older snapshots remain valid and gain the flag after access refresh. No SQL migration is required. ## What Changed - Add an All accounts view, account filter, search, and empty states. - Link both configuration controls to GitHub's app account chooser. - Place an accessible refresh icon beside the configuration button. - Keep the repository heading and list in one section. - Add GitHub icons and private-repository locks. - Cap the scrollable list at ten rows using a design token. - Persist GitHub's private flag only when the provider returns a boolean. - Recover missing legacy app configuration from GitHub installation metadata. - Update tests and the GitHub connection runbook. ## Verification - Focused tests passed: 54 permissions-page tests and four GitHub metadata tests. - UI and server typechecks passed before submission. Token gates passed. - Browser checks verified account filtering, search, empty results, and the configuration destination. - The live list contained 40 repositories. Its final height was 272 pixels, which fits ten single-line rows with gaps. Scrolling retained all rows. - A live access refresh populated 30 private-repository lock icons from GitHub metadata. - Full workspace typecheck and build passed. The broad local suite stopped in the general-server group with 18 failed files. Failures include macOS temporary-path handling and embedded PostgreSQL startup. That run also overlapped the legacy fix and retained a stale GitHub module; the final focused run passed all 58 tests. Clean-runner CI is tracked separately. - Latest-head review is 5/5 with the legacy chooser finding resolved. All CI checks passed on commit `0ff2b63f348f5c87d8b7df6e43388f60f5d872d9`, including build, typecheck, all test shards, browser tests, and canary dry run. ## Risks - Older repository snapshots lack visibility metadata until refreshed. Unknown visibility does not display a lock. - The account filter lists authorized installation owners. Users add other organizations through GitHub's chooser. - Filtering changes only the displayed list. GitHub remains authoritative for repository access. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) via Codex. Reasoning, code execution, and browser tools were used. The exact context window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bac60d9d31 |
fix: preserve GitHub sign-in and show connected repository access (#12993)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - GitHub connections give agents an account with selected repository access. > - Fresh local instances enroll with production Paperclip Cloud. > - Enrollment could finish while the GitHub OAuth profile remained disabled. > - Setup then switched to a personal access token form without explanation. > - This change preserves sign-in intent and shows the connected account and repositories. ## Linked Issues or Issue Description Related: #12907, #12943, #12947. Existing open GitHub connection work was checked. No duplicate was found. **What happened?** After Cloud enrollment, a fresh test-drive asked for a GitHub key. Production did not advertise the managed GitHub profile. Staging did. The permissions page also omitted the authenticated username and repository names. **Expected behavior** Continue with GitHub OAuth when available. Explain unavailable sign-in and allow retry otherwise. Show the GitHub username and complete accessible repository list. **Steps to reproduce** Start a fresh test-drive. Choose GitHub and complete instance enrollment while the Cloud GitHub profile is disabled. Open an existing GitHub connection's permissions page. ## What Changed - Preserve managed sign-in intent when the gallery omits its profile. - Refresh the selected gallery entry on retry without resetting the audience. - Fetch all pages of GitHub installations and repositories. - Store only repository IDs, full names, and installation IDs in grant metadata. - Show the GitHub username, repository list, management link, and refresh action. - Discard the repository snapshot after newer installation lifecycle events. Preserve snapshots verified after delayed events. - Lock and re-read grant metadata when applying installation events or saving refreshed access. Patch only webhook fields for other events. Reject snapshots if access changed during the external fetch, using unique access revisions even when timestamps collide. - Show repository installation recovery for managed OAuth even when the app also offers an advanced PAT method. - Update tests and the GitHub connection runbook. No SQL migration is required. ## Verification - Local typecheck, build, and token gates passed. All latest-head CI gates passed, including the complete test matrix and browser suites. Greptile is 5/5 with no unresolved findings. - All 382 focused setup, permissions, metadata, service, and webhook tests passed across final runs. One socket-hang-up test passed on rerun with the full service suite. Final service, metadata, and webhook checks passed all 230 tests. - The broad local suite was stopped after failures. Seven workspace-runtime exposure and control-conflict failures reproduce on base commit `54a99d884`. The broad run also overlapped local iteration; final focused tests and clean-checkout CI are tracked separately. - Browser: a fresh production-backed instance completed enrollment, retried after profile enablement, reached GitHub consent, recovered from a missing installation, and completed OAuth. - Browser: the permissions page showed the authenticated username and the selected private test repository. A real `get_me` call returned the same account. Reading the selected repository passed; reading an unselected private repository failed with 404. - Browser: a second fresh instance completed enrollment and OAuth without a PAT form or unavailable state. Its username and repository list survived reload and refresh. A real get_me call on the final code returned the displayed account. ## Risks - Repository names are now stored in company-scoped grant metadata and shown with that credential. They are display data, not authorization data. - Large selections require more GitHub API calls. A failed later page rejects the refresh rather than reporting a partial list. - Older grants and webhook-invalidated snapshots require Refresh access to load the list. - Cloud profile enablement is separate deployment configuration. This PR does not change OAuth scopes or GitHub App permissions. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) via Codex. Reasoning, code execution, and browser tools were used. The exact context window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fee8d8dc39 |
fix(runner): repair direct live provider bootstrap (#12932)
## Thinking Path > - Paperclip runs AI agents through qualified provider backends. > - The direct live eval workflow builds one immutable Runner runtime for every matrix cell. > - The workflow reinstalled the packed Runner with npm. > - That install discarded pnpm patches and selected provider dependencies outside the qualified lock. > - The first pnpm deployment model also placed its virtual-store marker at the wrong level; a real deployment keeps `.pnpm` beside the scoped Runner package. > - AgentCore enforced the current context-aware harness but the direct eval CLI did not supply the production v3 runtime context that harness requires. > - This pull request preserves the qualified dependency graph, resolves the real deployment layout, and makes direct evals exercise the production runtime-context contract. > - The benefit is that live eval cells reach their provider turn with the same artifacts and context contract that Paperclip qualified. ## Linked Issues or Issue Description Refs: #12931 **What happened?** The full direct live eval campaign failed every ACPX cell during `session.open`. The portable runtime had an incorrect dependency root. Its npm install also discarded the qualified ACP server patches. AgentCore cells first failed because Runner enforced `aws-agentcore-harness-v1` while the provisioned stack and eval profile use `aws-agentcore-harness-context-v2`; after aligning that revision, the direct eval CLI still omitted the required v3 runtime context. **Expected behavior** The direct eval runtime must preserve the frozen pnpm dependency graph and patched provider bytes. Runner, server validation, OpenAPI, and the deployed AgentCore stack must use one qualification revision. Direct eval attempts must supply the same immutable native runtime-context contract as production. **Steps to reproduce** 1. Dispatch `Runner Direct Live Protocol Evals` from `master`. 2. Select an ACPX Claude, ACPX Codex, or AgentCore roster. 3. Observe a pre-turn provider bootstrap failure. **Paperclip version or commit** `d96452db059338b329b458ba8fe359fef72f1363` **Deployment mode** GitHub Actions on the RunsOn Linux x64 fleet. ## What Changed - Build the reusable direct-eval runtime with `pnpm deploy --prod`. - Resolve ACPX dependencies from the actual scoped-package layout of a self-contained pnpm deployment. - Align AgentCore configuration and qualification checks on `aws-agentcore-harness-context-v2`. - Materialize a minimal immutable v3 runtime context for each isolated direct eval attempt. - Add workflow, package-authority, runtime-context, Rust, and server regression coverage. - Document the qualified packaging, runtime-context, and AgentCore revision contracts. ## Verification - `pnpm --filter @paperclipai/paperclip-runner typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/live/runnerd-codex-transport.test.ts` (70 tests) - `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/cli/eval-session-contract.test.ts` (14 tests) - Focused Runner contract tests (36 tests) - Focused server profile tests (47 tests) - Focused Rust managed-provider and native-selector tests (19 tests) - `node --test packages/paperclip-runner/scripts/runner-protocol-eval-workflow-security.test.mjs` - `actionlint .github/workflows/runner-protocol-live-evals.yml` - A local `pnpm deploy --prod` produced both qualified ACP server digests. - A Linux reproduction of the first follow-up smoke identified the real deployment root and the missing AgentCore runtime context. ## Risks The AgentCore revision change rejects profiles that still use the obsolete v1 value. This is intentional because the provisioned context-aware harness and current eval profile use v2. Direct eval prompts now receive the same fixed runtime-context preamble as production, so behavior scores may move; that is the intended qualification surface. The workflow package layout changes, but tests assert the new entrypoint and dependency root. This change does not modify the browser full-stack E2E workflow. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5.6. The context-window size is not exposed in this session. The model used extended reasoning, repository tools, code execution, Docker-based Linux reproduction, and GitHub Actions diagnostics. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (for example, `docs/...` or `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0c1e7504c0 |
fix(runner): persist warm Daytona workspaces across turns (#12904)
## Thinking Path > - Daytona preserves a stopped sandbox filesystem, but deleting or replacing a sandbox removes its only remote copy. > - Warm reuse therefore improves latency but cannot be Paperclip's durability boundary. > - The host execution workspace must remain authoritative after every successful turn, while same-run recovery must avoid overwriting unexported remote work. > - Result proposal, workspace export/merge, and terminal completion need a durable, replayable ordering so a crash never starts a duplicate provider turn. > - A paid browser acceptance suite must exercise both legacy Codex and Runner Codex for three real turns on one continuously warm Daytona sandbox. ## Linked Issues or Issue Description Refs #12901. Runner Codex did not previously export successful Daytona workspace changes back to the authoritative host workspace. That made warm reuse depend on Daytona's remote filesystem and left deleted/replacement sandboxes without a reliable reconstruction path. The existing paid fixture also lacked a focused three-turn continuity case for both Codex adapters. ## What Changed - Persist versioned, atomic native workspace-sync descriptors and durable seeds in `PAPERCLIP_HOME`, without credentials or a database migration. - Classify fresh, warm, replacement, and same-run-recovery workspace preparation explicitly; ambiguous lease/root/digest evidence fails closed. - Finalize native workspace export/merge after semantic result proposal and before run completion, with idempotent replay that never submits a second provider turn. - Surface legacy Codex workspace restoration failures instead of masking them, while preserving an earlier provider error when both fail. - Keep healthy reusable Daytona leases warm for legacy and native adapters, stamp finalized workspace generations, and retain existing cleanup behavior for per-turn or unhealthy leases. - Preserve Runner Codex's provider process/session across warm turns, including bounded post-terminal tail draining and exact authority rotation. - Add the exact paid `daytona-warm-continuity` matrix: - `legacy-codex × daytona × warm-three-turn` - `runner-codex × daytona × warm-three-turn` - Drive all three turns through the browser, verify ordered file continuity and stable lease/workspace/runtime identities, capture per-turn timings, and delete the sandbox immediately after assertions. - Document `pnpm test:e2e:runner -- --suite daytona-warm-continuity`; no package script was added. ## Verification - `pnpm typecheck` — passed, including migration safety (no migration added) - Focused server/runner Vitest coverage — 144 passed - `pnpm test:e2e:runner:unit` — 114 passed - `pnpm test:e2e:runner:typecheck` — passed - `pnpm --filter @paperclipai/paperclip-runner test:codex` — 66 passed, 1 helper ignored - `native-session-executor.test.ts` — 139 passed, including safe fail-closed cleanup after remote runner identity capture failure - Paid local browser acceptance, exact post-rebase Linux/amd64 runner binary: - Runner Codex — passed in 1.7m; 3 runs; lease outcomes `created, resumed, resumed`; 10/10 matchers; cleanup passed - Legacy Codex — passed in 2.7m; 3 runs; lease outcomes `created, resumed, resumed`; 10/10 matchers; cleanup passed - [Protected paid GitHub Actions campaign](https://github.com/paperclipai/paperclip/actions/runs/34026735033) against `7da42a91b95fa7fb2df126668ef7e37afb3b2b9d` — passed 2/2: - Runner Codex — 3 runs; lease outcomes `created, resumed, resumed`; evidence and cleanup passed - Legacy Codex — 3 runs; lease outcomes `created, resumed, resumed`; evidence and cleanup passed - Merge/enforcement, S3 history, and Pages publication jobs passed - Paid result artifacts were scanned for both provider credentials; neither secret was present. - Current PR checks — 31 passed, 1 expected Storybook skip; Greptile 5/5; Superagent security scan passed - `git diff --check origin/master...HEAD` — passed - Confirmed no `package.json`, lockfile, migration, or SQL changes. ## Risks - Workspace synchronization now sits on the terminal-success path, so a remote export failure deliberately prevents false success. Retryable state retains its lease/seed; loss of the only unexported remote copy fails closed. - Warm provider reuse has strict identity and quiescence checks. Mismatched or ambiguous evidence blocks reuse rather than risking concurrent provider work. - The paid suite incurs Daytona and Codex cost only in the existing protected scheduled/manual workflow and explicitly destroys its sandbox after each cell. ## Model Used OpenAI Codex with GPT-5 agentic reasoning, repository inspection, real browser E2E execution, Rust/TypeScript test execution, and GitHub Actions diagnostics. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked an existing issue or described the issue in-PR - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name contains no internal ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have documented the dedicated suite invocation without adding a package script - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green on the current revision - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups on the current revision - [x] I will address all reviewer comments before requesting merge |
||
|
|
3796c6f259 |
fix(connections): project GitHub identity into sandbox runners (#12907)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed GitHub connections resolve a responsible user's or dedicated agent's identity into an audited, run-scoped credential projection. > - An enrolled instance could retain a hidden managed setup method after Cloud stopped advertising it, producing a blank, disabled setup step. > - Native runner processes also dropped the resolved GitHub projection before the provider shell, so `gh` and Git could not use the selected identity in Daytona. > - Daytona already provides the outer isolation boundary. Applying Codex's inner Linux sandbox there both duplicated containment and failed because nested user namespaces are unavailable. > - This change repairs setup fallback, carries only the bounded GitHub projection across each runner boundary, and allows only a controller-selected managed sandbox transport to act as the outer sandbox. ## Linked Issues or Issue Description **What happened?** An enrolled self-hosted instance could show a blank GitHub setup step when its managed profile was unavailable. Separately, a native Codex run in Daytona could resolve a managed GitHub connection on the Paperclip host but lose it before the provider shell. Once projected, Codex's nested sandbox failed before commands could run because Daytona does not expose the user-namespace operation used by the inner sandbox. **Expected behavior** Setup must select an advertised customer method when the managed method is unavailable. A Daytona run must receive the exact managed GitHub identity selected for that run, support `gh` and HTTPS Git, and rely on Daytona as its outer sandbox without weakening local or SSH execution. **Steps to reproduce** 1. Enroll a self-hosted instance while Cloud does not advertise the managed GitHub profile and open GitHub setup. 2. Observe the blank second step and disabled action. 3. Configure a native Codex agent with a Daytona environment and a responsible-user GitHub grant. 4. Run `gh api user` or HTTPS Git from the agent shell. 5. Observe missing GitHub environment projection or nested-sandbox startup failure. **Paperclip version or commit** The setup bug reproduces on `1dceee9a4`; the runner proof was developed from the same branch and verified at the latest head below. **Deployment mode** Self-hosted Paperclip enrolled with Paperclip Cloud, using the Daytona sandbox-provider plugin and native Paperclip runner. ## What Changed - Wait for connector enrollment hydration, retain a hidden managed method only while enrollment is needed, and otherwise select an advertised customer fallback. - Add a single bounded GitHub credential-environment projection for `GH_TOKEN`, `GITHUB_TOKEN`, the process-only Git helper token, GitHub commit identity, and at most 32 controller-generated Git config entries. - Forward that projection through the durable controller, Codex app-server transport, and Rust provider child without placing token values in arguments or config. - Allow Codex shell inheritance only for the exact projected GitHub keys and enable provider network access only when the managed credential exists. - Derive outer-sandbox authority exclusively from a managed `sandbox` transport; strip the same flag from configured, host, local, and SSH environments. - Define a named external-sandbox permission profile that Codex resolves to `dangerFullAccess` for default-mode Daytona turns while plan mode remains read-only. - Add regression tests for setup fallback, credential projection, local/SSH/sandbox authority separation, provider forwarding, and permission-profile selection. ## Verification - `pnpm exec vitest run ui/src/pages/apps/AppsConnect.test.tsx` — 96 passed. - `pnpm exec vitest run src/drivers/codex/codex-security-config.test.ts src/drivers/codex/app-server-transport.test.ts src/control-plane/durable-prp-control-plane.test.ts` from `packages/paperclip-runner` — 31 passed. - Focused native-session executor tests — 3 passed. - `pnpm --filter @paperclipai/paperclip-runner typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml --locked -p paperclip-runner-core --lib` — 194 passed. - Real Codex app-server configuration probe accepted `paperclip-runner-external-sandbox` and reported `sandbox.type=\"dangerFullAccess\"` while using the named profile. - [Signed Daytona image workflow](https://github.com/paperclipai/paperclip/actions/runs/33995270328) built commit `86571f7997e7100e47bd131aac1f1e773112a0ce`; the isolated environment was pinned to `sha256:ecef21105f8de382d75787e59439d936be239b77ae74a31c8ed3a17cde39b023`. - Live isolated Daytona proof passed: the three projected token variables were non-empty and equal; the host-scoped Git credential helper returned the same token without printing it; `gh api user` resolved `cryppadotta`; authenticated `git ls-remote https://github.com/paperclipai/paperclip.git HEAD` returned `1dceee9a4e75b13456760bb54c752deb2dba1d79`; no repository mutation occurred. - The persisted 28,476-byte run log contains no GitHub token shape, bearer header, credential-bearing URL, or private-key marker. - Latest-head pull-request CI and reviews provide the remaining full-suite gate. ## Risks - This deliberately gives shell Git and `gh` access to the run's resolved GitHub identity. It is the audited class-3 behavior required by the GitHub connection design and is outside per-tool Ask-first controls. - The credential source is the trusted broker projection, which overwrites configured environment values. The helper is scoped to HTTPS `github.com`, revalidates protocol and host, and never places its token in command arguments, URLs, or files. - Managed Daytona sandboxes become the containment boundary for default-mode provider commands. Local and SSH targets retain the inner Codex workspace sandbox, and plan mode remains read-only everywhere. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5, with reasoning, browser control, shell access, and code execution. The product did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1dceee9a4e |
fix(runner): persist warm Daytona workspaces (#12901)
## Thinking Path > - Paperclip manages AI agent work and the execution state for each task. > - Remote agents run in sandbox environments such as Daytona. > - Daytona keeps files while a sandbox is stopped, but deletion removes those files. > - Runner Codex did not copy successful remote workspace changes back to the host workspace. > - A warm sandbox could therefore hide data loss until Daytona replaced or deleted the sandbox. > - This pull request makes the host workspace durable after every successful turn and keeps verified reusable sandboxes warm. > - The benefit is reliable multi-turn work across warm reuse, restart, stop, and sandbox replacement. ## Linked Issues or Issue Description **What happened?** A successful native Codex turn in Daytona could leave workspace changes only in the remote sandbox. A later warm turn appeared to work because it reused that filesystem. A replacement sandbox could start from stale host data and lose the successful changes. **Expected behavior** Paperclip must merge each successful remote turn into the authoritative host workspace before it completes the run. A verified warm lease may reuse its remote files. A replacement lease must reconstruct the exact durable workspace seed. **Steps to reproduce** 1. Run Codex in a reusable Daytona environment. 2. Write a file during one successful turn. 3. Replace the Daytona sandbox before the next turn. 4. Observe that the next turn can start without the prior file on the unpatched code. Related remote workspace foundation: #10070. ## What Changed - Added explicit `host_current`, `durable_seed`, and `adopt_remote` workspace preparation modes. - Added atomic, versioned native workspace descriptors and seed archives under `PAPERCLIP_HOME`. - Added real native sandbox export and three-way host merge before terminal result completion. - Added workspace-only recovery after a proposed result. Recovery does not submit another provider turn or consume the provider retry budget. - Added fail-closed handling when a sandbox with unexported changes is gone. - Kept healthy reusable Daytona sandboxes started for legacy Codex and Runner Codex. - Kept the Runner Codex process and provider session across verified warm turns. - Added the paid `daytona-warm-continuity` browser suite. It contains exactly the legacy Codex and Runner Codex cells. Each cell performs three measured turns. - Documented `pnpm test:e2e:runner -- --suite daytona-warm-continuity`. No package script was added. - Added no database migration. The metadata format is backward compatible and idempotent. ## Verification - `pnpm typecheck` - `pnpm test:e2e:runner:unit` — 114 passed - Native workspace, finalizer, session, and environment tests — 232 passed - Daytona provider tests — 150 passed - Workspace staging and merge tests — 98 passed - Runner transport tests — 63 passed - Legacy Codex restore tests — 5 passed - Rust format and compile checks pass through root typecheck - The paid Daytona suite was not run locally because the required Daytona, OpenAI, and immutable image credentials are not present. ## Risks - The main risk is an incorrect workspace identity or merge after a crash. Durable descriptors bind the run, workspace, lease, provider lease, local root, remote root, and baseline digest. Ambiguous evidence fails closed. - The host merge may conflict with concurrent host edits. The existing three-way merge and exclusion rules handle this case and surface failures. - A deleted sandbox cannot recover unexported bytes. Paperclip now blocks with `workspace_sync_out_unrecoverable` instead of reporting success or rerunning the provider. - There is no database migration. Descriptor writes and recovery are atomic and idempotent. ## Model Used OpenAI Codex with GPT-5. The run used agentic reasoning, repository inspection, code execution, test execution, Git, and GitHub CLI tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
60469a08e0 |
feat(agent-login): resume an active login session and permit concurrent login terminals (#12861)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent authentication uses server sessions, plugin workers, and browser login panels. > - A page reload loses an active login session, and one worker permits only one login terminal. > - These limits cause lost work and prevent two owners from logging in through one worker. > - This pull request lets the browser resume active sessions and lets workers serve concurrent login terminals. > - The benefit is reliable login recovery with a bounded process-wide route limit. ## Linked Issues or Issue Description **What existing behavior does this improve?** It improves agent credential login recovery and concurrent login terminal handling. **Subsystem affected** Cross-cutting (multiple of the above). **Current behavior** A page reload loses the active login session. A shared plugin worker rejects a second login terminal. **Proposed behavior** The browser reads and resumes the owner's active session. A worker supports multiple login terminal routes under a process-wide ceiling. **Reason and benefit** Owners keep login progress after a reload. Two owners can log in through one worker without removing the route limit. **Breaking changes** None. The change adds owner-scoped read routes and changes login terminal concurrency. ## What Changed - Replace the single worker login route with maps keyed by host route and worker session identifiers. - Add a process-wide login route ceiling and release each reserved slot on every exit path. - Add owner-scoped active-session reads with consistent negative responses and private cache control. - Keep the device-login prompt while the session has an active public status. - Add a durable setup-token cancel fallback for a lost in-memory session. - Resume active sessions when the agent configuration or onboarding panel mounts. - Remove routine unmount cancellation and keep explicit Cancel behavior. ## Verification - `pnpm --filter @paperclip/server test` — server route, service, and plugin-worker-manager suites. - `pnpm --filter @paperclip/plugin-sdk test` — worker RPC host suite. - `cd ui && npx vitest run src/components/AgentConfigForm.render.test.tsx src/components/OnboardingWizard.test.tsx`. - `cd ui && npx tsc -b`. - `tests/e2e/onboarding.spec.ts` — reload during login. - CI must pass on this pull request. ## Risks The change affects agent authentication and the sandbox-to-host boundary. Route cleanup must release every reserved slot. Owner checks must prevent cross-owner session access. Tests cover route cleanup, owner scope, reload recovery, and concurrent worker routes. ## Model Used Codex, OpenAI GPT-5, tool use and code review support. The implementation author owns the exact model details for the code changes. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR with the relevant issue-template fields - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3ed5b7c5c8 |
refactor(server): extract the active-run output watchdog into a feature module (#12853)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server recovery service monitors active runs and applies watchdog decisions > - The watchdog rules and database operations lived in one large recovery service > - This structure made the rules harder to test and made company scoping harder to inspect > - This pull request moves the watchdog into domain, application, and adapter layers > - The benefit is a smaller recovery service, pure policy tests, and clear company-scoped ports ## Linked Issues or Issue Description **What existing behavior does this improve?** The active-run output watchdog that detects silence, suppression, and terminal evidence. **Current behavior** The recovery service contains the watchdog policy, use cases, database operations, and process control in one file. **Proposed behavior** A feature module separates pure policy, use cases and ports, and Postgres and process adapters. The recovery service delegates its public watchdog methods to this module. **Reason and benefit** The separation makes policy decisions easy to test. Company identifiers on every reader and writer port make tenant scope clear. Smaller service methods reduce change risk. **Breaking changes** None. The recovery service keeps its public methods and call sites. Related public watchdog work includes [#7043](https://github.com/paperclipai/paperclip/pull/7043) and [#7770](https://github.com/paperclipai/paperclip/pull/7770). ## What Changed - Add the `server/src/modules/active-run-watchdog/` feature module with domain, application, and adapter layers. - Move watchdog policy, use cases, Postgres access, and local process control into the module. - Keep the recovery service public methods and delegate them to the module. - Add 53 pure module test cases and retain 8 Postgres integration cases. - Add company scoping and transaction rollback coverage. ## Verification - Run `vitest run --config vitest.config.ts src/modules` and confirm 3 files and 53 cases pass. - Run `vitest run --config vitest.config.ts src/__tests__/heartbeat-active-run-output-watchdog.test.ts` and confirm 1 file and 8 cases pass. - Run the full server suite in pull request CI. - Compare the type-check result with a fresh baseline on the same checkout. ## Risks The main risk is a behavior change in recovery decisions during the move across layers. The pure policy tests cover the moved rules. The integration tests cover database behavior, company scope, and transaction rollback. Pull request CI runs the full server suite. ## Model Used OpenAI Codex, GPT-5, runtime-managed context window, tool use, code execution, and repository review support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d2d647c34b |
fix(connections): honor identity after reconnect (#12897)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - Connections give agents access to external services with a selected credential identity. > - A removed connection keeps its database row so Paperclip can retain its history. > - A fresh GitHub setup can select a different identity from the removed connection. > - The retained row incorrectly kept its old credential policy after that new selection. > - The GitHub callback then could not save the new grant and returned the user to setup. > - This pull request applies the explicit identity selection when Paperclip revives an archived row. > - The benefit is a successful GitHub reconnect after the user changes from a dedicated agent account to a personal account. ## Linked Issues or Issue Description **What happened?** After a user removed a dedicated-agent GitHub connection, a fresh setup with “My GitHub account” returned to the setup page with `oauth=failed`. The Cloud claim succeeded, but the local connection still used the old `per_agent` policy. **Expected behavior** A fresh setup must apply the explicit identity choice. An interrupted draft or an explicit reconnect must keep its existing identity. **Steps to reproduce** 1. Connect GitHub with a dedicated agent identity. 2. Remove the connection. 3. Start a fresh GitHub connection with “My GitHub account.” 4. Complete GitHub OAuth. 5. Observe that Paperclip returns to the setup page instead of the permissions page. **Paperclip version or commit** `342c01fee` **Deployment mode** Local dev (`pnpm dev`) with embedded Postgres and the staging managed connector. **Additional context** This follows the GitHub access UI change in #12893. ## What Changed - Apply an explicit Access identity when a fresh gallery setup revives an archived connection row. - Preserve the identity for interrupted drafts and explicit reconnects. - Do not carry credential material across an identity change. - Apply the omitted organization default during a fresh archived-row recovery. - Restore the prior grants and credential policy transactionally if a revived setup rolls back. - Disable the connection and surface a specific failure if that restoration cannot complete. - Preserve newer concurrent grant changes with a row lock and optimistic version check. - Preserve newer concurrent connection identity/configuration changes with a locked state fingerprint. - Add regressions for dedicated-to-personal OAuth, organization-default recovery, rollback, rollback failure, and concurrent grant/connection changes. ## Verification - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts` — 222 tests passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - Browser proof on an isolated local instance: dedicated connection removed, personal setup selected, staging GitHub OAuth completed, permissions page opened, connection reported active and healthy, personal grant active, old agent grant revoked. ## Risks - Low migration risk. This change has no schema migration. - The behavior changes only when a fresh setup explicitly selects an identity for an archived connection row. - Existing draft resume and explicit reconnect behavior stays unchanged. - Connection-manager checks still protect changes to a retained credential identity. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5, with reasoning, browser control, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5da6499860 |
fix(connections): reuse one-time cloud enrollment (#12891)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Managed connections let agents use provider credentials without exposing those credentials to the control plane UI > - A self-hosted instance must first establish a trusted credential destination with Paperclip Cloud > - The GitHub connection flow repeated that trust decision before provider consent > - The local setup route also lost step 2 after enrollment and could display the PAT identity defaults before enrollment > - This pull request makes enrollment a one-time instance decision and sends later provider starts directly to provider consent > - The benefit is a shorter flow with one clear Paperclip approval and no required service restart ## Linked Issues or Issue Description Refs #12843. Companion Cloud change: [paperclipai/paperclip-cloud#391](https://github.com/paperclipai/paperclip-cloud/pull/391). ## What Changed - Made `stage=setup` authoritative during initial route hydration and enrollment return. - Added a contained one-time enrollment screen with provider-specific copy. - Accepted a provider `authorizationUrl` from Paperclip Cloud only when it matches the exact GitHub or Google OAuth endpoint. - Preferred the direct provider URL while retaining the legacy confirmation URL fallback. - Preserved the company-bound identity and agent-access draft across the full-page enrollment callback, including cold company-context hydration. - Kept GitHub defaulted to “My GitHub account” and “Any agent,” including before Cloud advertises the managed method. - Updated GitHub identity and agent-access copy for responsible-person and dedicated-agent behavior. - Labeled the provider action “Continue to GitHub.” - Added parser, routing, cold-hydration, access-restoration, visibility, fallback, defaults, and copy tests. ## Verification - `pnpm exec vitest run server/src/services/paperclip-cloud-connector.test.ts ui/src/pages/apps/AppsConnect.test.tsx` (114 tests passed) - `pnpm check:token-gates` - `pnpm -r typecheck` - `pnpm build` - Live browser proof used a new data directory on `127.0.0.1:3117` and the exact Cloud PR revision on staging. - The fresh flow selected “My GitHub account” and “Any agent,” showed one enrollment approval, returned to local step 2, and connected GitHub without a second Paperclip confirmation or login. - The connected screen showed one selected repository, a long-lived token, installation metadata, a successful access refresh, and healthy webhook delivery. - Gmail on the same instance went directly to Google consent without another Paperclip approval. - Restarting the same data directory preserved enrollment. A second new data directory required exactly one new approval. - A final fresh-data-dir rerun selected a dedicated GitHub identity for Ada before enrollment, approved the instance once, returned to step 2, retained Ada after a Back check, connected directly through GitHub, and finished with “Used only by Ada,” one selected repository, and a long-lived token. - Port 3100 remained untouched throughout the proof. ## Risks - The new Cloud field is additive and restricted to the exact GitHub and Google OAuth origins and paths, with no embedded credentials or URL fragment. - An older Cloud response still works through `confirmationUrl`. - A self-hosted instance still requires one signed Cloud enrollment. Managed Cloud instances do not render the enrollment screen. - Provider authentication and consent remain mandatory after instance enrollment. - No schema migration is included in this pull request. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, `gpt-5.6-sol`, extended reasoning, tool use, code execution, browser control, and multi-file repository editing. The context window size was not provided. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
bcc6fe7a44 |
fix(runner): restore multi-turn remote sessions (#12840)
## Thinking Path > - Paperclip manages AI agents and their work. > - The runner executes agent turns on local and remote providers. > - A remote per-turn session must save its state before Paperclip releases its sandbox. > - The session runtime returned after 100 milliseconds while the remote checkpoint still ran. > - The next turn also checked the local state path instead of the verified remote backup. > - This pull request waits for the bounded remote close and accepts only a verified suspended backup. > - The benefit is reliable multi-turn execution without weaker identity checks. ## Linked Issues or Issue Description **What happened?** A successful remote agent turn released its sandbox before the runner saved the verified continuation backup. The next turn failed with `runner_state_identity_mismatch`. **Expected behavior** Paperclip must finish the bounded remote checkpoint before it releases the sandbox. A later turn must validate and restore the digest-matched suspended backup. **Steps to reproduce** 1. Run a native ACPX Claude Plan test in a non-reusable Daytona sandbox. 2. Reject the first plan to start a second turn. 3. Observe that the second turn fails before provider execution. **Paperclip version or commit** The failure reproduced at `13775a90b078ff64872f50961ea1b83d575e7bc6`. **Deployment mode** GitHub Actions with a Daytona sandbox. ## What Changed - Wait for the internally bounded remote runner close and checkpoint before the host returns. - Preserve the existing short cleanup bound for other providers. - Validate remote continuation lifecycle from a complete digest-verified backup when local runner state is absent. - Keep corrupt, non-suspended, mismatched, and unverified state fail-closed. - Make native Plan completion and accepted-Plan wake prompts deterministic. ## Verification - A prior 45-cell local campaign passed 44 cells. The only failure was the OpenCode Plan prompt variance fixed here. - A focused OpenCode local Plan rerun passed. - ACPX Claude Daytona message and question cells passed. - Focused regressions cover delayed checkpoint close and verified remote backup lifecycle. - GitHub Build and the focused ACPX Claude Daytona Plan cell will validate this exact head. ## Risks Remote runnerd sessions now wait for their internally bounded close/checkpoint path before returning; generic provider cleanup retains the existing 100 millisecond bound. Durable run success still cannot be reversed. The environment release guard still blocks sandbox destruction when no verified backup stamp exists. ## Model Used OpenAI Codex, GPT-5.6, extended reasoning, with code execution and GitHub Actions inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal task id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open findings - [ ] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0ffc091473 |
feat(connections): add durable GitHub identities and webhooks (#12843)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents need source control access for repository work > - A shared token cannot preserve the responsible person's identity or an agent's dedicated identity > - GitHub App tokens also need durable refresh, repository access checks, and webhook delivery > - Paperclip already has managed connections, encrypted grants, run secret leases, and merge-confirmation behavior > - This pull request extends those systems with GitHub identities instead of adding a parallel credential system > - The benefit is durable GitHub access with explicit identity, repository, runtime, and webhook boundaries ## Linked Issues or Issue Description No public GitHub issue describes this connection change. This description follows the feature request template. **Subsystem affected** Connected Apps, connection grants, secret resolution, native Git runtime setup, webhook processing, and the Apps UI. **Problem or motivation** Users need to connect GitHub once and let agents use the correct GitHub identity. A run should use a dedicated agent account when one exists. Otherwise, it should use the responsible person's account. The connection must survive token expiry, repository access changes, and temporary instance downtime. **Proposed solution** Add user-owned and agent-owned GitHub grants to the existing connection model. Resolve one identity for MCP, Git, `gh`, health checks, and webhook bindings. Store provider tokens in the existing encrypted secret system. Refresh expiring token pairs under the existing lease and compare-and-swap path. Register signed Cloud webhook bindings and process normalized pull request and installation events through a durable local inbox. **Alternatives considered** An organization-wide GitHub token would lose person and agent attribution. Environment variables alone would bypass the managed connection and grant model. A new GitHub-only credential store would duplicate the existing secret and access systems. GitHub App installation tokens and private-key custody remain outside this first version. **Roadmap alignment** This change implements the Connected Apps direction. It also extends the shipped MCP Tool Gateway, per-agent secret access, and action-attribution systems. It does not add a repository catalog. The open repository catalog work in [#11234](https://github.com/paperclipai/paperclip/pull/11234) is related and complementary. ## What Changed - Added agent-owned connection grants and a per-agent credential policy with company and subject constraints. - Added a managed GitHub App method while keeping the personal access token method as an advanced fallback. - Added durable access-token and refresh-token handling with proactive rotation and one automatic recovery after a provider `401`. - Added GitHub identity and installation summaries without storing repository-name lists. - Added signed Cloud webhook binding, event lease, acknowledgement, local idempotency, pull request merge processing, and installation access handling. - Added one identity resolver for MCP, native Git, `gh`, checkout, health checks, and webhook bindings. - Added a class-3 run projection for `GH_TOKEN`, `GITHUB_TOKEN`, a `github.com`-only credential helper, SSH-to-HTTPS rewrite, and GitHub noreply commit attribution. - Added personal and dedicated-agent setup choices plus identity, repository, continuity, and webhook status in the Apps UI. - Added schema migrations, tests, and connection documentation. ## Verification - The current head is fully green in GitHub CI, including build, typecheck, all serialized/general server shards, all browser shards, policy, canary dry run, review, and security checks. - Live staging proof completed with a non-expiring GitHub App user token, selected-repository installation, repository add/remove refresh, managed MCP, native `gh`, HTTPS clone/push/delete, GitHub noreply commit attribution, signed merged-PR webhook acceptance, durable Cloud-to-instance delivery, and installation-access event processing. Temporary branches and temporary repository access were removed afterward. - `pnpm check:token-gates` passed. - `pnpm -r typecheck` passed before and after the rebase onto `origin/master`. - `pnpm build` passed. - The focused connector suite passed 285 tests after the rebase. - The full stable suite passed 5,790 tests and failed 22 tests across 8 general server files. The failures reproduced as shared-runner environment issues. They included `/tmp` versus `/private/tmp`, closed database connections, and invalid high ephemeral ports. The focused connection tests pass in isolation. ## Risks - Migrations add agent grant subjects and a durable connection-event inbox. Migration numbering and safety checks pass. - A raw GitHub user token enters the agent process for Git and `gh`. Per-tool Ask-first controls cannot limit those shell operations. The UI warns users about this boundary. - GitHub App user tokens can be non-expiring. Paperclip performs a continuity check every 30 days, but provider revocation still requires a reconnect. - The webhook path accepts only signed and bounded payloads. It stores a minimal normalized record and no raw provider payload. - GitHub repository permissions remain authoritative. Removed access can make a cached repository count temporarily stale, but runtime access fails immediately. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, `gpt-5.6-sol`, extended reasoning, tool use, code execution, browser control, and multi-file repository editing. The context window size was not provided. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
263f181fed |
fix(runner): complete live hot restart adoption (#12852)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner owns durable provider sessions and streams their work to the control plane. > - Pull request #12845 added native restart recovery for live and dead local runners. > - A real browser test found three live-adoption gaps after that pull request merged. > - Lazy runner process ownership was not always stored before restart. > - The old controller did not release its PRP authority without closing the provider turn. > - Reconnect events could arrive before the active provider turn was restored. > - This pull request closes those gaps and proves the same turn completes after a UI hot restart. ## Linked Issues or Issue Description Refs #12845 Related search results: #12646 covers indeterminate command results after a runner restart. It does not cover controller adoption or active-turn rebinding. No open duplicate pull request was found. ## What Changed - Store lazy runnerd process ownership after provider session creation, read, and resume. - Detach native PRP controller authority during coordinated hot shutdown. Keep the live provider turn running. - Restore the exact checkpointed provider session when bounded PRP identity events have been compacted. - Restore the active provider turn before reconnect events are replayed. This prevents `turn_binding_mismatch`. - Keep exact live ownership by the current controller out of generic orphan recovery. - Add driver, transport, and server regression tests for these paths. ## Verification - Ran 12 Codex driver lifecycle tests. - Ran 53 runnerd transport tests. - Ran 143 recovery and orphan-reaper server tests. - Ran all 8 real-process restart recovery scenarios. - Ran all 96 existing runner E2E unit tests. - Ran runner TypeScript typecheck. - Ran server TypeScript typecheck. - Ran the migration replay test and migration safety checks. - Tested the board UI on an isolated local instance. A real local Codex-backed turn entered a 120-second terminal wait. The UI `Restart now` action replaced the server and kept the same runner PID, process start time, run ID, native session ID, runner ID, provider session ID, and active turn. The original turn then completed. - Confirmed one heartbeat run, no retry row, one result, one proposed-result event, one terminal event, no protocol errors, no active recovery state, and no surviving runner or provider process. ## Risks - A live runner can continue provider work while no server owns the control route. Recovery fails closed when the process fingerprint or durable identity is ambiguous. - Provider identity can be restored from the database only for an exact verified adoption claim. An authenticated live `session.snapshot` validates that identity before the driver can resume. - The new detach path applies only to native sessions that expose restart detachment. Other adapters keep their existing shutdown behavior. - This follow-up does not change the database migration or `package.json`. The migration in #12845 remains replay-safe through `ADD COLUMN IF NOT EXISTS` and its embedded-Postgres idempotence test. The dedicated real-process command remains in `doc/DEVELOPING.md`. ## Model Used - OpenAI Codex based on GPT-5. The exact serving build and context-window size are not exposed. The run used extended reasoning, repository tools, shell execution, and in-app browser automation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
7b094724e6 |
fix(runner): recover native sessions across restarts (#12845)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner keeps durable run and provider state outside one server process. > - A server restart can leave that runner alive or can interrupt it after a provider checkpoint. > - The old startup path used handoff intent and PID evidence, but it did not reconstruct native ownership. > - That gap could block the issue, create a replacement run, or start duplicate provider work. > - This pull request adds durable same-run recovery for coordinated and uncoordinated restarts. > - The benefit is exact recovery of the run, runner, session, provider, steering, and finalization state. ## Linked Issues or Issue Description Refs #9628. That pull request added earlier local-adapter hot-restart work. This change adds native PRP authority reconstruction and same-run provider resume. Refs #10935. That pull request handles missing hot-restart snapshots. This change also supports hard restarts with no snapshot. Refs #11624. That pull request prevents unsafe retry after an adopted legacy process exits. This change reconciles native terminal evidence before provider recovery. Refs #12070. That pull request improves process liveness checks. This change also binds recovery to a process-start fingerprint and fails closed on ambiguity. **What happened?** The server could record hot-restart intent, but startup did not rebuild native runner ownership. A live runner could not re-register its PRP authority. A dead runner could not resume the exact native and provider session on the same heartbeat run. Generic recovery could then block the issue or create replacement work. **Expected behavior** A live native runner must reconnect with the same PID and logical identities. A dead runner must resume the same durable session and heartbeat run with only a new operating-system PID. A proposed or terminal result must finalize once before any provider turn starts. Ambiguous process or session evidence must stay blocked without a signal or duplicate spawn. **Steps to reproduce** 1. Start a Paperclip Runner heartbeat and wait for an active provider turn. 2. Restart only the Paperclip server, with or without a hot-restart marker. 3. Observe that the old startup path does not reconstruct the native control-plane authority. 4. Kill both the server and runner after a provider checkpoint. 5. Observe that the old path cannot resume the exact native session on the original heartbeat run. **Paperclip version or commit** The defect was reproduced from commit `1991f31fd53e7f7794d5c2e4b93be384ade2b41d`. This branch is rebased onto the current `master`. **Deployment mode** Local development and self-hosted server deployments that use the local Paperclip Runner. ## What Changed - Added correlated hot-restart requests and version-compatible native handoff fields. - Added controller boot identity, process-start identity, controller generation, recovery state, request id, and bounded history to the native finalization ledger. - Added transactional recovery claims for live-runner reattach, dead-runner resume, and incomplete bootstrap. - Added fail-closed ownership takeover rules and process identity validation. - Added live runner adoption to the local runner transport without a duplicate spawn. - Added same-run provider checkpoint resume and legacy retry-row compatibility. - Reconciled proposed and terminal results before runner or provider recovery. - Bound the HTTP and PRP listener before startup recovery and delayed scheduling and generic reapers until classification completes. - Added restart-aware health diagnostics, run-log recovery transitions, durable runner diagnostics, and bounded shutdown finalizer draining. - Moved restart-survivable diagnostics into runner-owned, pre-redacted bounded writes; raw stdout and stderr are never persisted. - Added process-start fencing for controller, runner, and provider PIDs; startup classifies every candidate without an implicit cap. - Added crash-recoverable, contention-safe development restart-request coordination and failed-startup listener cleanup. - Added a credential-free real-process restart suite for eight restart, scale, and identity scenarios. - Documented native restart operation, persistence, diagnostics, and verification. ## Verification - The documented native restart commands passed. They ran eight real-process/database recovery scenarios and the live runner adoption transport test. - Native executor tests passed: 111 tests. - Heartbeat recovery tests passed: 124 tests. - Hot restart, health, and shutdown tests passed: 52 tests. - The broader affected server suite passed: 350 tests. - Focused native recovery and startup tests passed: 49 tests. - Runner transport and control-plane tests passed: 63 tests. - Runner-owned diagnostic tests passed for write-time bounding, credential redaction, private file modes, and raw stream non-persistence. - Development restart coordination tests passed: 11 tests. - Database migration checks and the partial-application/replay regression test passed. - Server, database, and Paperclip Runner typechecks passed. - `git diff --check` passed. - Full Paperclip PR CI passed, including build, canary, all five general server shards, all five serialized server shards, all three browser E2E shards, workspace suites, and release-registry verification. - Greptile completed at 5/5 with no outstanding findings, recommendations, follow-ups, or open review threads. ## Risks - Moderate risk. This changes startup ordering and ownership transfer for active native runs. - The migration adds nullable columns and does not rewrite existing rows. - Recovery fails closed when process or durable session identity is incomplete or contradictory. - The first implementation supports the local Paperclip Runner. Remote targets keep their existing behavior. - The real-process suite covers cleanup and asserts that no runner or provider process survives each test. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The runtime did not expose a more specific model revision or context-window size. Repository editing, shell execution, database tests, and real-process test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
bf95a7eae2 |
fix(ui): stabilize active-run steering queue (#12834)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The issue detail page shows a live agent run and accepts follow-up instructions. > - A follow-up must stay in a stable queue until the user sends, reorders, or removes it. > - Native runners can receive a steering event in the active run. > - Legacy runners must interrupt the active run and start a follow-up run. > - The current UI moved comments between the queue and the transcript and could show duplicate text or ambiguous chronology. > - This pull request makes the queue projection durable, keeps each message in one clear place, and labels when queued input was actually steered or delivered. > - The benefit is predictable steering with stable ordering, no duplicate messages, and visible causal timing. ## Linked Issues or Issue Description Refs #11374. Refs #12591. **What happened?** During an active run, a new follow-up could first appear as a transcript bubble and then move into the steering queue. After a steer or remove action, it could appear again. Progress text could also repeat the final response text. Once consumed, a queued bubble displayed only its original submission time even though it moved to its later causal slot, and a native run split by steering looked like two unrelated runs. **Expected behavior** An active-run follow-up must appear in the queue immediately. A native steer must move it once into the active run. A legacy interrupt must move it once into the follow-up run. A removed item must stay removed. Progress text that is identical to the final response must appear once. Consumed follow-ups must show both queue and steer/delivery times, and post-steer native segments must identify themselves as continuations of the same run. **Steps to reproduce** 1. Start a long-running task. 2. Send two or more follow-up messages while the agent is active. 3. Reorder the messages and remove one message. 4. Send the first queued message as steering. 5. Observe the queue and transcript during and after both runs. **Paperclip version or commit** The problem reproduced on commit `da1e40302`. **Deployment mode** Local development with the embedded database. ## What Changed - Project queued comments into the steering well for native and legacy live runners. - Send native steering to the active run and use interrupt-and-follow-up for legacy runners. - Keep optimistic queue order stable across refreshes and roll back failed actions. - Remove discarded comments from the transcript cache and keep them removed when the queue becomes empty. - Collapse only the final progress occurrence matching the durable response, including across steered transcript segments. - Show `Queued … · Steered …` for same-run input and `Queued … · Delivered …` for successor-run input at their causal positions. - Label settled and live post-steer segments `Continued after steering` and time them from the steer boundary. - Add regression tests for queue display, steering, fallback interrupt, reorder, remove, rollback, duplicate text, causal timestamps, and live/settled continuation headers. ## Verification - Ran the final focused steering/chronology UI suite with 233 passing tests. - Ran the activity-service regression suite with 5 passing tests. - Ran the broader queue-focused UI suite with 298 passing tests before the final chronology refinement. - Ran `pnpm -r typecheck` successfully. - Ran `pnpm build` successfully. - Ran `pnpm check:token-gates` successfully. - Tested native steering in a real browser with a 90-second baseline wait and a three-second steering correction. - Confirmed that the old final response did not appear before the steered response. - Tested three queued messages in a real browser. - Confirmed that reorder changed delivery order and that the removed message was never sent or shown again. - Tested a legacy runner in a real browser. - Confirmed that it used the interrupt fallback and showed the follow-up once. - Reloaded a saved mixed-steer/successor-run thread and confirmed the causal timestamps and continuation header render in the correct positions. - The complete macOS suite reaches five unrelated platform assertions in workspace-runtime tests. Two compare `/var` with `/private/var`. Three require Linux `/proc` listener data. GitHub Actions provides the authoritative Linux run. ## Risks - Low risk. The change is limited to issue-chat queue projection and transcript presentation. - The server run-history API adds only a read-only `contextIssueId` projection; the database schema does not change. - Optimistic actions restore the prior UI state when a request fails. ## Model Used - OpenAI Codex with GPT-5, extended reasoning, browser automation, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
184b014c25 |
feat(telemetry): add the agent.task_run event and emit it at every terminal run transition (#12809)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip records agent run outcomes through telemetry and run lifecycle services > - Terminal run transitions need one consistent event for outcome analysis > - The current paths do not report every terminal transition through one event > - This pull request adds the agent.task_run event and emits it at each terminal transition > - The benefit is complete run outcome data without exposing raw task identifiers ## Linked Issues or Issue Description **What existing behavior does this improve?** Paperclip telemetry reports agent activity, but it does not report every terminal task run through one event. **Subsystem affected** Cross-cutting (multiple of the above): packages/shared telemetry and server run lifecycle services. **Current behavior** Several run paths write a terminal status without a matching agent.task_run telemetry event. **Proposed behavior** Each terminal run transition emits one agent.task_run event. The event records the terminal state and uses the existing pseudonym helper for the optional task identifier. **Reason and benefit** Complete terminal-run data helps operators measure agent outcomes. The pseudonym helper prevents the raw task identifier from leaving the installation. **Breaking changes** None. The change adds an event and keeps existing event behavior compatible. ## What Changed - Add the agent.task_run telemetry contract and client helper. - Reuse the existing pseudonym helper for the task identifier. The helper hashes the identifier with a per-installation salt and returns 16 hexadecimal characters. The raw identifier never leaves the installation. Existing identifiers do not move. - Emit one event from each legacy, native, recovery, and issue terminal transition. - Keep emissions outside database transactions and make delivery best-effort. - Add regression tests for event shape, hashing, terminal transitions, and emission failures. - Document the event and its privacy rule in the telemetry data contract. ## Verification - `npx tsc --noEmit` in `server/` passes at the submitted commit. - The pull-request CI suite must pass. CI is the authority because local Vitest has a known dependency artifact. - The added regression tests cover event output shape, per-installation hash divergence, raw identifier handoff, omitted identifiers, and non-throwing emits. ## Risks - A missed terminal path could reduce event coverage. - Telemetry delivery remains best-effort and cannot change run finalization. - The pseudonym helper uses installation-specific state, so identifiers differ between installations. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. Context window details were not provided. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I have addressed all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
af3023f1e3 |
fix(runner): repair paid provider startup paths (#12769)
## Thinking Path > - Paperclip manages AI agents that perform work. > - Paperclip Runner connects durable task runs to local provider processes. > - The full-stack paid matrix exposed failures after the runner integrity repair. > - Verified JavaScript entrypoints lost their relative module graph when Linux executed them through descriptor paths. > - Returned provider startup errors also remained pending and became indeterminate after recovery. > - Sparse Codex tool lifecycle events lost the `write_document` identity before task transcript projection. > - This pull request repairs those three boundaries and makes the structured-question fixture deterministic. > - The benefit is repeatable provider startup, exact failure replay, and correct inline Plan placement. ## Linked Issues or Issue Description Refs #12721 and #12700. **What happened?** The paid runner matrix failed ACPX and OpenCode startup before provider session creation. The runner journal then replaced the original startup error with an indeterminate recovery result. Native Codex saved a Plan but rendered it only as a fallback card. A legacy Claude waiting reply could also echo the reserved terminal marker before the answer arrived. **Expected behavior** Verified JavaScript providers must start from immutable descriptor-backed artifacts. Returned startup failures must persist as terminal failed command results. Native tool lifecycle updates must preserve the `write_document` boundary. Pre-answer fixture output must not contain the reserved terminal marker. **Steps to reproduce** 1. Run the local provider cells in the Runner Full-Stack E2E workflow. 2. Observe ACPX and OpenCode fail during `session.open` before provider execution. 3. Observe recovery report `execution_indeterminate` instead of the original startup error. 4. Run the native Codex Plan cell and observe the fallback Plan card after the tool activity row. 5. Run the legacy Claude structured-question resume cell and observe an early marker echo in waiting prose. **Paperclip version or commit** `0f9452101740835ce0b1488a204bf48acd5bafc3` **Deployment mode** Local development with the paid GitHub Actions acceptance workflow. ## What Changed - Bundle the ACPX sidecar and OpenCode proxy as self-contained Node ESM entrypoints before hashing and verified descriptor launch. - Anchor ACPX dynamic provider package resolution at a controller-derived provider-pack root and keep that root out of the provider child environment. - Persist executor-returned startup errors as redacted durable failed command results while retaining indeterminate recovery for true process death. - Coalesce sparse native tool items by stable ID so a late `write_document` name, input, and result reach the transcript boundary once. - Forbid the structured-question fixture from spelling or announcing its reserved terminal marker before the user answers. ## Verification - Rust and TypeScript regression tests cover durable failed replay, true crash ambiguity, bundle closure, package-root derivation, environment filtering, exact Codex tool lifecycle coalescing, and prompt determinism. - Local execution is intentionally limited to formatters and static diff checks. GitHub Actions will run tests, type checks, builds, and security checks. - After ordinary CI is green, scoped paid cells will validate one ACPX launch, one OpenCode launch, native Codex Plan projection, and legacy Claude structured resume before a complete matrix rerun. - Prior failing matrix: https://github.com/paperclipai/paperclip/actions/runs/33682434315 ## Risks - Bundling changes the bytes covered by provider launch hashes. Provider-pack generation already hashes the final built files. - ACPX still loads qualified provider packages dynamically. The controller supplies a normalized package root, while existing version, digest, path, and descriptor checks remain active. - Durable `failed` is terminal. Replays return the same redacted result and do not execute the provider effect twice. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5 with agentic reasoning, repository inspection, code editing, Git, parallel subagents, and GitHub Actions coordination. The exact deployed snapshot and context-window size are not exposed to this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked related public work or described the bug in this PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [ ] I have run tests locally and they pass (intentionally deferred to GitHub Actions) - [x] I have added or updated tests where applicable - [x] No documentation change is required for this runtime repair - [x] I have considered and documented the risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
54dd0f4868 |
feat(agents): grant new agents hire permission by default (#12814)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent permissions control which agents can create or hire other agents (`canCreateAgents`) > - Today only CEO-role agents get this permission by default; every other agent starts without it > - Teams that want agents to delegate and build out their own teams must flip the toggle on each hire, and most operators want delegation to work out of the box > - This pull request makes `canCreateAgents` default to enabled for new standard-trust agents, while low-trust agents keep a disabled default > - The benefit is that agent teams can grow without per-agent permission toggling, while low-trust containment and checkout protection stay intact ## Linked Issues or Issue Description Related (not fixed by this PR): #8064 also decouples an authority from `agents:create`. **Subsystem affected** Server agent permissions (`server/src/services/agent-permissions.ts`), authorization (`server/src/services/authorization.ts`), the shared `agentPermissionsSchema` validator, and the UI trust-preset helper. **Problem or motivation** New agents cannot hire other agents unless an operator enables `canCreateAgents` on each one. Only CEO-role agents get the permission by default. This blocks delegation-by-default workflows. Operators must toggle the permission for every hire. **Proposed solution** Default `canCreateAgents` to `true` for newly created agents. Apply and persist the default at creation only. Stored rows without an explicit value stay fail-closed at read and enforcement time. Keep the default at `false` when the agent's permissions record marks it low-trust (the `low_trust_review` preset or a trust boundary). Explicit values always win. Decouple `tasks:manage_active_checkouts` from `canCreateAgents` so the default-on flag does not let a peer agent write over another agent's checked-out issue. **Alternatives considered** Granting the default only at the route layer would leave stored rows and enforcement out of sync. Keeping the checkout authority coupled to `canCreateAgents` would void the active-checkout write protection once the flag is default-on. A per-company setting adds configuration surface without a clear need; explicit per-agent overrides already exist. **Roadmap alignment** Governance and trust-preset work already separates standard-trust from low-trust agents. This change follows that line: capability by default for standard trust, containment by default for low trust. ## What Changed - `normalizeAgentPermissions` now takes a `create`/`stored` context. Creation writes get the new default: enabled unless `permissionsImplyLowTrust()` detects the low-trust review preset or a trust boundary. Stored rows without an explicit value normalize to disabled (fail-closed). The role parameter is gone. - `agentPermissionsSchema` no longer injects `canCreateAgents: false` when the field is omitted. The server-side default applies instead. - `authorization.ts` normalizes raw agent rows for `agents:create`, so enforcement matches what the API reports for legacy rows. - `tasks:manage_active_checkouts` no longer rides on `canCreateAgents`. CEO role, explicit grants, and the manager chain remain the paths. - `agents:create` is denied outright inside any resolved low-trust execution context (agent, project, issue, or run policy). The default-on flag can never reach the legacy creator allow there. - The UI trust-preset helper sets `canCreateAgents: false` when an agent is switched to the low-trust preset, instead of carrying the old value forward. - `doc/CLI.md` describes the new default for `teams install`. - Tests pin the default matrix (standard, low-trust, explicit overrides) on the server and in the UI helper. ## Verification - `cd server && npx vitest run src/__tests__/agent-permissions-service.test.ts src/__tests__/agent-permissions-routes.test.ts src/__tests__/low-trust-red-team-routes.test.ts src/__tests__/authorization-service.test.ts` — 143 tests pass. - Broader sweep: 18 suites that touch `canCreateAgents` (hire, pending-approval, teams catalog, portability, built-in agents, plugin-managed agents) pass locally. - `cd ui && npx vitest run src/lib/trust-policy-ui.test.ts src/components/TrustPresetSection.test.tsx src/pages/NewAgent.test.tsx src/pages/Agents.test.tsx` — passes. - Typecheck is clean for the changed files in `packages/shared`, `server`, and `ui`. ## Risks - Behavioral shift: agents created after this change persist `canCreateAgents: true` unless low-trust. Pre-existing agents keep their stored value. Legacy or malformed permission records without an explicit value stay fail-closed at read and enforcement time; they never gain the authority retroactively. - Low-trust runs can no longer create agents at all, even when the agent carries an explicit `canCreateAgents: true`. Before this change, that combination could hire. The red-team suite and a new authorization test pin the denial. - Narrowing: a non-CEO agent with `canCreateAgents: true` loses implicit `tasks:manage_active_checkouts`. The manager chain and explicit grants still provide it. This narrowing is deliberate; without it, the default-on flag would let any peer bypass active-checkout write protection. - No migrations. No API shape changes. Low-trust defaults are covered by the red-team regression suite. ## Model Used - Claude Fable 5 (`claude-fable-5`), Anthropic — via Claude Code CLI with extended thinking and tool use (code search, editing, local test execution). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
b98badb246 |
fix(recovery): exclude hidden issues from stranded recovery and continuation wakes (#5648)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The recovery subsystem watches assigned issues and re-wakes an agent whose run ended without finishing the work > - Intake can hide a duplicate issue by setting `hiddenAt` while leaving its status and assignee in place > - The stranded-issue query and the terminal-run cleanup both ignore `hiddenAt`, so a hidden issue is re-woken on every cycle > - Nothing on the board shows the hidden issue, so the repeated wakes have no visible cause > - This pull request adds a hidden-issue guard to both predicates and a test for each > - The benefit is that hiding an issue stops recovery work on it, with no other change in behavior for visible issues ## Linked Issues or Issue Description **What happened?** When intake marks an issue as a duplicate it sets `hiddenAt` but leaves the status at `todo` or `in_progress` with the agent still assigned. The stranded-issue recovery timer selects that issue on every tick and queues an `issue_continuation_needed` wake for it. The agent's run on the hidden issue fails or is cancelled, the terminal-run cleanup queues immediate recovery for the same issue, and the cycle repeats indefinitely. Hidden issues are invisible on the board, so nothing a person can see explains the wakes. **Expected behavior** A hidden issue is never a recovery candidate. Stranded-issue reconciliation skips it, and a failed, timed-out or cancelled run on it releases the issue without queuing a continuation. **Steps to reproduce** 1. Assign an issue to an agent and leave it `in_progress`. 2. Hide the issue (set `hiddenAt`, for example by marking it a duplicate through intake) without changing its status or assignee. 3. Let a run on that issue fail, or wait for the stranded-issue recovery timer. 4. Observe a new `issue_continuation_needed` heartbeat run queued for the hidden issue on every cycle. **Paperclip version or commit** Reproduced on `master` when this PR was opened (May 2026). The two predicates are unchanged on current `master`; this branch is rebased onto it. **Deployment mode** Not deployment-specific: both guards are in the server's recovery and heartbeat services and apply in every mode. ## What Changed - `server/src/services/recovery/service.ts`: `isNull(issues.hiddenAt)` added to the `reconcileStrandedAssignedIssues` candidate query, so hidden issues never enter the stranded set. - `server/src/services/heartbeat.ts`: `!issue.hiddenAt` added to `issueNeedsImmediateRecovery`, so terminal-run cleanup releases a hidden issue instead of queuing a continuation. - `server/src/__tests__/heartbeat-process-recovery.test.ts`: one test per guard. A failed run on a hidden issue queues no recovery run, and a hidden stranded issue is left out of reconciliation. ## Verification - `heartbeat-process-recovery.test.ts` covers both guards; CI runs it against embedded Postgres. ## Risks Low. Both changes narrow an existing predicate to exclude rows that already carry `hiddenAt`; visible issues take exactly the path they take today. A hidden issue that genuinely needs recovery would have to be unhidden first, which matches how hidden issues behave everywhere else in the board. ## Model Used The original two-line fix was authored by @im0xMagnus. The rebase onto current `master`, the two regression tests, and this description were produced with Claude (claude-fable-5-1, extended thinking, tool use) driven by a Paperclip maintainer through Prospector's triage flow. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com> |
||
|
|
f449b05bc5 |
feat(apps): unify permissions and action testing (#12802)
## Thinking Path > - Paperclip is the control plane for companies that use AI agents. > - Apps give humans and agents controlled access to external services. > - The existing app detail flow split permissions, tests, setup, and activity across separate pages. > - The split made access rules harder to understand and made reconnect work hard to find. > - New write actions also defaulted to Ask first, which did not match the intended connection policy. > - This pull request combines permission control and action testing, removes the setup page, and moves connection activity into Audit. > - The benefit is one clear place to configure, test, reconnect, and review each app. ## Linked Issues or Issue Description **What existing behavior does this improve?** The installed app Permissions, Test, Setup, and Activity views. **Subsystem affected** Cross-cutting. This change updates the React UI, shared app defaults, server permission behavior, tests, smoke scripts, and connection documentation. **Current behavior** App access and action testing use separate pages. The app detail view also links to a setup page after installation. Connection activity uses a separate tab. New write actions default to Ask first. **Proposed behavior** Permissions uses the connection access language from the initial flow. It includes searchable Read and Write sections, a three-state permission control, and a Test dialog for each action. Reconnect appears below a Needs attention header on Permissions and Review. Old Setup and Test links redirect to Permissions. Old Activity links redirect to the filtered company Audit feed. New write actions default to Allowed. **Reason and benefit** A person can understand and test app access without moving between several pages. Reconnect work stays visible where the person reviews the connection. Audit events use one consistent feed and filter model. New connections have the intended default policy. **Breaking changes** The Setup, Test, and app Activity tabs are removed. Existing deep links redirect to their replacement pages. Existing saved action permissions do not change. Only defaults for new write actions change. **Additional context** This builds on the managed app connection work in #12728. A search found no duplicate open pull request or issue. ## What Changed - Combined action testing with Permissions. - Added searchable Read and Write action groups. - Added Off, Ask first, and Allowed controls with tooltips. - Added an action Test dialog with agent selection, arguments, and formatted results. - Removed the installed-app Setup and Activity tabs. - Added reconnect guidance to Permissions and Review when a connection needs attention. - Routed connection activity into the company Audit feed and preserved the Apps & tools filter in streamlined Audit. - Moved connection removal to the Connectors-page management menu. - Made new write actions default to Allowed across connection creation paths. - Updated regression tests, browser suites, smoke scripts, and connection documentation. ## Verification - `pnpm check:token-gates` - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts server/src/__tests__/generic-mcp-connection.test.ts server/src/__tests__/tool-access-service.test.ts ui/src/components/AppConnectionSidebar.test.tsx ui/src/pages/apps/AppDetail.test.tsx ui/src/pages/apps/AppNotConnected.test.tsx ui/src/pages/apps/AppsConnect.test.tsx ui/src/pages/apps/Browse.test.tsx ui/src/pages/apps/Connections.test.tsx ui/src/pages/apps/composio-services.test.ts ui/src/pages/audit/AuditFeed.test.tsx ui/src/pages/tools/PasteConfigTab.test.tsx` (517 tests passed) - `pnpm exec vitest run ui/src/pages/apps/app-detail/TestPanel.test.tsx ui/src/pages/audit/AuditHub.test.tsx ui/src/pages/audit/AuditFeed.test.tsx ui/src/pages/apps/AppDetail.test.tsx ui/src/pages/apps/Browse.test.tsx` (96 tests passed) - Targeted Playwright verification for connection removal, rename on Permissions, inline action testing, and Smoke Lab Audit evidence (5 flows passed) - `pnpm -r typecheck` - `pnpm build` - `pnpm test:run` completed with 5,755 passing tests and 20 unrelated macOS harness failures. The failures use `/tmp` versus `/private/tmp`, invalid ports above 65535, and workspace fixtures outside this change. ## Risks - Low migration risk. This change has no database migration. - Old app-detail URLs depend on redirect compatibility. - New connections grant write actions by default. Finalization remains configure-authorized and audited, Ask first and Off remain available per action, and existing connections keep their saved policy. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected - check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, exact model ID `gpt-5`. The client does not expose the context-window size. The model used reasoning, repository tools, code execution, and browser verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
505e7b40fc |
fix: guard listComments against non-UUID afterCommentId to prevent 500 errors (#8695)
## Thinking Path > - Paperclip is an open-source app for managing AI agents > - The issue history subsystem stores comments per issue, with cursor-based pagination via the `after` query parameter > - `GET /issues/:id/comments?after=<commentId>` looks up the anchor comment by UUID to get its created_at timestamp > - When agents store an incorrect or truncated comment ID (e.g. `670427ab` instead of `670427ab-e0ae-4a54-959e-2b13a2e33d14`), Postgres throws `invalid input syntax for type uuid` before the anchor-not-found guard can execute > - This surfaces as an unhandled 500 and causes agents to fail when doing incremental comment reads on any issue > - This pull request adds a UUID validation guard in `listComments` using the already-imported `isUuidLike` helper > - The benefit is that invalid cursors get a clean empty-array response instead of a 500, matching what already happens when a valid UUID simply isn't found ## Linked Issues or Issue Description Refs #2612 (a different 500 on the same `after=` cursor path, fixed earlier; this PR covers the malformed-cursor case that remains). **What happened?** `GET /issues/:id/comments?after=<value>` returns a 500 when `after` is not a UUID. The route trims the query value and passes it straight to the anchor lookup, so Postgres raises `invalid input syntax for type uuid: "670427ab"` before the anchor-not-found guard can run. Any agent that stored a truncated or malformed comment ID as its pagination cursor gets stuck in a 500 loop on that issue. **Expected behavior** A cursor that cannot name a comment behaves like a cursor that names a missing comment: the endpoint returns `[]`. **Steps to reproduce** 1. Pick any issue id on a running instance. 2. Call `GET /api/issues/<issue-id>/comments?after=670427ab` (8 hex characters instead of a full UUID). 3. Observe a 500 with `PostgresError: invalid input syntax for type uuid: "670427ab"`, where a full-but-unknown UUID such as `00000000-0000-0000-0000-000000000000` returns `[]`. **Paperclip version or commit** `master` at the time this PR was opened (June 2026). The `listComments` anchor lookup in `server/src/services/issues.ts` is unchanged on current `master`, so the failure still reproduces there. **Deployment mode** Local dev (`pnpm dev`). Not deployment-specific: the failure is in the server's comment-listing service, so it reproduces in every mode. ## What Changed - `server/src/services/issues.ts` — added `if (!isUuidLike(afterCommentId)) return [];` guard in `listComments` before the DB anchor lookup, using the already-imported `isUuidLike` helper ## Verification ```bash # Start the dev server pnpm dev # Pass a truncated UUID — should return [] instead of 500 curl -s "http://localhost:3100/api/issues/<any-valid-issue-id>/comments?after=670427ab" # Expected: [] # Pass a valid full UUID that doesn't exist — should also return [] curl -s "http://localhost:3100/api/issues/<any-valid-issue-id>/comments?after=00000000-0000-0000-0000-000000000000" # Expected: [] # Pass a valid full UUID that exists — should return comments after that cursor curl -s "http://localhost:3100/api/issues/<any-valid-issue-id>/comments?after=<real-comment-uuid>" # Expected: array of comments ``` ## Risks Low risk. The change only adds an early-return guard for values that are provably invalid UUIDs. The code path for valid UUIDs is unchanged. The existing behavior for anchor-not-found (returning `[]`) is preserved for invalid UUIDs, which is the correct semantic (cursor not found → no comments after it). ## Model Used Claude Sonnet 4.6 (`claude-sonnet-4-6`) via Paperclip CTO agent, tool use + code execution mode, 200K context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [ ] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip CTO <cto@paperclip.ai> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com> |
||
|
|
5f87090894 |
Make managed Cloud OAuth handoffs invisible (#12790)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Apps let people give agents governed access to external providers > - Paperclip Cloud brokers shared provider authorization for managed stacks > - The managed flow sent the browser through a confirmation page after the tenant had already prepared sign-in > - A lost confirmation response could also show an expired-session error before the provider page opened > - This pull request adds an opaque handoff contract and one shared tenant coordinator > - The benefit is a direct and recoverable transition from Paperclip to every Cloud-brokered provider ## Linked Issues or Issue Description **What happened?** A managed Paperclip Cloud connection opened the Cloud confirmation route. A response-loss race could show an expired-session error while the authorization still continued. **Expected behavior** The current Paperclip loading state must stay visible while the tenant exchanges an opaque session. The browser must then open the provider directly. Self-hosted and direct OAuth must keep their existing behavior. **Steps to reproduce** 1. Open Apps on a Paperclip Cloud stack. 2. Start a managed provider connection. 3. Select Continue to sign in. 4. Observe that the browser visits the Cloud confirmation route before it reaches the provider. **Paperclip version or commit** `b872cd3d1b404bdaff70af493a2973ceb7e5d6ec` **Deployment mode** Paperclip Cloud hosted stack. No related open issue or pull request was found in the repository search. ## What Changed - Add a backward-compatible opaque Cloud handoff to the shared OAuth start contract. - Validate the Cloud descriptor on the server and expose no browser-selected endpoint. - Exchange managed handoffs through one fixed same-origin route in every Apps OAuth launcher. - Keep dialog popups reserved before asynchronous work and retain the tenant loading state. - Add recent-login resume storage, bounded retry behavior, terminal tenant errors, tests, and Storybook states. ## Verification - `pnpm check:token-gates` - `pnpm -r typecheck` - Focused connector and UI suites: 184 passed and 202 skipped. - `pnpm build` - `pnpm build-storybook` - The full local suite reached one unrelated macOS path-alias failure. The untouched test expected `/var/...` and received the equivalent `/private/var/...`. The same test reproduces in isolation. ## Risks - A malformed managed descriptor now fails closed in Paperclip instead of opening a URL. - A legacy Cloud deployment can omit the descriptor. Paperclip then uses the existing validated confirmation URL. - Direct provider OAuth and self-hosted flows do not receive a handoff and remain unchanged. - Rollback is a normal revert of this commit because the contract is optional and backward compatible. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5.6, reasoning mode, tool use, code execution, and browser verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
db4eeb1688 |
fix(server): validate project goal ids exist and belong to the company (#12779)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Projects can link to goals, through the `goalIds` list or the legacy
`goalId` field. The project service writes those links on create and
update.
> - The service never checked the goal ids. A nonexistent id died at the
`projects.goal_id` foreign key as an opaque 500, and the caller got no
actionable feedback — observed live on 2026-09-03, where one caller
retried the same bad id four times.
> - The foreign key also only proves a goal exists, not who owns it. A
goal id from another company linked silently on a multi-company
instance.
> - This pull request asserts every resolved goal id exists under the
caller's company before any write, and rejects with a 422 that names the
unknown ids.
> - The benefit is a clear, actionable client error instead of a 500,
and no cross-company goal links.
## Linked Issues or Issue Description
**What happened?**
`POST /companies/:companyId/projects` with a `goalIds` entry that does
not exist fails with an internal error: `insert or update on table
"projects" violates foreign key constraint
"projects_goal_id_goals_id_fk"`. The caller sees a 500 and retries. A
goal id that exists but belongs to a different company is accepted and
linked.
**Expected behavior**
The request fails fast with a 422 that names the unknown goal id(s).
Goals from other companies are rejected the same way. Valid links behave
exactly as before.
**Steps to reproduce**
1. Create a company and no goals.
2. `POST /companies/:companyId/projects` with `{ "name": "Rocket",
"goalIds": ["<any-uuid>"] }`.
3. Before this change: 500 from the foreign key. After: 422 naming the
id.
**Deployment mode**
Any; observed on an authenticated public deployment.
## What Changed
- `assertGoalsBelongToCompany` in the project service: one query for the
resolved ids scoped to the company; unknown ids produce `unprocessable`
(422) with the ids in the message and details
- called on create (before the project row insert, so no partial writes)
and on update (scoped to the existing project's company); both `goalIds`
and the legacy `goalId` field flow through the same resolution
- new embedded-Postgres test file: valid link, nonexistent id on create
with no partial insert, legacy field, another company's goal on create,
and a foreign-goal update that leaves existing links unchanged
## Verification
- `pnpm vitest run src/__tests__/project-goal-validation.test.ts` — 5
passed
- adjacent suites (`project-icon-persistence`,
`project-shortname-resolution`, `issue-goal-fallback`,
`project-goal-telemetry-routes`, `heartbeat-referenced-projects`,
`projects-list-archived-routes`) — 35 passed
## Risks
- Low risk. One extra indexed select per create/update that carries goal
ids. Requests that previously 500ed now 422; requests that silently
linked a foreign goal now fail — both are corrections, not regressions.
- Existing rows with foreign links (written before this check) are
untouched; only new writes validate.
## Model Used
Claude Fable 5 (claude-fable-5) via Claude Code, extended thinking with
tool use.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (none found for goal-id validation)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (doc
comments; no user-facing docs affected)
- [x] I have considered and documented any risks above
|
||
|
|
9dd6526b47 |
fix(security): harden privileged server boundaries (#12776)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server controls secrets, host files, outbound requests, and workspace commands > - A red-team review found cases where restricted callers could cross these trust boundaries > - These cases could expose credentials or let untrusted input reach privileged resources > - This pull request applies least-privilege checks at each affected server boundary > - The benefit is safer agent execution without changing the private-instance bootstrap contract ## Linked Issues or Issue Description **What happened?** Several server paths used authorization, redaction, or content-delivery rules that were too broad. Restricted agent keys could obtain company-level operational data. Some adapter and instruction paths could reach server-owned network or file resources without the required owner approval. **Expected behavior** Paperclip must redact credential values, enforce restricted-key scopes, guard outbound network access, prevent same-origin script execution, and reserve host-level file and command controls for authorized operators. **Steps to reproduce** 1. Configure an authenticated development instance at the parent commit. 2. Exercise the affected APIs with a restricted agent key or a non-instance-admin company user. 3. Observe that the parent commit returns privileged data or accepts a privileged operation. 4. Repeat on this branch and observe a redacted response, a safe download, or an HTTP 403 response. **Paperclip version or commit** The findings reproduce from commit `39898ab22` and are fixed by this pull request. **Deployment mode** Authenticated self-hosted server and local development modes. **Installation method** Built from source with pnpm. ## What Changed - Redact generic secret `value` and `token` fields recursively in structured logs. - Classify exact and separator-suffixed `KEY` environment names as secrets in company exports. - Limit restricted self-identity responses and protect company run, log, and secret catalog APIs. - Route HTTP adapter requests through DNS-pinned SSRF protection with exact private-origin allowlisting. - Download HTML, SVG, and other script-capable assets with `nosniff` and a sandbox CSP. - Require instance-admin access for external instruction roots and exports that read them. - Block agent-authenticated host command persistence across supported workspace runtime shapes. - Apply the central runtime-management decision before workspace command controls. - Keep the documented first-user instance-admin claim contract unchanged. - Add regression tests and server-owner configuration documentation. ## Verification - `pnpm -r typecheck` passes. - The Node 24 remediation suite passes with 365 tests. It skips 25 environment-gated tests. - `pnpm build` passes under Node 24. - `git diff --check` passes. - The full local runner reaches known macOS-only general-server harness failures before the serialized route lane. The Linux PR matrix is the authoritative full-suite gate. ## Risks - Restricted agent keys now receive HTTP 403 responses from company-wide run, log, and secret catalog endpoints. - Script-capable assets now download instead of rendering inline. - External instruction roots now require instance-admin access. - Private HTTP adapter endpoints now require an exact origin in `PAPERCLIP_HTTP_ADAPTER_PRIVATE_ENDPOINT_ALLOWLIST`. - Public HTTP adapter endpoints remain enabled. Redirects and metadata or link-local targets remain blocked. - No database migration is required. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5. The exact serving snapshot and context-window size are not exposed. The model used tool-enabled reasoning, repository access, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0798c77fde |
Secure Cloud canonical runtime identity (#12766)
Accept and persist Cloud-signed canonical runtime identity before activation, then route absolute self-URLs through the durable runtime identity provider. Co-Authored-By: Codex <codex@openai.com> |
||
|
|
597fd63b61 | feat(ui): add streamlined navigation foundation (#12746) | ||
|
|
9064cfd09e |
feat(codex-local): give each Codex account its own home and path secret (#12709)
## Thinking Path > - Paperclip is the control plane for companies that use AI agents for work > - Local adapters connect Paperclip agents to provider command line tools > - The Codex adapter stores login data in a shared company home > - A shared home cannot keep credentials for more than one Codex account > - This pull request gives each account a safe home and a matching company secret > - The benefit is that one company can use multiple Codex accounts at the same time ## Linked Issues or Issue Description **Problem or motivation** A company can hold only one Codex subscription credential because device login uses one shared home. A second account cannot log in without replacing or conflicting with the first credential. **Proposed solution** This change validates the vendor account identifier, stores each credential in its own home, and creates a company secret that points to that home. Repeat login calls return success when the matching secret already exists. **Roadmap alignment** The change supports the roadmap goal for centrally managed secrets with scoped access and audited resolution. **Additional context** The security review returned approve with no blocking finding. The branch adds shared account-handle validation and tests for device login and the Codex local adapter. ## What Changed - Add strict allowlist validation for Codex account handles. - Store each Codex account credential in a separate home under the Codex cache root. - Verify that the resolved account home stays inside the cache root. - Create the `CODEX_HOME_<handle>` company secret for each account. - Keep repeat and concurrent login calls safe and idempotent. - Add shared helper and route, adapter, and validation tests. ## Verification - `pnpm --filter @paperclipai/adapter-codex-local test` passes with 343 tests. - `pnpm --filter @paperclipai/server test src/__tests__/agent-device-login-routes.test.ts` passes with 25 tests. - The adapter suite passes with 23 tests. - The shared package and Codex adapter typechecks pass. - Continuous integration must pass on every check before merge. ## Risks The account handle becomes part of a directory path and secret name. The strict allowlist and root containment check reduce path traversal risk. Existing single-account homes remain unchanged unless a new device login creates an account-specific home. ## Model Used OpenAI GPT-5 (exact runtime model ID: gpt-5), with tool use and code execution. The runtime context window is not exposed in this run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0f94521017 |
fix(runner): restore local session and task integrity (#12721)
## Thinking Path > - Paperclip is the control plane for agents that perform work. > - Paperclip Runner connects durable provider sessions to individual task runs through PRP. > - Provider continuity and per-run authority are different lifetimes. > - The existing implementation mixed those lifetimes and lost event metadata between provider frames, runnerd, persistence, API sanitization, and the task thread. > - That caused failed continuation, missing progress and Plans, duplicate replies, hidden failures, and unsafe recovery. > - This repair gives every heartbeat fresh authority, preserves qualified provider-session continuity, and restores one lossless presentation path without changing direct adapters. ## Linked Issues or Issue Description **What happened?** A second native heartbeat could reuse tickets, leases, command receipts, sequence state, and run identity from the first heartbeat. Provider phase and item identity could be lost before the UI read them. Redaction could corrupt protocol discriminators while still missing malformed credential tails. The task thread could fold progress into the final response, hide failures, or show more than one final answer. Native Codex also exposed approval modes that do not yet have a durable approval bridge. **Expected behavior** Each heartbeat uses a new PRP authority epoch. Codex and OpenCode preserve exact qualified provider sessions; ACPX emits an explicit continuity event when its qualified process-replacement policy is used. Every accepted provider event is presented, classified as internal, or surfaced as unsupported. The task page shows chronological progress, reasoning summaries, activity, Plans, interactions, terminal failures, and exactly one final reply. Direct adapters retain their existing path. **Steps to reproduce** 1. Enable the unified experimental Paperclip Runner setting. 2. Create a local native Codex, OpenCode, ACPX Claude, or ACPX Codex agent. 3. Run response, Plan, structured-question/resume, restart, cancellation, and failure scenarios. 4. Reload the task while active, waiting, failed, and settled. 5. On the old implementation, observe stale run authority, missing classifications, incomplete output, or duplicated/folded replies. **Paperclip version or commit** The repair is based directly on `master` at `87d05e194b643810d16d20612115acd01d735d43`. **Deployment mode** Local development with the embedded database. Related work: Refs #12616, #12646, #12666, #12685, and #12700. ## What Changed - Rotates PRP control-plane, outbox, ticket, lease, command, receipt, and sequence authority for each heartbeat while carrying forward only a validated provider-session identity. - Reads `control-plane-state.json`, validates both durable schemas and lifecycle values, resumes coherent current runs, archives qualified settled authority, and quarantines malformed or mismatched scoped state without moving ambiguous live legacy state. - Preserves Codex provider phase and stable item identities so commentary remains progress and only `final_answer` becomes final. - Adds raw OpenCode HTTP/SSE boundary coverage and canonical reasoning lifecycle mapping. - Makes ACPX normalization lossless for visible reasoning, tool lifecycle metadata, stable bounded identities, Plan revisions, structured requests, failures, and qualified process replacement. Only the compatible terminal assistant message is promoted as final. - Applies schema-aware redaction before generic JWT-shaped detection and scans every diagnostic string leaf. Malformed raw/escaped quoted credential tails are redacted in both server and durable Rust state. - Restores snapshot-style chronological task presentation, expandable tool activity, inline Plan cards, visible waiting/resume/cancel/failure states, and exactly one final answer. - Makes `never` the only qualified native Codex permission mode and rejects unsupported persisted native modes with remediation. OpenCode and ACPX policies remain intact. - Keeps the unified experimental Runner setting as the only enablement flag. Onboarding and direct Codex, Claude, and OpenCode stay on their legacy execution/finalization paths. - Adds cross-language goldens, authority/recovery/fault coverage, exact response/count assertions, and native plus legacy acceptance scenarios. ## Verification - Pull-request GitHub Actions run Rust formatting/tests, TypeScript checks, server/UI tests, builds, protocol drift checks, browser E2E, and security scans. - A separate workflow-only validation ref is pinned directly on this PR head and runs the 35-cell paid local matrix: three core scenarios plus structured-question resume and restart/resume for native Codex, native OpenCode, ACPX Claude, ACPX Codex, and direct Codex/Claude/OpenCode. Run: https://github.com/paperclipai/paperclip/actions/runs/33682434315 - Acceptance requires exact single visible replies, monotonic sequences, matching envelope discriminators, one semantic terminal, one run terminal, no unresolved interaction, no duplicate mutation, no secret leakage, provider continuity, and zero native rows for direct adapters. - Per maintainer direction, tests are running in GitHub Actions rather than on the slower local host. Only formatters and static diff checks were run locally. ## Risks - Recovery from old or partial filesystem state is sensitive. The repair fails closed, preserves active or unverifiable authority, and quarantines only state whose scoped ownership is safe to move. - Provider event formats can change. Closed validators and boundary goldens turn new or malformed events into visible diagnostics instead of silent drops. - Shared task presentation could affect direct adapters. Runtime-fact gating plus the direct-adapter matrix protect the existing path. - Managed and remote providers are not qualified here. Shared code continues to compile and fail safely, but live qualification is deferred. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5. The exact deployed snapshot and context-window size are not exposed to this task. It used agentic reasoning, repository inspection, code editing, Git, parallel subagents, and GitHub Actions. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass (intentionally deferred to GitHub Actions) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] The paid local-provider matrix is green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
87d05e194b |
feat(work-products): add rich cards and run artifact inventory (#12717)
## Thinking Path > - Paperclip is the control plane for AI-agent companies. > - Agent outputs must remain visible after a run and easy to inspect from a task. > - The thread and artifact inventory need one consistent rich-card vocabulary. > - Run uploads also need durable artifact registration and producing-run context. > - Reviewers need deterministic examples for each rich-card kind and state. > - This pull request adds the shared presentation, registration, inventory, and Storybook review coverage. > - The benefit is a complete output path that reviewers can inspect without seeded data. ## Linked Issues or Issue Description **What existing behavior does this improve?** This change improves work-product presentation in task threads and the task Artifacts tab. **Subsystem affected** The change affects shared work-product contracts, the runner diff path, server attachment and work-product services, GitHub metadata refresh, the React board UI, and Storybook. **Current behavior** The thread used generic cards. Some files uploaded by a run existed only as message attachments. The Artifacts tab showed a flat list without run context or filters. Storybook showed only one resting card per kind. **Proposed behavior** The thread uses rich cards for supported work-product types. Each run-produced file registers one attachment-backed artifact work product. The Artifacts tab groups outputs by run and supports filters. Storybook shows every kind and requested state, PR lifecycle states, stats variants, truncation, mobile layout, and message-tail media. **Reason and benefit** Users can identify outputs quickly. Reviewers can inspect all card permutations without creating task data. **Breaking changes** None. The metadata fields and automatic artifact registration are additive. Existing attachments and work products keep their current behavior. ## What Changed - Added a shared rich work-product card with kind-specific content and a compact inventory variant. - Added pull-request and commit diff metadata plus bounded GitHub state refresh. - Added media strips and typed file chips to message-tail attachments. - Registered each run-produced attachment as an artifact work product in the same server transaction. - Grouped task artifacts by run with agent and timestamp headings. - Added type and run filters, image thumbnails, compact cards, and a company Artifacts link. - Added a Storybook kind-by-state matrix with stats variants for all eight visual kinds. - Added PR open, draft, merged, and closed examples, long-title truncation, an exact 375-pixel viewport, and message-tail overflow coverage. - Closed reconciled runtime work products when the linked runtime stops or disappears, so the card shows `Stopped` instead of `Unhealthy`. ### Screenshots Before: one resting card per kind.  After: the kind and state matrix.  After: message-tail media at 375 pixels.  [Open the Storybook evidence viewer](https://pages.paperclip.ing/rich-work-product-storybook-20260902/). The earlier artifact inventory comparison remains available in the [artifact inventory viewer](https://pages.paperclip.ing/rich-artifacts-inventory-proof-20260902/). ## Verification - `pnpm --filter @paperclipai/ui typecheck` passed. - `pnpm check:token-gates` passed. - `pnpm build-storybook` passed. - `pnpm exec vitest run server/src/__tests__/work-product-runtime-reconciliation.test.ts` passed with 5 tests. - Chromium visual checks passed at desktop and 375-pixel widths. - All 30 latest-head GitHub checks passed. One unrelated annotation test was flaky and passed on its single retry. - Greptile passed at 5/5 with zero unresolved threads. ## Risks - Low risk. The Storybook change adds review fixtures only. The runtime fix changes read-time reconciliation without database writes. - The matrix is intentionally large so every permutation stays visible in one review surface. > I checked `ROADMAP.md`. This work does not duplicate planned core work. ## Model Used - OpenAI Codex with GPT-5 and GPT-5.6-sol across this pull request. Reasoning, tool use, and code execution were enabled. The context-window size is not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My public branch name describes the change and contains no internal task id - [x] I have run tests locally and the changed-path tests pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8c3b8c432a |
Simplify app connections and enable managed Google access (#12728)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Apps subsystem gives humans and agents governed access to external tools. > - The current connection flow hides Apps behind an experimental gate and repeats setup text. > - Google sharing choices and generic MCP permissions do not use one consistent opening model. > - Self-hosted installs also need a safe default origin for managed OAuth without a manual config file. > - This pull request makes Apps available, simplifies connection setup, and applies one governed permissions model. > - The benefit is a shorter connection flow that works on a clean self-hosted install. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the Apps connection setup flow, managed Google connection flow, generic MCP connection flow, navigation, and runtime origin discovery. **Subsystem affected** Cross-cutting. This changes `ui/`, `server/`, `packages/shared/`, connector documentation, and browser tests. **Current behavior** Apps require an experimental switch. Setup pages repeat titles and explanatory copy. Connection names require manual input. Google credential sharing does not always offer both personal and organization access. Generic MCP providers do not start with the same permission choices. Managed OAuth needs a public URL setting even when the request already has a safe HTTPS origin. **Proposed behavior** Apps are available by default. Setup asks only for required permissions and sharing choices. Paperclip creates conflict-free connection names. Google apps and generic MCP providers use the same human and agent access model. Managed OAuth derives a validated same-origin HTTPS URL when no explicit public URL is set. **Reason and benefit** A clean self-hosted install can connect a managed Google app without hidden setup. Humans can share a service account with their organization. The shorter flow reduces duplicated choices and setup errors. **Breaking changes** The Apps experimental switch is removed. Existing connection APIs remain compatible. New connections can receive a numeric suffix when a name already exists. No duplicate or related public issue was found. ## What Changed - Removed the Apps experimental gate and the breadcrumb that leaves the Apps section. - Simplified all connection setup pages and moved optional provider requirements into one small link. - Added consistent human and agent access choices for Google apps, Zapier, and generic MCP connections. - Added organization sharing to Google Workspace credentials while keeping personal access available. - Generated connection names automatically and resolved name conflicts with numeric suffixes. - Derived a validated public HTTPS origin from the request for config-free managed OAuth. - Updated connector contracts, tests, browser coverage, and authoring documentation. ## Verification - `pnpm check:token-gates` - `pnpm -r typecheck` - `pnpm build` - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts server/src/__tests__/generic-mcp-connection.test.ts` (273 passed) - Targeted UI/service regression suite (308 passed) - Six targeted Playwright connection journeys on a fresh onboarding instance (6 passed) - Fresh-install browser proof through Tailscale HTTPS: enrolled with Paperclip Cloud, connected managed Google Drive, and completed a real read operation. - [Exact-head CI run](https://github.com/paperclipai/paperclip/actions/runs/33669760711): all 23 matrix jobs passed, including build, typecheck, server, serialized, canary, and all browser shards. - Greptile 5/5 on `0ae2a859f269984ee950d0af231a5b09a06f3dfd`, with no unresolved review threads. ## Risks Apps are now visible to all operators. The removed experimental flag no longer hides unfinished app definitions. Managed Google availability still depends on the Cloud profile rollout and active instance enrollment. Automatic conflict handling changes only the display name of a newly conflicting connection. > I checked [`ROADMAP.md`](ROADMAP.md). MCP Tool Gateway and Apps are shipped. Connected Apps is planned, and this change improves the existing shipped connection flow. ## Model Used OpenAI Codex, GPT-5, with reasoning, browser control, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fdf8c8464d |
feat(runner): add managed provider backends (#12699)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner provides durable, provider-neutral agent execution. > - The current stack supports qualified local providers but omits the managed provider paths from the integration branch. > - Claude Managed Agents and AWS AgentCore need explicit profile qualification, durable recovery, usage accounting, and cleanup controls. > - This pull request adds those managed backends as the third part of the Runner parity stack. > - The benefit is managed execution without weakening the default-off Runner rollout gate. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: Runner, server orchestration, database profiles, CLI, and adapter configuration UI. **Problem or motivation** The current Runner stack cannot select or execute the managed Claude Agents API or AWS Bedrock AgentCore Harness backends. It also lacks qualified profile storage and recovery checks for those remote resources. **Proposed solution** Add qualified managed and remote profiles, API and CLI management, exact provider selection, durable lifecycle handling, cumulative usage accounting, bounded cleanup, and retention acknowledgement. Keep `enableNativeRunner` default-off. **Alternatives considered** A direct copy of the old integration branch was rejected because its provider contracts, model values, credential flow, and migration history no longer match the current base. A single large parity pull request was also rejected because stacked review keeps each subsystem bounded. **Roadmap alignment** This continues the existing Runner architecture and rollout work. It does not introduce a separate execution system. **Additional context** This pull request is based on the merged #12691 and #12685 stack. It also closes the delayed security-review findings reported on #12691 by binding qualified ACPX and OpenCode launch artifacts to the bytes actually executed. A GitHub search for managed agent, AgentCore, and Claude managed work found no duplicate public issue or pull request. ## What Changed - Add Claude Managed Agents and AWS AgentCore provider executors to runnerd. - Add qualified managed and remote profile storage, routes, OpenAPI contracts, CLI commands, and migration 0237. - Validate profile ownership, enabled state, exact qualified revision, model, agent version, and secret binding before persistence and recovery. - Persist durable provider session and owned skill state for restart-safe cleanup. - Reconcile uncertain create responses and delete remote sessions before owned skills. - Track cumulative provider usage and enforce positive session spend caps. - Recover interrupted AgentCore usage at the next turn boundary by charging the prior invocation ceiling exactly once; keep the session gated until an explicit monotonic budget raise. - Isolate AgentCore AWS configuration from host profiles and credential-process/SSO configuration while preserving workload identity. - Require OpenCode 1.18.17 and fixed build-owned provider-pack artifact paths; remove the ambient executable override. - Snapshot and content-verify ACPX and OpenCode commands, scripts, and provider executables before launch. Linux executes sealed inherited descriptors; macOS uses authenticated private snapshots with retry-safe rematerialization at the spawn boundary. - Persist canonical ACPX and OpenCode launch-profile digests, reject drift across fresh recovery, and make recovery failures sticky. - Close and journal unsafe ACPX active-turn recovery before any provider bootstrap or reconnect. - Add managed provider fields to the Runner configuration UI and permission projection. - Preserve the default-off `enableNativeRunner` experimental flag. ## Verification - `pnpm -r typecheck` - `pnpm build` - Focused managed server, database, CLI, Runner TypeScript, Rust, Claude, AgentCore, ACPX, OpenCode, process-supervisor, and durable-recovery tests passed. - `cargo test -p paperclip-runner-core --lib --locked` (160 tests) - `cargo check --workspace --all-targets --locked` - Native Codex integration tests passed (60 tests); native provider tests passed (7 tests); server native-runtime tests passed (87 tests). - Verified-launch replacement, nested-spawn retry, exact-version, profile-drift, sticky-failure, and no-bootstrap active-recovery tests passed. - `git diff --check` - The PR changes 91 files. `pnpm-lock.yaml` is unchanged. The Rust workspace lockfile adds the approved `rustix` dependency used for safe descriptor handling while `#![forbid(unsafe_code)]` remains enabled. ## Risks - The provider APIs can change while they are in beta. Exact qualification and fail-closed recovery checks limit drift. - Remote cleanup can fail after a partial create. Durable ownership inventories and retry-safe deletion preserve recovery state. - Migration 0237 adds profile tables. The generated migration and snapshot pass the repository migration checks. - Managed execution can incur provider cost. Positive default spend caps and explicit retention acknowledgement limit accidental use. - An interrupted AgentCore invocation without final metadata is conservatively charged to its active session ceiling. This can overstate cost, but cannot undercount it; later work requires an explicit budget increase. - Linux qualified launches use sealed memory descriptors. macOS lacks executable-descriptor APIs, so the runner uses owner-only private snapshots and minimizes linked-path lifetime; hostile same-UID processes remain outside the documented local-host trust boundary. - The global Runner feature remains default-off. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, with tool use, code execution, and subagent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
84bedd4ca1 |
feat(runner): activate qualified OpenCode and ACPX providers (#12691)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner is the experimental native runtime for governed agent work. > - The runtime contracts already describe Codex, OpenCode, and ACPX providers. > - The merged control plane still rejected OpenCode and ACPX for new runner agents. > - Runnerd also selected only the Codex provider implementation. > - This pull request activates the qualified OpenCode and ACPX paths from the form to runnerd. > - The benefit is one durable runner path with provider-specific permissions and recovery. ## Linked Issues or Issue Description Refs #12685 **Subsystem affected** This change affects the runner package, server orchestration, adapter configuration, and UI configuration. **Problem or motivation** Paperclip Runner stores provider contracts for OpenCode and ACPX. New agents cannot select those providers. Runnerd cannot execute those stored provider descriptors. The UI also shows only Codex. **Proposed solution** Accept the qualified OpenCode 1.18.17 profile and the fixed ACPX Claude and Codex profiles. Route them through runnerd. Keep provider selection, model selection, permissions, credentials, events, and recovery inside closed provider-specific boundaries. **Alternatives considered** One option was to keep the contracts dormant. That option leaves stored configuration and runtime behavior out of sync. Another option was to enable every ACPX agent. That option is not safe because Pi does not yet have the same verified launch path. **Roadmap alignment** This change supports the completed cloud and sandbox agent milestone. It also supports self-healing runs and governed agent execution. It does not add a new roadmap surface. ## What Changed - Add one server profile resolver for Codex, OpenCode, and qualified ACPX descriptors. - Keep `adapterConfig` as the provider and permission authority for fresh runs. - Add Paperclip Runner provider, ACPX agent, and provider-specific permission controls to the UI. - Reset the model to a compatible qualified value when the provider changes. - Route Codex, OpenCode, and ACPX through the durable runnerd provider selector. - Add a durable ACPX executor with bounded state, recovery, events, tool receipts, and identity checks. - Remove Codex labels from OpenCode events, results, evidence, and recovery diagnostics. - Pass only provider-specific credential names to child processes. - Keep ACPX Pi unavailable and reject it before process launch. - Keep the existing Paperclip Runner experimental flag unchanged. ## Verification - `pnpm exec vitest run packages/paperclip-runner/src/backends/native-backend-factory.test.ts packages/paperclip-runner/src/live/runnerd-codex-transport.test.ts packages/adapters/codex-local/src/ui/build-config.test.ts ui/src/adapters/codex-local/config-fields.test.tsx server/src/__tests__/adapter-registry.test.ts server/src/__tests__/adapter-routes.test.ts server/src/__tests__/agent-adapter-validation-routes.test.ts server/src/__tests__/company-portability.test.ts server/src/services/native-runtime/runtime-mode.test.ts server/src/services/native-runtime/native-session-executor.test.ts server/src/services/heartbeat-runner-provider-config.test.ts` - The focused TypeScript, server, and UI suites passed 274 tests. - `cargo test -p paperclip-runner-core --test native_provider_backend` - The executable native provider integration suite passed 4 tests. - `cargo test -p paperclip-runner-core --lib` - The Rust unit suite passed 91 tests. - `pnpm -r typecheck` - `pnpm check:token-gates` - `pnpm build` - `git diff --check codex/runner-parity-task-runtime...HEAD` ## Risks - This changes provider process selection and durable recovery. The experimental flag still gates every fresh Paperclip Runner run. - OpenCode requires a model in `provider/model` form and stays pinned to version 1.18.17. - ACPX accepts only exact Claude and Codex profile versions and models. Pi stays unavailable. - ACPX steering stays unavailable and reports that limit through the driver capabilities. - Child processes receive explicit environment allowlists. They do not inherit the full server environment. - This pull request has no database migration. ## Model Used OpenAI Codex, GPT-5, with tool use, code execution, and subagent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
b4f302d040 |
feat(grok-local): copy a refreshed sandbox credential back to the host (#12696)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs agents through provider-specific adapters in local and remote environments > - A remote Grok run can refresh its credential inside its sandbox > - The host copy can become stale when teardown discards that refreshed credential > - This pull request copies the refreshed credential back through a locked, fail-closed teardown path > - The benefit is that later Grok runs can use the refreshed host credential without another login ## Linked Issues or Issue Description Refs: #12618 **Agent or provider** Grok local adapter. **Why this adapter is useful** A remote Grok run can refresh its access token during a run. Copying the refreshed credential back to the host keeps later runs ready to use. **How the agent is invoked** Paperclip invokes the Grok local adapter through its remote subscription run path. The adapter stages the company Grok home as a sandbox asset. The change adds a copy-out step on the teardown path. ## What Changed - `grok-auth-merge-decision.cjs` adds a host predicate in its own process. It compares the whole `<issuer>::<uuid>` identity key of the two files. It reads `expires_at` as an ISO-8601 string, an epoch-seconds number, or an epoch-milliseconds number. It exits 10 to use the source, 20 to keep the destination, 21 when the expiry shape is unreadable, and 22 when the source expiry sits more than 400 days after the host clock. It fails closed in every unclear case: an unusable side, a different identity, an absent expiry, a tie, an unreadable expiry, and an implausible expiry all keep the destination. - `grok-auth-merge-decision.ts` adds a wrapper that runs the predicate and maps the exit code to a typed result. - `grok-auth-copyback.ts` adds `copyBackGrokAuth({ hostHomeDir, readSandboxAuth, log, env })`. It locks on `hostHomeDir` with `withDirectoryMergeLock`, stages the sandbox bytes into a private `0600` temporary file, runs the predicate, and installs the file with an atomic rename in the same directory. It keeps no backup of the displaced credential. It leaves no temporary file on the success path, the keep path, or an error path. On an error it logs the `errno` code only, then re-throws. - `execute.ts` adds a `restore` callback to the Grok `home` asset. The callback takes the destination from `resolveManagedGrokHomeDir(process.env, agent.companyId)`, never from `env.GROK_HOME`. A copy-out failure does not fail the run. - `package.json` updates the `build` script to copy `grok-auth-merge-decision.cjs` into `dist/server/`, because `tsc` does not copy a `.cjs` file. **The credential shape this predicate reads** A redacted sample of a real vendor credential answered four structural questions. The answers hold no credential bytes, no account identifier, no file path, and no timestamp value. 1. `expires_at` is present. 2. `expires_at` sits inside the value object, under the `<issuer>::<uuid>` key. It is not a top-level field. 3. `expires_at` is an ISO-8601 string. It carries UTC time with a trailing `Z` and six fractional-second digits. 4. A normal run rewrites `auth.json`. The value object carries a `refresh_token` next to `expires_at`, so the client refreshes the access token and rewrites the file. ## Verification - [x] `pnpm vitest run packages/adapters/grok-local` — 114 tests in 12 files pass. - [x] `pnpm --filter @paperclipai/adapter-grok-local typecheck` — clean. - [x] `pnpm --filter @paperclipai/adapter-grok-local build` — succeeds, and `dist/server/grok-auth-merge-decision.cjs` exists after the build. - [x] Continuous integration is green on every check. ## Risks The predicate keeps the host credential when identity, expiry, file access, or freshness data is unclear. The copy-out path can log an error and leave the run successful when it cannot install the refreshed credential. The atomic rename and directory lock protect the host file from partial writes and concurrent copy-out actions. ## Model Used OpenAI GPT-5, current deployment. The exact runtime version and context window are not exposed to this agent. The model used tool calls and code inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8f9f850c20 |
fix: limit plan-to-auto transition to plan confirmation (#12695)
<!-- This pull request uses ASD-STE100 Simplified Technical English. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The issue thread controls plan review and agent work modes. > - A user can accept a full plan or confirm a smaller checkbox action. > - Only full plan acceptance must start automatic agent work. > - The current transition did not check the interaction kind. > - This pull request limits the transition to an accepted plan confirmation. > - The benefit is a safe and clear start of agent work after plan approval. ## Linked Issues or Issue Description **What happened?** An accepted confirmation that targeted a plan could change an issue from planning mode to standard mode. This included a checkbox confirmation. A checkbox action is not approval of the full plan. **Expected behavior** Only acceptance of a current full-plan confirmation starts automatic agent work. Other interaction kinds and rejected confirmations keep the current work mode. **Steps to reproduce** 1. Put an issue in planning mode. 2. Create a checkbox confirmation that targets the current plan revision. 3. Accept the checkbox confirmation. 4. Observe that the issue enters standard mode before this fix. **Paperclip version or commit** The problem was present on `master` before this change. **Deployment mode** The problem is in the core server logic and is not deployment-specific. ## What Changed - Require a full `request_confirmation` interaction before plan acceptance starts automatic work. - Add service tests for acceptance, rejection, stale interaction kinds, and unchanged standard-mode behavior. - Check the route activity log for the planning-to-standard mode change. - Document the plan acceptance transition in the V1 contract. ## Verification - `pnpm exec vitest run server/src/__tests__/issue-thread-interactions-service.test.ts server/src/__tests__/issue-thread-interaction-routes.test.ts` passes 140 tests. - `pnpm -r typecheck` passes. - `pnpm build` passes. - `pnpm test:run` was also started. Unrelated workspace-runtime tests failed because fixed local runtime ports were occupied or offset on the shared host. The same failures reproduce alone. The changed test files pass alone. ## Risks - Risk is low. The change adds one interaction-kind guard to the existing transition. - A full accepted plan confirmation still changes planning mode to standard mode and an eligible review issue to todo in one transaction. - Checkbox confirmations, questions, rejection, and standard-mode issues keep their previous behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5, reasoning, tool use, and code execution. The runtime does not expose the exact model suffix or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4b6de5327e |
Remove cheap model profiles (#12683)
## Thinking Path > - Paperclip manages agents that use different model providers and adapters. > - Paperclip must keep agent execution rules clear and predictable. > - The cheap-model profile added a second execution mode across adapters, task recovery, APIs, and the UI. > - That mode increased configuration and recovery complexity. > - This pull request removes the cheap-model profile as a product feature. > - The benefit is one model-selection path for normal work and recovery work. ## Linked Issues or Issue Description **What existing behavior does this improve?** This change simplifies model selection across agent configuration, task execution, recovery, and adapter capabilities. **Current behavior** Paperclip exposes cheap-model profiles in adapter metadata, agent runtime configuration, task overrides, recovery rules, APIs, and the board UI. Recovery work can select a different model profile from the agent's configured model. **Proposed behavior** Paperclip uses the agent's configured model for normal work and recovery work. Status-only recovery stays limited to coordination work. The API rejects legacy model-profile configuration. A migration removes stored model-profile values from existing agent, issue, and historical revision records. **Reason and benefit** One model path reduces configuration, API, UI, and recovery complexity. It also prevents status recovery from becoming a separate product-level model-routing feature. **Breaking changes** This change removes model-profile fields and adapter capability metadata. Existing stored model-profile values are removed by an idempotent migration. The validators reject new legacy profile values with clear errors. ## What Changed - Removed model-profile types, adapter capabilities, API fields, and model selection logic. - Removed cheap-model controls from agent and task UI surfaces. - Kept status-only recovery limited to coordination context while normal continuations use the configured agent model. - Added an idempotent migration that removes stored model-profile values from agents, issues, and configuration revisions without changing issue update timestamps. - Updated tests and product documentation for the single-model behavior. ## Verification - `pnpm check:token-gates` passes. - `pnpm -r typecheck` passes. - `pnpm build` passes. - `pnpm test:run` completed with 5,607 passing tests and 8 environment-sensitive failures in unrelated fixed-port and database-deadlock suites. The same failures repeated in an isolated rerun. CI is the final clean-room result. ## Risks - This is an intentional breaking change for clients that send model-profile fields. - The migration changes legacy agent, issue, and configuration-revision JSON. It is idempotent and preserves unrelated fields and issue update timestamps. - The change is cross-cutting because the removed feature existed in adapters, shared contracts, the server, plugins, and the UI. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5`. Reasoning and tool use were enabled. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1ab159d3a7 |
feat(apps): consolidate connector management (#12684)
Completes the post-managed-OAuth connector lifecycle, Paperclip Cloud provisioning defaults, governed test flows, and consolidated Apps UI.\n\nCo-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
141f202e40 |
Clean up experimental settings features (#12681)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - Instance settings control optional product features and developer tools. > - The experimental settings page mixed active experiments, internal tools, and old recovery controls. > - Some workspace links also used the selected company instead of the workspace owner. > - These problems made settings hard to scan and could send users to the wrong company route. > - This pull request removes old controls, groups developer settings, and resolves workspace links from workspace data. > - The benefit is a smaller settings surface and correct workspace navigation. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the instance experimental settings page, task watchdog controls, dependency wake recovery, and execution workspace routes. **Current behavior** The settings page shows old recovery controls and mixes product experiments with internal developer settings. Task watchdogs require an extra feature flag. Some direct workspace links use the current company prefix instead of the company that owns the workspace. **Proposed behavior** Remove the old task recovery experiment and its unused API surface. Make task watchdog controls available without the removed flag. Put worktree execution and managed environment controls in the developer section. Resolve direct workspace links from the workspace owner and reject a company prefix that does not own the workspace. **Reason and benefit** The smaller settings page is easier to understand. The server keeps only the dependency wake backstop that it still uses. Workspace links open under the correct company route. **Breaking changes** This removes the experimental issue graph recovery preview and run endpoints. It also removes the task watchdog feature flag. Task watchdog data and dependency wake behavior remain available. ## What Changed - Removed the old task watchdog and issue graph recovery feature flags. - Removed the old issue graph recovery preview, run controls, API contracts, and unused recovery implementation. - Kept resolved dependency wakes as the scheduler backstop. - Grouped product experiments and Paperclip developer settings on the instance settings page. - Made task watchdog controls available without an extra experimental flag. - Added owner-aware redirects and company checks for execution workspace routes. - Hid the false stopped-state badge while a workspace has no active runtime state. - Updated focused server and UI tests for the new behavior. ## Verification - `pnpm check:token-gates` - `pnpm -r typecheck` - `pnpm build` - `pnpm test:run` completed with 5,620 passing tests and four failures in unchanged workspace runtime port tests. The same four failures repeat when the two files run alone. - The complete GitHub CI matrix passed, including all server, serialized server, build, canary, and end-to-end jobs. ## Risks - Clients that call the removed experimental recovery endpoints must stop calling them. - The route checks depend on workspace detail access. An unknown or cross-company workspace returns the global not-found page. - There are no database migrations, lockfile changes, workflow changes, or design image changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5. The exact deployment ID and context window are not exposed. Reasoning, tool use, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
86ebdf842e |
fix(runner): keep agents running when app connections expire (#12670)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents can receive governed access to connected apps through the runtime MCP gateway. > - A connected app can become unavailable when its sign-in expires or its health state needs attention. > - The native runner treated that optional app state as a fatal runtime setup error. > - One unavailable app could therefore stop all unrelated agent work. > - This pull request removes the fatal dependency and keeps the available app assignment immutable. > - The benefit is that an agent can continue its work while the stream tells the user which app needs reconnection. ## Linked Issues or Issue Description **What happened?** An agent could not start a native run when one assigned app connection was disabled, degraded, failed, or missing its secret. Runtime context creation or MCP delivery threw an error before the agent could do unrelated work. **Expected behavior** The run must continue without the unavailable app. Healthy assigned apps must remain available. The stream must explain which app needs reconnection. A changed assignment must not give a native run new access after its immutable context is captured. **Steps to reproduce** 1. Assign an MCP app connection to a Paperclip Runner agent. 2. Set the connection to a state that needs attention, such as `degraded`. 3. Start a task run for that agent. 4. Observe that native runtime setup fails before the agent starts. **Paperclip version or commit** Reproduced from `ee2a19062`. The branch is rebased on `dda4dff64`. **Deployment mode** Local development from source with embedded Postgres. No matching public issue or open pull request was found in the GitHub search. ## What Changed - Filter unavailable assigned app connections from the immutable native runtime MCP snapshot. - Keep healthy assigned connections and their tools in the snapshot. - Replace the fatal native MCP availability check with an optional stream warning callback. - Withhold MCP delivery when the current assignment digest does not match the captured native context. - Prevent a warning delivery failure from stopping the agent run. - Add regression tests for disabled, degraded, mixed healthy and unavailable, and assignment-drift cases. ## Verification - `pnpm exec vitest run server/src/services/native-runtime/runtime-context.test.ts server/src/__tests__/heartbeat-runtime-mcp-servers.test.ts` passes with 8 tests. - `pnpm -r typecheck` passes. - `pnpm check:token-gates` passes. - `pnpm build` passes. - `pnpm test:run` was attempted. Unrelated workspace runtime and port-exposure tests failed on this macOS host. The same files also failed when run without the changed MCP tests. The changed MCP tests remained green. Clean GitHub CI is the final full-suite check. ## Risks - Low migration risk. This change has no schema or API contract migration. - An unavailable app is absent from the run MCP surface until it is reconnected and a later run captures it again. - Assignment drift fails closed. The agent keeps running, but the changed gateway is not delivered. - This pull request does not auto-block the issue before the agent decides that the app is required. It emits reconnect guidance in the stream. The existing connection-request interaction remains the path for a required app. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, `gpt-5.6-sol`, with high reasoning, repository tools, code execution, and browser automation. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
14c7efa068 |
fix(workspaces): enable UI hot reload by default (#12612)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Managed worktrees can run a Paperclip development server for each task > - The managed runtime used the built UI when its service did not set the UI development middleware option > - This made new UI source changes require a manual build instead of a hot reload > - The runtime must supply the development default while it must keep an explicit operator choice > - This pull request enables the UI development middleware for new managed Paperclip development services > - The benefit is that UI edits appear in the managed worktree browser without a manual build ## Linked Issues or Issue Description **What happened?** A new managed Paperclip development worktree served the built UI by default. An operator had to set `PAPERCLIP_UI_DEV_MIDDLEWARE=true` before UI source changes could hot reload. **Expected behavior** New managed Paperclip development worktrees must enable the UI development middleware by default. An explicit `PAPERCLIP_UI_DEV_MIDDLEWARE=false` value must continue to disable it. **Steps to reproduce** 1. Start a managed Paperclip development service without `PAPERCLIP_UI_DEV_MIDDLEWARE`. 2. Open its UI. 3. Change a UI source file. 4. Observe that the browser does not receive the change until the UI is built again. **Paperclip version or commit** This was reproduced on `317394456` from `master`. **Deployment mode** Local development with a managed worktree runtime. ## What Changed - Set `PAPERCLIP_UI_DEV_MIDDLEWARE=true` for managed `paperclip-dev` services when the service does not set a value. - Keep explicit service values, including `false`. - Add a regression test and document the default and the opt-out. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/workspace-runtime.test.ts -t "enables UI dev middleware by default"` - `pnpm -r typecheck` - `pnpm build` - `pnpm test:run` completed with 5,397 passing tests. Four existing runtime-port tests could not use ports `42000` and `52000` because a live managed runtime owns those ports on this host. The new regression test passed separately. ## Risks - Risk is low. The change applies only to managed services named `paperclip-dev`. - A service can keep the built UI by setting `PAPERCLIP_UI_DEV_MIDDLEWARE=false`. - There is no database or API contract change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, `gpt-5.6-sol`, hosted Codex context window, high reasoning, tool use, code execution, and multi-file repository editing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ee2a190626 |
Unify Paperclip Runner experimental controls (#12666)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner is an experimental execution adapter. > - The adapter and its required sandbox ingress had separate settings. > - A user could enable one setting and still have an unusable runner configuration. > - The runtime already makes one durable native or legacy decision for each run. > - This pull request uses that runtime decision for ingress authorization. > - The benefit is one clear opt-in with safe recovery for existing native runs. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the experimental settings and transport authorization for Paperclip Runner. **Subsystem affected** Cross-cutting. This change affects the React settings UI, shared settings contracts, adapter utilities, and server runtime selection. **Current behavior** Settings shows separate Paperclip Runner and Runner Preview Ingress controls. A user can enable the runner but leave required sandbox ingress disabled. **Proposed behavior** Settings shows only Paperclip Runner. Its native runtime decision also authorizes provider WebSocket ingress when the execution target requires it. A persisted native run keeps its recovery transport after the setting is disabled. **Reason and benefit** Paperclip Runner is one experimental capability. One opt-in removes an invalid partial configuration and makes the rollout boundary easier to understand. **Breaking changes** The Runner Preview Ingress card is removed. The old `enableRunnerPreviewIngress` key remains accepted in stored settings and managed configuration, but it has no server runtime effect. The public adapter-utils input remains compatible through a deprecated alias. **Additional context** Refs: #12638, #12641, #12656. ## What Changed - Removed the separate Runner Preview Ingress card from Experimental Settings. - Made resolved native runtime selection authorize required provider ingress. - Preserved ingress recovery for persisted native runs after the rollout flag is disabled. - Kept the old settings key and adapter-utils input as deprecated compatibility contracts. - Added focused UI, runtime policy, transport, stored-settings, and managed-config regression tests. - Updated deployment documentation and feature descriptions. ## Verification - GitHub Actions will run typecheck, tests, build, policy, and browser shards. - Focused tests cover the single settings control, runtime authorization, fail-closed transport selection, the deprecated public input, and old managed configuration. - No local tests were run, per the maintainer request to use GitHub Actions for verification. - `git diff --check` passes. ## Risks Low to moderate risk. The effective ingress gate changes from a separate stored flag to the resolved native run decision. Fresh runs still require `enableNativeRunner`. Persisted native runs remain recoverable. Legacy adapters never receive ingress authorization. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, with reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |