mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
ff245787658df5bb90ee89eac0cc87e583fef650
4228
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ff24578765 |
fix(runner): initialize Rustls crypto provider (#13023)
## Thinking Path > - Paperclip runs AI agents through local and remote execution adapters. > - The native runner daemon uses Rustls for secure transport. > - The production dependency graph enables the `ring` and `aws-lc-rs` crypto backends. > - Rustls cannot select a default backend when both backends are active. > - The runner daemon did not select a backend before it built a TLS client configuration. > - This pull request installs the workspace-selected `ring` provider during process startup. > - The benefit is that the runner can start reliably with the production feature graph. ## Linked Issues or Issue Description **What happened?** `paperclip-runnerd` exited with code 101 before it opened a provider session. Rustls reported that it could not select a process-level `CryptoProvider` because the binary included two crypto backends. **Expected behavior** The runner daemon must select its configured crypto provider before it creates a TLS client configuration. The daemon must start and open the provider session. **Steps to reproduce** 1. Build `paperclip-runnerd` with the locked production dependency graph. 2. Start the daemon through the local loopback transport. 3. Observe the Rustls provider-selection panic before this change. **Paperclip version or commit** `d8b958053` on `master`. **Deployment mode** Self-hosted server. **Installation method** Built from source. **Agent adapter(s) involved** Codex through `paperclip_runner`. **Operating system** Linux 7.0.0-1010-aws on aarch64. ## What Changed - Install the Rustls `ring` provider before runner setup reaches TLS initialization. - Accept an existing process-level provider as an initialized state. - Add a focused regression test for TLS builder creation and repeated initialization. ## Verification - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml --locked -p paperclip-runner-core --bin paperclip-runnerd startup_installs_a_crypto_provider_before_tls_initialization` - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml --locked -p paperclip-runner-core --test local_runner runnerd_startup_reports_build_metadata_without_panicking` - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml --locked -p paperclip-runner-core --test local_runner happy_path_emits_one_result_and_one_terminal` - `cargo fmt --manifest-path packages/paperclip-runner/runner/Cargo.toml --all -- --check` - Built the debug daemon and ran `paperclip-runnerd --build-metadata` successfully. - `pnpm -r typecheck` - `pnpm build` ## Risks - Low risk. The change selects the crypto provider that the workspace already declares. - A host process can install a provider first. The runner accepts that initialized state. - No database, API, UI, telemetry, observability, or run-log contract changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with model `gpt-5`. The context-window size is not exposed. The model used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.908.0-canary.0 |
||
|
|
d8b9580531 |
fix(ui): debounce recent task ordering (#13007)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The recent tasks sidebar helps operators return to active and
completed work.
> - Task detail refreshes update the activity timestamps used to sort
that list.
> - Older responses can undo newer activity. Concurrent updates can move
rows repeatedly.
> - This pull request preserves the newest observed activity and
debounces row moves.
> - Operators can select a stable row while task text and live state
stay current.
## Linked Issues or Issue Description
**What happened?**
Recent task rows can repeatedly swap positions when two tasks receive
updates. Older detail responses can also move a task below another task
after a newer comment promoted it.
**Expected behavior**
Rows remain stable during an update burst. The final activity order
appears after one second of quiet. Old responses never reduce the
recorded activity time.
**Steps to reproduce**
1. Enable the streamlined UI and open multiple tasks.
2. Send alternating updates to two tasks in the recent list.
3. Observe row order while those updates arrive and while older detail
data refreshes.
**Paperclip version or commit**
Base commit:
canary/v2026.907.0-canary.5
|
||
|
|
1cc45086d3 |
feat: use the responsible person's GitHub for shared agent operations (#13005)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Several people can send instructions to the same agent and task. > - A fixed GitHub token in the provider process can keep the first person's access after another person's message is accepted. > - Task ownership cannot select credentials for each accepted instruction or preserve the identity of an operation already in progress. > - This pull request records ordered execution identity contexts and resolves credentials when managed Git, gh, or GitHub tools start. > - The benefit is automatic personal GitHub access for shared agents, with durable continuation rules and no teammate credential fallback. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: orchestration, connection grants, database, runtime adapters, native runners, and run details. **Problem or motivation** A shared agent must use the person whose instructions it has accepted. A queued message must retain its author. A retry or approval without new instructions must retain the originating identity. GitHub must remain optional for ordinary work. **Proposed solution** Persist execution identity separately from task ownership. Give new processes a run-scoped broker capability and token-free managed launchers. Capture identity at operation start. Keep an explicit dedicated-agent grant as an override. Show redacted diagnostics in run details. **Alternatives considered** Per-task ownership, fixed provider tokens, and mutable repository author configuration do not handle accepted steering or concurrent operations. A manual account-selection action would add unnecessary setup to each turn. **Roadmap alignment** This completes the existing Multiple Human Users, MCP Tool Gateway & Apps, Secrets Manager, and Self-healing Runs capabilities. The implementation follows the maintainer-approved plan. Related work: Refs #12843, Refs #12907. Existing proposals #4618 and #8945 cover per-agent or per-worktree author configuration. This change instead follows the accepted human instruction across runtime types. Refs #11831 for governed personal connection delegation; this change preserves connection audience checks and does not use standing delegation as a personal credential fallback. ## What Changed - Add durable, ordered identity contexts and active run references. Preserve message authors through consolidation, steering, retries, delegation, approvals, routines, and restart. - Add an authenticated operation-time GitHub credential broker and local/remote managed git and gh launchers. Keep personal tokens out of the long-lived provider process. - Resolve GitHub gateway and server-side Git operations through the same responsible-person or dedicated-grant selection rules. - Make absent and unavailable GitHub credentials non-blocking at generic startup. Clear host and prior-person credentials. Keep anonymous Git access where supported. - Add run-detail identity history and the dedicated-account warning. Keep task ownership and queue-versus-steer decisions unchanged. - Preserve personal OAuth declarations through connection edits. Retain exact selected grants in the gateway. - Fix continuation races found during real acceptance: verify a warm owner before credential rotation, and wait for bounded durable runner suspension before the next run starts. - Make migrations replay-safe. Retain identity through agent/run deletion, remove it with its company, and clean terminal launcher directories before releasing execution environments. Document coordinated release and rollback. ## Verification - Full workspace typecheck, build, and token gates passed. The complete local suite passed in its normal test groups: 17,120 passing tests, including all 143 serialized server suites. After integrating the newly merged runner API work, full local typecheck and build passed again, along with 890 focused integration tests. All 31 checks on the integrated revision passed, including build, browser E2E, release registry, canary dry run, typecheck, security and all test suites. Greptile is 5/5 with all review threads resolved. - Current focused checks passed: 142 native executor tests, 67 runtime lifecycle tests, 9 durable identity tests, 75 credential/routine tests, 19 low-trust/resumption tests, and the executable migration replay test. - Authenticated browser acceptance with two Paperclip users and two GitHub accounts on one shared native agent passed. Real commits and pushes followed A → B accepted steering → queued A continuation in the same saved conversation. GitHub commit author and committer identities matched all three operations. Both runs succeeded and task ownership stayed unchanged. - Real GitHub MCP calls switched from A to B after accepted steering. A delegated subtask retained its originating identity across a server restart. - Disabling B's GitHub connection left ordinary work successful. Managed gh was unauthenticated and the provider had no inherited GH_TOKEN or GITHUB_TOKEN. - The browser displayed run-detail diagnostics and the exact dedicated-account warning. A final controller-restart check followed by another-person continuation retained the conversation, selected the correct GitHub login and Git author, and removed each terminal launcher directory. - Company-lifetime migration and all five previously failing CI suites passed locally (167 tests). Same-token gateway A → B → A and six broker/launcher boundary tests passed. - Remote callback, launcher, sandbox, and runtime contract tests passed. Both native and legacy Codex completed actual Daytona executions on the integrated revision ([campaign results](https://github.com/paperclipai/paperclip/actions/runs/34155056509)). The remote package-manager shim staging regression also passed locally. ## Risks - Deploy the migrations, server broker, launchers, and runner artifacts together. Existing processes finish with their original contract. New managed processes need the broker endpoint for GitHub operations. - Finish or stop new managed executions before rolling application code back. Keep the additive schema and identity history during rollback. - Scripts that require a persistent raw GH_TOKEN must use managed git, gh, or GitHub gateway tools. Run capabilities authorize code executing within that run to acquire its current identity; this is not hostile-code isolation within one execution principal. Managed commands prevent automatic credential carryover; arbitrary code deliberately copying a credential is outside that boundary. - Uncertain steering acknowledgement deliberately holds new credential acquisition until reconciliation. Already-started operations retain their captured identity. - GitHub private access and provider outages can still fail the specific operation that needs them. Dedicated grant failure does not fall back to personal access. ## Model Used OpenAI GPT-6 through Codex assisted implementation, review, shell execution, and browser acceptance. The exact model variant and context-window size are not exposed in this session. Tool use included TypeScript and Rust tests, database integration tests, GitHub CLI, and authenticated browser control. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5bddff0920 |
feat(runner): add guarded API search and call fallback (#13003)
## Thinking Path > - Paperclip manages AI agents and their work. > - The new runner gives agents dedicated tools for common tasks. > - Some API operations and parameters have no dedicated tool. > - Agents need a controlled way to find and use those operations. > - This pull request adds API search and calls through the real server routes. > - Existing tools remain the preferred path. The new tools are disabled by default. > - Paired tests measure correctness, tool choice, cost and time. ## Linked Issues or Issue Description **Subsystem affected** Paperclip Runner contracts, production tool authority and the server API catalog. **Problem or motivation** The runner cannot use much of the API described by the old Paperclip skill. A generic HTTP client would also let agents bypass runner control rules. **Proposed solution** Add `search_api` and `call_api`. Resolve calls from the mounted API catalog. Use server-held, run-bound credentials. Preserve route checks and runner lifecycle rules. Keep the tools disabled until an operator enables selected companies. **Alternatives considered** A dedicated tool for every endpoint would add a large initial prompt. An unrestricted HTTP tool would weaken authorization and replay controls. **Roadmap alignment** This extends the native runner tooling. The repository owner requested this design and implementation. The roadmap and related open PRs were checked. No duplicate API escape-hatch PR was found. ## What Changed - Register two compact fallback tools in canonical contracts and provider projections. - Build deterministic API discovery from OpenAPI, mounted experimental routes and the old skill reference. - Execute bounded JSON, text, file and download requests through authenticated HTTP routes. - Recheck active runs, company access and work modes. Block runner lifecycle, scheduling, credential and approval bypasses. Keep routine annotation collaboration available. - Retain mutation receipts. Report uncertain outcomes without blindly repeating writes. - Add a company rollout gate and a durable eval worker with complete cost accounting checks. - Record child-task creation in the activity log with the agent and run. - Add contract, authorization, file, replay and real runnerd/PRP/HTTP tests. - Document rollout gates and paid coverage limits. The companion eval repository retains immutable attempts and reports. ## Verification - Final app commit `da58370524c3626a744eec20164397c5fb6ba9ef`: all 32 checks passed; the unrelated Storybook visual check was skipped. Greptile 5/5; no unresolved review threads. - Full Linux build and recursive typecheck passed. Repository tests were run by project and serialized shard; all 143 serialized server suites passed. - Runner TypeScript: 1,599 passed, two skipped. Rust release: 451 passing test reports. Conformance and replay parity passed. The required API check passed 837 tests, including runnerd → PRP → authority → real HTTP. - Bindings cannot enable API tools without the explicit deployment flag. Unit and real-authority tests prove the default-off boundary. - The standalone API check builds and stages its own binary. It passed after existing staged and debug binaries were removed from the test container. - UI and CLI tests passed. Initial environment failures (missing jq, Docker overlay file identity, and parallel linker memory pressure) and focused passing reruns are retained. The macOS full runner suite has platform-specific failures; Linux is the qualified full-check platform. - Eval harness: 27 tests passed; existing CI discovery ran 86 tests with two unrelated skips. Credential export rejection is tested against the actual report command. - Luna and OpenRouter Sonnet each passed 60 common-workflow runs: ten workflows, three repetitions per arm, zero unnecessary API fallback. - Sonnet passed 11 selected capability/contract cases after fixes. Gemini passed three smoke cases. DeepSeek exceeded the 120-second limit and remains unqualified. - Luna's two cost flags received focused follow-up. The original flags and a later n=1 latency flag remain visible. Sonnet had no cost or latency increase above 20%. - The catalog contains 785 entries; 58 were exercised across all stages. Most operation probes remain unrun and some need additional fixtures. Authored probes do not establish successful coverage. - Total conservative accounted cost: $9.875960. Active paid-campaign time: 88.16/90 minutes. No missing accounting. Later security and harness fixes have provider-free verification; no paid validation is claimed for those revisions. - Inspect the [qualification report](https://github.com/paperclipai/paperclip-evals/blob/codex/seach-call-api-tools/evals/runner-api-tools/reports/2026-09-07-production/READINESS.md) and [verification record](https://github.com/paperclipai/paperclip-evals/blob/codex/seach-call-api-tools/evals/runner-api-tools/reports/2026-09-07-production/verification.json). ## Risks - This is a broad authenticated API surface. Keep the default-off gate until an operator selects initial rollout companies. - Paid coverage is incomplete. Small regression samples do not prove all workflows are unchanged. - A timeout or server failure can follow a committed mutation. The result reports an unknown outcome and requires state inspection. - The new definitions add prompt tokens. The report retains cost flags and cache variation. - No database migration is required. - Repository rules require code-owner approval before merge. Technical CI and automated review are complete. ## Model Used OpenAI Codex based on GPT-6 assisted with code, tests and review. The exact serving model ID and context window are not exposed in this session. It used reasoning, tool calls and code execution. Eval models: `gpt-5.6-luna` with low reasoning, `openrouter/anthropic/claude-sonnet-5`, `openrouter/google/gemini-3.8-flash`, and `openrouter/deepseek/deepseek-v4-flash-0731`. Attempts retain runtime versions, model identity, usage and source provenance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.907.0-canary.4 |
||
|
|
392ab26b1e |
chore(lockfile): refresh pnpm-lock.yaml (#12999)
Record the CI-generated dependency graph for #12994. Co-Authored-By: lockfile-bot <lockfile-bot@users.noreply.github.com> Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.907.0-canary.3 |
||
|
|
f6a211479f |
fix: share current CLI runtimes across sandbox adapters (#12994)
## Thinking Path - Paperclip Runner needs its runtime preinstalled for fast sandbox startup. - Native and local adapters should launch one current CLI installation per provider. - An older global copy can shadow that installation, and exact native compatibility pins must match it. - Update the qualified releases and binary digests, expose shared CLI entrypoints from the provider pack, and prefer the image-owned bin directory. - Keep dependency installation in the image build; task startup only discovers, links, and verifies artifacts. ## Linked Issues or Issue Description **What happened?** Remote native startup rejected a stale global Codex, while CLI-only images lacked runnerd entirely. **Expected behavior** An image-baked runtime starts without uploading binaries or installing packages. All adapters share the same current provider CLI. **Steps to reproduce** Start a native remote task with the old global Codex and the updated runtime available only under `/opt/paperclip-runner/bin`. **Paperclip version or commit** Discovery behavior at `54a99d884`. **Deployment mode** Docker with a remote sandbox. ## What Changed - Prefer `/opt/paperclip-runner/bin`, then the user's local bin directory, then PATH. Existing metadata and version validation remains in force. - Qualify Codex 0.153.4, OpenCode 1.18.29, and Claude SDK 0.3.263 / CLI 2.1.263. Update binary digests, TypeScript/Rust checks, registry defaults, and the displayed OpenCode version together. - Share Codex and Claude's native executable with the ACP bridges through exact dependency overrides. Preserve the separately qualified ACP bridge implementations and their security patches. - Expose shared provider-pack CLI launchers; fail the pack build if Codex ACP resolves a separate Codex installation. Update the eval image's other agent CLIs to current stable releases and remove duplicate global provider installs. - Document the single-current-CLI policy in source comments and development guidance. Latest stable releases are resolved at review/build preparation and pinned; task startup never auto-updates. ## Verification - Native-session and adapter-registry suites: 158 tests passed. - Provider suites: 88 tests passed, 7 Linux-only checks skipped on macOS. One existing macOS temporary-path alias assertion passed when rerun with canonical `TMPDIR=/private/tmp`. - Package-contract and OpenCode materialization tests: 11 passed. - Full typecheck, build, and token gates passed. Rust native-provider/recovery tests: 19 passed. - Broad local suite: 5,974 passed, 23 failed, 41 skipped. Failures are in unchanged macOS workspace/path/port and connection suites; focused runtime tests pass. All latest-head Linux PR checks passed, including the full test shards, typecheck, build, runner verification, browser suites, and canary dry run. - The standalone fleet image built with one current provider CLI each and passed native Codex/Claude binary-integrity checks. A disposable Daytona sandbox reported ready in 798 ms; its baked runner completed an API-key `gpt-5.6-luna` turn in 2,430 ms and returned the expected marker with a usage receipt. No runtime artifacts were uploaded or installed. - The normal shared `codex exec` entrypoint also completed an API-key `gpt-5.6-luna` turn in 2,321 ms. - Both image builds verify the complete generated lockfile against a reviewed SHA-256 before package installation or lifecycle execution. Root lockfile changes remain CI-owned. Merge and rollout remain on hold for operator review. ## Risks - Updating provider CLIs changes their behavior for all adapters; version probes and live native smoke testing are required before image promotion. - The image-owned directory takes precedence. Its entries must launch the same shared CLI as the global PATH, not a private older/newer copy. - Application qualification pins and the deployed image must move together. No startup fallback installation is added. - No schema or authentication-policy changes. ## Model Used OpenAI GPT-6 (Codex). The session does not expose a more specific model ID or context-window size. Used reasoning, repository inspection, code execution, and browser verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
932ddb7b37 |
feat: browse GitHub repository access across organizations (#12998)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - GitHub connections give agents access to approved repositories. > - One GitHub identity can use installations across several organizations. > - The permissions page linked to one installation and showed an unfiltered list. > - Users could not easily find another organization or inspect a large selection. > - This change adds account filtering, search, and access configuration links. > - Users can inspect repository access in one compact view. ## Linked Issues or Issue Description Refs #12993. Related repository-catalog work in #11228 and #11234 was checked. This change only improves the existing GitHub connection permissions page. **What existing behavior does this improve?** The GitHub connection permissions page and its repository display metadata. **Current behavior** The page links directly to an existing installation. The repository list has no account filter, search, height limit, or private-repository marker. Refresh access occupies a separate section. **Proposed behavior** Show all authorized repositories by default. Filter by account or organization and search by name. Open GitHub's account chooser to configure access across organizations. Show GitHub icons and private-repository locks. Keep refresh beside configuration and limit the visible list to about ten rows. **Reason and benefit** Users can find repositories across organizations and configure missing access without creating another GitHub identity. Large repository lists no longer fill the page. **Breaking changes** None. Repository display metadata gains an optional private flag. Older snapshots remain valid and gain the flag after access refresh. No SQL migration is required. ## What Changed - Add an All accounts view, account filter, search, and empty states. - Link both configuration controls to GitHub's app account chooser. - Place an accessible refresh icon beside the configuration button. - Keep the repository heading and list in one section. - Add GitHub icons and private-repository locks. - Cap the scrollable list at ten rows using a design token. - Persist GitHub's private flag only when the provider returns a boolean. - Recover missing legacy app configuration from GitHub installation metadata. - Update tests and the GitHub connection runbook. ## Verification - Focused tests passed: 54 permissions-page tests and four GitHub metadata tests. - UI and server typechecks passed before submission. Token gates passed. - Browser checks verified account filtering, search, empty results, and the configuration destination. - The live list contained 40 repositories. Its final height was 272 pixels, which fits ten single-line rows with gaps. Scrolling retained all rows. - A live access refresh populated 30 private-repository lock icons from GitHub metadata. - Full workspace typecheck and build passed. The broad local suite stopped in the general-server group with 18 failed files. Failures include macOS temporary-path handling and embedded PostgreSQL startup. That run also overlapped the legacy fix and retained a stale GitHub module; the final focused run passed all 58 tests. Clean-runner CI is tracked separately. - Latest-head review is 5/5 with the legacy chooser finding resolved. All CI checks passed on commit `0ff2b63f348f5c87d8b7df6e43388f60f5d872d9`, including build, typecheck, all test shards, browser tests, and canary dry run. ## Risks - Older repository snapshots lack visibility metadata until refreshed. Unknown visibility does not display a lock. - The account filter lists authorized installation owners. Users add other organizations through GitHub's chooser. - Filtering changes only the displayed list. GitHub remains authoritative for repository access. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) via Codex. Reasoning, code execution, and browser tools were used. The exact context window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bac60d9d31 |
fix: preserve GitHub sign-in and show connected repository access (#12993)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - GitHub connections give agents an account with selected repository access. > - Fresh local instances enroll with production Paperclip Cloud. > - Enrollment could finish while the GitHub OAuth profile remained disabled. > - Setup then switched to a personal access token form without explanation. > - This change preserves sign-in intent and shows the connected account and repositories. ## Linked Issues or Issue Description Related: #12907, #12943, #12947. Existing open GitHub connection work was checked. No duplicate was found. **What happened?** After Cloud enrollment, a fresh test-drive asked for a GitHub key. Production did not advertise the managed GitHub profile. Staging did. The permissions page also omitted the authenticated username and repository names. **Expected behavior** Continue with GitHub OAuth when available. Explain unavailable sign-in and allow retry otherwise. Show the GitHub username and complete accessible repository list. **Steps to reproduce** Start a fresh test-drive. Choose GitHub and complete instance enrollment while the Cloud GitHub profile is disabled. Open an existing GitHub connection's permissions page. ## What Changed - Preserve managed sign-in intent when the gallery omits its profile. - Refresh the selected gallery entry on retry without resetting the audience. - Fetch all pages of GitHub installations and repositories. - Store only repository IDs, full names, and installation IDs in grant metadata. - Show the GitHub username, repository list, management link, and refresh action. - Discard the repository snapshot after newer installation lifecycle events. Preserve snapshots verified after delayed events. - Lock and re-read grant metadata when applying installation events or saving refreshed access. Patch only webhook fields for other events. Reject snapshots if access changed during the external fetch, using unique access revisions even when timestamps collide. - Show repository installation recovery for managed OAuth even when the app also offers an advanced PAT method. - Update tests and the GitHub connection runbook. No SQL migration is required. ## Verification - Local typecheck, build, and token gates passed. All latest-head CI gates passed, including the complete test matrix and browser suites. Greptile is 5/5 with no unresolved findings. - All 382 focused setup, permissions, metadata, service, and webhook tests passed across final runs. One socket-hang-up test passed on rerun with the full service suite. Final service, metadata, and webhook checks passed all 230 tests. - The broad local suite was stopped after failures. Seven workspace-runtime exposure and control-conflict failures reproduce on base commit `54a99d884`. The broad run also overlapped local iteration; final focused tests and clean-checkout CI are tracked separately. - Browser: a fresh production-backed instance completed enrollment, retried after profile enablement, reached GitHub consent, recovered from a missing installation, and completed OAuth. - Browser: the permissions page showed the authenticated username and the selected private test repository. A real `get_me` call returned the same account. Reading the selected repository passed; reading an unselected private repository failed with 404. - Browser: a second fresh instance completed enrollment and OAuth without a PAT form or unavailable state. Its username and repository list survived reload and refresh. A real get_me call on the final code returned the displayed account. ## Risks - Repository names are now stored in company-scoped grant metadata and shown with that credential. They are display data, not authorization data. - Large selections require more GitHub API calls. A failed later page rejects the refresh rather than reporting a partial list. - Older grants and webhook-invalidated snapshots require Refresh access to load the list. - Cloud profile enablement is separate deployment configuration. This PR does not change OAuth scopes or GitHub App permissions. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) via Codex. Reasoning, code execution, and browser tools were used. The exact context window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.907.0-canary.2 |
||
|
|
54a99d8840 |
fix(evals): make the chat viewer the default published Evalbook (#12952)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Direct Runner evals retain evidence across model configurations. > - Evalbook already has a grid and a read-only Runner Lab chat viewer. > - Public projection stripped the view and selected a second plain result page. > - This change uses the existing viewer for public and private results. > - The data access differs, but the presentation does not. ## Linked Issues or Issue Description Refs #12931, #12945. Related open runtime-contract PR #11634 does not contain this report-only change. **What happened?** The public direct-eval campaign opened plain result pages. The access-controlled artifact used the chat viewer. Users could not follow the same recorded interaction from the published grid. **Expected behavior** Every newly generated Runner Evalbook opens the existing chat viewer. The grid and durable run history remain. Public evidence has explicit redactions. **Steps to reproduce** Open campaign gha-34062394019-1 from the direct-eval history. Click a result, then compare its plain page with the corresponding Actions artifact. **Paperclip version or commit** Reproduced atcanary/v2026.907.0-canary.1 |
||
|
|
ae03465ad4 |
fix(onboarding): drop the sign-in card's Cancel, and resume after Back (#12958)
The card carried a Cancel beside its instruction, directly above the step's own Back. Two ways out of one screen is one too many. Removing it was only safe if returning genuinely resumed the running session, and writing the test for that disproved the premise: coming back started a second login rather than adopting the first, which the server's per-owner lease rejects. Cancel was not duplicative — it was the only way out of a login the customer could no longer reach. The cause was the cache. The active-session read is a useQuery with no staleTime, so on the second mount it answered instantly from the first mount's cached null: isFetched was true, the resume effect concluded there was nothing to adopt, and autoStart began a new login while the refetch was still in flight. Both panels did it, and it was pre-existing on master. Both reads now refuse the cache, so the resume works and the removal is safe because of it. Removing the button also stranded the chain behind it: handleCancel in both panels, the onCancel prop on AdapterLoginPanel, and the wizard's wiring. The panel chrome's own Cancel is untouched. An abandoned login is collected on a five-minute timer (DEVICE_LOGIN_TIMEOUT_MS) with the reaper as the restart-safe backstop, and the lease is keyed on adapterType, so a source switch does not contend. Both covered by tests.canary/v2026.907.0-canary.0 |
||
|
|
856813ba3a |
fix(connections): distinguish local setup from provider handoff (#12947)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give those agents access to external services. > - Connection setup first asks who may use the credential. > - Some providers need a local method selection before OAuth starts. > - The Access button said it would open GitHub even when it opened another local step. > - This pull request names the actual next action and shows the external arrow only for a provider handoff. > - Users can distinguish local setup from leaving Paperclip. ## Linked Issues or Issue Description Refs: #12943 The browser audit found a second, separate clarity problem. GitHub showed "Continue to GitHub" twice: once to open its local method selection and once to start OAuth. The first label promised the wrong action. This change fixes that label without adding or removing a setup step. ## What Changed - Use "Continue" when Access opens another local OAuth setup step. - Keep "Continue to <provider>" and the external arrow when Access starts OAuth directly. - Test GitHub defaults, pre-enrollment setup, direct Notion OAuth, and dialog behavior. - Document the two GitHub button actions. - Preserve the external-handoff arrow in the direct-OAuth Storybook fixture. ## Verification - Targeted connection tests: 114 passed. - `pnpm --filter @paperclipai/ui typecheck`: passed. - `pnpm --filter @paperclipai/ui build`: passed. - `node scripts/check-token-gates.mjs`: passed. - `git diff --check`: passed. - After the Storybook review fix: 100 connection-flow tests, UI typecheck, token gates, and Storybook build passed; visually confirmed the Gmail handoff arrow and local GitHub step labels. - Actual test-drive browser: confirmed default personal identity and Any agent; clicked Continue to the local method screen; switched PAT and OAuth; returned to Access with selections intact; cancelled without a duplicate connection. - Live OAuth, local MCP identity, and a real staging sandbox task passed during the audit. This patch does not change those paths. - Full local suite has known unrelated macOS path/runtime test failures from the audit. All CI passed on head `936ce09f4e1faaf09f9fb778c2cd4c65c2a41972`, including typecheck, build, all test shards, browser tests, and canary dry run. Greptile: 5/5; no unresolved comments. ## Risks - Low risk. This changes button text and an icon cue only. There is no new schema, permission, consent, or credential behavior. - Providers with multiple methods now say Continue before their method screen. Direct OAuth providers keep their existing wording. ## Model Used - OpenAI Codex assisted with code, tests, terminal checks, and actual browser verification. The runtime does not expose the exact model ID or context-window size, so those details are unavailable. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted tests; full-suite caveats above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.906.0-canary.5 |
||
|
|
83987210d6 |
fix(runner): align direct eval provider setup with qualified runtimes (#12945)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Runner direct live evals test semantic tools against a mock control
plane.
> - The first complete AWS campaign exercised 358 cells.
> - It exposed setup differences from the working full-stack harness.
> - This pull request corrects those direct-harness differences.
> - It preserves production permission defaults and the full-stack
workflow.
## Linked Issues or Issue Description
**What happened?**
Native Codex cells could not find a global codex executable. ACPX denied
unattended tool requests and lost valid provider usage receipts.
AgentCore hit a 30-second facade timeout while its worker allows 120
seconds for delivery. Several models reported a native run result
without updating the separate mock task state.
**What did you expect?**
The direct harness should use the pinned executable, explicit test
permissions, and a timeout compatible with the provider delivery
contract. Its instructions should explain which operation changes mock
task state.
**Steps to reproduce**
Run the full Runner Direct Live Protocol Evals workflow. Baseline
campaign:
https://github.com/paperclipai/paperclip/actions/runs/34059009921.
**Paperclip version**
Master at
canary/v2026.906.0-canary.4
|
||
|
|
d293dd3d14 |
fix(connections): refresh expired pending enrollment links (#12943)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Managed connections deliver credentials to an approved instance.
> - A self-hosted instance starts with a time-limited approval link.
> - The setup page cached that link and reused it after expiry.
> - This pull request lets the existing server choose a valid enrollment
link.
> - Users can recover without another setup step or a service restart.
## Linked Issues or Issue Description
Related: #12891 (one-time enrollment) and #12907 (sandbox GitHub
identity).
No duplicate open enrollment-recovery PR or issue was found.
**What happened?**
After a pending Cloud enrollment expired, Continue opened the same
expired
approval link. Cloud returned ENROLLMENT_NOT_AVAILABLE. Returning to
setup and
selecting Continue repeated the failure.
**Expected behavior**
Continue must ask the server for a valid enrollment. The server must
reuse a
live pending enrollment and replace an expired one. Existing approved
instances
must not need another approval.
**Steps to reproduce**
1. Start GitHub setup on a fresh self-hosted instance.
2. Start Cloud enrollment but leave it unapproved until the link
expires.
3. Return to step 2 and select Continue.
4. Before this change, the browser opens the expired link again.
**Paperclip version or commit**
Reproduced from master commit
canary/v2026.906.0-canary.3
|
||
|
|
b55ce03e86 |
chore(lockfile): refresh pnpm-lock.yaml (#12931)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com>canary/v2026.906.0-canary.2 |
||
|
|
fee8d8dc39 |
fix(runner): repair direct live provider bootstrap (#12932)
## Thinking Path > - Paperclip runs AI agents through qualified provider backends. > - The direct live eval workflow builds one immutable Runner runtime for every matrix cell. > - The workflow reinstalled the packed Runner with npm. > - That install discarded pnpm patches and selected provider dependencies outside the qualified lock. > - The first pnpm deployment model also placed its virtual-store marker at the wrong level; a real deployment keeps `.pnpm` beside the scoped Runner package. > - AgentCore enforced the current context-aware harness but the direct eval CLI did not supply the production v3 runtime context that harness requires. > - This pull request preserves the qualified dependency graph, resolves the real deployment layout, and makes direct evals exercise the production runtime-context contract. > - The benefit is that live eval cells reach their provider turn with the same artifacts and context contract that Paperclip qualified. ## Linked Issues or Issue Description Refs: #12931 **What happened?** The full direct live eval campaign failed every ACPX cell during `session.open`. The portable runtime had an incorrect dependency root. Its npm install also discarded the qualified ACP server patches. AgentCore cells first failed because Runner enforced `aws-agentcore-harness-v1` while the provisioned stack and eval profile use `aws-agentcore-harness-context-v2`; after aligning that revision, the direct eval CLI still omitted the required v3 runtime context. **Expected behavior** The direct eval runtime must preserve the frozen pnpm dependency graph and patched provider bytes. Runner, server validation, OpenAPI, and the deployed AgentCore stack must use one qualification revision. Direct eval attempts must supply the same immutable native runtime-context contract as production. **Steps to reproduce** 1. Dispatch `Runner Direct Live Protocol Evals` from `master`. 2. Select an ACPX Claude, ACPX Codex, or AgentCore roster. 3. Observe a pre-turn provider bootstrap failure. **Paperclip version or commit** `d96452db059338b329b458ba8fe359fef72f1363` **Deployment mode** GitHub Actions on the RunsOn Linux x64 fleet. ## What Changed - Build the reusable direct-eval runtime with `pnpm deploy --prod`. - Resolve ACPX dependencies from the actual scoped-package layout of a self-contained pnpm deployment. - Align AgentCore configuration and qualification checks on `aws-agentcore-harness-context-v2`. - Materialize a minimal immutable v3 runtime context for each isolated direct eval attempt. - Add workflow, package-authority, runtime-context, Rust, and server regression coverage. - Document the qualified packaging, runtime-context, and AgentCore revision contracts. ## Verification - `pnpm --filter @paperclipai/paperclip-runner typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/live/runnerd-codex-transport.test.ts` (70 tests) - `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/cli/eval-session-contract.test.ts` (14 tests) - Focused Runner contract tests (36 tests) - Focused server profile tests (47 tests) - Focused Rust managed-provider and native-selector tests (19 tests) - `node --test packages/paperclip-runner/scripts/runner-protocol-eval-workflow-security.test.mjs` - `actionlint .github/workflows/runner-protocol-live-evals.yml` - A local `pnpm deploy --prod` produced both qualified ACP server digests. - A Linux reproduction of the first follow-up smoke identified the real deployment root and the missing AgentCore runtime context. ## Risks The AgentCore revision change rejects profiles that still use the obsolete v1 value. This is intentional because the provisioned context-aware harness and current eval profile use v2. Direct eval prompts now receive the same fixed runtime-context preamble as production, so behavior scores may move; that is the intended qualification surface. The workflow package layout changes, but tests assert the new entrypoint and dependency root. This change does not modify the browser full-stack E2E workflow. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5.6. The context-window size is not exposed in this session. The model used extended reasoning, repository tools, code execution, Docker-based Linux reproduction, and GitHub Actions diagnostics. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (for example, `docs/...` or `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
05735b3d87 |
fix(ui): show continuation actions in confirmation receipts (#12939)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task interactions let an operator approve completion or ask an agent to continue. > - The server stores the continue choice as a rejected completion request so it can resume the same task. > - The task feed ignored the configured action label and showed the generic text `Declined request`. > - The generic text made a successful three-turn continuation look like a failed request. > - This pull request keeps the server state and shows the action that the operator selected. > - The benefit is an accurate task feed for legacy Codex and Runner Codex. ## Linked Issues or Issue Description **What happened?** A rejected confirmation always appeared as `Declined request`. The feed did not use a custom rejection action such as `Continue work`. **Expected behavior** The resolved receipt and success toast must show the selected custom action. Confirmations without a custom action must keep the current fallback text. **Steps to reproduce** 1. Create a completion confirmation with `rejectLabel` set to `Continue work`. 2. Select `Continue work` and enter a continuation note. 3. Open the completed task feed. 4. Observe that the old UI says `Declined request` instead of the selected action. **Paperclip version or commit** `539c9212f4b98e37643f5a8e3b603f1b5845b5d7` **Deployment mode** GitHub Actions Runner E2E with warm Daytona sandboxes. ## What Changed - Show `Selected “Continue work”` when a rejected confirmation has that custom action label. - Use the same action-aware text in the success toast. - Keep `Declined request` as the fallback for confirmations without a custom rejection label. - Make the legacy warm-turn prompt ask if the task is ready to complete. - Require both warm Daytona matrix cells to show two continuation receipts and no generic decline receipt. ## Verification - `pnpm exec vitest run ui/src/pages/IssueDetail.test.tsx ui/src/lib/issue-thread-interactions.test.ts ui/src/components/task-chat/TaskChatInteractionCard.test.tsx` - `pnpm test:e2e:runner:unit` - `pnpm test:e2e:runner:typecheck` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/ui build` - `pnpm check:token-gates` - Paid `daytona-warm-continuity` campaign: both three-turn cells passed on attempt 1 ([workflow run](https://github.com/paperclipai/paperclip/actions/runs/34052047946)); the downloaded evidence aggregates locally as 2/2. The trusted report job did not publish because default branch `pnpm-lock.yaml` was transiently behind its manifest. - `pnpm typecheck` reached an unrelated `plugin-workspace-diff` dependency type error after the current master manifest and lockfile resolved different versions. The focused UI and runner checks pass. ## Risks - Low risk. The database state and continuation behavior do not change. - A custom rejection label now appears in resolved receipts and success toasts. - The paid warm Daytona suite has a stricter browser assertion. ## Model Used - OpenAI Codex, GPT-5, with reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in this PR with the bug template fields - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal task identifier - [x] I have run focused tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated the relevant E2E fixture and assertions - [x] I have considered and documented the risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
539c9212f4 |
fix(onboarding): keep the pasted code on screen, and make the auto-copy real (#12921)
The Claude card cleared its field the instant the paste was submitted, so the code appeared and vanished in the same breath and the seconds that followed had no feedback at all - it read as the paste having been dropped. The code stays now; the field is disabled and the step's button carries the status. Scoped to the onboarding chrome: the panel in the agent form stays open long after its login and has its own status area, so clearing is still right there. The OpenAI card's auto-copy took its latch before the attempt rather than on success, so the first refusal was permanent - and the first attempt is the one most likely to be refused, since it fires while the card is still arriving. It also ran with an empty code on the paths where the row renders before the server's one-time prompt has a value, which wrote nothing and then said Copied! over whatever the clipboard already held. It now waits for the document to have focus, retries when focus returns, guards against a second attempt while one is in flight, and only latches once a write has actually landed. The arc's question is a quarter larger, 36px to 45px, as a token with its own leading. paperclipai/paperclip-cloud#400 moves the signup wizard's matching heading by the same quarter - the two flows are pinned to each other on purpose. |
||
|
|
d96452db05 |
fix(runner): restore Vite 6 viewer compatibility (#12929)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner includes an issue-thread viewer for direct Evalbook reports. > - The runner package uses Vite 6.4.3. > - Dependabot changed the React plugin from version 4.7.0 to version 6.1.1. > - React plugin 6.1.1 requires Vite 7 or Vite 8 package internals. > - The live evaluation workflow found this mismatch when it built the viewer on a clean Linux worker. > - This pull request restores the compatible plugin and adds the viewer build to pull request CI. > - The benefit is that CI detects this class of report-viewer build failure before a paid evaluation campaign starts. ## Linked Issues or Issue Description **What happened?** The direct live evaluation workflow failed before model execution. The `build:issue-thread` command could not load `@vitejs/plugin-react@6.1.1` with Vite 6.4.3. The plugin imported the unavailable `vite/internal` package path. **Expected behavior** The Evalbook issue-thread viewer must build from a clean frozen-lockfile installation before the live evaluation matrix starts. **Steps to reproduce** 1. Check out commit `165ca56a22adb60e5fda56045442d9c8498116a8`. 2. Run `pnpm install --frozen-lockfile --ignore-scripts`. 3. Run `pnpm --filter @paperclipai/paperclip-runner build:issue-thread`. 4. Observe the `ERR_PACKAGE_PATH_NOT_EXPORTED` error for `vite/internal`. **Paperclip version or commit** `165ca56a22adb60e5fda56045442d9c8498116a8` **Deployment mode** Built from source in GitHub Actions on Ubuntu. ## What Changed - Restore `@vitejs/plugin-react` 4.7.0 in the Vite 6 runner package. - Follow the repository policy: trusted PR CI regenerates and verifies the lockfile artifact, and the master refresh workflow commits the lock-only update after merge. - Build the Runner Evalbook viewer in pull request CI. ## Verification - `ci / policy`: regenerated the dependency lock artifact successfully - `pnpm --filter @paperclipai/paperclip-runner build:issue-thread` - `actionlint .github/workflows/pr-trusted.yml .github/workflows/runner-protocol-live-evals.yml` - `node --test packages/paperclip-runner/scripts/runner-protocol-eval-workflow-security.test.mjs` - `git diff --check` ## Risks Low risk. This change restores the previous React plugin major version for one package. The selected version declares support for Vite 6. Pull request CI now builds the affected viewer directly. > This bug fix does not add or change a roadmap feature. ## Model Used - OpenAI Codex with GPT-5. Tool use and code execution were enabled. The Codex app managed the context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
9ecd93a54d |
test(e2e): link runner campaign summaries (#12927)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip uses a paid full-stack campaign to verify runner behavior across providers and environments. > - The campaign already creates an interactive report, workflow logs, and retained evidence artifacts. > - The merge job summary shows result totals but does not link to those resources. > - Reviewers must search several workflow jobs and artifacts to find the executed cells. > - This pull request adds direct and safe links to the exact campaign, each cell, the workflow logs, and the artifacts. > - The benefit is that a reviewer can inspect a result from the Actions summary with one click. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the `Merge and enforce campaign result` summary in the `Runner Full-Stack E2E` workflow. **Subsystem affected** The runner E2E report generator and its GitHub Actions workflow are affected. **Current behavior** The summary lists each selected cell and its result. It does not link to the published campaign report, the workflow logs, or the evidence artifacts. **Proposed behavior** The summary includes a `View results` section. It links to the exact immutable campaign report, the workflow logs, and the artifacts. Each cell name links to its stable section in the campaign report. **Reason and benefit** The current summary does not show reviewers where to inspect the run. Direct links make the result evidence discoverable without manual URL construction or artifact searches. **Breaking changes** None. This change only adds links and stable HTML anchors to existing report output. **Additional context** Related: #12904. The cited successful campaign is [run 34026735033](https://github.com/paperclipai/paperclip/actions/runs/34026735033). ## What Changed - Add a safe URL builder for public campaign, workflow, and artifact links. - Add a `View results` section to the GitHub Actions campaign summary. - Link each summary table cell to its exact section in the immutable campaign report. - Add stable execution anchors to the generated dashboard. - Reject non-HTTPS, credential-bearing, malformed, and ambiguous link destinations. - Document the new links and their retention or publication timing. ## Verification - `pnpm test:e2e:runner:unit` — 116 tests passed. - `pnpm test:e2e:runner:typecheck` — passed. - `pnpm typecheck` — passed, including migration safety. - `pnpm build` — passed. - `pnpm exec prettier --check ...` for all changed files — passed. - `git diff --check origin/master...HEAD` — passed. - The full local server suite also ran. One unrelated macOS workspace-runtime file passed 157 tests and failed 4 existing path and port assumptions. Two failures compare `/var` with `/private/var`. Two failures cannot reserve a port outside a hard-coded range. This PR does not change that file or its dependencies. ## Risks - The immutable campaign link becomes available after the history publisher completes. The workflow and artifact links remain available while publication runs. - The artifact link requires GitHub access and follows the existing 30-day retention period. - Invalid configured URLs are omitted instead of being rendered into the summary. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex desktop agent with GPT-5. The runtime does not expose the context-window size. The agent used repository inspection, agentic reasoning, code execution, and GitHub CLI tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/no-internal-issue-references`, `fix/sandbox-secret-resolution`, `feat/adapter-retry-backoff`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
165ca56a22 |
fix(runner): scope live eval tokens to eval repo (#12911)
## Thinking Path > - The merged direct live eval workflow must read the private `paperclip-evals` repository at an exact commit. > - Its first hosted dispatch failed before provider execution because the GitHub App token was minted from the `paperclip` repository installation. > - GitHub returned 404 while resolving the private eval commit, proving that token did not have the required repository scope. > - Minting each short-lived token from the exact private eval repository installation supplies only the cross-repository read boundary the workflow needs. > - A workflow regression now verifies every eval-token block keeps that exact scope. ## Linked Issues or Issue Description The first default-branch run of Runner Direct Live Protocol Evals failed in its immutable eval-commit verification step with HTTP 404. No provider jobs ran and no provider spend occurred. **What existing behavior does this improve?** It allows the protected direct live eval workflow to verify and check out the private `paperclipai/paperclip-evals` repository. **Current behavior** All four eval-token blocks set `GH_REPO` to `paperclipai/paperclip`, selecting a token installation that cannot read the private eval repository. **Proposed behavior** Set `GH_REPO` to the exact `paperclipai/paperclip-evals` repository in authorization, catalog, matrix, and report jobs. **Reason and benefit** The app mints a short-lived token from the correct repository installation while the main repository continues to use its ordinary read-only workflow token. **Breaking changes** None. ## What Changed - Scoped all four private-eval installation tokens to `paperclipai/paperclip-evals`. - Added a regression requiring that exact scope in every token block. ## Verification - `pnpm --filter @paperclipai/paperclip-runner test:runner-protocol-eval-publish` — 15 passed. - `node --test .github/scripts/tests/get-bot-token.test.mjs` — 3 passed. - `actionlint .github/workflows/runner-protocol-live-evals.yml` — passed. - `git diff --check` — passed. ## Risks - The workflow reads a private repository. The token is still short-lived, repository-specific, masked immediately, and used only by the protected default-branch workflow. - This changes no provider execution, Runner behavior, report content, S3 publishing, or browser E2E behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex on GPT-5. The exact deployment ID and context-window size are not exposed. The model used reasoning, repository inspection, code editing, GitHub Actions diagnostics, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation where applicable - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.906.0-canary.1 |
||
|
|
0c1e7504c0 |
fix(runner): persist warm Daytona workspaces across turns (#12904)
## Thinking Path > - Daytona preserves a stopped sandbox filesystem, but deleting or replacing a sandbox removes its only remote copy. > - Warm reuse therefore improves latency but cannot be Paperclip's durability boundary. > - The host execution workspace must remain authoritative after every successful turn, while same-run recovery must avoid overwriting unexported remote work. > - Result proposal, workspace export/merge, and terminal completion need a durable, replayable ordering so a crash never starts a duplicate provider turn. > - A paid browser acceptance suite must exercise both legacy Codex and Runner Codex for three real turns on one continuously warm Daytona sandbox. ## Linked Issues or Issue Description Refs #12901. Runner Codex did not previously export successful Daytona workspace changes back to the authoritative host workspace. That made warm reuse depend on Daytona's remote filesystem and left deleted/replacement sandboxes without a reliable reconstruction path. The existing paid fixture also lacked a focused three-turn continuity case for both Codex adapters. ## What Changed - Persist versioned, atomic native workspace-sync descriptors and durable seeds in `PAPERCLIP_HOME`, without credentials or a database migration. - Classify fresh, warm, replacement, and same-run-recovery workspace preparation explicitly; ambiguous lease/root/digest evidence fails closed. - Finalize native workspace export/merge after semantic result proposal and before run completion, with idempotent replay that never submits a second provider turn. - Surface legacy Codex workspace restoration failures instead of masking them, while preserving an earlier provider error when both fail. - Keep healthy reusable Daytona leases warm for legacy and native adapters, stamp finalized workspace generations, and retain existing cleanup behavior for per-turn or unhealthy leases. - Preserve Runner Codex's provider process/session across warm turns, including bounded post-terminal tail draining and exact authority rotation. - Add the exact paid `daytona-warm-continuity` matrix: - `legacy-codex × daytona × warm-three-turn` - `runner-codex × daytona × warm-three-turn` - Drive all three turns through the browser, verify ordered file continuity and stable lease/workspace/runtime identities, capture per-turn timings, and delete the sandbox immediately after assertions. - Document `pnpm test:e2e:runner -- --suite daytona-warm-continuity`; no package script was added. ## Verification - `pnpm typecheck` — passed, including migration safety (no migration added) - Focused server/runner Vitest coverage — 144 passed - `pnpm test:e2e:runner:unit` — 114 passed - `pnpm test:e2e:runner:typecheck` — passed - `pnpm --filter @paperclipai/paperclip-runner test:codex` — 66 passed, 1 helper ignored - `native-session-executor.test.ts` — 139 passed, including safe fail-closed cleanup after remote runner identity capture failure - Paid local browser acceptance, exact post-rebase Linux/amd64 runner binary: - Runner Codex — passed in 1.7m; 3 runs; lease outcomes `created, resumed, resumed`; 10/10 matchers; cleanup passed - Legacy Codex — passed in 2.7m; 3 runs; lease outcomes `created, resumed, resumed`; 10/10 matchers; cleanup passed - [Protected paid GitHub Actions campaign](https://github.com/paperclipai/paperclip/actions/runs/34026735033) against `7da42a91b95fa7fb2df126668ef7e37afb3b2b9d` — passed 2/2: - Runner Codex — 3 runs; lease outcomes `created, resumed, resumed`; evidence and cleanup passed - Legacy Codex — 3 runs; lease outcomes `created, resumed, resumed`; evidence and cleanup passed - Merge/enforcement, S3 history, and Pages publication jobs passed - Paid result artifacts were scanned for both provider credentials; neither secret was present. - Current PR checks — 31 passed, 1 expected Storybook skip; Greptile 5/5; Superagent security scan passed - `git diff --check origin/master...HEAD` — passed - Confirmed no `package.json`, lockfile, migration, or SQL changes. ## Risks - Workspace synchronization now sits on the terminal-success path, so a remote export failure deliberately prevents false success. Retryable state retains its lease/seed; loss of the only unexported remote copy fails closed. - Warm provider reuse has strict identity and quiescence checks. Mismatched or ambiguous evidence blocks reuse rather than risking concurrent provider work. - The paid suite incurs Daytona and Codex cost only in the existing protected scheduled/manual workflow and explicitly destroys its sandbox after each cell. ## Model Used OpenAI Codex with GPT-5 agentic reasoning, repository inspection, real browser E2E execution, Rust/TypeScript test execution, and GitHub Actions diagnostics. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked an existing issue or described the issue in-PR - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name contains no internal ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have documented the dedicated suite invocation without adding a package script - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green on the current revision - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups on the current revision - [x] I will address all reviewer comments before requesting merge |
||
|
|
af8439a70b |
feat(runner): restore direct live eval campaigns and reports (#12909)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Runner executes agents through native and managed provider drivers. > - The direct live eval layer had drifted from the current Runner contracts. > - The old local workflow did not provide a complete parallel campaign or durable report history. > - The Runner also needed current native OpenCode and OpenRouter qualification. > - This pull request restores the direct campaign, corrects the runtime gaps that the campaign found, and adds safe hosted Evalbook history. > - The benefit is repeatable model comparison against an immutable Runner and eval source revision. ## Linked Issues or Issue Description Refs #11297 Refs #11634 **What existing behavior does this improve?** This improves the direct live `paperclip-runner` eval workflow, provider execution contract, and static Evalbook reporting path. **Current behavior** The direct evals do not have one maintained full campaign on current `master`. OpenCode has no qualified multi-model OpenRouter roster. Parallel provider bursts can compact committed events before the transport observes them. Local reports do not have a separate safe S3 history index. **Proposed behavior** Run one immutable roster-plus-case matrix. Use the shared paid AWS runner fleet. Keep raw artifacts access-controlled. Publish a sanitized canonical Evalbook report under the separate `runner-protocol-evals` S3 prefix. Keep immutable campaign directories plus root history, latest, and latest-green pointers. **Reason and benefit** Maintainers can compare native Codex, native OpenCode, ACPX, Claude Managed, and AWS AgentCore behavior over time. They can inspect failures without mixing this direct protocol layer with browser full-stack E2E. **Breaking changes** None. The new workflow and S3 prefix are additive. The existing Runner full-stack E2E workflow and report remain separate. ## What Changed - Added a trusted two-shard direct live workflow for up to 393 roster-plus-case cells. - Reused the numeric actor allowlist, protected paid environment, and RunsOn fleet controls from Runner full-stack E2E. - Added immutable Runner and eval revision resolution, exact credential boundaries, bounded retries, and cost ceilings. - Added a public report projection that removes sessions, transcripts, tool payloads, state, traces, raw failures, remote profile identities, and credential-shaped values. - Added additive S3 history under `runner-protocol-evals`, with immutable campaigns and mutable root index pointers. - Added native OpenCode model injection and current OpenRouter pricing contracts. - Fixed direct eval completion, workflow execution, semantic discovery, warm-attach state reset, executable binding, and event-burst handling. - Kept Runner browser full-stack E2E behavior and publication separate. - Documented local and hosted direct eval operation. ## Verification - `pnpm --filter @paperclipai/paperclip-runner test:runner-protocol-eval-publish` — 15 passed. - `pnpm --filter @paperclipai/paperclip-runner build:typescript` — passed. - `actionlint .github/workflows/runner-protocol-live-evals.yml .github/workflows/runner-full-stack-e2e.yml` — passed. - Local current matrix at the revision in [paperclip-evals#17](https://github.com/paperclipai/paperclip-evals/pull/17) — 323 cells across 10 enabled configurations completed. - Final local current matrix — 269 passed, 11 behavior failures, and 43 expected macOS-only ACPX platform failures. - Targeted Runner checks — 13/13 eval-session tests, 15/15 publisher/security tests, and package typecheck passed; complete PR CI is green, including all browser E2E shards. ## Risks - Paid live campaigns can consume provider budget. Actor authorization, exact per-cell ceilings, protected environments, and explicit schedule enablement bound this risk. - Public reports can leak provider data. The workflow publishes only a separately projected report and validates every file before upload. - The new workflow cannot publish until it is present on the default branch. This pull request does not change the existing `runner-full-stack-e2e` publication path. - The campaign is large. It uses two GitHub matrices and caps combined concurrency at the shared fleet limit. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex on GPT-5. The exact deployment ID and context-window size are not exposed. The model used reasoning, code editing, browser inspection, repository tools, and live provider execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.906.0-canary.0 |
||
|
|
3796c6f259 |
fix(connections): project GitHub identity into sandbox runners (#12907)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed GitHub connections resolve a responsible user's or dedicated agent's identity into an audited, run-scoped credential projection. > - An enrolled instance could retain a hidden managed setup method after Cloud stopped advertising it, producing a blank, disabled setup step. > - Native runner processes also dropped the resolved GitHub projection before the provider shell, so `gh` and Git could not use the selected identity in Daytona. > - Daytona already provides the outer isolation boundary. Applying Codex's inner Linux sandbox there both duplicated containment and failed because nested user namespaces are unavailable. > - This change repairs setup fallback, carries only the bounded GitHub projection across each runner boundary, and allows only a controller-selected managed sandbox transport to act as the outer sandbox. ## Linked Issues or Issue Description **What happened?** An enrolled self-hosted instance could show a blank GitHub setup step when its managed profile was unavailable. Separately, a native Codex run in Daytona could resolve a managed GitHub connection on the Paperclip host but lose it before the provider shell. Once projected, Codex's nested sandbox failed before commands could run because Daytona does not expose the user-namespace operation used by the inner sandbox. **Expected behavior** Setup must select an advertised customer method when the managed method is unavailable. A Daytona run must receive the exact managed GitHub identity selected for that run, support `gh` and HTTPS Git, and rely on Daytona as its outer sandbox without weakening local or SSH execution. **Steps to reproduce** 1. Enroll a self-hosted instance while Cloud does not advertise the managed GitHub profile and open GitHub setup. 2. Observe the blank second step and disabled action. 3. Configure a native Codex agent with a Daytona environment and a responsible-user GitHub grant. 4. Run `gh api user` or HTTPS Git from the agent shell. 5. Observe missing GitHub environment projection or nested-sandbox startup failure. **Paperclip version or commit** The setup bug reproduces on `1dceee9a4`; the runner proof was developed from the same branch and verified at the latest head below. **Deployment mode** Self-hosted Paperclip enrolled with Paperclip Cloud, using the Daytona sandbox-provider plugin and native Paperclip runner. ## What Changed - Wait for connector enrollment hydration, retain a hidden managed method only while enrollment is needed, and otherwise select an advertised customer fallback. - Add a single bounded GitHub credential-environment projection for `GH_TOKEN`, `GITHUB_TOKEN`, the process-only Git helper token, GitHub commit identity, and at most 32 controller-generated Git config entries. - Forward that projection through the durable controller, Codex app-server transport, and Rust provider child without placing token values in arguments or config. - Allow Codex shell inheritance only for the exact projected GitHub keys and enable provider network access only when the managed credential exists. - Derive outer-sandbox authority exclusively from a managed `sandbox` transport; strip the same flag from configured, host, local, and SSH environments. - Define a named external-sandbox permission profile that Codex resolves to `dangerFullAccess` for default-mode Daytona turns while plan mode remains read-only. - Add regression tests for setup fallback, credential projection, local/SSH/sandbox authority separation, provider forwarding, and permission-profile selection. ## Verification - `pnpm exec vitest run ui/src/pages/apps/AppsConnect.test.tsx` — 96 passed. - `pnpm exec vitest run src/drivers/codex/codex-security-config.test.ts src/drivers/codex/app-server-transport.test.ts src/control-plane/durable-prp-control-plane.test.ts` from `packages/paperclip-runner` — 31 passed. - Focused native-session executor tests — 3 passed. - `pnpm --filter @paperclipai/paperclip-runner typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml --locked -p paperclip-runner-core --lib` — 194 passed. - Real Codex app-server configuration probe accepted `paperclip-runner-external-sandbox` and reported `sandbox.type=\"dangerFullAccess\"` while using the named profile. - [Signed Daytona image workflow](https://github.com/paperclipai/paperclip/actions/runs/33995270328) built commit `86571f7997e7100e47bd131aac1f1e773112a0ce`; the isolated environment was pinned to `sha256:ecef21105f8de382d75787e59439d936be239b77ae74a31c8ed3a17cde39b023`. - Live isolated Daytona proof passed: the three projected token variables were non-empty and equal; the host-scoped Git credential helper returned the same token without printing it; `gh api user` resolved `cryppadotta`; authenticated `git ls-remote https://github.com/paperclipai/paperclip.git HEAD` returned `1dceee9a4e75b13456760bb54c752deb2dba1d79`; no repository mutation occurred. - The persisted 28,476-byte run log contains no GitHub token shape, bearer header, credential-bearing URL, or private-key marker. - Latest-head pull-request CI and reviews provide the remaining full-suite gate. ## Risks - This deliberately gives shell Git and `gh` access to the run's resolved GitHub identity. It is the audited class-3 behavior required by the GitHub connection design and is outside per-tool Ask-first controls. - The credential source is the trusted broker projection, which overwrites configured environment values. The helper is scoped to HTTPS `github.com`, revalidates protocol and host, and never places its token in command arguments, URLs, or files. - Managed Daytona sandboxes become the containment boundary for default-mode provider commands. Local and SSH targets retain the inner Codex workspace sandbox, and plan mode remains read-only everywhere. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5, with reasoning, browser control, shell access, and code execution. The product did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1dceee9a4e |
fix(runner): persist warm Daytona workspaces (#12901)
## Thinking Path > - Paperclip manages AI agent work and the execution state for each task. > - Remote agents run in sandbox environments such as Daytona. > - Daytona keeps files while a sandbox is stopped, but deletion removes those files. > - Runner Codex did not copy successful remote workspace changes back to the host workspace. > - A warm sandbox could therefore hide data loss until Daytona replaced or deleted the sandbox. > - This pull request makes the host workspace durable after every successful turn and keeps verified reusable sandboxes warm. > - The benefit is reliable multi-turn work across warm reuse, restart, stop, and sandbox replacement. ## Linked Issues or Issue Description **What happened?** A successful native Codex turn in Daytona could leave workspace changes only in the remote sandbox. A later warm turn appeared to work because it reused that filesystem. A replacement sandbox could start from stale host data and lose the successful changes. **Expected behavior** Paperclip must merge each successful remote turn into the authoritative host workspace before it completes the run. A verified warm lease may reuse its remote files. A replacement lease must reconstruct the exact durable workspace seed. **Steps to reproduce** 1. Run Codex in a reusable Daytona environment. 2. Write a file during one successful turn. 3. Replace the Daytona sandbox before the next turn. 4. Observe that the next turn can start without the prior file on the unpatched code. Related remote workspace foundation: #10070. ## What Changed - Added explicit `host_current`, `durable_seed`, and `adopt_remote` workspace preparation modes. - Added atomic, versioned native workspace descriptors and seed archives under `PAPERCLIP_HOME`. - Added real native sandbox export and three-way host merge before terminal result completion. - Added workspace-only recovery after a proposed result. Recovery does not submit another provider turn or consume the provider retry budget. - Added fail-closed handling when a sandbox with unexported changes is gone. - Kept healthy reusable Daytona sandboxes started for legacy Codex and Runner Codex. - Kept the Runner Codex process and provider session across verified warm turns. - Added the paid `daytona-warm-continuity` browser suite. It contains exactly the legacy Codex and Runner Codex cells. Each cell performs three measured turns. - Documented `pnpm test:e2e:runner -- --suite daytona-warm-continuity`. No package script was added. - Added no database migration. The metadata format is backward compatible and idempotent. ## Verification - `pnpm typecheck` - `pnpm test:e2e:runner:unit` — 114 passed - Native workspace, finalizer, session, and environment tests — 232 passed - Daytona provider tests — 150 passed - Workspace staging and merge tests — 98 passed - Runner transport tests — 63 passed - Legacy Codex restore tests — 5 passed - Rust format and compile checks pass through root typecheck - The paid Daytona suite was not run locally because the required Daytona, OpenAI, and immutable image credentials are not present. ## Risks - The main risk is an incorrect workspace identity or merge after a crash. Durable descriptors bind the run, workspace, lease, provider lease, local root, remote root, and baseline digest. Ambiguous evidence fails closed. - The host merge may conflict with concurrent host edits. The existing three-way merge and exclusion rules handle this case and surface failures. - A deleted sandbox cannot recover unexported bytes. Paperclip now blocks with `workspace_sync_out_unrecoverable` instead of reporting success or rerunning the provider. - There is no database migration. Descriptor writes and recovery are atomic and idempotent. ## Model Used OpenAI Codex with GPT-5. The run used agentic reasoning, repository inspection, code execution, test execution, Git, and GitHub CLI tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.905.0-canary.11 |
||
|
|
be407f3456 |
fix(onboarding): hide the probe's diagnostics while the hire is in flight (#12902)
After a successful sign-in, a block of amber diagnostics flashed up and vanished as the step advanced. They are the identity and target INFO checks every environment test reports: they make the result a warn without blocking anything, and the step rendered any non-pass result, so they owned the screen for the window between the probe returning and setStep(5). Gated on loading rather than the connect phase. A blocking result stops the hire and handleGiveHeartbeat clears loading in the finally after its early return, so a genuine block still shows its checks while a hire in flight shows none. The button needed no change: step 4 forces the footer's loading to false and takes its label from connectCta, which reads Connecting for that whole window.canary/v2026.905.0-canary.10 |
||
|
|
60469a08e0 |
feat(agent-login): resume an active login session and permit concurrent login terminals (#12861)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent authentication uses server sessions, plugin workers, and browser login panels. > - A page reload loses an active login session, and one worker permits only one login terminal. > - These limits cause lost work and prevent two owners from logging in through one worker. > - This pull request lets the browser resume active sessions and lets workers serve concurrent login terminals. > - The benefit is reliable login recovery with a bounded process-wide route limit. ## Linked Issues or Issue Description **What existing behavior does this improve?** It improves agent credential login recovery and concurrent login terminal handling. **Subsystem affected** Cross-cutting (multiple of the above). **Current behavior** A page reload loses the active login session. A shared plugin worker rejects a second login terminal. **Proposed behavior** The browser reads and resumes the owner's active session. A worker supports multiple login terminal routes under a process-wide ceiling. **Reason and benefit** Owners keep login progress after a reload. Two owners can log in through one worker without removing the route limit. **Breaking changes** None. The change adds owner-scoped read routes and changes login terminal concurrency. ## What Changed - Replace the single worker login route with maps keyed by host route and worker session identifiers. - Add a process-wide login route ceiling and release each reserved slot on every exit path. - Add owner-scoped active-session reads with consistent negative responses and private cache control. - Keep the device-login prompt while the session has an active public status. - Add a durable setup-token cancel fallback for a lost in-memory session. - Resume active sessions when the agent configuration or onboarding panel mounts. - Remove routine unmount cancellation and keep explicit Cancel behavior. ## Verification - `pnpm --filter @paperclip/server test` — server route, service, and plugin-worker-manager suites. - `pnpm --filter @paperclip/plugin-sdk test` — worker RPC host suite. - `cd ui && npx vitest run src/components/AgentConfigForm.render.test.tsx src/components/OnboardingWizard.test.tsx`. - `cd ui && npx tsc -b`. - `tests/e2e/onboarding.spec.ts` — reload during login. - CI must pass on this pull request. ## Risks The change affects agent authentication and the sandbox-to-host boundary. Route cleanup must release every reserved slot. Owner checks must prevent cross-owner session access. Tests cover route cleanup, owner scope, reload recovery, and concurrent worker routes. ## Model Used Codex, OpenAI GPT-5, tool use and code review support. The implementation author owns the exact model details for the code changes. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR with the relevant issue-template fields - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3ed5b7c5c8 |
refactor(server): extract the active-run output watchdog into a feature module (#12853)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server recovery service monitors active runs and applies watchdog decisions > - The watchdog rules and database operations lived in one large recovery service > - This structure made the rules harder to test and made company scoping harder to inspect > - This pull request moves the watchdog into domain, application, and adapter layers > - The benefit is a smaller recovery service, pure policy tests, and clear company-scoped ports ## Linked Issues or Issue Description **What existing behavior does this improve?** The active-run output watchdog that detects silence, suppression, and terminal evidence. **Current behavior** The recovery service contains the watchdog policy, use cases, database operations, and process control in one file. **Proposed behavior** A feature module separates pure policy, use cases and ports, and Postgres and process adapters. The recovery service delegates its public watchdog methods to this module. **Reason and benefit** The separation makes policy decisions easy to test. Company identifiers on every reader and writer port make tenant scope clear. Smaller service methods reduce change risk. **Breaking changes** None. The recovery service keeps its public methods and call sites. Related public watchdog work includes [#7043](https://github.com/paperclipai/paperclip/pull/7043) and [#7770](https://github.com/paperclipai/paperclip/pull/7770). ## What Changed - Add the `server/src/modules/active-run-watchdog/` feature module with domain, application, and adapter layers. - Move watchdog policy, use cases, Postgres access, and local process control into the module. - Keep the recovery service public methods and delegate them to the module. - Add 53 pure module test cases and retain 8 Postgres integration cases. - Add company scoping and transaction rollback coverage. ## Verification - Run `vitest run --config vitest.config.ts src/modules` and confirm 3 files and 53 cases pass. - Run `vitest run --config vitest.config.ts src/__tests__/heartbeat-active-run-output-watchdog.test.ts` and confirm 1 file and 8 cases pass. - Run the full server suite in pull request CI. - Compare the type-check result with a fresh baseline on the same checkout. ## Risks The main risk is a behavior change in recovery decisions during the move across layers. The pure policy tests cover the moved rules. The integration tests cover database behavior, company scope, and transaction rollback. Pull request CI runs the full server suite. ## Model Used OpenAI Codex, GPT-5, runtime-managed context window, tool use, code execution, and repository review support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.905.0-canary.9 |
||
|
|
d2d647c34b |
fix(connections): honor identity after reconnect (#12897)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - Connections give agents access to external services with a selected credential identity. > - A removed connection keeps its database row so Paperclip can retain its history. > - A fresh GitHub setup can select a different identity from the removed connection. > - The retained row incorrectly kept its old credential policy after that new selection. > - The GitHub callback then could not save the new grant and returned the user to setup. > - This pull request applies the explicit identity selection when Paperclip revives an archived row. > - The benefit is a successful GitHub reconnect after the user changes from a dedicated agent account to a personal account. ## Linked Issues or Issue Description **What happened?** After a user removed a dedicated-agent GitHub connection, a fresh setup with “My GitHub account” returned to the setup page with `oauth=failed`. The Cloud claim succeeded, but the local connection still used the old `per_agent` policy. **Expected behavior** A fresh setup must apply the explicit identity choice. An interrupted draft or an explicit reconnect must keep its existing identity. **Steps to reproduce** 1. Connect GitHub with a dedicated agent identity. 2. Remove the connection. 3. Start a fresh GitHub connection with “My GitHub account.” 4. Complete GitHub OAuth. 5. Observe that Paperclip returns to the setup page instead of the permissions page. **Paperclip version or commit** `342c01fee` **Deployment mode** Local dev (`pnpm dev`) with embedded Postgres and the staging managed connector. **Additional context** This follows the GitHub access UI change in #12893. ## What Changed - Apply an explicit Access identity when a fresh gallery setup revives an archived connection row. - Preserve the identity for interrupted drafts and explicit reconnects. - Do not carry credential material across an identity change. - Apply the omitted organization default during a fresh archived-row recovery. - Restore the prior grants and credential policy transactionally if a revived setup rolls back. - Disable the connection and surface a specific failure if that restoration cannot complete. - Preserve newer concurrent grant changes with a row lock and optimistic version check. - Preserve newer concurrent connection identity/configuration changes with a locked state fingerprint. - Add regressions for dedicated-to-personal OAuth, organization-default recovery, rollback, rollback failure, and concurrent grant/connection changes. ## Verification - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts` — 222 tests passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - Browser proof on an isolated local instance: dedicated connection removed, personal setup selected, staging GitHub OAuth completed, permissions page opened, connection reported active and healthy, personal grant active, old agent grant revoked. ## Risks - Low migration risk. This change has no schema migration. - The behavior changes only when a fresh setup explicitly selects an identity for an archived connection row. - Existing draft resume and explicit reconnect behavior stays unchanged. - Connection-manager checks still protect changes to a retained credential identity. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5, with reasoning, browser control, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3da58b185e |
fix(cli): restore test-drive credential inputs (#12898)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The CLI provides a test-drive command for a ready local test instance. > - That command must accept a provider credential before it creates the first agent. > - The command did not accept a literal key and could lose an exported key during server startup. > - This pull request accepts both credential paths and captures the CLI environment before startup. > - The benefit is a reliable one-command test drive from an existing shell. ## Linked Issues or Issue Description Refs #12894 **What happened?** `paperclipai test-drive --api-key <value>` failed because the option did not exist. An exported canonical provider variable could also become unavailable before the post-listen bootstrap read it. **Expected behavior** The command must accept a literal key when the operator requests it. The command must also use a provider variable that was present when the CLI started. **Steps to reproduce** 1. Export `ANTHROPIC_API_KEY` in the shell. 2. Run `pnpm paperclipai test-drive`. 3. Observe that bootstrap can report that no credential exists. 4. Run `pnpm paperclipai test-drive --api-key test-value`. 5. Observe that Commander reports an unknown option on the prior implementation. **Paperclip version or commit** The problem exists on `master` after #12894. **Deployment mode** Local development with `pnpm` and the embedded database. ## What Changed - Add the `--api-key <value>` test-drive option. - Keep `--api-key` and `--api-key-env` mutually exclusive. - Capture provider variables before in-process server startup changes the process environment. - Redact literal and environment-backed credentials from Paperclip errors, including custom environment-variable names with surrounding whitespace. - Scrub split and joined literal-key forms from the JavaScript `process.argv` view before telemetry, diagnostics, API work, or server startup. - Warn about process argument and shell history exposure without printing the key. - Add tests for literal keys, option conflicts, environment snapshots, argv handling, and error redaction. - Update the CLI and development documentation, including the remaining external argv exposure tradeoff. ## Verification - `pnpm exec vitest run cli/src/__tests__/test-drive.test.ts` passed 32 tests. - `pnpm -r typecheck` passed. - `pnpm build` passed. - A live literal-key smoke test created one company and one CEO agent without a provider call. - A live exported-variable smoke test created the same clean instance without credential flags. - `pnpm test:run` completed locally with 5,870 passing tests and 19 host-dependent failures in six unrelated suites. The failures came from macOS `/tmp` aliases, exhausted test ports, and existing workspace-runtime fixture assumptions. - GitHub CI passed the full build, typecheck, canary dry run, general tests, serialized server tests, and e2e matrix on clean Linux runners. - Greptile rated the exact latest head 5/5, Superagent passed, and all review threads are resolved. ## Risks - A raw value passed through `--api-key` can appear in operating-system process listings, shell history, or parent-wrapper output before Paperclip can scrub its own JavaScript argv view. The command warns about this risk and Paperclip does not print the value. - The environment snapshot contains the process environment only in memory for the life of the foreground command. - This change adds no schema migration and no REST endpoint. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5. The exact deployment version and context window are not exposed to the agent. Extended reasoning, tool use, code execution, and GitHub access were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
70c9ca7410 |
fix(server): stop mock leakage between interaction-route tests (#12807)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server test suite checks issue-thread interaction routes > - Shared Vitest mocks can keep queued one-shot values between tests > - A leftover value can change the issue returned to a route and cause a false authorization failure > - This pull request resets all mocks and restores the plain run-attribution value before each test > - The benefit is stable interaction-route tests that do not depend on test order ## Linked Issues or Issue Description This pull request has no public issue link. The bug details follow. **What happened?** The interaction-route server test suite failed intermittently in continuous integration. The test named `lets a watchdog-scoped assignee withdraw through ordinary containment` sometimes failed because `withdrawInteraction` received no call. The test suite used `vi.clearAllMocks()`, which clears call history but does not clear queued one-shot mock values. A queued value from `mockIssueService.getById` could change a later test's issue and make the route return `403`. **Expected behavior** Each test must start with empty mock queues and the default run-attribution value. Test results must not depend on test order. **Steps to reproduce** 1. Run the interaction-route test file many times in sequence. 2. Run the same file with shuffled test seeds. 3. Observe the intermittent containment failure before this change. **Paperclip version or commit** Current `master` plus commit `8cd56b38a77f1feecac495f57a48d3f0a1b3b01c`. **Deployment mode** Built from source. The failure occurs in the server test suite. ## What Changed - Replace the partial mock reset with `vi.resetAllMocks()`. - Reset `mockRunAttribution.value` before each test. - Keep the change within the interaction-route test file. ## Verification - The target suite passes 77 of 77 runs. - The test count remains 64 `it(...)` sites, including five `it.each(...)` blocks that expand to 77 runs. - The file contains no skipped or focused tests. - The diff changes no production code. - A defect-detection round trip reproduces the failure when the containment guard is broken and passes after the guard is restored. - Ten sequential runs and five shuffled-seed runs pass 77 of 77. ## Risks Low risk. The change affects test setup only. It does not change production code or route behavior. ## Model Used OpenAI GPT-5, exact runtime model `gpt-5`, tool use and code execution, context window not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
45725ad820 |
fix(server): hoist the comment-cancel route test suite's module graph (#12877)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server test suite checks issue comment and cancellation routes.
> - The comment-cancel route suite reloaded about 40 modules before
every test.
> - Repeated module reloads created a race between service mocks and
real services.
> - The race caused an HTTP 500 when the test expected HTTP 200.
> - This pull request loads the mocked module graph once for the suite.
> - The benefit is stable route tests with clearer diagnostics for
future failures.
## Linked Issues or Issue Description
**What happened?**
The comment-cancel route test suite failed intermittently in continuous
integration with an HTTP 500 where the test expected HTTP 200. The suite
reset modules and re-imported the route graph before every test. A
re-import could bind the real service module to the test's minimal fake
database and cause a `TypeError`.
**Expected behavior**
The suite must run all seven route tests without intermittent HTTP 500
responses. A future server error must show its underlying cause in the
test output.
**Steps to reproduce**
1. Run `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/issue-comment-cancel-routes.test.ts`.
2. Repeat the test command under continuous integration load.
3. Compare the result with a run that reloads the route module graph
before every test.
**Paperclip version or commit**
`0a6a7087ed6c3bb1cadf59fcfec362d6fc9a6d14`
**Deployment mode**
Built from source with the server test runner.
**Installation method**
Built from source.
**Agent adapter(s) involved**
Not adapter-specific (core test issue).
**Database mode**
Not database-related.
**Access context**
Unclear / not applicable.
**Additional context**
The route module and error-handler middleware remain real. The service
layer remains mocked. Authorization and non-leakage assertions remain
unchanged.
## What Changed
- Register mocks once and load the route module graph once through
`hoistModuleGraph`.
- Remove the per-test module reset and re-import.
- Add a `res.on("finish")` diagnostic listener for server error context.
- Keep all seven test titles and the existing authorization assertions.
## Verification
- Run `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/issue-comment-cancel-routes.test.ts`.
- Confirm that all seven tests pass.
- Confirm that the full continuous integration suite passes on this pull
request.
## Risks
Low risk. This change modifies one test file and does not change
production code or test coverage. The local worktree cannot start this
suite because it lacks `packages/adapters/droid-local`; continuous
integration must verify the complete repository dependency set.
## Model Used
OpenAI Codex, GPT-5. The model used repository tools, GitHub tools, and
code review reasoning. The execution context window is not exposed by
the runtime.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
58ed2b64ea |
test(server): remove a concurrent-import race in the approval routes suite (#12876)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Approval routes use server tests to protect access and idempotency behavior > - The approval routes test suite loaded mocked modules at the same time > - Concurrent module loading could lose a service mock and produce a false test failure > - This pull request loads the shared module graph once and reuses it across the suite > - The benefit is stable approval route tests with unchanged coverage ## Linked Issues or Issue Description **What happened?** The approval routes test suite loaded two mocked modules in one concurrent import. A module interleaving could remove the approval service mock. The route then returned HTTP 404 instead of the expected HTTP 403. **Expected behavior** The suite must keep the approval service mock when it loads the route modules. The access test must return HTTP 403 on every run. **Steps to reproduce** 1. Run `npx vitest run server/src/__tests__/approval-routes-idempotency.test.ts`. 2. Repeat the test command while the test runner loads the module graph. 3. Observe a false HTTP 404 result when the module mock interleaves. **Paperclip version or commit** Current `master` at the base commit of this pull request. **Deployment mode** Built from source with the server test runner. ## What Changed - Reused the existing `hoistModuleGraph` helper for the approval route modules. - Loaded the route modules once in sequence instead of in one concurrent import. - Kept per-test mock behavior, Express app setup, database doubles, test names, and assertions unchanged. ## Verification - `npx vitest run server/src/__tests__/approval-routes-idempotency.test.ts` — 11 of 11 tests passed. - Ten repeat runs passed. - `npx tsc --noEmit -p server` produced no new errors against the base branch. - All Paperclip CI checks passed. - Greptile reported 5/5 with no open findings. ## Risks This change affects test module setup only. It does not change production code or test coverage. Risk is low. ## Model Used OpenAI Codex with the `gpt-5` model family. The serving snapshot and context-window size are not exposed. The agent used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and found no duplicate for this test race - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ca9c17df60 |
test(ui): flush passive effects with React act in CompanySkills tests (#12878)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The user interface tests verify company skill installation flows.
> - The install-preview dialog sets state through passive React effects.
> - The local test helper flushed render work but did not flush passive
effects.
> - Timer hops did not guarantee that React completed those effects
before user actions.
> - This pull request delegates the helper to React act and removes the
timer hops.
> - The benefit is deterministic CompanySkills tests without a
product-code change.
## Linked Issues or Issue Description
**What happened?**
The CompanySkills install-preview dialog tests used a timer hop after
opening the dialog. The timer could run before React completed passive
effects. A later click then used stale dialog state.
**Expected behavior**
The test helper must flush passive effects before the test interacts
with the dialog. The tests must pass without a race between timer
callbacks and React scheduler tasks.
**Steps to reproduce**
1. Run the CompanySkills test file many times from the ui directory.
2. Observe intermittent failures that report an empty agent list or a
null slug.
3. Run the same tests with React act to flush passive effects.
**Paperclip version or commit**
Commit
|
||
|
|
342c01fee8 |
fix(connections): simplify GitHub access details (#12893)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Connections show the identity and access that agents use > - The GitHub permissions page showed several operational fields in one large status card > - That card made the repository actions harder to scan > - A dedicated GitHub identity also named its agent without linking to the agent > - The permanent shell Git warning also competed with the controls it explained > - This pull request replaces the card with two compact action rows, links the agent label, and reveals the warning only when an action permission is changed > - The benefit is a shorter page with direct navigation to the controls that matter ## Linked Issues or Issue Description Related: #12891 **What existing behavior does this improve?** The GitHub connection permissions view. **Subsystem affected** `ui/` — React and Vite board UI. **Current behavior** The view shows installation, token, webhook, event, and refresh metadata in a large card. The dedicated-agent text is not interactive, and a shell Git warning is always visible even before the user interacts with action permissions. **Proposed behavior** Show one repository-management row and one access-refresh row. Link the dedicated-agent text to that agent. Hide the shell Git warning until the user changes an action permission, then show it in the Actions section. **Reason and benefit** The two actions are easier to find. Users can open the dedicated agent directly, and see the shell Git limitation at the moment it becomes relevant. **Breaking changes** None. This change removes secondary display fields from this view. It does not change GitHub credentials, grants, or API data. ## What Changed - Replaced the GitHub status card with repository and refresh rows. - Kept the token-backed all-repositories warning inside the repository row. - Added explicit labels for selected, all, mixed, and empty repository access. - Added a direct link from “Used only by” to the dedicated agent. - Moved the shell Git/`gh` warning into the Actions section and reveal it only after a permission-change attempt. - Added render coverage for the rows, removed fields, links, actions, and contextual warning. ## Verification - `pnpm exec vitest run ui/src/pages/apps/AppDetail.test.tsx` — 50 tests passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm --filter @paperclipai/ui build` — passed. - `pnpm check:token-gates` — passed. - Verified the live page initially hides the shell warning, then shows it after changing an action permission; restored the test permission afterward. ## Risks Low risk. This is a display-only change. The existing management URL, refresh action, and action-permission mutations are unchanged. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, `gpt-5.6-sol`, extended reasoning, tool use, code execution, browser control, and multi-file repository editing. The context window size was not provided. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.905.0-canary.8 |
||
|
|
87832c48fd |
feat(runner-e2e): publish declared screenshots (#12895)
## Thinking Path > - Paperclip uses runner end-to-end reports to compare agent profiles and execution environments > - The report dashboard shows each reviewed final-state screenshot as a thumbnail and gallery item > - The public history publisher removed all per-attempt images before it regenerated the dashboard > - Therefore the public dashboard had the new layout but could not show the screenshots from the run > - The publisher needs a narrow rule that keeps only screenshots from the exact live fixture issue route > - This pull request keeps those trusted PNG files in every future public S3 and GitHub Pages report > - The benefit is that each future report can show its screenshot gallery without exposing logs, traces, videos, archives, arbitrary images, or generated report trees ## Linked Issues or Issue Description **What happened?** The runner E2E job captured final-state screenshots in its private artifact. The public S3 and GitHub Pages publication step removed those screenshots before it regenerated the dashboard. As a result, the public report showed the new dashboard controls but no screenshot thumbnails or gallery items. **Expected behavior** Each future public runner E2E report must include reviewed PNG screenshots from the live fixture issue. Other captures and active or unsafe evidence must stay private. **Steps to reproduce** 1. Run the runner full-stack E2E workflow on `master` before this change. 2. Open the private `runner-e2e-report-*` artifact and confirm that it contains per-attempt PNG screenshots. 3. Open the public campaign URL and confirm that the dashboard has no screenshot gallery items. **Paperclip version or commit** The issue was reproduced on commit `64d8929`, after the report design change in PR #12889. **Deployment mode** GitHub Actions with the public S3 and GitHub Pages report publishers. Related design work: Refs #12889. ## What Changed - Mark screenshots from the exact server-created live fixture issue route with `public-runner-fixture`. - Keep marked PNG files in both the S3 history bundle and the GitHub Pages bundle. - Keep captures from other issue routes, sensitive routes, and external origins private. - Bind public files to the normalized execution ID, attempt, and safe PNG base name. - Validate every retained image with the existing PNG signature and 12 MiB size checks. - Skip missing-artifact sentinel results with attempt `0` when they have no public screenshots. - Continue to remove unmarked images, videos, traces, archives, generated HTML reports, and other private evidence. - Update publisher tests, workflow checks, report copy, and the public-evidence security documentation. ## Verification - `pnpm exec vitest run --config tests/runner-e2e/vitest.config.ts tests/runner-e2e/history.test.ts tests/runner-e2e/report.test.ts` - `pnpm exec vitest run --config tests/runner-e2e/vitest.config.ts tests/runner-e2e/workflow-security.test.ts -t "uses environment-scoped OIDC"` - `pnpm test:e2e:runner:typecheck` - `PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/runner-e2e-dashboard.spec.ts` - `pnpm -r typecheck` - `pnpm build` - Regenerated the dashboard from retained evidence for Actions run `33968240659` without a paid matrix rerun. The public-stage proof contained 121 screenshot gallery items and thumbnail frames, with zero generated HTML report files. The trusted-fixture marker and route gate have separate focused tests. - All pull request CI checks pass on commit `ccd2b49e1`. ## Risks - This change intentionally makes marked fixture screenshots public at the campaign URL. A screenshot can show data that a raw-byte secret scan cannot detect. - The capture helper marks a screenshot only on the exact loopback issue route for the fixture that the harness created. A different issue, sensitive page, or external origin stays private. - The publisher also requires the marker, a safe normalized path, a valid PNG signature, and the size limit. - The change does not publish videos, logs, traces, archives, arbitrary images, or generated browser report trees. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, model `gpt-5.6-sol`, with high reasoning, repository tool use, shell execution, browser inspection, and GitHub CLI access. The working context was the Codex desktop task context. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8f0c1d4548 |
feat(cli): add isolated test-drive command (#12894)
Add a foreground-only test-drive workflow with isolated data, provider-backed CEO bootstrap, OpenCode/OpenRouter support, worktree execution setup, reuse safeguards, and delayed browser opening. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.905.0-canary.7 |
||
|
|
5da6499860 |
fix(connections): reuse one-time cloud enrollment (#12891)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Managed connections let agents use provider credentials without exposing those credentials to the control plane UI > - A self-hosted instance must first establish a trusted credential destination with Paperclip Cloud > - The GitHub connection flow repeated that trust decision before provider consent > - The local setup route also lost step 2 after enrollment and could display the PAT identity defaults before enrollment > - This pull request makes enrollment a one-time instance decision and sends later provider starts directly to provider consent > - The benefit is a shorter flow with one clear Paperclip approval and no required service restart ## Linked Issues or Issue Description Refs #12843. Companion Cloud change: [paperclipai/paperclip-cloud#391](https://github.com/paperclipai/paperclip-cloud/pull/391). ## What Changed - Made `stage=setup` authoritative during initial route hydration and enrollment return. - Added a contained one-time enrollment screen with provider-specific copy. - Accepted a provider `authorizationUrl` from Paperclip Cloud only when it matches the exact GitHub or Google OAuth endpoint. - Preferred the direct provider URL while retaining the legacy confirmation URL fallback. - Preserved the company-bound identity and agent-access draft across the full-page enrollment callback, including cold company-context hydration. - Kept GitHub defaulted to “My GitHub account” and “Any agent,” including before Cloud advertises the managed method. - Updated GitHub identity and agent-access copy for responsible-person and dedicated-agent behavior. - Labeled the provider action “Continue to GitHub.” - Added parser, routing, cold-hydration, access-restoration, visibility, fallback, defaults, and copy tests. ## Verification - `pnpm exec vitest run server/src/services/paperclip-cloud-connector.test.ts ui/src/pages/apps/AppsConnect.test.tsx` (114 tests passed) - `pnpm check:token-gates` - `pnpm -r typecheck` - `pnpm build` - Live browser proof used a new data directory on `127.0.0.1:3117` and the exact Cloud PR revision on staging. - The fresh flow selected “My GitHub account” and “Any agent,” showed one enrollment approval, returned to local step 2, and connected GitHub without a second Paperclip confirmation or login. - The connected screen showed one selected repository, a long-lived token, installation metadata, a successful access refresh, and healthy webhook delivery. - Gmail on the same instance went directly to Google consent without another Paperclip approval. - Restarting the same data directory preserved enrollment. A second new data directory required exactly one new approval. - A final fresh-data-dir rerun selected a dedicated GitHub identity for Ada before enrollment, approved the instance once, returned to step 2, retained Ada after a Back check, connected directly through GitHub, and finished with “Used only by Ada,” one selected repository, and a long-lived token. - Port 3100 remained untouched throughout the proof. ## Risks - The new Cloud field is additive and restricted to the exact GitHub and Google OAuth origins and paths, with no embedded credentials or URL fragment. - An older Cloud response still works through `confirmationUrl`. - A self-hosted instance still requires one signed Cloud enrollment. Managed Cloud instances do not render the enrollment screen. - Provider authentication and consent remain mandatory after instance enrollment. - No schema migration is included in this pull request. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, `gpt-5.6-sol`, extended reasoning, tool use, code execution, browser control, and multi-file repository editing. The context window size was not provided. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.905.0-canary.6 |
||
|
|
64d8929ce9 |
fix(runner-e2e): bound completed cell teardown (#12890)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner E2E suite verifies complete agent tasks against real providers. > - A Codex Plan test finished in 76 seconds, but its Playwright process stayed alive for 25 more minutes. > - The launcher accepted the saved passing result after its watchdog killed the process. > - The existing Plan limits also allowed much more time than recent successful runs need. > - This pull request adds a bounded result-to-exit check and safe process evidence. > - It also reduces the Plan limits while it keeps large headroom over measured success times. > - The benefit is faster diagnosis and no false green result after a teardown stall. ## Linked Issues or Issue Description **Pre-submission checklist** I searched open pull requests for runner E2E timeout and Playwright cleanup changes. I found no duplicate. The problem reproduces on `master`. **What happened?** The local Codex Plan cell completed its test in 76 seconds. Playwright then stayed alive for about 25 minutes. The launcher watchdog killed it after 26.5 minutes, but the launcher still accepted the saved passing result. **Expected behavior** The launcher must stop a process that stays alive after all results exist. It must report a cleanup failure instead of a pass. Plan tests must also use limits that match measured successful runs. **Steps to reproduce** 1. Run `core-compatibility.runner-codex.local.plan-revise-accept`. 2. Observe a valid result and the Playwright pass output. 3. Observe that the process can stay alive until the old launcher watchdog stops it. **Paperclip version or commit** The evidence came from `bcc6fe7a442dae74ab0321ad472f7536ffa58f04` in [Actions run 33963318820](https://github.com/paperclipai/paperclip/actions/runs/33963318820). ## What Changed - Reduce the Plan attempt limit from 20 to 8 minutes for local execution. - Reduce the Plan attempt limit from 35 to 12 minutes for Daytona execution. - Stop Playwright after it stays alive for 120 seconds after every result exists. - Record only allowlisted process kinds in the stall diagnostic. - Validate process identities before cleanup and retain continuously live process groups through member replacement. - Treat watchdog, post-result, cleanup, and nonzero-exit conflicts as cleanup failures. - Keep interactive `--ui` and `--debug` sessions exempt from the result-to-exit check. ## Verification - Prettier completed for all changed files. - `git diff --check` passed. - Static review confirmed the timeout derivation and cleanup boundaries. - An independent review found no blocking issue in the final patch. - I did not run local tests, builds, or type checks because this workstation must use the lightweight workflow. - GitHub CI and the exact paid Codex Plan cell will verify this commit. ## Risks The main risk is a false cleanup failure when Playwright needs more than 120 seconds after it writes all results. The allowance is separate from the task limit. Interactive modes are exempt. The diagnostic does not print command arguments or environment values. ## Model Used OpenAI Codex with GPT-5.6, reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.905.0-canary.5 |
||
|
|
4a7172f5ac |
feat(runner-e2e): improve matrix report browsing (#12889)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - The runner E2E system verifies agent profiles across supported execution environments. > - Its generated report is the main surface for inspecting those results and their visual evidence. > - Expanded matcher content could change the matrix column widths and make comparisons difficult. > - Screenshots also required extra navigation, and the report did not have current-result search and filters. > - This pull request stabilizes the matrix layout and makes visual evidence directly browsable. > - The benefit is faster inspection of retained test evidence without another paid matrix run. ## Linked Issues or Issue Description **What happened?** The runner E2E report changed matrix column widths when a matcher table expanded. The report also made screenshot comparison and current-result discovery slower than necessary. **Expected behavior** The matrix columns must remain stable. Each retained screenshot must appear as a thumbnail. The gallery must support keyboard navigation and show the relevant execution metadata. The report must support client-side search and filters. **Steps to reproduce** 1. Open a runner E2E matrix report that contains retained screenshots. 2. Expand the matcher details in a matrix cell. 3. Observe the matrix column movement in the old report. **Paperclip version or commit** `8430bd897` **Deployment mode** Generated static runner E2E report. ## What Changed - Keep matrix and matcher table widths stable when details expand. - Show retained screenshot thumbnails in each test card. - Add a full-screen evidence gallery with mouse, keyboard, and swipe navigation. - Show agent, environment, runtime, status, duration, token, and matcher data in the gallery header. - Add client-side search and profile, environment, suite, and status filters below the report section tabs. - Keep the filters in normal document flow while the report tabs remain sticky. - Add report generator assertions for the new layout and controls. - Add Playwright coverage for filtering, stable matcher expansion, and filtered gallery navigation. ## Verification - `pnpm test:e2e:runner:typecheck` - `pnpm exec vitest run --config tests/runner-e2e/vitest.config.ts tests/runner-e2e/report.test.ts` - `PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/runner-e2e-dashboard.spec.ts` - `node --test ./scripts/__tests__/e2e-shard.test.mjs` - `pnpm check:token-gates` - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` - Regenerated the report from GitHub Actions run `33963318820` without rerunning the matrix. - Verified 121 retained thumbnails across 66 results in a local browser. - Verified that expanded matchers keep matrix widths at 260, 486, and 486 pixels in a 1280-pixel viewport. - Verified search, filters, modal metadata, and arrow-key gallery navigation. ## Risks Low risk. This change only modifies the static runner E2E report generator and its tests. It does not change runner execution or retained evidence data. > This work is focused report polish. It does not duplicate planned core work in `ROADMAP.md`. ## Model Used OpenAI Codex with `gpt-5.6-sol`. Codex desktop managed the context window. The model used reasoning, filesystem tools, code execution, GitHub CLI access, and in-app browser verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8430bd897f |
ci: reuse trusted cache for Daytona images (#12862)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The full-stack runner campaign checks local and Daytona runner behavior. > - A Daytona image content miss starts a cold multi-stage Docker build. > - Stable dependency and agent CLI layers take most of the image build time. > - Development targets must not write shared cache state. > - This pull request adds a registry cache with a default-branch write gate. > - It also puts volatile source inputs after stable install layers. > - The benefit is a shorter Daytona image build without weaker secret isolation. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the Daytona runner image stage in the full-stack E2E workflow. **Subsystem affected** The GitHub Actions runner E2E workflow and its Daytona Docker image are affected. **Current behavior** Each new Daytona image content ID starts with an empty BuildKit cache. A runner source change also invalidates dependency and agent CLI install layers because volatile inputs occur before those layers. **Proposed behavior** All authorized campaigns can read one GHCR BuildKit cache. Only a campaign whose target ref is the repository default branch can update that cache. The Dockerfile installs dependencies and agent CLIs before it consumes volatile runner source or revision metadata. **Reason and benefit** The paid runner matrix spends several minutes building the image before any selected cell can start. Cache reuse removes repeated stable setup work and makes focused Daytona iterations faster. **Breaking changes** None. The immutable content tag, digest inspection, Cosign signature, image labels, pinned base images, and provider credential boundary stay unchanged. ## What Changed - Read a registry-backed BuildKit cache for Daytona image content misses. - Export the cache only when the resolved target ref is the default branch. - Keep provider credentials outside the image build and cache. - Install provider-pack dependencies before runner source is copied. - Keep expensive agent CLI installs before source revision metadata. - Add workflow and Docker layer-order contract checks. ## Verification - `prettier --write .github/workflows/runner-full-stack-e2e.yml tests/runner-e2e/daytona-image.test.ts tests/runner-e2e/workflow-security.test.ts` - `actionlint .github/workflows/runner-full-stack-e2e.yml` - `git diff --check` - I did not run a test suite or Docker image build locally. The requested iteration policy reserves those checks for GitHub Actions. ## Risks Low risk. BuildKit can use a cache record only when its content key matches the build instruction and input. Development targets have read-only cache access. The cache contains public source and build outputs, but it does not receive provider credentials or the GitHub token as Docker build inputs. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.905.0-canary.4 |
||
|
|
bcc6fe7a44 |
fix(runner): restore multi-turn remote sessions (#12840)
## Thinking Path > - Paperclip manages AI agents and their work. > - The runner executes agent turns on local and remote providers. > - A remote per-turn session must save its state before Paperclip releases its sandbox. > - The session runtime returned after 100 milliseconds while the remote checkpoint still ran. > - The next turn also checked the local state path instead of the verified remote backup. > - This pull request waits for the bounded remote close and accepts only a verified suspended backup. > - The benefit is reliable multi-turn execution without weaker identity checks. ## Linked Issues or Issue Description **What happened?** A successful remote agent turn released its sandbox before the runner saved the verified continuation backup. The next turn failed with `runner_state_identity_mismatch`. **Expected behavior** Paperclip must finish the bounded remote checkpoint before it releases the sandbox. A later turn must validate and restore the digest-matched suspended backup. **Steps to reproduce** 1. Run a native ACPX Claude Plan test in a non-reusable Daytona sandbox. 2. Reject the first plan to start a second turn. 3. Observe that the second turn fails before provider execution. **Paperclip version or commit** The failure reproduced at `13775a90b078ff64872f50961ea1b83d575e7bc6`. **Deployment mode** GitHub Actions with a Daytona sandbox. ## What Changed - Wait for the internally bounded remote runner close and checkpoint before the host returns. - Preserve the existing short cleanup bound for other providers. - Validate remote continuation lifecycle from a complete digest-verified backup when local runner state is absent. - Keep corrupt, non-suspended, mismatched, and unverified state fail-closed. - Make native Plan completion and accepted-Plan wake prompts deterministic. ## Verification - A prior 45-cell local campaign passed 44 cells. The only failure was the OpenCode Plan prompt variance fixed here. - A focused OpenCode local Plan rerun passed. - ACPX Claude Daytona message and question cells passed. - Focused regressions cover delayed checkpoint close and verified remote backup lifecycle. - GitHub Build and the focused ACPX Claude Daytona Plan cell will validate this exact head. ## Risks Remote runnerd sessions now wait for their internally bounded close/checkpoint path before returning; generic provider cleanup retains the existing 100 millisecond bound. Durable run success still cannot be reversed. The environment release guard still blocks sandbox destruction when no verified backup stamp exists. ## Model Used OpenAI Codex, GPT-5.6, extended reasoning, with code execution and GitHub Actions inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal task id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open findings - [ ] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
a7ed22e3dd |
refactor(onboarding): reconcile the arc column's width comment with its width (#12875)
The shell carried two comments arguing opposite things: the older made the case for a 64px inset and said 40px was too wide, the newer made the case for the 40px the code uses. The older one had also drifted from the code independently - it described 68px sides and a 424px column while the file used --sz-64px, a 432px column. One comment now: 40px sides, a 480px column, the measure the connect sequence is drawn to and which the arc shares. The earlier objection is kept and marked untested, with a note that if step 1 or step 3 reads loose the fix belongs in those steps' content rather than the shared shell. Comment-only; no behaviour change.canary/v2026.905.0-canary.2 |
||
|
|
f2349990cc |
feat(onboarding): the connect step's sign-in as one continuous sequence (#12863)
Picking a source starts the sign-in: the row collapses to the answer, the card opens where the credential link was, and the footer button walks Sign in -> Waiting for code -> Connecting before the step advances. Back unwinds it a beat at a time. Nothing mounts to change layout - the card and the link are always rendered and their heights animate, with inert holding the a11y line - because a mount changes the page in one frame and no easing can smooth a step already taken. Review fixes in the same branch: the displayed-code panel now reports its prompt upward (the OpenAI path could not leave the loading beat without it), the two-second hold is a cancellable beat rather than a dropped timer, unwinding a sequence that never opened a card no longer starts a login to cancel it, and the key field regains focus-on-open.canary/v2026.905.0-canary.1 |
||
|
|
4b0e324c63 |
ci: activate the Docker context integrity gate for PRs (#12860)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Pull request CI runs through `pr.yml`, which pins the reusable `pr-trusted.yml` workflow by commit SHA, so `pr-trusted.yml` changes take effect only when the pin moves. > - PRs #12855 and #12858 added the `docker_context_integrity` job and wired it into the `verify` aggregate, but the pin still points at a commit from before them. > - Until the pin moves, a pull request that strips a committed Docker build input still merges green and breaks every post-merge image build. > - This pull request bumps the pin to the #12858 merge commit, the standard second step of every `pr-trusted.yml` change. > - The benefit is that the Docker context integrity gate now blocks merges, which closes out the 2026-09-04 image-publishing incident end to end. ## Linked Issues or Issue Description Refs #12855 and #12858 (the gate this activates) and #12769 (the incident that motivated it). **What happened?** The `docker_context_integrity` job exists on master but does not run on pull requests, because `pr.yml` pins `pr-trusted.yml` at `a0a78ee6`, which predates it. **Expected behavior** Pull requests run the gate, and the `verify` required check fails when a change strips a committed Docker build input. **Steps to reproduce** 1. Open any pull request before this change: the ci run shows no "Docker context integrity" job. 2. After this change, the job runs on every full-CI pull request and `verify` requires its result. **Paperclip version or commit** Pin moves from `a0a78ee60946a5f79f85b2bd0584fc766fae43bb` to `03609aa6` (the #12858 merge commit). **Deployment mode** GitHub Actions pull request CI. ## What Changed - `.github/workflows/pr.yml`: the `pr-trusted.yml` pin moves to the #12858 merge commit. One line. ## Verification - The pinned commit is master's current tip and contains the job, the `verify` wiring, and both probe revisions; its own CI (on #12855 and #12858) is fully green. - This PR's ci run itself executes the newly pinned workflow, so the gate's first live run is visible on this very pull request. ## Risks - Low. Identical mechanism to every previous pin bump. If the new lane misbehaves on some runner, reverting this one line restores the previous pin. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Claude Fable 5 (Anthropic, model id `claude-fable-5`), extended thinking, agentic tool use in Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (one-line pin bump; the pinned workflow's tests ran green on #12855/#12858) - [x] I have added or updated tests where applicable (covered by the pin tests updated in #12855) - [x] I have updated relevant documentation to reflect my changes (not applicable to a pin bump) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.905.0-canary.0 |
||
|
|
03609aa6ec |
ci: keep traceability regression tests in the Docker build context (#12858)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - GitHub Actions builds the Docker images that ship Paperclip, and the image build re-runs the runner's committed-artifact checks. > - PR #12855 restored the capability-contract files that PR #12769's context slimming stripped, and image builds then progressed one step further in the chain. > - The next check, `check:runner-workflow-traceability`, access()es every regression test its spec names — `src/**/*.test.ts` files that the same slimming block also strips. > - Every image build since #12855 merged now fails there with ENOENT, so image publishing is still down. > - This pull request restores those files with one more narrow exception and teaches the context probe to derive the required paths from the spec itself. > - The benefit is that image publishing recovers, and the probe now covers this input class without a hand-maintained path list that could rot. ## Linked Issues or Issue Description Refs #12855 (first restoration from the same incident) and #12769 (the context-slimming change). **What happened?** After #12855 merged, every `Docker` workflow run on master still failed, now inside `check:runner-workflow-traceability`: `Error: ENOENT ... access '/app/packages/paperclip-runner/src/contracts/native-execution.test.ts'`. The check access()es all 29 regression tests named by `spec/evals/stress-workflow-traceability.json`; they are `src/**/*.test.ts` files, and the `packages/paperclip-runner/**/*.test.ts` ignore rule strips them from the build context. **Expected behavior** The Docker build context must contain every file the image build reads. The context-integrity probe must catch this class on the pull request, including inputs named dynamically by a spec. **Steps to reproduce** 1. Check out master after #12855. 2. Run `docker buildx build -f .github/docker-context-checks.Dockerfile .` with this PR's probe, or the real `Docker` workflow build. 3. Observe the ENOENT above; with this PR's `.dockerignore` exception, both pass. **Paperclip version or commit** `bb920fb8` (first post-#12855 failing image build) through master tip. **Deployment mode** GitHub Actions image builds (`docker.yml`), consumed by managed cloud deployments. ## What Changed - `.dockerignore`: re-include `packages/paperclip-runner/src/**/*.test.ts` and `.tsx` — the traceability spec references only files under `src`, so the remaining test exclusions stay. - `.github/docker-context-checks.Dockerfile`: new spec-driven existence walk that replicates the traceability check's own access() loop against the exact build context. The path list comes from the spec at probe time, so a future spec change is covered automatically; the check itself still runs only inside the real image build, where `dist/` exists. ## Verification - `docker buildx build -f .github/docker-context-checks.Dockerfile .` without the `.dockerignore` exception: fails with the exact production ENOENT (`src/contracts/native-execution.test.ts`). - Same command with the exception: passes end to end (all probe stages, including the drift checks from #12855). - Static re-sweep of the remaining image-build chain steps (`build:binary`, replay goldens, semantic-action catalog) against the ignore rules: their inputs are all in the context; cargo needs no `tests` directories (no crate declares an explicit `[[test]]` target). ## Risks - Low. The exception re-adds source test files to the build context only; image contents do not change (tests are neither compiled into the production output nor run in the image build — the check only requires that the referenced files exist). - The probe addition is one dependency-free Node one-liner. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Claude Fable 5 (Anthropic, model id `claude-fable-5`), extended thinking, agentic tool use in Claude Code: GitHub Actions log forensics, spec-driven path inventory, and local docker buildx verification in both failing and fixed states. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (the probe in both failing-before and passing-after states) - [x] I have added or updated tests where applicable (the spec-driven probe walk is the regression test) - [x] I have updated relevant documentation to reflect my changes (inline comments explain the invariant) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
5b56d430e9 |
feat(ui): refine core navigation and task detail (#12854)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The main navigation and task detail view are core operator surfaces. > - Several controls used different hover states, popover layouts, and spacing rules. > - Recent task actions also needed a compact menu and correct inbox archive behavior. > - These differences made the interface feel inconsistent and caused some content to look crowded or clipped. > - This pull request aligns these surfaces with the Paperclip design tokens and current interaction patterns. > - The benefit is a simpler and more consistent operator experience in light and dark modes. ## Linked Issues or Issue Description **What happened?** Profile and organization popovers used inconsistent layouts. Navigation controls used different hover and selected backgrounds. Task warnings and the composer could crowd nearby content. Archiving a recent task could also remove it from more than the inbox. **Expected behavior** Popover menus should use the same compact visual language. Navigation controls should share readable hover and selected tokens. Task detail content should keep consistent spacing. Archiving should hide a task from the inbox while keeping it in the task list. **Steps to reproduce** 1. Open the main sidebar in light or dark mode. 2. Open the profile and organization menus. 3. Hover navigation items, the organization trigger, the profile trigger, and the feedback flag. 4. Open a task with a warning banner and a long thread. 5. Use the recent task overflow menu and archive a task. **Paperclip version or commit** Reproduced on `master` before this branch. **Deployment mode** Local dev (`pnpm dev`). ## What Changed - Rebuilt the profile and organization popovers with compact token-based layouts. - Matched organization popover width and alignment to the profile popover. - Unified sidebar hover and selected states in light and dark modes. - Added a recent task overflow menu with rename, archive, and pause or restart actions. - Kept archived tasks in the task list while removing them from the inbox. - Improved warning banner and composer spacing in task detail views. - Added and updated focused UI tests for the changed behavior. ## Verification - `pnpm check:token-gates` passed. - `pnpm --filter @paperclipai/ui typecheck` passed. - The seven affected UI test files passed with 216 tests. - `pnpm --filter @paperclipai/ui build` passed. - GitHub CI passed the full build, typecheck and release registry, general test, serialized server, canary dry-run, and end-to-end matrices. - Greptile reviewed commit `6e296be85` at 5/5 with no outstanding actionable findings. ## Risks - Risk is limited to sidebar presentation, recent task actions, and task detail layout. - The recent task archive action now follows inbox-only archive semantics. - No database schema or public API contract changed. > I checked [`ROADMAP.md`](ROADMAP.md). This pull request does not duplicate planned core work. ## Model Used - OpenAI Codex, GPT-5.6. The model used high reasoning, tool use, and code execution. The context window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Scott Tong <scott@scottsmbpm5max.lan> Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.904.0-canary.15 |
||
|
|
0ffc091473 |
feat(connections): add durable GitHub identities and webhooks (#12843)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents need source control access for repository work > - A shared token cannot preserve the responsible person's identity or an agent's dedicated identity > - GitHub App tokens also need durable refresh, repository access checks, and webhook delivery > - Paperclip already has managed connections, encrypted grants, run secret leases, and merge-confirmation behavior > - This pull request extends those systems with GitHub identities instead of adding a parallel credential system > - The benefit is durable GitHub access with explicit identity, repository, runtime, and webhook boundaries ## Linked Issues or Issue Description No public GitHub issue describes this connection change. This description follows the feature request template. **Subsystem affected** Connected Apps, connection grants, secret resolution, native Git runtime setup, webhook processing, and the Apps UI. **Problem or motivation** Users need to connect GitHub once and let agents use the correct GitHub identity. A run should use a dedicated agent account when one exists. Otherwise, it should use the responsible person's account. The connection must survive token expiry, repository access changes, and temporary instance downtime. **Proposed solution** Add user-owned and agent-owned GitHub grants to the existing connection model. Resolve one identity for MCP, Git, `gh`, health checks, and webhook bindings. Store provider tokens in the existing encrypted secret system. Refresh expiring token pairs under the existing lease and compare-and-swap path. Register signed Cloud webhook bindings and process normalized pull request and installation events through a durable local inbox. **Alternatives considered** An organization-wide GitHub token would lose person and agent attribution. Environment variables alone would bypass the managed connection and grant model. A new GitHub-only credential store would duplicate the existing secret and access systems. GitHub App installation tokens and private-key custody remain outside this first version. **Roadmap alignment** This change implements the Connected Apps direction. It also extends the shipped MCP Tool Gateway, per-agent secret access, and action-attribution systems. It does not add a repository catalog. The open repository catalog work in [#11234](https://github.com/paperclipai/paperclip/pull/11234) is related and complementary. ## What Changed - Added agent-owned connection grants and a per-agent credential policy with company and subject constraints. - Added a managed GitHub App method while keeping the personal access token method as an advanced fallback. - Added durable access-token and refresh-token handling with proactive rotation and one automatic recovery after a provider `401`. - Added GitHub identity and installation summaries without storing repository-name lists. - Added signed Cloud webhook binding, event lease, acknowledgement, local idempotency, pull request merge processing, and installation access handling. - Added one identity resolver for MCP, native Git, `gh`, checkout, health checks, and webhook bindings. - Added a class-3 run projection for `GH_TOKEN`, `GITHUB_TOKEN`, a `github.com`-only credential helper, SSH-to-HTTPS rewrite, and GitHub noreply commit attribution. - Added personal and dedicated-agent setup choices plus identity, repository, continuity, and webhook status in the Apps UI. - Added schema migrations, tests, and connection documentation. ## Verification - The current head is fully green in GitHub CI, including build, typecheck, all serialized/general server shards, all browser shards, policy, canary dry run, review, and security checks. - Live staging proof completed with a non-expiring GitHub App user token, selected-repository installation, repository add/remove refresh, managed MCP, native `gh`, HTTPS clone/push/delete, GitHub noreply commit attribution, signed merged-PR webhook acceptance, durable Cloud-to-instance delivery, and installation-access event processing. Temporary branches and temporary repository access were removed afterward. - `pnpm check:token-gates` passed. - `pnpm -r typecheck` passed before and after the rebase onto `origin/master`. - `pnpm build` passed. - The focused connector suite passed 285 tests after the rebase. - The full stable suite passed 5,790 tests and failed 22 tests across 8 general server files. The failures reproduced as shared-runner environment issues. They included `/tmp` versus `/private/tmp`, closed database connections, and invalid high ephemeral ports. The focused connection tests pass in isolation. ## Risks - Migrations add agent grant subjects and a durable connection-event inbox. Migration numbering and safety checks pass. - A raw GitHub user token enters the agent process for Git and `gh`. Per-tool Ask-first controls cannot limit those shell operations. The UI warns users about this boundary. - GitHub App user tokens can be non-expiring. Paperclip performs a continuity check every 30 days, but provider revocation still requires a reconnect. - The webhook path accepts only signed and bounded payloads. It stores a minimal normalized record and no raw provider payload. - GitHub repository permissions remain authoritative. Removed access can make a cached repository count temporarily stale, but runtime access fails immediately. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, `gpt-5.6-sol`, extended reasoning, tool use, code execution, browser control, and multi-file repository editing. The context window size was not provided. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
263f181fed |
fix(runner): complete live hot restart adoption (#12852)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner owns durable provider sessions and streams their work to the control plane. > - Pull request #12845 added native restart recovery for live and dead local runners. > - A real browser test found three live-adoption gaps after that pull request merged. > - Lazy runner process ownership was not always stored before restart. > - The old controller did not release its PRP authority without closing the provider turn. > - Reconnect events could arrive before the active provider turn was restored. > - This pull request closes those gaps and proves the same turn completes after a UI hot restart. ## Linked Issues or Issue Description Refs #12845 Related search results: #12646 covers indeterminate command results after a runner restart. It does not cover controller adoption or active-turn rebinding. No open duplicate pull request was found. ## What Changed - Store lazy runnerd process ownership after provider session creation, read, and resume. - Detach native PRP controller authority during coordinated hot shutdown. Keep the live provider turn running. - Restore the exact checkpointed provider session when bounded PRP identity events have been compacted. - Restore the active provider turn before reconnect events are replayed. This prevents `turn_binding_mismatch`. - Keep exact live ownership by the current controller out of generic orphan recovery. - Add driver, transport, and server regression tests for these paths. ## Verification - Ran 12 Codex driver lifecycle tests. - Ran 53 runnerd transport tests. - Ran 143 recovery and orphan-reaper server tests. - Ran all 8 real-process restart recovery scenarios. - Ran all 96 existing runner E2E unit tests. - Ran runner TypeScript typecheck. - Ran server TypeScript typecheck. - Ran the migration replay test and migration safety checks. - Tested the board UI on an isolated local instance. A real local Codex-backed turn entered a 120-second terminal wait. The UI `Restart now` action replaced the server and kept the same runner PID, process start time, run ID, native session ID, runner ID, provider session ID, and active turn. The original turn then completed. - Confirmed one heartbeat run, no retry row, one result, one proposed-result event, one terminal event, no protocol errors, no active recovery state, and no surviving runner or provider process. ## Risks - A live runner can continue provider work while no server owns the control route. Recovery fails closed when the process fingerprint or durable identity is ambiguous. - Provider identity can be restored from the database only for an exact verified adoption claim. An authenticated live `session.snapshot` validates that identity before the driver can resume. - The new detach path applies only to native sessions that expose restart detachment. Other adapters keep their existing shutdown behavior. - This follow-up does not change the database migration or `package.json`. The migration in #12845 remains replay-safe through `ADD COLUMN IF NOT EXISTS` and its embedded-Postgres idempotence test. The dedicated real-process command remains in `doc/DEVELOPING.md`. ## Model Used - OpenAI Codex based on GPT-5. The exact serving build and context-window size are not exposed. The run used extended reasoning, repository tools, shell execution, and in-app browser automation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
bb920fb859 |
ci: keep the Docker build context complete and guard it on every PR (#12855)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - GitHub Actions builds the Docker images that ship Paperclip, and downstream deployments consume the `-cloud` image variant on every master merge. > - PR #12769 slimmed the Docker build context with a broad `.dockerignore` block for `packages/paperclip-runner`, and the block also removed three files the image build itself reads. > - The image build re-runs the runner's generated-file drift checks, so it found no committed capability contract in the context and failed on every master commit after the merge. > - PR CI never runs those checks against the Docker context, so the pull request stayed green and the breakage only appeared post-merge, on every image build. > - This pull request restores the three files with narrow `.dockerignore` exceptions and adds a PR CI job that runs the drift checks against the exact Docker build context. > - The benefit is that image publishing works again now, and the next context-slimming regression fails the pull request instead of every post-merge image build. ## Linked Issues or Issue Description Refs #12769 (the context-slimming change that exposed this) and #12608 (which committed the generated contract outputs the image build checks). **What happened?** Every `Docker` workflow run on master failed from 2026-09-04 12:58Z onward, in both the `build-and-push` and `build-and-push-cloud` jobs. The failing step reported `Generated contract drift: generated/capability/capability-contract.md` from `check:capability-contract` inside `pnpm --filter @paperclipai/server build`. The committed contract file is current — regeneration on a full checkout is a no-op. The file was simply absent from the build context: the new `packages/paperclip-runner/**/*.md` ignore rule strips the committed drift-check outputs (`generated/capability/capability-contract.md`, `generated/capability/downstream-handoff.md`), and the `packages/paperclip-runner/docs` rule also strips `docs/capability-contract.md`, which `check:capability-inventory` reads next in the chain. No cloud image published for eight hours, which stalled every downstream deployment that consumes the canary images. **Expected behavior** The Docker build context must contain every file the image build reads, and a change that removes one must fail the pull request that introduces it, not every image build after the merge. **Steps to reproduce** 1. Check out master at any commit from `af3023f1` onward. 2. Run `docker buildx build -f .github/docker-context-checks.Dockerfile .` (the probe added by this PR), or start the real `Docker` workflow build. 3. Observe `Generated contract drift: generated/capability/capability-contract.md` — while `node packages/paperclip-runner/scripts/generate-capability-contract.mjs --check` passes on the same checkout outside Docker. **Paperclip version or commit** `d593463ab` (master tip at diagnosis time; first failing commit `af3023f1`). **Deployment mode** GitHub Actions image builds (`docker.yml`), consumed by managed cloud deployments. ## What Changed - `.dockerignore`: narrow exceptions (last match wins) re-include the committed drift-check outputs (`!packages/paperclip-runner/generated/**`) and the inventory check's documentation input (`!packages/paperclip-runner/docs/capability-contract.md`). Every other exclusion from #12769 stays: no crate declares an explicit `[[test]]` target, so cargo builds without the `tests` directories, and the image build chain never runs the excluded smoke scripts. - `.github/docker-context-checks.Dockerfile` (new): a small probe that COPYs the real build context — identical `.dockerignore` semantics — and runs the dependency-independent drift checks inside it (`generate-capability-contract.mjs --check`, `check-capability-inventory.mjs`). ajv installs in an isolated directory for schema validation only; codegen checks such as `generate-protocol-schema-module` stay out because their emitted bytes vary with the ajv release and would raise false drift alarms outside the locked dependency tree. - `.github/workflows/pr-trusted.yml`: new `docker_context_integrity` job builds the probe on every full-CI pull request, and the existing `verify` aggregate now requires its result, so the guard gates merges through the same required check as the other lanes. - Activation note: `pr.yml` pins `pr-trusted.yml` by commit SHA, so the new job starts gating pull requests after the usual follow-up `ci: activate ...` pin bump once this merges. The `.dockerignore` fix needs no activation — `docker.yml` reads it directly, so image builds recover on the first master commit after this merges. ## Verification - `docker buildx build -f .github/docker-context-checks.Dockerfile .` on master (before the `.dockerignore` fix): fails with the exact production error, `Generated contract drift: generated/capability/capability-contract.md`. - Same command with the `.dockerignore` exceptions applied: passes, which also proves BuildKit honors the `!` exceptions, including the file inside the excluded `docs` directory. - `node scripts/generate-capability-contract.mjs --check` on a full checkout: passes both before and after, which confirms the committed contract was never stale — only missing from the context. - Static sweep of every script in the image build chain (`build`, `build:typescript` and their `check:*` steps) against the ignore rules: the three restored files are the only build inputs the #12769 block strips. - YAML for `pr-trusted.yml` lints clean. ## Risks - Low. The `.dockerignore` exceptions only re-add three committed files to the build context; image contents do not change otherwise. - The probe job adds one context transfer and two Node scripts per full-CI pull request run (about one to two minutes, no dependency install beyond one isolated ajv package). - The `verify` aggregate now also requires the new job, mirroring the existing pattern for the other lanes; on non-full-CI runs the job skips and `verify` asserts the skip, unchanged from how the other lanes behave. - The new job only takes effect for pull requests after a follow-up pin bump in `pr.yml` (same two-step flow as every `pr-trusted.yml` change). > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Claude Fable 5 (Anthropic, model id `claude-fable-5`), extended thinking, agentic tool use in Claude Code: GitHub Actions log forensics to isolate the failing check, static analysis of the build-chain scripts against the ignore rules, and local docker buildx runs to reproduce the failure and verify the fix. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (the docker probe, both failing-before and passing-after; the drift checks themselves on a full checkout) - [x] I have added or updated tests where applicable (the probe IS the regression test for this class) - [x] I have updated relevant documentation to reflect my changes (inline comments in `.dockerignore` and the probe explain the invariant) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.904.0-canary.14 |
||
|
|
77312ee2d9 |
feat(codex): add GPT-6 Astra support (#12851)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - The Codex local adapter supplies model metadata to the server and the user interface. > - OpenAI now lists `gpt-6-astra` as a supported Codex model. > - Paperclip did not list this model or its model-specific controls. > - This pull request adds the model through the existing adapter metadata path. > - The benefit is that agents and task overrides can use the exact model ID and supported controls. ## Linked Issues or Issue Description **Subsystem affected** `packages/adapters` and `ui` **Problem or motivation** Paperclip does not expose `gpt-6-astra` in Codex model selectors. Operators cannot select and save the model through the normal agent and task forms. **Proposed solution** Register the exact model ID in the Codex local adapter. Use the adapter as the source for the model-specific reasoning options. Preserve the current default model. Forward the saved model, reasoning effort, and fast-mode controls through both Codex execution lanes. **Alternatives considered** A user-interface-only model list would duplicate adapter metadata. A model alias would not match the official model ID. Both options were rejected. **Roadmap alignment** This is a small adapter compatibility update. It does not duplicate a planned item in `ROADMAP.md`. ## What Changed - Added `gpt-6-astra` to the Codex local adapter model registry and fast-mode support list. - Added the official Astra reasoning efforts: `low`, `medium`, `high`, `xhigh`, `max`, and `ultra`. - Used the adapter metadata in agent and task model selectors. - Preserved supported effort choices when the model changes. Cleared an effort only when the new model does not support it. - Added tests for registration, user-interface selection, configuration persistence, and CLI and ACP forwarding. ## Verification - `pnpm exec vitest run packages/adapters/codex-local/src/index.test.ts packages/adapters/codex-local/src/server/acp.test.ts packages/adapters/codex-local/src/server/codex-args.test.ts packages/adapters/codex-local/src/ui/build-config.test.ts ui/src/lib/codex-reasoning-effort.test.ts ui/src/components/AgentConfigForm.render.test.tsx ui/src/components/IssueProperties.test.tsx ui/src/components/NewIssueDialog.test.tsx ui/src/lib/issue-assignee-overrides.test.ts` passed 245 tests. - `pnpm -r typecheck` passed. - `pnpm check:token-gates` passed all four gates across 939 files. - `pnpm --filter @paperclipai/ui build` passed and supplied isolated user-interface build proof. - `pnpm build` passed. - `pnpm test:run` passed 5,812 tests and failed 24 workspace-runtime tests in this isolated host. The failures use invalid generated ports above 65,535, incomplete nested-worktree fixture configuration, or `/tmp` path aliases. The focused tests for this change all pass. GitHub CI must pass before review handoff. - GitHub CI run `33918372718` passed all required checks and the aggregate verify gate on exact head `6ac6be2cee0a5996c82bdf674fcb7f46cb4c5fde`. - Independent engineering review approved the exact remediation head after 170/170 reviewer tests passed. - Greptile reported 5/5 with no open review threads on exact head `6ac6be2cee0a5996c82bdf674fcb7f46cb4c5fde`. - The model ID and capabilities were checked against the [official OpenAI Codex model list](https://developers.openai.com/codex/models). ## Risks - Low risk. The change adds one model and model-specific selector options. It does not change the default model. - OpenAI can change model capabilities later. The adapter metadata must stay aligned with the official Codex metadata. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with model ID `gpt-5.6-sol`, a 272,000-token context window, reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |