mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
fff410dfe777ae0427385e8297df992ba9aed4ce
4561
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b64469e403 |
feat(workspaces): add an operator default for isolated execution workspaces (#13444)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The execution workspace subsystem decides if a task run uses the
shared project checkout or an isolated per-task git worktree
> - The mode comes from the project policy, then the task settings. A
project that stores no policy always falls back to the shared checkout
> - An operator who wants every project to use isolated workspaces must
therefore edit each project one at a time, and must repeat this for each
new project
> - There is no instance-level control, so a fleet operator cannot set
this default at all
> - This pull request adds a managed experimental flag that moves the
default for projects that store no policy of their own
> - The benefit is that an operator sets the workspace default one time,
and every current and future project follows it
## Linked Issues or Issue Description
No public issue exists. The description follows the feature request
template.
**Subsystem affected**
Execution workspaces. The files are
`server/src/services/execution-workspace-policy.ts` and the run dispatch
path in `server/src/services/heartbeat.ts`.
**Problem or motivation**
`resolveExecutionWorkspaceMode` reads only the project policy, the task
settings, and a legacy field. Its last statement returns
`shared_workspace`. A project that stores no policy always gets the
shared project checkout.
An operator has no way to change this default for many projects at the
same time. The operator must edit each project, and must edit each new
project again later. Tasks in one project therefore share one checkout,
and they run one at a time when the environment driver makes the
scheduler serialize them.
**Proposed solution**
Add the managed experimental flag `enableIsolatedWorkspacesByDefault`.
When the flag is on, a project that stores no policy of its own resolves
as if it selected isolated workspaces. A project that stores a policy
keeps that policy.
The new helper substitutes a project policy. It does not move the last
statement of `resolveExecutionWorkspaceMode`. Two behaviors make this
necessary:
- A task that has no project must keep its current behavior. An isolated
workspace needs a repository to cut a worktree from.
`isUnrunnableWorktreeCombo` blocks an isolated task that has no
`projectId` and no `projectWorkspaceId`. A moved fallback would resolve
isolated for project-less tasks, such as agent chat, and stop them
before dispatch.
- The mode and the strategy must agree.
`buildExecutionWorkspaceAdapterConfig` supplies the default
`git_worktree` strategy only when one layer asserts workspace control. A
moved fallback would leave isolated mode with a `project_primary`
strategy.
**Alternatives considered**
- Change the last statement of `resolveExecutionWorkspaceMode` to
`isolated_workspace`. This is one line, but it changes the default for
every deployment. It is also not gated, so it would apply where isolated
workspaces are off.
- Write the policy to each project row with a script. This does not
cover new projects, and it does not cover new instances.
- Add an instance defaults section to the managed-config document. This
needs a new document key, new validation, and new delivery code. A
boolean flag reuses the delivery machinery that exists today.
**Roadmap alignment**
`ROADMAP.md` does not list execution workspace defaults. This change
adds an operator control to an existing capability. It does not add a
new capability.
**Additional context**
The flag is `tier: "managed"`. A cloud operator can therefore deliver it
with the managed-config machinery that exists today. No new delivery
code is needed.
## What Changed
- Add `enableIsolatedWorkspacesByDefault` to the feature catalog with
`tier: "managed"`. Both defaults are off.
- Add the flag to the experimental settings schema, the type, and both
branches of `normalizeExperimentalSettings`.
- Add `applyDefaultIsolatedExecutionWorkspacePolicy` to
`execution-workspace-policy.ts`. It substitutes `{ enabled: true,
defaultMode: "isolated_workspace" }` only when the flag is on, the task
has a project, and the project stores no policy.
- Apply the helper in the run dispatch path in `heartbeat.ts`, after the
existing `gateProjectExecutionWorkspacePolicy` call. The `hasProject`
argument reads the resolved project row, not the raw `projectId` of the
task.
- Gate the new flag behind `enableIsolatedWorkspaces` at the call site.
The new flag does nothing on its own.
- Add a toggle card to the instance experimental settings page. The card
shows only when isolated workspaces are on.
- Add eight tests for the new helper.
## Verification
Commands:
```
pnpm --filter @paperclipai/shared typecheck
pnpm --filter @paperclipai/ui typecheck
cd server && ../node_modules/.bin/tsc --noEmit
```
The server typecheck script also builds a Rust binary. I ran `tsc`
directly because this machine has no `cargo`. The server package reports
no type errors.
Tests:
```
./node_modules/.bin/vitest run \
server/src/__tests__/execution-workspace-policy.test.ts \
server/src/__tests__/instance-settings-service.test.ts \
server/src/__tests__/instance-settings-cloud-defaults.test.ts \
server/src/__tests__/instance-settings-managed-overlay.test.ts \
server/src/__tests__/instance-settings-routes.test.ts \
server/src/__tests__/managed-config.test.ts \
server/src/__tests__/heartbeat-workspace-busy.test.ts \
server/src/__tests__/heartbeat-workspace-session.test.ts \
server/src/__tests__/heartbeat-workspace-ready-comment.test.ts \
server/src/__tests__/execution-workspaces-service.test.ts \
server/src/__tests__/issue-runtime-workspace-binding.test.ts \
server/src/__tests__/run-trust-preset.test.ts \
packages/shared/src/feature-catalog.test.ts \
packages/shared/src/settings-visibility.test.ts \
packages/shared/src/validators/instance.test.ts \
ui/src/pages/InstanceExperimentalSettings.test.tsx \
ui/src/components/Sidebar.test.tsx
```
All of these files pass. The new tests cover each of these cases:
- The helper substitutes an isolated policy for a project that stores
none.
- The helper changes nothing while the flag is off.
- The helper changes nothing for a task that has no project.
- The helper keeps a stored policy, including a policy with `enabled:
false`.
- The resolver returns `isolated_workspace` for an unpolicied project.
- An explicit task setting still wins over the operator default.
- The substituted policy produces the `git_worktree` strategy.
- A project-less task does not become an unrunnable worktree.
To confirm the behavior by hand:
1. Turn on Isolated Workspaces, then turn on Use Isolated Workspaces By
Default.
2. Open a project that has no execution workspace policy.
3. Start a task in that project.
4. The run gets its own worktree. Tasks in that project no longer wait
for each other.
## Risks
Low to medium. The details:
- The flag defaults to off, and it is inert unless
`enableIsolatedWorkspaces` is also on. An instance that does not turn on
both flags sees no change.
- A project that stores a policy keeps it. This includes a policy with
`enabled: false`, which the helper reads as a decision to stay on the
shared checkout.
- When an operator turns the flag on, the workspace configuration
fingerprint changes for projects that store no policy. Their next run
creates a new workspace. This is correct, because the mode did change,
but the first run after the change does more setup work.
- A task that is in flight when the flag changes resumes with a
different workspace path than the path its session remembers. An
operator should let current runs finish before turning the flag on.
- Isolated workspaces use more disk, because each task gets its own
worktree.
## Model Used
Claude Opus 5 (`claude-opus-5`) in Claude Code, with extended thinking
and tool use. The model read the repository, made the change, and ran
the typechecks and tests above.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
canary/v2026.915.0-canary.1
|
||
|
|
ee81cee76d |
fix: allow concurrent agent runs on one Anthropic subscription connection (#13445)
## Thinking Path > - Paperclip manages AI agents and their work. > - Agent runs use provider connections and subscription credentials. > - File-backed subscriptions need a lease while a run writes new credentials. > - Anthropic subscriptions pass a token and do not write a credential file. > - The current lease blocks concurrent Anthropic runs without protecting shared state. > - This pull request limits the lease to file-backed subscriptions and tests both paths. > - The result lets Anthropic runs share one connection safely while file-backed credentials remain serialized. ## Linked Issues or Issue Description **What existing behavior does this improve?** This change improves concurrent use of one subscription connection. Anthropic runs no longer wait for a lease that protects no shared file. **Subsystem affected** server/ — REST API and orchestration services, with related server tests and connection documentation. **Current behavior** The server takes a connection lease for every subscription invocation. The lease blocks a second Anthropic run while the first runtime remains open. **Proposed behavior** The server takes the lease only when the subscription writes a provider authentication file back to the grant. Anthropic runs can proceed at the same time. **Reason and benefit** Anthropic passes its token through an environment variable and writes no file. Removing this unnecessary wait improves concurrency and preserves credential safety for OpenAI and xAI. **Breaking changes** None. OpenAI and xAI keep the existing lease and retry behavior. ## What Changed - Limit the AI credential lease to subscriptions that write a provider authentication file. - Add coverage for concurrent Anthropic runs and hired-agent runs. - Keep the existing OpenAI contention assertions as a sensitivity control. - Update the AI connection documentation to describe the lease boundary. ## Verification - Run `server/src/__tests__/ai-connections.test.ts`. - Run `server/src/__tests__/agent-hire-ai-connections.test.ts`. - Run `server/src/__tests__/heartbeat-ai-subscription-contention.test.ts`. - Run the server type check. - Review the pull request checks after GitHub starts continuous integration. ## Risks The main risk is an incorrect provider classification. The write-back condition remains unchanged, and OpenAI contention tests keep the file-backed lease behavior under test. No schema or API contract changes occur. ## Model Used OpenAI GPT-5. Exact runtime model ID and context window are not exposed in this execution. The model used tool calls, shell commands, and GitHub operations. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
52d120f68d |
fix(release): wait 30 minutes for npm to expose a published version (#13436)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Every release publishes a batch of npm packages and then waits for each one to become visible before continuing > - npm accepts a publish immediately but exposes it later, so that wait exists to keep a release from continuing past a package nobody can install yet > - Today that wait was too short twice in a row, and each timeout aborted the release with the batch half-published > - Version numbers are derived from what is already on npm, so every retry moves to a new number and meets the same lag > - Two canary attempts burned two versions this way and shipped nothing > - This pull request raises the per-package budget from ten minutes to thirty > - The benefit is that ordinary registry lag costs waiting instead of a failed, half-published release ## Linked Issues or Issue Description No existing issue. The problem, in the bug report format: **What happened** `publish_canary` failed twice in a row with the batch half-published: ``` Warning: npm accepted @paperclipai/server@2026.914.0-canary.2, but the version did not become registry-visible. Error: stopping release: npm did not publish and expose @paperclipai/server@2026.914.0-canary.2 ``` The version was accepted at 17:49:09 and became visible at 18:04:29 — about five minutes after the poll gave up. `shared`, `db` and `adapter-utils` published at that version; `server`, `paperclip-runner` and the root package did not. **Expected behavior** Ordinary registry propagation delay costs the release some waiting, not a failure. A package that becomes visible after 15 minutes must not abort the batch because of the old 10-minute wait. Longer registry outages can still leave a partial batch. **Steps to reproduce** 1. Publish any channel while npm is propagating slowly. 2. A package takes longer than `NPM_PUBLISH_VERIFY_ATTEMPTS * NPM_PUBLISH_VERIFY_DELAY_SECONDS` to become visible. 3. The release aborts, that version is half-published, and the retry picks a new version number and meets the same lag. **Paperclip version or commit** Present on master. Observed on 2026-09-14 across canary runs in workflow run 34869494325. ## What Changed - Increase npm visibility checks from 60 to 180, retaining the 10-second delay: about 30 minutes per package. - Increase canary, nightly, beta, and stable publish job timeouts from 90 to 150 minutes. - Add offline regression tests using the workflow's actual settings. They cover the observed 15-minute 20-second delay, immediate visibility, exhausted retries, and job timeout sizing. - Load the access router in test setup so its cold transform does not consume the first permission test's 10-second timeout. The permission assertions are unchanged. ## Verification - `node --test scripts/release-lib.test.mjs`: 14 passed. - `pnpm run test:release-registry`: 129 passed. - `pnpm exec vitest run server/src/__tests__/access-routes-permissions-upgrade.test.ts`: 3 passed. - Regression proof in temporary fixtures: restoring 60 attempts fails the observed-delay test; restoring 90-minute jobs fails the timeout-budget test. - [CI run 34916632804](https://github.com/paperclipai/paperclip/actions/runs/34916632804): all jobs passed, including typecheck/release registry, build, all general and serialized server shards, browser tests, runner verification, and canary dry run. The PR has 31 successful checks and two expected Storybook skips at `a27f5e896`. - [Previously failing serialized shard](https://github.com/paperclipai/paperclip/actions/runs/34916632804/job/104215809616): all three access-route permission tests passed in CI after preloading the router. - Greptile's final review is 5/5 with no outstanding findings. Both review threads are resolved. - Local limits: `pnpm -r typecheck` and `pnpm build` stop at the runner package because this machine has no Rust `cargo` executable. The duplicate full local `pnpm test:run` was interrupted while the complete CI matrix ran. Targeted local results are listed above; full validation is from CI. The previous CI failures were unrelated to npm propagation: - [PR review](https://github.com/paperclipai/paperclip/actions/runs/34891770396) required a test file for this fix. - [Serialized server shard 3](https://github.com/paperclipai/paperclip/actions/runs/34891773736/job/104136392196) timed out in the first access-route permission test at 10 seconds. The other two tests in that file passed. - The canary dry run passed in that same CI run. ## Risks Low risk. Production behavior changes only in release waiting budgets. - Polling exits as soon as npm exposes the version, so healthy publishes do not wait longer. - An unavailable version now takes about 30 minutes to report. Polls remain bounded and still fail the release on exhaustion. - The 150-minute jobs leave roughly 30 minutes for setup/build plus four full polling windows. npm command runtime and later release steps also consume that budget; a broader outage can still interrupt a batch. - The permission-test change moves module loading into a bounded setup hook; it does not relax authorization assertions. - No schema changes or operational migrations. ## Model Used - Claude Fable 5 (`claude-fable-5`), 1M context, extended thinking, run through Claude Code with tool use and code execution. - OpenAI GPT-6 (Codex), with reasoning, tool use, and code execution, for the CI follow-up. Context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — targeted checks listed above; full validation passed in CI - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes — workflow comments and verification details - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a2e7ffdc34 |
fix(runtime): validate sandbox paths and preserve live controller leases (#13432)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents can run in remote sandboxes. > - Connection checks must use the selected execution target. > - Recovery must respect the controller that owns an active run. > - A host path or PID does not describe a remote sandbox. > - This pull request checks sandbox paths on the target and preserves live controller leases. ## Linked Issues or Issue Description **What happened?** Selecting an AI account for a sandbox agent could fail because the Claude ACP environment check tried to create the sandbox directory on the Paperclip host. The recovery sweep could also interrupt a sandbox run while its controller lease was still valid. It treated a PID absent from the local host as proof that the run had stopped. **Expected behavior** ACP checks directories on the selected execution target. Recovery leaves a run with a live controller lease alone. Its final database write rejects a stale snapshot after renewal, a claim, a controller change, or a runtime change. **Steps to reproduce** 1. Test a Claude ACP sandbox agent with a directory that cannot be created on the host. The check fails before this fix. 2. Give a running sandbox task a valid controller lease and a PID absent from the host. Run the stale-lock sweep without an in-memory handle. The sweep interrupts the run before this fix. 3. Renew or replace the controller between the sweep's read and write. The old snapshot must not end that controller's run. **Paperclip version or commit** Rebased onto `origin/master` at `0e9b24c8216171c26c8358ba387d77858e02c7a9`. All seven regression cases still fail against this base. Refs #13438, which supplies the managed hiring and task-connection behavior, and #13433, which preserves non-assignee subscription comment wakes. This PR preserves both upstream changes and addresses the two remaining sandbox failures. ## What Changed - Resolve and create Claude ACP test directories through the execution-target helpers. - Preserve active legacy controller leases during stale-lock recovery, including finalization after a task becomes terminal. - Recheck the controller, lease, runtime mode, and native ownership in the terminal database write. - Add two sandbox-directory cases and five database-backed controller-lease cases. - Document the target used for ACP directory checks. ## Verification - Red: all seven new cases fail against `0e9b24c82` without these two implementation changes. - Before the final upstream sync, 192 focused tests passed. Full `pnpm test:run` coverage completed using the repository's group/shard runner: all general server and workspace groups passed, and all 147 serialized server suites passed across the initial run and isolated continuations. Five cold-import timeout suites passed with `--experimental.fsModuleCache`; their assertions and deadlines were unchanged. - The hiring routes, default-selection service, and upstream hiring tests match `origin/master` exactly. The two remaining fixes are unchanged by the final rebase. - On final head `bd28d5cefbdf7084acc3759199cd9661907e7a26`, all 372 focused tests pass across 16 suites covering both upstream changes and these fixes. Two timeouts in the combined run (database setup and an existing ACP case) pass in isolated reruns with fresh test homes and temporary directories. `pnpm -r typecheck` and `pnpm build` also pass on this head. - All 32 active checks pass on final head `bd28d5cef`, including the full test matrix and browser shards, in [CI run 34904204849](https://github.com/paperclipai/paperclip/actions/runs/34904204849). Two Storybook checks are skipped by path filters. - Greptile reviewed final head `bd28d5cef` at 5/5 with no findings or unresolved review threads. ## Risks - A failed remote directory check still blocks connection adoption. - A live controller retains finalization authority after its task becomes terminal. Cleanup waits for ownership to expire and must pass the final ownership check. - No schema or credential-storage changes. ## Model Used OpenAI GPT-6 in Codex, with reasoning, repository inspection, code execution, and API tools. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8f1905d34d |
fix: provision all project repositories for local and sandbox tasks (#13442)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Projects can now attach several source repositories. > - Task preparation still treated these sources as alternative workspaces. > - Sandbox sync preserved Git history only for the selected repository. > - A task needs every attached repository to complete work across the project. > - This pull request prepares all distinct project repositories and preserves their separate Git histories through sandbox restore. ## Linked Issues or Issue Description **What happened?** A user reported that a project with two repositories received only the first repository in Daytona. Repository-only project rows also reached the agent with null local paths. Managed checkouts with matching repository names could resolve to the same directory. **Expected behavior** Local and sandbox tasks receive every distinct repository attached to their project. Repository-only sources work without preconfigured local folders. Each repository keeps its own Git history and working files. **Steps to reproduce** 1. Create a project with two repository sources and no local folder paths. 2. Assign a task to the project and run it in Daytona. 3. Inspect the task workspace and the repository paths exposed to the agent. 4. Observe that the original implementation supplies only the selected checkout. Related change: #13010 added multiple repository selection. The open repository-catalog proposals #11234 and #11228 cover a different data model. This fix uses the existing project workspaces. ## What Changed - Materialize each additional distinct repository as an editable checkout inside the task root. Seed configured local sources with their current working files and retain task edits across runs. - Pass materialized repository paths to local agents and native sandbox task prompts. Apply existing run-scoped Git credentials to each remote clone. - Preserve each repository's Git history, dirty files, and restore baseline during sandbox staging and durable recovery. Apply each repository's ignore rules and the operator's workspace exclusions. - Keep same-name managed repositories in separate directories. Report additional clone failures before the task starts. - Add task-level, checkout, sandbox round-trip, environment-hint, and recovery-descriptor regression coverage. Document checkout and restore behavior. ## Verification - Red: the original implementation fails the sandbox test because the second repository has no Git directory. It also fails the same-name checkout test and both real-database task tests because repository hints have no local path. - Green: focused tests pass for one and two repository-only sources, local source edits, clone failures, per-repository credentials, separate Git histories, ignored files, and recovery from remote or durable seed state. - Live Daytona smoke passed with two disposable repositories through the production provider sync functions. Both repositories arrived with Git history. Commits from both restored locally. Ignored files stayed excluded. The disposable sandbox was deleted. - Passed on final commit `93ab76763`: `pnpm -r typecheck` and `pnpm build`. - Final focused coverage: 254 assertions across the six changed test areas passed across the serial run and an isolated rerun of the existing process-kill timing test. The live Daytona smoke also passed. - The local `pnpm test:run` overlapped source edits and retained stale transformed code. Its first phase reported 12,240 passed assertions, nine failed assertions, three hook failures, and one worker error; later phases did not run locally. This run is not claimed as green. Fresh focused tests verify the changes, and every general/workspace and serialized-server CI shard passes on the final commit. - Final CI is green on `93ab76763`: all test shards, all three browser shards, typecheck, build, runner verification, canary dry run, and security checks. The initial unrelated chat-delivery browser timing failure passed in the final CI run. Optional Storybook visual checks were skipped. - Greptile is 5/5 on the final commit with no unresolved review threads. Its checkout-race finding was reproduced with a failing test, fixed, and rechecked. ## Risks - Additional repositories need disk space and clone time. Access failure for an attached repository stops preparation. - Additional checkouts live under `.paperclip-repositories/` and keep independent histories. Changes stay in those task copies; they do not overwrite configured source folders. - Detached or reconfigured repository copies are retained under `.paperclip-runtime/detached-repositories/`. Sandbox recovery retains per-repository merge baselines. - No database migration, UI contract change, or new credential delegation is required. Referenced projects retain their separate read-only behavior. ## Model Used OpenAI Codex, based on GPT-6, with repository inspection, code execution, and tool use. The runtime does not expose a more specific model deployment ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>nightly/v2026.915.0-nightly.1 canary/v2026.914.0-canary.7 |
||
|
|
0e9b24c821 |
fix: resume subscription comment wakes and preserve retry status (#13433)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Concurrent task runs can share a managed AI subscription with one credential lease. > - #13438 added durable retries for busy subscriptions and task-lock checks. > - A comment wake can run for an agent who is not the task assignee. Such a run never owns the task lock, so the new checks suppress its retry. > - This pull request preserves those comment wakes while keeping the lock checks for assignee runs. > - It also shows the subscription wait in task status and preserves the failure count through the database projection. ## Linked Issues or Issue Description Refs #13438. **What happened?** When a subscription is busy, a non-assignee comment wake is cancelled without a successor. Repeated subscription waits also display an inflated attempt count because the execution query omits the preserved failure count. **Expected behavior** An eligible comment wake waits and resumes without claiming the assignee's task lock. The task shows “Waiting for AI subscription”. Waiting does not consume provider-failure retries. Assignee retries still stop when the task lock is cleared or transferred. **Steps to reproduce** 1. Configure an agent with a managed subscription and concurrent runs. 2. Hold its credential lease in one run. 3. Mention the agent on a task assigned to another actor. 4. Release the lease and inspect whether the comment wake has a scheduled retry. The test uses a real embedded Postgres database and a real credential lease. Provider execution uses a fixture. The reassignment test applies a database mutation after the real checkout. No live provider account is needed. **Paperclip version or commit** Based on `f912ecaac` from #13438. Before the reconciliation fix, the rebased regression at `2063cfd13` failed because the comment wake had no scheduled retry. The original configuration failure was reproduced on `5282cabde` before #13438 merged. **Deployment mode** Built from source with embedded Postgres for local verification. ## What Changed - Record non-assignee comment-wake authority at admission while holding the task and run locks. A later reassignment cannot grant this exception. Preserve it through repeated subscription waits. - Keep master’s lock checks for assignee retries, including the check inside the scheduling transaction. - Show the subscription wait and preserve its failure count in the execution projection query. - Limit pre-provider wait receipts to fresh executions. A persisted native execution input must retain its recovery path. - Add lease, retry, projection, cancellation, reassignment, pause, revocation, service-recreation, and ownership regression coverage. - Use master’s 60–120 second retry interval and document the resulting behavior. ## Verification - Red: after rebase, the comment-wake test failed with no scheduled retry; the other 50 contention and retry-scheduling tests passed. - Red/green: reassigning the task immediately after real checkout reproduced an unwanted successor for both assignment and comment wakes. Both tests pass after recording authority at admission. - Green: all 285 targeted tests pass across 12 suites, including #13438’s four cancellation-race phases, its hiring/connection suite, run-dispatch integration tests, and both reassignment regressions. - `pnpm -r typecheck` passes after rebase. - `pnpm build` passes on the final revision. - The full general and serialized test suites pass in [CI run 34899514681](https://github.com/paperclipai/paperclip/actions/runs/34899514681) on `99f1e4c1d`. Full-suite verification ran in CI; local verification used the 285 targeted tests. - All 32 active PR checks pass, including build, typecheck, native runner verification, canary dry run, and all browser shards. Two Storybook checks are skipped by path filters. - Greptile reviewed `99f1e4c1d` at 5/5 with no outstanding findings. All review threads are resolved. - `git diff --check` and `node scripts/check-module-boundaries.mjs` pass. ## Risks - The non-assignee exception requires authority recorded under admission locks or its server-created subscription retry. Tests cover caller-supplied flags, assignment and comment reassignment races, repeated comment waits, and assignee lock protections. - A wait uses master’s 60–120 second interval. A long-held lease can produce multiple wait records. - Existing native sessions can have provider effects. They must not receive a fresh-execution receipt that permits replay. - No schema migration or credential permission change is required. ## Model Used OpenAI Codex, model `gpt-6-astra`, with repository search, code editing, shell execution, and test tools. The session does not expose its context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f912ecaacf |
fix: carry AI connections through hiring and unblock task execution (#13438)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents hire other agents and assign tasks to them. > - Managed AI connections must follow those hires across legacy and native runners. > - Missing accounts should pause task execution and let the user connect from the task. > - Subscription contention must wait without asking for new credentials. > - This pull request fixes these paths and the native tool and Daytona staging failures found during live tests. > - The result is a working hire, subtask, and connection setup flow on local and remote runners. ## Linked Issues or Issue Description **What happened?** A managed Claude or Codex agent could hire a teammate without a usable AI binding. Cross-provider hiring could fail before the user had a chance to connect the new provider. First-time task setup did not show the existing AI credential form inline. A busy subscription could request a new connection. Native API replies could stop the parent after a hire had already committed. Fresh Daytona sandboxes could fail to extract read-only skill directories created on macOS. **Expected behavior** Compatible hires inherit the managed connection choice. A hire for another provider uses the responsible user's default. If that account is missing, the hire succeeds and the task asks for a connection. Completing setup in the task resumes work automatically. Explicit child auth settings and existing unmanaged login paths keep precedence. Shared-account access checks remain in force. **Steps to reproduce** 1. Connect a Claude or Codex parent with a managed AI account. 2. Ask it to hire one agent of each provider and create a self-assigned subtask. 3. Assign work to both hires without connecting the second provider first. 4. Connect the missing provider from its task card. 5. Check that all tasks finish and same-provider work uses the original account. 6. Repeat with native runners and fresh Daytona sandboxes. The opt-in browser suite in `tests/hiring-ai-connections/README.md` performs these steps. **Paperclip version or commit** The live failures were reproduced from `f2c5e54dc`. The branch is rebased onto `5282cabde`. **Deployment mode** Isolated local development instance. Legacy CLI and native runners. Local execution and ephemeral Daytona sandboxes. Related work: Refs #13247 for managed AI connections. Refs #13268 for legacy credential-reference inheritance, which this branch preserves. Refs #13432 for a concurrent managed-inheritance fix. This PR also covers cross-provider task setup, subscription waits, native API replies, and Daytona extraction. It permits missing responsible-user defaults at hire time; restricted shared selections still fail. ## What Changed - Apply managed connection defaults to both agent creation routes. Preserve explicit auth choices and legacy credential-reference inheritance. - Allow hires before their responsible user connects the provider. Keep approval gates, company boundaries, and shared-account access checks. - Reuse the production AI credential form inside the pending task card. Resume the task after setup. - Retry subscription lease contention without consuming the provider-failure allowance or creating a connection request. - Require task execution-lock ownership when scheduling, promoting, and dispatching subscription retries. Recheck ownership under the issue row lock. - Rename the HTTP operation identity at the native tool boundary so it cannot override the runner's operation identity. - Delay directory permission restoration during Daytona extraction. Preserve the final read-only modes. - Add database-backed regressions, real browser acceptance tests, and Storybook states. Document setup and run-log behavior. ## Verification - Six real browser scenarios passed: both parent providers on legacy local and legacy Daytona; native Codex locally; native Claude on Daytona. Each scenario hires both providers, completes a self-subtask and assigned work, and connects the missing provider inline with automatic continuation. - Successful runs verify the account, responsible user, runner mode, and Daytona lease. All 18 test sandboxes were deleted. - Live authentication used API keys. Subscription inheritance, lease contention, and retry have integration coverage. Fresh subscription OAuth sign-in was not automated. - Red/green tests reproduced missing bindings, missing inline forms, subscription contention, native API reply failure, and GNU tar permission failure. - Seven Storybook browser checks passed. They cover both providers, method selection, narrow layout, completion, cancellation, and invalid credentials. - Full local suite coverage completed before rebase. Initial timing and fixture startup failures passed unchanged on isolated reruns. The first full command did not exit cleanly; the remaining workspace and serialized groups were completed separately. - After rebase, 107 hiring/auth/retry tests and 59 native API, task-card, and Daytona tests passed. The full workspace typecheck, production build, and token gates passed again. Storybook build passed before rebase. - Review fixes: 169 hiring/retry/dispatch tests, 37 adjacent tests, and four explicit cancellation-race cases passed. Eight cross-provider cases cover stale auth keys on both creation routes and both runner types. Server typecheck and build passed. - Final CI on `ee4890837a8a4913e07453392b9a75969580dae1`: 32 checks passed. Two optional Storybook jobs were skipped. The full server, workspace, browser, native runner, build, typecheck, and release checks passed. - Three unchanged tests initially failed on a busy port, a chat row-lock race, and preview-server readiness. Each affected job passed after one CI rerun. Isolated local checks also passed: 41 credential tests, the chat-concurrency case, and 25 preview-runtime tests. - Greptile reviewed the final commit at 5/5. Both review threads are resolved. GitHub reports no merge conflicts. ## Risks - A missing personal account now defers authentication to the first task. Explicit incompatible bindings and restricted shared accounts still fail at hire time. - An inherited personal default uses the responsible user's existing authorization to install access for the new agent. It never copies credentials or another user's identity. - Subscription contention retries after a delay and rechecks task eligibility. It does not consume the provider-failure budget. - Native hiring uses the existing managed API-tools opt-in. Remote native runners require a matching Linux binary and provider pack, as documented in the acceptance README. - No schema changes. Live tests make paid provider calls and remain opt-in. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) in Codex, with reasoning, repository inspection, code execution, browser automation, and API tools. The context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5282cabde8 |
fix(ui): make every agent reachable from Chats (#13420)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Experimental Agent Chat gives each person a persistent conversation
with each agent.
> - The sidebar only showed starred or previously visited agents.
> - A person could not start a conversation with an agent absent from
that list.
> - This pull request adds a searchable agent picker and a default
sidebar entry.
> - People can now reach every company agent while keeping their
frequent conversations close.
## Linked Issues or Issue Description
Refs #13283, which introduced the existing task-backed agent chat
surface. This change adds discovery to that surface after maintainer
review of the Storybook design.
**What happened?**
With Agent Chat enabled, a person with no starred or recent agents had
no agent shortcut. The sidebar also lacked a direct way to start a chat
with another agent.
**Expected behavior**
Show the first-created company agent by default. Let the person search
all company agents and open a persistent conversation.
**Steps to reproduce**
1. Enable Agent Chat in a company with agents.
2. Use an account with no starred agents or recent chats.
3. Look for a chat entry in the sidebar, or try to chat with an agent
absent from its shortcuts.
**Paperclip version or commit**
The existing behavior is present on master at
|
||
|
|
5054c9ef9b |
fix(grok): stage the environment test from a host directory, and survive an absent workspace (#13416)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - An agent runs through an adapter. Before you use an environment, the
adapter's environment test probes it and reports structured checks
> - The grok adapter's environment test stages the managed account into
a remote environment first. It gave the runtime its resolved `cwd` as
the host workspace directory
> - For a remote target that `cwd` is a path inside the sandbox. It does
not exist on the host that runs the test
> - A new workspace ignore scan reads that host directory. The scan
failed on the missing directory, and the error escaped the environment
test
> - The user got a server error instead of a result with checks
> - This pull request stages from a host temporary directory, sends the
remote path separately, and reports a staging failure as a check
> - The benefit is that the grok environment test always answers with
checks, and a workspace directory that does not exist no longer stops a
runtime preparation
## Linked Issues or Issue Description
No existing issue. The problem, in the bug report format:
**What happened**
The grok environment test failed with `Error: Workspace ignore scan
failed: git-ignore-scan-failed`. The route returned a server error, so
the user saw no checks at all.
**Expected behavior**
The environment test always returns `{status, checks}`. Every other
problem it finds (an invalid working directory, a command it cannot
resolve, a probe that times out) becomes a check with a level. A
credential staging problem must do the same.
**Steps to reproduce**
1. Configure a grok agent with a managed AI connection.
2. Point the agent at a remote (sandbox) environment.
3. Run the environment test for that environment.
**Paperclip version or commit**
Present on master. Both halves landed on 2026-09-12: the call site in
#13247, and the scan that it trips in #13353.
## What Changed
- `packages/adapters/grok-local/src/server/test.ts` stages the managed
account through a fresh host temporary directory and passes the remote
path as `workspaceRemoteDir`. This is the split `codex-local` and
`opencode-local` already use.
- A failure while staging becomes a `grok_environment_unprepared` check
with level `error`, instead of an exception that escapes the function.
The probe does not run after it, because there is no prepared
environment to probe.
- The temporary directory is removed in the existing `finally` block.
- `packages/adapter-utils/src/sandbox-managed-runtime.ts` treats a
`workspaceLocalDir` that does not exist as "nothing to sync". A
directory that is not there has no files for ignore rules to govern,
nothing to stage, and nothing to restore.
- Tests: the grok environment test now covers the managed-connection
remote branch, which had no coverage. `sandbox-file-sync.test.ts` covers
a preparation whose workspace directory does not exist.
## Verification
```
npx vitest run packages/adapters/grok-local/src/server/test.test.ts # 9 passed
npx vitest run packages/adapter-utils/src/sandbox-file-sync.test.ts \
packages/adapter-utils/src/sandbox-managed-runtime.test.ts # 90 passed
cd packages/adapter-utils && npx tsc --noEmit # clean
cd packages/adapters/grok-local && npx tsc --noEmit # clean
```
The two new grok tests fail against the old call site: the first asserts
the runtime never receives the remote path as its host workspace
directory, and the second asserts a staging failure becomes a check
instead of an exception.
## Risks
Low risk, and limited to environment preparation.
- The adapter change only affects the managed-connection remote branch
of one adapter's environment test.
- The `adapter-utils` change makes a preparation that used to throw now
continue with no workspace sync. A real workspace is unaffected, because
the directory exists in that case and the scan runs exactly as before.
- No migration. No API change.
## Model Used
- Claude Fable 5 (`claude-fable-5`), 1M context, extended thinking, run
through Claude Code with tool use and code execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes — no
documented behavior changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green — pending first run
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
pending first review
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
46bcc0863c |
test(runner): stop two CI-contention flakes that starve the canary lane (#13418)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Each master commit is published as an npm canary, and the nightly release gate installs the newest canary and runs the onboarding smoke against it > - A canary only publishes after the Cloud readiness workflow passes for that exact commit > - Cloud readiness fails often, and almost every failure stops that commit from ever becoming a canary > - The nightly gate then keeps testing an old canary, so a fix that landed after it cannot reach the gate, and the gate stays red for reasons no longer in the code > - Most of these failures are short timeouts that only expire because the runner is loaded > - This pull request gives two such timeouts the budget the surrounding tests already use > - The benefit is that more master commits publish a canary, so the nightly gate tests current code ## Linked Issues or Issue Description No existing issue. The problem, in the bug report format: **What happened** Cloud readiness failed on 25 of 85 runs over three days (29%). 24 of those failures stopped the commit from publishing a canary. Two timeout clusters account for 10 of the 25: - `packages/paperclip-runner/src/control-plane/durable-prp-control-plane.test.ts` — `Error: Test timed out in 5000ms.` in "promptly fails the real transport request and notification paths on authenticated bad semantic input". Seen in runs 34652752663, 34654215468, 34656885170, 34657111560, 34661910120. - `packages/paperclip-runner/src/live/live-session.test.ts` — `Turn settled before the intentional SIGKILL: Error: Capability Codex turn timed out after 2000ms`. Seen in runs 34632264036, 34658026290, 34658106386, 34727991195, 34760283177. **Expected behavior** A commit that is good publishes its canary. A test fails because the behavior is wrong, not because the runner was busy. **Steps to reproduce** 1. Run the Cloud readiness workflow on master repeatedly. 2. Watch the Runner protocol and chaos jobs. 3. The two tests above fail intermittently, and the commit then publishes no canary. **Paperclip version or commit** Present on master. Measured over runs from 2026-09-11 to 2026-09-14, which is the workflow's whole history. ## What Changed - `durable-prp-control-plane.test.ts`: the "promptly fails the real transport request..." `it.each` now takes the same 15s budget as the test directly above it. It ran on the 5s default, although it builds a real control plane over a socket. The neighbouring test already asks for 15s, so this reads as an omission. - `live-session.test.ts`: the real-runner SIGKILL test gives the turn a 60s timeout instead of 2s. The test kills the runner process and then asserts the turn was still running at that moment, so the timeout is only a backstop. At 2s a loaded machine settles the turn first and the assertion fails. Neither change weakens an assertion. No test in this PR asserts that a timeout happens, and the tests that do assert one (`turnTimeoutMs: 500` and `turnTimeoutMs: 20`) are untouched. ## Verification ``` cd packages/paperclip-runner npx vitest run src/control-plane/durable-prp-control-plane.test.ts # 87 passed npx vitest run src/control-plane/durable-prp-control-plane.test.ts \ -t "promptly fails the real transport" # 2 passed (both it.each cases) ``` The live-session change is verified by the readiness lane itself, because that suite needs a real runner and a provider. ## Risks Low risk, and limited to test budgets. - A genuine hang in either test now takes longer to report: 15s instead of 5s, and 60s instead of 2s. - A real regression still fails, because the assertions are unchanged. - Timeouts are not a complete fix for this lane. Port races (`EADDRINUSE`), lease races, and lost runner VMs cause about another 40% of its failures. A retry on the `verify` matrix would cover all of those classes, and none of them reproduce on a second attempt. That is a larger policy change, so it is not in this PR. ## Model Used - Claude Fable 5 (`claude-fable-5`), 1M context, extended thinking, run through Claude Code with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes — no documented behavior changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green — pending first run - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — pending first review - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
e1d2e279a3 |
fix(db): replay a query whose socket write failed on a recycled pooled connection (#13417)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - A hosted instance keeps its data in Postgres, and it reaches that
database through a connection pooler
> - A pooler recycles the server side of a connection that sat idle. The
client still holds the socket and believes it is open
> - The next query fails while the driver writes it to that dead socket,
with `write CONNECTION_CLOSED host:5432`
> - The request that happened to draw the recycled connection fails,
although nothing is wrong with the query or the database
> - Only one code path guards against this today, so the same error
keeps reaching users and error reporting from every other path
> - This pull request replays a query when the write itself failed,
because those bytes never reached the server
> - The benefit is that a recycled connection costs one retry instead of
one failed request
## Linked Issues or Issue Description
No existing issue. The problem, in the bug report format:
**What happened**
Requests fail with `write CONNECTION_CLOSED <host>:5432` (driver code
`CONNECTION_CLOSED`). It is most visible on the first queries after an
idle period, when the pool holds connections the pooler already
recycled.
**Expected behavior**
A connection that the pooler recycled while it was idle is not a
user-visible failure. The client notices the dead socket and uses a live
connection.
**Steps to reproduce**
1. Run the server against a pooled Postgres endpoint (for example a Neon
`-pooler` host).
2. Leave the instance idle until the pooler recycles the server side of
the pooled connections.
3. Issue any request that queries the database.
**Paperclip version or commit**
Present on master.
## What Changed
- `packages/db/src/transient-write-retry.ts` (new) wraps the root
`postgres.js` client. When a query fails while the driver writes it to
the socket, the wrapper runs it again on a fresh connection. Three
attempts, 50 ms then 100 ms backoff.
- `packages/db/src/client.ts` gives Drizzle the wrapped client. The
teardown registry keeps the real client, because shutdown must end the
actual pool.
- The wrapper is deliberately narrow:
- It matches only the write phase (`code === "CONNECTION_CLOSED"` and a
message that starts with `write CONNECTION_CLOSED`). The write failed,
so the server never saw the query, and a replay cannot run anything
twice. That makes it safe for reads and writes alike.
- An error after the write propagates untouched, because the server may
have acted on the query.
- `CONNECTION_ENDED` and `CONNECTION_DESTROYED` propagate untouched,
because they mean a deliberate shutdown.
- Queries inside `db.transaction()` run on the scoped client that
`sql.begin()` returns, which the wrapper does not touch. A transaction
that loses its connection must abort, not replay.
- A query runs once however many handlers attach to it, and it still
executes lazily, like `postgres.js` itself.
## Verification
```
npx vitest run packages/db/src/transient-write-retry.test.ts # 7 passed
npx vitest run packages/db/src/client-teardown-registry.test.ts \
packages/db/src/client-options.test.ts # existing db suites pass
npx vitest run server/src/__tests__/cloud-tenant-transient-db-retry.test.ts # 6 passed
cd packages/db && npx tsc --noEmit # clean
```
New tests cover the replay, the `.values()` form Drizzle uses, the
attempt budget, an error that must not replay, and one execution per
pending query. A Drizzle round trip runs through a fake wire-protocol
server, which shows the wrapper is transparent to ordinary queries.
## Risks
Low risk.
- The retry only fires for a failure during the socket write, where the
server never received the query. A replay therefore cannot duplicate an
effect.
- A permanently unreachable database costs two extra attempts and 150 ms
before the same error surfaces.
- Transactions keep exactly their current behavior.
- Existing `retryOnTransientDbConnectionError` in the auth middleware
stays. It wraps a broader set of codes for one path, and it is
unaffected.
- No migration. No configuration change. No API change.
## Model Used
- Claude Fable 5 (`claude-fable-5`), 1M context, extended thinking, run
through Claude Code with tool use and code execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes — no
documented behavior or configuration changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green — pending first run
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
pending first review
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
02c7175e72 |
feat(agents): hired agents inherit provider credential references (#13268)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters need provider credentials to start work > - A hired agent can lack the credential reference that its hiring agent already uses > - The child agent then cannot authenticate, even when the company has a valid credential > - This pull request copies matching credential references from the hiring agent to the hired agent > - The benefit is that hired agents can start with the provider access that the hiring agent already uses ## Linked Issues or Issue Description This change relates to [PR #9920](https://github.com/paperclipai/paperclip/pull/9920), which covers credential inheritance for other agent creation paths. This pull request covers hiring-specific inheritance and fixed Claude OAuth binding checks. ### What existing behavior does this improve? The agent hire route builds the child adapter configuration from the hire request only. ### Current behavior A hired agent does not receive matching provider credential references from its hiring agent. The child agent cannot run when the request omits the credential. ### Proposed behavior The hire route inherits matching credential references from the hiring agent. The request keeps priority. A Claude hire that supplies any Claude credential inherits none. ### Reason and benefit The child agent can use the provider access that the hiring agent already uses. The change copies references only and never copies raw token values. ### Breaking changes None. The change affects only hires that need an inherited reference. ## What Changed - Copy matching credential references from the hiring agent into the hired agent adapter configuration. - Preserve pinned versions, `required`, and `allowMissingOverride` fields on each copied reference. - Keep hire-request credentials ahead of inherited credentials. - Reject inherited fixed Claude OAuth bindings unless the parent agent passes company, adapter, and exact-binding checks inside the same transaction. - Add route and service tests for inheritance, precedence, and binding validation. ## Verification - `pnpm exec vitest run --project @paperclipai/server src/__tests__/agent-hire-auth-inheritance-routes.test.ts src/__tests__/agents-claude-oauth-binding.test.ts` — 60 passed. - `pnpm exec vitest run --project @paperclipai/server src/__tests__/agent-hire-idempotency-routes.test.ts src/__tests__/agents-service-secret-bindings.test.ts src/__tests__/secrets-service-user-secret-owner-scoped.test.ts` — 28 passed. - `tsc --noEmit` in `server/` — the error count matches the merge base, with no error in either changed source file. - `git diff --check` — clean. ## Risks Low risk. The route copies references, not raw tokens. The request keeps precedence. Company, adapter, and exact-binding checks protect the inherited Claude OAuth path. ## Model Used Codex, OpenAI GPT-5, with code execution and review support. The implementation commit predates this pull request handoff. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: nickyleach <331803+nickyleach@users.noreply.github.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9a5a364170 |
chore(lockfile): refresh pnpm-lock.yaml (#13350)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com> |
||
|
|
78ce96a48b |
fix(runner): restore native Claude context and read permissions (#13422)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner supplies each agent with instructions, assigned skills, and tools. > - Native Claude lost its skill snapshot before provider launch. It also changed exact model IDs to aliases. > - The read permission mode denied ordinary Paperclip reads because this runner has no interactive approval handler. > - This pull request restores the missing context and permits only assigned tools that Paperclip defines as reads. > - Other requests that need approval stop with a clear action for the operator. They do not retry automatically. > - Agents can complete read tasks. Users can see why a restricted task stopped. ## Linked Issues or Issue Description **What happened?** Native Claude could not find assigned skills. The ACP layer could replace an exact model ID with an alias and then fail identity verification. The default `approve-reads` mode denied Paperclip read tools. A denied write showed a generic transport failure. The tool descriptions also called live operations “mock” operations. The completion prompt said to call a completion tool once, although the protocol can reject a claim and require a corrected call. **Expected behavior** Load assigned skills before launch. Keep the selected model ID. Allow assigned Paperclip reads. Stop an operation that needs approval with clear instructions when no approval handler exists. **Steps to reproduce** 1. Configure a native Paperclip Runner agent with ACPX Claude and an exact model ID. 2. Assign a skill and ask the agent to use it. 3. Select the read permission mode and ask the agent to read task context and list documents. 4. Ask the agent to write a document. Check the task state and recovery message. **Paperclip version or commit** Reproduced from `d351e08deee1b49d3467a950d1a3f01131943441`. **Deployment mode** Local native runner with real Claude, Rust runnerd, the Paperclip server, and embedded PostgreSQL. Related: #13196 fixed remote skill staging in the legacy `claude_local` adapter. The native runner uses a separate path, which this pull request fixes. ## What Changed - Carry the runtime context through Rust and the ACPX sidecar. Load assigned Claude skills after the provider lifetime lease is held. Refresh the files on each open. - Write the exact requested model ID into the isolated Claude settings. - Grant exact MCP permissions for the intersection of assigned tools and Paperclip's read catalog. Provider hints cannot grant access. Existing task-control permissions stay in place. - Stop requests that need an unavailable approval handler. Preserve the typed error through the server. Show “Approval required” on the task and require operator action without automatic retry. - Label the setting “Allow Paperclip reads.” Remove “mock” from live tool descriptions and regenerate the contracts. - Change one completion-prompt sentence to require one accepted result. Add regression tests and update the runner documentation. ## Verification - Red/green regression tests cover read admission, unavailable approval handling, server recovery, and the task error message. - A real Claude read trial failed before the fix and completed after it: [red trace](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=95e6a326e24437f1), [green trace](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=85a1f347db5604d1). - Full browser tests used real Claude and the production tool authority. Reads completed. A write stopped with “Approval required.” No document was created and no automatic retry was scheduled. All nine checks passed: [events and screenshots](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=1c4ff10be0d779cc). - A separate full browser test assigned a skill, invoked it, and completed with a marker absent from the task prompt: [skill evidence](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=ba5fbcb29a60be52). - Post-rebase checks passed: 128 targeted runner tests, 46 Rust tests, 27 sidecar/protocol contract tests, `pnpm -r typecheck`, and `pnpm build`. The server/UI red-green checks and `pnpm check:token-gates` also passed. The full suite passed in CI, including all general and serialized test shards, all three browser shards, and runner `check:all`. The duplicate unsharded local full-suite run was stopped after CI passed. Braintrust links require project access. The traces contain provider/runner events and app outcomes. They do not contain raw model HTTP requests. ## Risks - `approve-reads` now allows assigned Paperclip reads. Other operations that need approval stop the turn. Users must review the operation and change permissions before retrying. - The Claude settings depend on the pinned ACP and SDK behavior. Automated tests and real Claude trials cover this boundary. - The native Codex path is unchanged. The ACPX Codex fallback shares the clearer approval failure handling. - Local execution was tested end to end. Remote execution was not run. This change has no database migration. ## Model Used OpenAI GPT-6 through Codex. The agent used reasoning, code execution, and browser tools. The exact deployment ID and context-window size were not exposed in the session. Live acceptance tests used `claude-sonnet-5`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
728f7185f6 |
feat: add native in-app announcements with persistent dismissal (#13403)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Self-hosted boards need a way to show occasional product announcements. > - An app release should not be required to publish or withdraw a card. > - Native card controls keep publishing consistent; the hero can use a static image or isolated HTML/CSS animation. > - This pull request renders a validated JSON feed with native components. > - It stores dismissals per account on each instance, so a closed card stays closed across companies and browsers. > - Named staging feeds let authors test content before production publication. ## Linked Issues or Issue Description **Subsystem affected** Board application shell, announcement delivery, and user preferences. **Problem or motivation** Operators need a small, optional announcement card. Users need reliable dismissal state. Authors need to test remote content without changing the production feed. **Proposed solution** Add one non-modal AnnouncementWell. Fetch validated JSON and content-addressed media through the instance server. Keep card controls native, with optional sandboxed HTML/CSS animation in the hero. Use stable announcement IDs for dismissal, an explicit empty manifest and quiet 404 handling. Provide a staged publishing helper and isolated test-drive guide. **Alternatives considered** Hosting the entire card as a page would move navigation and dismissal into remote content. This change limits HTML to a scriptless, isolated visual hero and keeps controls native. Browser-only storage would lose dismissals across browsers, so the instance stores account preferences. **Roadmap alignment** ROADMAP.md has no overlapping announcement feature. A GitHub title search found no related announcement pull requests. This work implements a maintainer-requested feature. ## What Changed - Add shared feed types, strict validation of every object, supported routes, expiration and version checks. - Add a board-only current-feed API, constrained media proxy, and idempotent dismissal API. Store the first dismissal and its company audit entry in one transaction. - Cache upstream data for one hour. Use conditional requests, request deduplication, response limits, public destination checks, and a three-second deadline. Treat a remote 404 as an empty feed with a fifteen-minute retry cooldown. - Keep announcement visibility stable when focus moves to browser chrome or another app pane; only tab visibility starts a return check. - Add a responsive native announcement card. Respect onboarding, dialogs and toast placement. Sync pending dismissals across tabs and retry after reconnect or return. - Add idempotent migrations for dismissals and validated publication IDs, design-guide examples, static and animated Storybook examples, and focused tests. The publication registry supports offline retries without accepting caller-invented IDs. - Add HTML/CSS animated heroes with static posters, automatic playback, reduced-motion handling, strict DOMPurify validation, an empty iframe sandbox and CSP that blocks scripts/network resources. - Add validated staging publication, content-addressed assets, an empty production manifest, preview fixtures, and authoring/operator documentation. ## Verification - The preceding implementation passed 98 targeted shared/server/publisher/route/OpenAPI/UI tests and 127 tests including the master rebase. The playback-control removal passes all 21 announcement UI tests, covering the rendered sandbox, fallback, reduced motion, dismissal and slow/stale state lookups. The preceding shared/server tests cover HTML validation and response sandbox headers. - The playback-control removal passes UI typecheck, production UI build, Storybook build and token gates locally. Browser verification confirms the animated card has only its dismiss button and two links, with no page errors. The full canonical CI matrix passed on current head `00e416431edb610861599d50490270bbd0f3c6b6`: 32 successful checks and two optional Storybook deployment checks skipped. This run needed no retries. Greptile reviewed this same head at 5/5 with no outstanding findings. - The local canonical general-server run passed 12,063 tests before reporting embedded-PostgreSQL startup failures in an unrelated fixture. All 31 tests in that fixture passed across isolated retries. The UI group passed 6,219 tests and other workspace groups passed 3,201; two CLI database-startup failures also passed individually. Serialized server suites were verified by the full CI matrix rather than repeating them locally. No source changes were needed for these environment failures. - The real S3/CloudFront staging manifest and both media asset headers were verified. Production remains empty/unpublished. The guide distinguishes the preview host's disabled edge cache from production cache requirements. - In the isolated test-drive, the animation visibly moves without playback controls. A 390×844 browser viewport keeps the card above navigation. Reduced motion makes no animation request. Both themes render correctly and browser page errors are empty. Browser fault injection verified that scripts cannot execute and CSS cannot make network requests; a missing animation leaves its poster and controls. - Refresh leaves the animated card visible. Closing it persists after reload and the API returns null. Earlier live checks verified dismissal across browsers, company-relative CTA navigation, modal deferral/restoration, and new-ID eligibility after restarting the same database. - The deployed empty feed and a real remote 404 return HTTP 200 with null from the board API, with a usable dashboard and no announcement popup or browser warnings. - Authoring documentation covers staging, animated HTML constraints, test-drive, withdrawal, ID reuse and cache-refresh steps. ## Risks - Animation supports self-contained visual HTML/CSS and inline SVG, without JavaScript or external resources. A static image is required. Older builds that do not recognize the optional animation field quietly hide that unsupported feed. - The default feed makes an outbound request from an instance when a board is used. Operators can disable it. Requests contain no account IDs, company data, cookies or interaction events. - Feed publication and withdrawal can take about 65 minutes to reach returning users because of CDN and instance caches. Expiration also removes visible cards locally. - Dismissals follow an account within one instance. No-login instances share the existing local-board identity. Separate installations do not share state. - Both tables are additive. A unique key prevents duplicate dismissals; the transaction prevents duplicate first-dismissal audit entries. The publication registry retains only validated IDs. AGENTS.md and the implementation spec document the required exception to company scope for these instance-level records. - Publication was limited to separate public staging prefixes on the existing preview host. Production remains empty/unpublished. No AWS policies or infrastructure were changed. ## Model Used OpenAI GPT-6 through Codex. The exact runtime model ID and context-window size are not exposed in this session. Capabilities used: reasoning, code editing, shell execution, tests, browser interaction, and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.914.0-canary.1 |
||
|
|
f1d57863d2 |
fix: make connection checks and task handoffs reliable (#13404)
Preserve connection-probe outcomes through cleanup, reduce unrelated startup work, and report selected Claude authentication accurately. Make artifact download actions match their labels. Route delegated feedback through its active child, retain accepted messages across completion, and avoid redundant worker runs for proven closing notes. Preserve explicit follow-ups, human input, company boundaries, source provenance, and mixed issue references. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2d47a8058d |
fix(apps): update connector artwork and theme fallback (#13361)
> Awaiting author review. Do not merge until the author explicitly approves. ## Thinking Path > - Paperclip helps people manage AI agents for work. > - Connector screens need recognizable app artwork. > - Several bundled marks have inconsistent artwork or dark-theme behavior. > - The shared logo component should retain the existing frame. > - This change replaces selected artwork and fixes local theme fallback. > - Users see consistent connector icons across shared component callers. ## Linked Issues or Issue Description **Current behavior** Some connector marks use outdated artwork or unsuitable theme variants. A remote dark logo can override a canonical local mark that works in both themes. **Proposed behavior** Use the selected bundled artwork with the existing gray rounded frame. Use local light artwork in both themes unless a distinct local dark variant exists. **Subsystem affected** Connector artwork, the shared AppLogo resolver, and its validation. **Breaking changes** No connector capability, permission, credential, or catalog activation changes. No duplicate PR was found in the earlier search. ## What Changed - Update 46 artwork files and only the app definitions whose logo paths need to change. - Keep a compact public manifest of identities, paths, visibility, and aliases. - Prefer local artwork in both themes and preserve the existing frame and padding. - Add locally runnable artwork safety checks and a light/dark Storybook gallery; leave PR workflows unchanged. - Document artwork conventions. Source research is kept outside the public manifest. ## Verification - Artwork check: 69 identities pass. - Node artwork validation tests: 13 pass (rechecked September 14). - Focused resolver, component, and catalog tests: 40 pass (rechecked September 14). - Maintainer decision: dedicated icon-validation CI is not required; the workflow remains unchanged. Greptile acknowledged 5/5 with no remaining code concerns on September 14. - UI typecheck, token gates, and Storybook build pass. - Earlier full build and repository typecheck passed. The broad local test run was stopped after workspace-runtime dependency fixture failures outside this change. Current GitHub CI remains the full-suite gate. - Visual approval remains outstanding. Review the canonical icon registry in both themes at 24–48px. ## Risks The SVG checker rejects common active features; it is not a general sanitizer for arbitrary uploads. Optical balance still requires human review. Remote fallback for unknown brands retains existing behavior. This change does not add an asset importer or new connector capabilities. ## Model Used OpenAI Codex, assisted by GPT-5 and GPT-6 with code execution and browser tooling. Exact hosted model IDs and context window sizes were not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Artwork comparison Before and after for every affected identity. Images follow your GitHub light/dark theme and are pinned to the base and PR commits. This compares artwork; the existing gray rounded product frame and padding are unchanged. | Connector | Before | After | |---|:---:|:---:| | AgentMail | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/agentmail-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/agentmail.svg" alt="AgentMail" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/agentmail.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/agentmail.svg" alt="AgentMail" width="48" height="48"></picture> | | Airtable | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/airtable.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/airtable.svg" alt="Airtable" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/airtable.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/airtable.svg" alt="Airtable" width="48" height="48"></picture> | | Asana | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/asana.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/asana.svg" alt="Asana" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/asana.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/asana.svg" alt="Asana" width="48" height="48"></picture> | | ClickHouse | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/clickhouse.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/clickhouse.svg" alt="ClickHouse" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/clickhouse-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/clickhouse.svg" alt="ClickHouse" width="48" height="48"></picture> | | Cloudflare | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudflare.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudflare.svg" alt="Cloudflare" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudflare.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudflare.svg" alt="Cloudflare" width="48" height="48"></picture> | | Cloudinary | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudinary-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudinary.svg" alt="Cloudinary" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudinary-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudinary.svg" alt="Cloudinary" width="48" height="48"></picture> | | Discord | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/discord.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/discord.svg" alt="Discord" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/discord-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/discord.svg" alt="Discord" width="48" height="48"></picture> | | GitHub | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/github-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/github.svg" alt="GitHub" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/github-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/github.svg" alt="GitHub" width="48" height="48"></picture> | | Gmail | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/gmail.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/gmail.svg" alt="Gmail" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/gmail.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/gmail.svg" alt="Gmail" width="48" height="48"></picture> | | Google Calendar | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-calendar.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-calendar.svg" alt="Google Calendar" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-calendar.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-calendar.svg" alt="Google Calendar" width="48" height="48"></picture> | | Google Chat | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-chat.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-chat.svg" alt="Google Chat" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-chat.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-chat.svg" alt="Google Chat" width="48" height="48"></picture> | | Google Docs | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-docs.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-docs.svg" alt="Google Docs" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-docs.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-docs.svg" alt="Google Docs" width="48" height="48"></picture> | | Google Drive | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-drive.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-drive.svg" alt="Google Drive" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-drive.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-drive.svg" alt="Google Drive" width="48" height="48"></picture> | | Google People | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-people.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-people.svg" alt="Google People" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-people.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-people.svg" alt="Google People" width="48" height="48"></picture> | | Google Sheets | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-sheets.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-sheets.svg" alt="Google Sheets" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-sheets.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-sheets.svg" alt="Google Sheets" width="48" height="48"></picture> | | Google Slides | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-slides.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-slides.svg" alt="Google Slides" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-slides.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-slides.svg" alt="Google Slides" width="48" height="48"></picture> | | Google Workspace Search | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-workspace-search.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-workspace-search.svg" alt="Google Workspace Search" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-workspace-search.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-workspace-search.svg" alt="Google Workspace Search" width="48" height="48"></picture> | | Hugging Face | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/hugging-face.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/hugging-face.svg" alt="Hugging Face" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/hugging-face.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/hugging-face.svg" alt="Hugging Face" width="48" height="48"></picture> | | Jam.dev (library only) | New library entry | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jam-dev-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jam-dev.svg" alt="Jam.dev" width="48" height="48"></picture> | | Jira | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/jira-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/jira.svg" alt="Jira" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jira.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jira.svg" alt="Jira" width="48" height="48"></picture> | | Linear | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/linear.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/linear.svg" alt="Linear" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/linear-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/linear.svg" alt="Linear" width="48" height="48"></picture> | | Manufact | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/manufact.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/manufact.svg" alt="Manufact" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/manufact-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/manufact.svg" alt="Manufact" width="48" height="48"></picture> | | Microsoft Teams | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/microsoft-teams.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/microsoft-teams.svg" alt="Microsoft Teams" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/microsoft-teams.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/microsoft-teams.svg" alt="Microsoft Teams" width="48" height="48"></picture> | | Miro | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/miro-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/miro.svg" alt="Miro" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/miro.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/miro.svg" alt="Miro" width="48" height="48"></picture> | | Mixpanel | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/mixpanel.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/mixpanel.svg" alt="Mixpanel" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/mixpanel-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/mixpanel.svg" alt="Mixpanel" width="48" height="48"></picture> | | Netlify | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/netlify.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/netlify.svg" alt="Netlify" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/netlify-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/netlify.svg" alt="Netlify" width="48" height="48"></picture> | | Notion | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/notion-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/notion.svg" alt="Notion" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/notion.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/notion.svg" alt="Notion" width="48" height="48"></picture> | | PagerDuty | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/pagerduty.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/pagerduty.svg" alt="PagerDuty" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/pagerduty.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/pagerduty.svg" alt="PagerDuty" width="48" height="48"></picture> | | PostHog | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/posthog-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/posthog.svg" alt="PostHog" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/posthog-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/posthog.svg" alt="PostHog" width="48" height="48"></picture> | | Postman | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/postman.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/postman.svg" alt="Postman" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/postman.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/postman.svg" alt="Postman" width="48" height="48"></picture> | | Shopify | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/shopify.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/shopify.svg" alt="Shopify" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/shopify.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/shopify.svg" alt="Shopify" width="48" height="48"></picture> | | Slack | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/slack.png"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/slack.png" alt="Slack" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/slack.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/slack.svg" alt="Slack" width="48" height="48"></picture> | | Stripe | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/stripe.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/stripe.svg" alt="Stripe" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/stripe.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/stripe.svg" alt="Stripe" width="48" height="48"></picture> | | Supabase | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/supabase.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/supabase.svg" alt="Supabase" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/supabase.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/supabase.svg" alt="Supabase" width="48" height="48"></picture> | | Telegram | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/telegram.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/telegram.svg" alt="Telegram" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/telegram.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/telegram.svg" alt="Telegram" width="48" height="48"></picture> | | Todoist | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/todoist.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/todoist.svg" alt="Todoist" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/todoist.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/todoist.svg" alt="Todoist" width="48" height="48"></picture> | | Wix | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/wix.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/wix.svg" alt="Wix" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/wix-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/wix.svg" alt="Wix" width="48" height="48"></picture> | | Zapier | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/zapier.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/zapier.svg" alt="Zapier" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/zapier.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/zapier.svg" alt="Zapier" width="48" height="48"></picture> | |
||
|
|
3052686b1e |
fix(ui): hide task chat attribution badge (#13251)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task chat shows messages from agents and users. > - An agent message can contain an internal on-behalf-of user value. > - The new task view showed this value as a badge next to the agent name. > - This badge added unwanted identity text to the task chat. > - This pull request removes the badge from the new task view. > - The benefit is a simpler agent identity row in the task chat. ## Linked Issues or Issue Description **What happened?** The new task view shows a `for <user>` badge next to an agent name when a message has an on-behalf-of user value. **Expected behavior** The task chat shows the agent name without the on-behalf-of badge. **Steps to reproduce** 1. Open the new task view. 2. Show an agent message that has an on-behalf-of user value. 3. See the attribution badge next to the agent name. **Paperclip version or commit** The issue occurs on `master` at commit `d0b7ba419`. **Deployment mode** Local dev and all other modes that use the web UI. ## What Changed - Removed the attribution chip from the task chat agent identity. - Added a regression test that supplies an on-behalf-of value and confirms that the badge is absent. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/task-chat/TaskChatBubble.test.tsx` passed with 21 tests. - `pnpm check:token-gates` passed. - `pnpm --filter @paperclipai/ui typecheck` passed. ## Risks - Low risk. The change only removes one badge from the new task chat view. - The message data remains unchanged. > This is a small UI fix. It does not add roadmap scope. ## Model Used - OpenAI Codex, GPT-5.4, with reasoning and tool use. The provider did not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/no-internal-issue-references`, `fix/sandbox-secret-resolution`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ef84f363b7 |
fix(ui): remove execution workspace access card (#13406)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Execution workspace pages show the tools and state for one isolated workspace. > - The workspace access card repeats runtime state that the page already shows. > - The card also adds open and repair actions that are not needed on this page. > - This pull request removes the complete workspace access card from execution workspace detail pages. > - The benefit is a smaller page that keeps attention on workspace controls and work results. ## Linked Issues or Issue Description **What existing behavior does this improve?** The execution workspace detail page shows a workspace access status card. **Current behavior** The page can show Ready, preparation, degraded, failed, or Workspace is not running states. The card can also show Open workspace and repair actions. **Proposed behavior** The page does not show the workspace access status card or its page-only actions. **Reason and benefit** The card adds status and controls that are not needed on this page. Its removal makes the isolated workspace page simpler. **Breaking changes** The workspace access card and its actions are no longer available from the execution workspace detail page. Workspace runtime controls remain available. ## What Changed - Removed the workspace access status card from execution workspace detail pages. - Removed the page-only login handoff and repair mutations that served the card. - Added a regression test for the removed card and text. ## Verification - `pnpm exec vitest run ui/src/pages/ExecutionWorkspaceDetail.test.tsx` passes with 8 tests. - `pnpm typecheck` passes. - `pnpm check:token-gates` passes. - `pnpm build` passes. - `pnpm test:run` found unrelated failures in `server/src/__tests__/workspace-runtime.test.ts` during the local broad run. The focused changed-area tests pass. - The complete GitHub Actions matrix passes on the latest run, including build, typecheck, server tests, runner verification, and e2e tests. ## Risks - Low risk. The change removes one UI card and the page-only code that supported it. - Users must use the remaining workspace runtime controls instead of this card. > This change is focused UI cleanup. It does not add a roadmap feature. ## Model Used - OpenAI Codex based on GPT-5, with reasoning, tool use, and code execution. The exact context window is not exposed in this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4cd7b40255 |
fix: recover new messages after historical native runs stop (#13405)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native recovery must distinguish a fresh user request from replay of
failed work.
> - Older runs can lose their process fields before a local stop receipt
exists.
> - A suspended durable session can still prove that the exact runner
and provider session are idle.
> - The message admission path ignored that evidence and kept new user
messages blocked.
> - This pull request uses the existing exact-state verifier for those
historical runs.
> - The user can start one fresh turn while the old history and unknown
outcomes remain intact.
## Linked Issues or Issue Description
**What happened?**
A user sent a new message after a native run exhausted recovery.
Paperclip saved the message but said the previous run had no verified
stop record. The old runner was suspended, with no active provider turn
or pending output. Its process fields had been cleared before stop
receipts were added.
**Expected behavior**
A new user message starts a fresh turn when the exact retained session
proves it is suspended and the other execution gates pass.
**Steps to reproduce**
1. Retain a failed native run with a terminal controller, cleared
process fields, and no process receipt events.
2. Retain its exact suspended runner state and idle provider state. Keep
its recovery hold.
3. Send a new user comment. Before this fix, admission returns no
successor.
**Paperclip version or commit**
Reproduced on master at
|
||
|
|
d351e08dee |
fix(ui): stop Recent Tasks storage feedback across tabs (#13402)
Version recent-task snapshots separately from activity, reject stale cache updates, and publish only on query or membership changes. Migrate history into an isolated storage namespace while preserving restart retries. Add regression coverage for conflicting caches, migration, comments, unchanged writes, and storage implementations on Node 24 and Node 26. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.914.0-canary.0 |
||
|
|
13368c5183 |
fix: unblock clean-machine onboarding for api_key AI connections (nightly smoke) (#13372)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The release pipeline gates each nightly on a Docker onboarding smoke. The smoke proves a clean machine can finish onboarding and hire the first agent. > - #13247, #13248, #13344, and #13351 changed the Connect step. Connect now creates an AI connection that the server verifies live with the provider. > - The managed adoption check also demanded a CLI hello probe. A clean machine has no provider CLI and cannot complete a subscription login. Onboarding dead-ends and the nightly gate fails. > - This pull request lets a live-verified API key adopt on the engine's own verdict, and re-verifies the key with the provider at adoption time. > - It also drives the release smoke through the API-key path against a provider mock that lives inside the test harness. > - The benefit is a green, deterministic release gate with no paid credential in CI, and a working first run for API-key users on clean installs. ## Linked Issues or Issue Description No public issue exists. The failure surfaced in the nightly release gate. Related PRs (no duplicates found): #13247, #13248, #13344, #13351 (the Connect changes), and #12423, #12135, #12151 (earlier release-smoke updates). **What happened?** The nightly Release cut failed its gate: [run 34749840498](https://github.com/paperclipai/paperclip/actions/runs/34749840498), job `smoke_nightly / smoke`, on published canary `2026.913.0-canary.2`. The wizard never left the "Connect a model" step. The subscription path waits for a human to run `claude auth login` on the server. The API-key path saves and live-validates the key, but the environment test then fails with `Command not found in PATH: "claude"` and `ai_connection_validation_incomplete`, and the wizard blocks the hire. **Expected behavior** A clean machine with a provider-accepted API key completes onboarding and hires the lead agent. The release smoke passes without a real paid credential in CI. **Steps to reproduce** 1. Run `scripts/docker-onboard-smoke.sh` with `PAPERCLIPAI_VERSION=2026.913.0-canary.2`. 2. Sign in, complete onboarding to "Connect a model", select "Use API key instead", pick Claude, enter a valid API key, and press Connect. 3. The environment test fails on the missing `claude` CLI and blocks the hire. **Paperclip version or commit** `2026.913.0-canary.2` (nightly candidate `c9e3bb7ca`). ## What Changed - `server/src/routes/agents.ts`: `testManagedEnvironment` no longer forces the CLI-lane hello probe for a resolved `api_key` binding. It re-verifies the key against the provider's live endpoint instead (the same `validateAiApiKey` check the save performed, which needs no CLI). A key the provider rejects fails adoption with `ai_connection_api_key_rejected`. Subscription adoption keeps the strict hello-probe requirement. - `scripts/docker-onboard-smoke.sh`: the harness now serves `api.anthropic.com` itself. A sibling container (the already-built smoke image) runs a small HTTPS mock. The app container gets `--add-host` for that one hostname and trusts the mock's certificate through `NODE_EXTRA_CA_CERTS`. The private key stays mode 600 in the mock container; the app container mounts only the certificate. The mock serves only `GET /v1/models` and returns 404 for every other path. `SMOKE_PROVIDER_MOCK=false` disables it. - `tests/release-smoke/docker-auth-onboarding.spec.ts`: the spec drives the API-key path — switch the credential mode before the source tile (the link hides when the row collapses), enter the key, and Connect. Loopback targets use a placeholder key that the mock accepts. Any other target must set `PAPERCLIP_RELEASE_SMOKE_ANTHROPIC_API_KEY`, and the test fails on arrival without it. - `server/src/__tests__/agent-test-environment-routes.test.ts`: three new route tests cover accepted keys (no CLI probe consulted), provider-rejected keys, and subscriptions that cannot complete a hello probe. ## Verification - `npx vitest run src/__tests__/agent-test-environment-routes.test.ts` — 26/26 pass. - `npx vitest run src/__tests__/ai-connections.test.ts src/__tests__/ai-legacy-compatibility.test.ts` — 42/42 pass. - `tsc --noEmit` reports no errors in the touched files. - Full local harness + suite run against the exact failing canary: the app container reaches the mock (request visible in the mock log), the placeholder key validates, and the connection saves as the default. The flow then stops at the forced CLI hello probe — the exact server check this PR removes, still present in the published canary. The next canary that includes this fix is the end-to-end proof. - Hardening check: from inside the app container, the mock answers with status 200 and `key.pem` is not visible. ## Risks - Behavior shift: `api_key` adoption no longer requires a CLI hello probe. It re-verifies the key with the provider at adoption instead. Subscription adoption is unchanged. - The mock returns 404 for unexpected provider calls, so a future onboarding change that calls a new endpoint fails the smoke loudly instead of passing silently. - Release-smoke runs against non-loopback targets now require an explicit key and fail fast without one. - No database migration. No dependency change. No provider routing change. ## Model Used Claude Fable 5 (`claude-fable-5`) through the Claude Code CLI, with extended thinking and tool use (shell, file edits, Playwright runs, GitHub CLI). No other models were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
f2c5e54dca |
fix(runner): preserve handoff work and publish requested files (#13355)
## Thinking Path > - Paperclip lets people manage AI agents and their tasks. > - A task keeps its instructions, progress, and files when its assigned agent changes. > - The replacement runner lost the interrupted run's context and could overwrite an existing draft. > - A saved message also stayed attached to the former agent and could reopen the task after the replacement finished. > - File tasks could report Done with only a local path that the user could not open. > - This pull request transfers handoff context and saved messages, and makes requested files accessible through the existing attachment contract. > - Users can change agents and collect completed work without repeating instructions or confirming bookkeeping. ## Linked Issues or Issue Description Refs #13338. Builds on merged #13354 for queue admission and #13353 for remote workspace retry. #10123 concerns restricted recovery-model escalation; this change instead covers ordinary native handoff and file completion. **What happened?** Codex wrote a draft before a user assigned the task to Claude. The replacement lacked continuation context and replaced the draft. A queued user message could later restart the former agent and reopen the completed task. Separately, a runner could finish a requested file but return only a machine-local path. Remote native runs had no bound file publication tool. **Expected behavior** The replacement reads and preserves existing work, receives saved messages once, and keeps each message's author. The former agent stays stopped. A requested file has a working attachment or accessible work product before Done. Text-only tasks do not require attachments. **Steps to reproduce** 1. Ask Codex to save three newsletter names and then wait. 2. Queue an instruction to keep those names and expand the draft. 3. Use Interrupt and assign to select Claude. 4. Verify the original names survive, the result has a working download, and only the source and replacement runs exist. 5. Ask either provider for a Markdown checklist and open the file from its completed response. ## What Changed - Carry the exact same-task interrupted run's summary, semantic receipts, and history into handoff context. Tell the replacement to inspect existing files before editing. - Adopt saved ordinary task comments into the successor's receipt under the task lock. Preserve authors and separate mention, chat, and interaction contracts. - Prevent a former-assignee comment wake from reopening a completed task or starting a stale execution. - Reject workspace-only, fabricated, and cross-task file completion references with actionable runner feedback. New file output also needs a matching current-run publication receipt and asset filename/size/hash, or an accessible work product registered by the current run. Prior output can remain context alongside a current file, or be verified and re-registered internally. Authorized chat attachment reuse retains its verified current-run clone receipt; older receipt shapes require an intact matching source. - Bind remote file reads to the active environment runner and reuse the existing attachment and work-product publication path. - Enforce workspace confinement, regular single-link files, stable identity, a 10 MiB limit, and exact size and SHA-256 checks. Rotate the native session fingerprint for the updated tool contract. - Contain rejected remote signals and protocol-failure cleanup, including logging failures. Preserve the original cleanup rejection for its owner; a rejected operation never supplies stop acknowledgement or cleanup proof. - Allow exactly one maximum-size base64 file through the native SSH command adapter, preserving a finite output cap. - Document handoff and accessible file completion rules. ## Verification - Each observed bug has a failing regression before its fix. Final post-rebase integration passed 732 tests across 13 files before the final receipt and signal guards; final affected results are below. - Publication provenance and compatibility: 8 provenance regressions and 2 compatibility regressions failed before their fixes; the final four affected suites pass 53 tests, including mixed old/new references and real authorized chat reuse. Controls cover old attachments and work products, filename/size/hash/origin mismatch, missing/wrong receipts, current-run publication, same-run durable proof, internally re-registering preserved bytes, and no-new-file follow-ups. - Remote signal rejection: the real Node subprocess previously exited 1 when the production launcher signalled a deleted sandbox. It now stays alive for both a failed signal and failed logging; all 349 executor tests pass. The failed signal still provides no termination proof. - Remote file reader and SSH command boundary: 33 tests passed, including real Linux descriptor reads and the actual SSH adapter subprocess output cap (network executable replaced by a deterministic fixture). Exact 10 MiB bytes pass, one byte beyond the encoded cap fails. - Live local Claude and Codex Stop journeys preserve the saved file, deliver queued instructions once, and reach Done with two total runs. The handoff journey preserves the original names and download with exactly two runs. Both local providers deliver exact checklist files without a completion confirmation. - Live combined Daytona verification passed: the original failed task's Retry reused its sandbox; a selected Git subfolder produced an exact downloadable file; warm and deliberately resumed Claude runs took about 33 seconds. Codex produced a 240-byte download in 33.1 seconds after 134 seconds of contention/backoff. Both cloud downloads retained exact bytes after the two owned sandboxes were deleted. - The sandbox-deletion retest identified a separate ignored promise in protocol-failure cleanup. Two real Node subprocess regressions failed under fatal unhandled-rejection policy before the fix; all 36 protocol, lifecycle, and integrity tests now pass. The original close promise still rejects to its owning runtime. The final live retest passed: a normal Claude Daytona task completed in 132.352 seconds, then its sandbox was deleted. Thirteen samples over 361 seconds confirmed the same controller stayed healthy, the task stayed Done with unchanged run IDs, and its attachment retained exact bytes. The post-deletion browser download passed with zero page errors; all five owned sandboxes are confirmed absent. - Full repository typecheck and build passed on final commit `de64d16f1`. Final-head Greptile is 5/5 with no unresolved threads. [Final-head CI](https://github.com/paperclipai/paperclip/actions/runs/34736623758) passed: 32 successful checks and two conditional skips. The earlier mixed-source full local test invocation was deliberately stopped before rebase, so no pristine green full local aggregate is claimed. Its known failures passed in later affected suites. ## Risks - Handoff may adopt only ordinary comments from its validated former owner. Other delivery contracts must remain independent. - File verification fails closed if a remote file changes during reading. The runner must retry publication or explain a blocker. - The updated session fingerprint starts a fresh provider process where needed to install the new tool contract. - No schema migration or historical status reconciliation is included. ## Model Used OpenAI `gpt-6-astra` through Codex, with reasoning, code execution, browser testing, and tool use. The context-window size is not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.913.0-canary.3 |
||
|
|
827ba8a434 |
fix: cancel stalled sandbox startup without waiting for setup (#13352)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandbox runs prepare credentials and files before an agent starts. > - Stop must work during that preparation. > - ACPX registered cancellation, but did not handle it while a setup command was waiting. > - Daytona cleanup waited for that same command before stopping the sandbox. > - This change stops the run's sandbox first and requires proof before abandoning setup. ## Linked Issues or Issue Description **What happened?** Stop left a sandbox run active when remote credential setup stalled. The run kept its connection lease until the sandbox was stopped separately. **Expected behavior** Stop terminates the selected run's sandbox, prevents later setup from launching the agent, and lets run cleanup finish. It must not report success without proof from the provider. **Steps to reproduce** 1. Start an ACPX agent in Daytona. 2. Hold a command during remote credential or file setup. 3. Select Stop before the agent starts. 4. Before this fix, cleanup waits for the held command and never reaches sandbox stop. Related: #13351 exposed this during connection acceptance testing. #12150 addresses scheduler load and session initialization limits, a separate startup problem. ## What Changed - Handle cancellation during ACPX sandbox preparation with a host-owned stop callback. - Pass an explicit active-work cancellation flag through environment cleanup. - Stop Daytona before draining setup commands. Keep normal graceful cleanup. - Require an exact run and lease termination receipt. Keep ownership of outstanding requests when stop cannot be verified. - Reject late setup work and defer sandbox resume until old requests settle. - Add regression tests and document the cancellation boundary. ## Verification - Red: both the stalled ACPX setup test and the Daytona cancellation test failed before the fix because Stop never reached the provider. - Green: adapter and Daytona suites passed, along with cancellation boundary and database-backed receipt/isolation tests. - Two real Daytona probes ran a five-minute setup command. Cancellation returned matching stopped receipts in 6.01 and 6.533 seconds. No later setup command ran. Both test sandboxes were deleted. - Repository typecheck and build passed locally. The full required CI suite passed, including all server/workspace tests, browser shards, Paperclip Runner verification, and canary packaging. The duplicate local full-suite run was interrupted after CI passed; it is not claimed as a completed local pass. - No UI changes. The live probe uses the actual adapter cancellation boundary and Daytona plugin; it is not a browser acceptance test. ## Risks - Provider stop failures remain unacknowledged. The adapter keeps ownership while its original requests remain active. - A retry can receive a settling-work error until old provider requests finish. - Cancelled sandboxes are stopped and retained under their existing provider expiry policy. - Local execution and cancellation after the agent turn starts keep their current behavior. - Other providers must return an exact termination receipt to permit early setup cancellation. No database migration. ## Model Used OpenAI GPT-6 through Codex, with code execution and tool use. The exact deployment identifier and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c9e3bb7ca4 |
fix: preserve queued work after native Stop and honor steering support (#13354)
## Thinking Path > - Paperclip lets people manage AI agents and their tasks. > - The runner owns execution, while the task keeps user instructions and status. > - Stop must stop the current response without losing instructions that the user already sent. > - The queue stored its original reason inside saved context. Recovery checked the outer deferred reason and left the message waiting. > - Claude also exposed Steer through a shared method even though its driver did not support it. A rejected request could remove its own error row. > - This pull request keeps queued work until execution has stopped, uses the driver's real capability, and preserves completion event order. > - Users can continue work without repairing task state or repeating messages. ## Linked Issues or Issue Description Refs #13338. Related recovery work: #13353. **What happened?** A message sent during a native run stayed queued after Stop. Claude exposed an unsupported Steer action. A steering failure could hide the queued row and its error. A terminal event could also precede the final provider result, and subtree Stop omitted the board actor. **Expected behavior** Stop ends the current execution. Once Paperclip proves that execution has stopped, it delivers the saved instruction once through normal admission. Pause and recovery holds still prevent dispatch. Unsupported controls stay disabled, and a rejected action leaves an actionable error visible. Final results precede terminal events. **Steps to reproduce** 1. Start a Claude or Codex task that writes a file and then waits. 2. Send a follow-up instruction while it runs. 3. Press Stop. Check that the queued instruction runs once and preserves the file. 4. Check Claude's Steer control and simulate a server rejection on the only queued message. ## What Changed - Recover saved native comments after acknowledged Stop using their original wake reason. - Require durable remote termination receipts or verified local process termination before dispatch. - Preserve actor identity, queued-message deduplication, Pause, and recovery gates. - Derive steering support from the driver descriptor and reject unsupported calls. - Keep the queue mounted until a steering request succeeds so its error remains visible. - Emit provider results before terminal events and pass the board actor into subtree Stop. - Document Stop and steering behavior. ## Verification - Final-head continuation suite: 104 passed, after failing regressions for saved wake reasons, cleanup proof, and deduplication of every queued message. Steering UI: 121 passed. Driver capability: 26 passed. Runner backend/transport: 205 passed; Rust library: 285 passed. - Real Claude and Codex browser journeys both preserved the saved file, delivered the queued instruction once after Stop, and reached Done with exactly two total runs. The process Stop browser fixture also passed. - Local full repository typecheck and build passed on `afaa35139`; server typecheck and the affected 104-test suite passed after the final queue changes. Token gates passed. Final-head CI verifies the complete integrated source. - Local aggregate evidence has explicit limits: the general-server invocation overlapped the queue fixes and finished with 11,989 passed, 2 failed, and 80 skipped; both failures are covered by the final 104-test pass. The UI and CLI then passed all 6,184 and 485 tests; the complete 145-file serialized rerun passed all 2470 tests. The shared-package lock fixture passed unchanged on rerun, but the package phase subsequently stopped at an embedded-Postgres bootstrap resource failure. No single pristine green local full aggregate is claimed. - Greptile reviewed `fa66e2bd5` at [5/5 with no unresolved findings](https://github.com/paperclipai/paperclip/pull/13354#issuecomment-5650334242). [Final-head CI completed successfully](https://github.com/paperclipai/paperclip/actions/runs/34733781888/attempts/2): 33 successful checks, 2 conditional skips, including all server, package, UI, browser, runner, typecheck, and build gates. The first attempt hit a preview-readiness/port-collision fixture; its unchanged local control passed 25 tests with 3 skips, and one supported unchanged CI retry passed the affected shard and aggregate gates. ## Risks - Queue recovery must never overlap an old execution. Unknown cleanup state remains blocked. - Driver descriptors are now authoritative for steering; a wrong descriptor disables the action instead of attempting it. - No schema migration or historical status reconciliation is included. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, browser testing, and tool use. The exact hosted model ID and context-window size are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.913.0-canary.2 |
||
|
|
6809314a3f |
fix: recover sandbox workspace setup and retries (#13353)
## Thinking Path > - Paperclip manages AI agents and their tasks. > - Sandbox tasks need a workspace and a provider session before work can start. > - A selected Git subfolder is a valid workspace, but it is not a repository fetch source. > - A resumed sandbox must keep one resource identity when the provider fills in its region. > - Workspace reuse does not prove that a provider session has started. > - A recorded continuation must not leave an old failure blocking Retry for a newer attempt. > - This pull request fixes those setup and recovery boundaries while preserving ownership checks. ## Linked Issues or Issue Description **What happened?** A Daytona task failed before provider startup when its project folder was inside a parent Git repository. After the folder was repaired, Retry resumed the same sandbox and verified its workspace sentinel, but workspace preparation reported that the lease was no longer active. A later attempt could fail because the resumed workspace had no provider checkpoint. The old run also kept a recovery projection after the server recorded its explicit successor, hiding Retry for the new failure. **Expected behavior** Sync the selected folder without importing parent files or history. Keep the existing sandbox identity stable. Allow a new provider session only with proof that its exact session has never started. Let the current failed attempt retain Retry once the old recovery has a recorded successor. **Steps to reproduce** 1. Select a subfolder of a Git repository as a project workspace and start a Daytona task. 2. Leave the target region unset. Release its reusable sandbox after setup fails, then resume and realize the workspace. 3. Retry a provider setup that failed before any session directory or checkpoint was created. 4. Record an explicit successor for a native failure, fail that successor during setup, and inspect Retry in the task thread. **Paperclip version or commit** Reproduced against `8d1f0c20a` with new failing regressions before each production fix. **Deployment mode** Local controller with a Daytona environment. Automated tests use deterministic provider fixtures and real filesystem operations. Related: #13338, #13163, #13264, #13349. The open work-folder stack in #13264 includes a broader fresh-session authority change. This patch addresses the independently reproduced startup failure with existing durable bootstrap proof and atomic directory creation. It does not include the work-folder migration or credential changes from that stack. ## What Changed - Classify only the selected Git repository root as a fetch source. Sync subfolders as directories and preserve their enclosing Git ignore rules on upload and restore. Recognize the shared scheduler’s completed non-repository result so ordinary folders still sync; timeout, cancellation, and output-limit failures remain closed. - Remove the placement target from Daytona account cache identity. Keep API endpoint, credential digest, company, environment, driver, and sandbox ID boundaries. - Preserve closed-lease admission until a sentinel-verified resume reopens that same resource. - Show the already-recorded explicit successor of a resolved recovery. Keep the old failure evidence and unresolved holds. No historical status writes occur. - Permit a resumed workspace to create a new session directory only with matching durable identity, zero connections and events, untouched bootstrap commands, no backup, and an absent remote session. Claim the directory atomically. Existing, partial, or ambiguous state still blocks startup. ## Verification - Final head `c43c7403f`: [CI completed successfully](https://github.com/paperclipai/paperclip/actions/runs/34733846820/attempts/2), with 32 successful checks and 2 conditional skips. This includes every server, UI, package, serialized-route, browser, native-runner, typecheck, and build gate. Greptile reviewed the same head at [5/5 with no unresolved findings](https://github.com/paperclipai/paperclip/pull/13353#issuecomment-5650270168). - New regressions failed before each of the four production fixes. Git/archive/restore suites: 148 passed. The scheduler-wrapped non-Git regression also failed before its fix; 135 affected Git/sync/Codex tests then passed. - Daytona plugin: 230 passed, 6 opt-in live tests skipped. Native executor, projection, and TaskChatThread: 494 passed, including existing/partial state, wrong identity, prior connections or turns, backups, unavailable proof, and unresolved recovery controls. - Local full-repository typecheck, build, and token gates passed on the final head. Local CLI: 485 passed. Complete single-worker package rerun: 3,224 passed, 19 skipped. - Local verification is an aggregate with recorded retries, not one pristine green invocation: the general-server run began on `e721a920a` and finished with 12,002 passed, 3 failed, 70 skipped. Its real Codex scheduler failure is fixed above; the socket and workspace-runtime timeout failures passed unchanged in focused reruns. Both Inbox failures passed unchanged in the full 27-test Inbox file; database/shared-package failures passed in the single-worker package rerun. The supplemental local serialized-route rerun remains in progress; all five corresponding final-head CI lanes passed. - The first final-head CI attempt hit a Daytona fixture-readiness race and a signoff-browser heartbeat receipt timeout. One supported unchanged failed-job rerun passed both and the aggregate gates. Live combined user-journey verification is tracked in the related follow-up; this PR's provider tests use deterministic fixtures and real filesystem checks. ## Risks - A selected subfolder uses directory sync and does not carry parent Git history. Its ignored files stay local. - The target region remains a creation setting and part of workspace reuse policy; it does not split the account identity of an existing sandbox. - Incomplete or conflicting provider state still fails closed. This change does not erase a session, infer completed work, bypass a user decision, or replay uncertain actions. - A later remote setup failure can leave a claimed partial session directory. It remains blocked rather than being overwritten. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact hosted model ID and context-window size are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8d1f0c20af |
fix: let responsible users choose either AI subscription or API key (#13351)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - AI Connections select the account used for each run. > - A responsible-user binding must follow the person whose work the agent performs. > - The saved sign-in method currently blocks users with another method for the same provider. > - This pull request resolves a personal default by company, user, and provider. > - Each user can use a subscription or API key with the same bot and model. ## Linked Issues or Issue Description Refs #13247, #13248, #13346, #13347. **What happened?** A bot configured with a Claude subscription rejects another responsible user’s Claude API key. Inline repair also limits that person to the original sign-in method. **Expected behavior** The same bot uses each responsible user’s default Claude account, whether it is a subscription or API key. Explicit shared account selections remain fixed. **Steps to reproduce** 1. User A connects a Claude subscription and creates a bot using the responsible user’s connection. 2. User C connects a personal Claude API key. 3. User C runs the same bot. Before this fix, credential resolution fails. ## What Changed - Add a personal provider-default table. Preserve legacy per-method preferences and backfill the most recently updated preference, including unavailable defaults. Repeated migration does not replace a selection. A database trigger propagates old-server default updates without treating new accounts as replacement defaults. - Resolve responsible-user bindings by provider. Retain the method as a wire compatibility hint for old servers. Explicit selections still require the exact method and grant. - Use the selected account’s method for credential isolation, refresh locking, and run attribution. - Update onboarding, agent setup, the picker, and inline task repair. Keep existing authentication components and harness/model settings. - Add mixed-method runtime, migration, repair, and Storybook coverage. Include upstream’s duplicate Anthropic option fix through the base branch. ## Verification - Focused resolver, migration, connection-intent, onboarding, agent setup, model, and connector UI suites: 331 tests passed. - Onboarding and new-agent regression suites passed during the initial focused run. - UI typecheck, token gates, and Storybook build passed. - Live browser checks passed for Claude and Codex API-default execution, switching both back to subscriptions, and both existing shared-account bots. Bot configuration remained unchanged. - One Daytona startup command stalled before Claude launched. The test run was cancelled, its sandbox stopped, and the same account/task passed on retry. Startup cancellation remains a separate environment finding; this PR does not change that command transport. - Browser review: all eight assertions passed in the new mixed-method story, including shared selection, return to responsible-user selection, and unchanged harness/model. - Repository build and typecheck passed after refreshing upstream dependencies. Final resolver and historical rollback verification: 38 tests passed. All latest-head CI gates passed, including server/workspace/serialized suites, browser E2E, build, typecheck, and runner verification. The extra serial local full-suite run was stopped after equivalent CI passed; focused local checks completed. ## Risks - Users with both historical method defaults get their most recently updated preference as the initial provider default. They can change it explicitly in Connections. - A revoked or unavailable default blocks. Connecting an additional account does not silently replace it. - Existing legacy authentication is unchanged. Managed responsible-user bindings intentionally stop pinning a method. - Live staging: the same Claude and Codex bots completed real API-key runs after changing only the personal default, then completed subscription runs after restoring the original defaults. Read-only database verification confirms unchanged bot configuration and actual method attribution. Distinct-user concurrency is covered by automated real-database tests with synthetic credentials, not two live human logins. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, code execution, and browser interaction. The runtime does not expose a more specific model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.913.0-canary.1 |
||
|
|
422287eecd |
fix: preserve runner recovery, warm sessions, and task outcomes (#13338)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner connects task messages, provider execution, and task outcomes. > - First-time user tests exposed gaps in recovery, completion permissions, message delivery, and Stop behavior. > - These gaps left usable output hidden, completed work waiting for bookkeeping, or safe work unable to continue. > - This pull request fixes the shared lifecycle and receipt paths while preserving process ownership and action checks. > - Users can continue work with accurate task state and durable messages. ## Linked Issues or Issue Description **What happened?** A stopped local Codex execution could remain blocked even after its processes had stopped and its complete transcript proved that no external action needed replay. Claude under Conservative permissions could fail to call task completion tools. Recovery could reuse an assistant item ID and overwrite prior output. A delivered comment could remain marked uncertain after navigation. Stop could look like Pause or a new recovery incident. Workspace contention could look like cancellation. A direct reply reopening Done could enter a clarification loop. **Expected behavior** Recover automatically only with verified termination and complete action receipts. Preserve answers and messages. Keep task completion available under Conservative permissions without broad tool access. Show crashes as Blocked, actual human decisions as In Review, and ordinary workspace contention as waiting. Stop the current response and allow a new direction. **Steps to reproduce** 1. Create ordinary response tasks with local Codex and Claude Code, then send follow-up messages through the task composer. 2. Interrupt a disposable local Codex runner during text-only work. Verify automatic continuation and retained output. 3. Stop a response, send a new request, answer a clarification, and reopen completed work with another message. 4. Navigate or reload while a comment submission is pending. Confirm the exact persisted request receipt settles it without removing newer draft text. 5. Run two tasks in a shared Daytona workspace. Confirm waiting does not appear as failure. **Paperclip version or commit** Initial acceptance baseline: `c9021c6721f91e2c74bd9fee9d3fd41c999d17b7`. Current integration base: `6cef9743c`. Both operator-interruption and workspace-waiting guards are preserved; native restart and legacy permission rules remain documented. **Deployment mode** An isolated source-built test-drive instance, with real local Codex and Claude Code providers and disposable Daytona environments. Related work: #13314, #13316, #13327, #13344, #13239, #13254, #13163. This PR addresses additional failures from ordinary task journeys, including controller restart handoff and repeated warm sandbox setup. Historical task status reconciliation is excluded. ## What Changed - Persist runner ownership immediately at spawn and resume an explicitly adopted runner even when the controller crashed before the first driver checkpoint. Detach the controller safely across graceful restarts, including session startup. Prevent an old finalizer from suspending or signaling an adopted runner. Checkpoint idle warm sessions before shutdown. Preserve the same run and queued follow-up messages. - Scope saved legacy queue successor checks to the queue owner while preserving ordinary task locks, operator identity, assignment gates, and exactly-once delivery. - Preserve managed Codex credential files when an old session is detached for restart; normal owned cleanup still copies refreshed auth back and removes the scoped copy. - Reuse the bound warm shared sandbox and fully verify an existing staged provider pack before using it. This avoids repeated uploads when the pack is already valid. - Add a narrow local Codex replacement path with stopped-process proof, a closed transcript inventory, exact completion receipts, and fresh-session lineage. Preserve no-replay holds when evidence is incomplete. Recovery may clear only the same run's recorded Blocked status version; manual re-blocking and dependency changes invalidate that receipt, while queued comments do not. Later blocks stop scheduled, queued, and final dispatch; queued/final checks re-read dependencies even when the task status stays In Progress. - Permit only task delivery and human-input tools through the isolated Claude runner's exact task bridge. - Scope assistant item identity to the provider turn and ignore only authority-free Codex skill-change notifications during startup. - Reconcile composer submissions by client request ID across response loss, navigation, and reload. Retain text typed during delivery. - Keep acknowledged run-only Stop neutral and show workspace contention as waiting. Project exhausted native failures as Blocked. - Restore the guarded task-page retry action for failed legacy runs, including the server-supported explicit new-attempt path for stopped conversation adapters. Preserve native/process recovery holds and avoid promising Retry while a decision or execution gate hides it. - Refresh delivered artifacts and handle direct user replies that reopen completed work without a clarification loop. - Check the embedded PostgreSQL PID, data directory, and actual port before connecting or migrating. - Document accepted behavior and add focused regressions at lifecycle, route, transcript, and UI boundaries. ## Verification - Final head `fece606ac2` passes the complete GitHub CI matrix: **34 green checks, two expected Storybook skips, no failures or pending checks**, including `ci / verify`, `ci / e2e`, full runner verification, typecheck, build, every server/workspace shard, and all browser shards. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34727183287). Greptile is **5/5 with no open findings**. The final two commits only refine test fixtures; both affected suites pass 24/24 locally and in CI, with server typecheck green. - Complete local Vitest coverage uses the canonical groups/shards: all 635 general server suites, all 145 serialized suites, and all workspace packages. The aggregate began on `0a8001c18` while the final queue fix arrived: 23,903 passed, five failed, 87 skipped. The five port/socket/timing failures passed unchanged in follow-ups (60 tests in the exposure/file suites and 412 tests covering the serialized failures and unrun tails). The final queue/operator-identity suites separately passed 52/52. This is aggregate coverage plus explicit reruns, not a pristine single-command final-head run. - After integration with current master, queue/operator-identity/continuation suites passed 162/162 and affected UI suites passed 140/140. ACP Stop/continuation and legacy task/Inbox/message browser suites passed 9/9, including both task recovery Retry and thread Try again, automatic saved-message delivery, exactly one new run, Done, and retained output after reload. The default process Stop/Pause/Resume browser case passed (the native-provider case is opt-in and skipped by default). The complete Board attachment/receipt browser suite passed 11/11 on a disposable instance, covering both composers, exact receipts after lost responses, no replay, bound attachments, and newer drafts after reload. - Blocking-intent regressions cover pre-existing Blocked, a mismatched run/cause, an explicit manual re-block, changed dependencies, a queued comment after failure, and a block arriving between scheduling and provider dispatch. The negative cases reproduced before the fix. All 478 affected executor/recovery/dispatch tests passed; both database suites ran separately after availability-probe skips in the first combined command. The final late-dependency check passed all 143 affected recovery/dispatch tests (zero skips) after two new negative cases reproduced the bug. - Focused runtime regressions cover awaited runner ownership publication, authenticated adoption before the first checkpoint, old-finalizer detachment, idle and busy warm-session shutdown, rejected checkpoint propagation, provider-pack verification, and managed-Codex credential preservation. Four managed credential detachment cases reproduced the bug before the fix; normal owned cleanup still succeeds exactly once. - Live local Claude: SIGKILL 2.6 seconds into startup recovered the same run automatically in 53 seconds, then a normal follow-up completed in 24 seconds. SIGTERM 2.5 seconds into startup preserved the same run (54 seconds) and its queued follow-up (21 seconds). Answers remained visible and the task reached Done. - Live Claude Daytona: a warm follow-up retained its sandbox and fell from 121 seconds to 44 seconds. A separate cold turn took 127 seconds; after controller shutdown and checkpointing, its follow-up completed in 33 seconds with the same sandbox, workspace, native session, and runner. Both answers remained visible and the task was Done. - Other live journeys covered task completion and follow-up with local and Daytona Codex, local Codex crash recovery, Stop then new direction, clarification response, live artifact refresh, and shared-workspace waiting. - Validation limits: the opt-in native composer Stop/Pause→subtree Resume fixture exposes terminal/result ordering and subtree-cancellation attribution bugs that can leave a child task blocked; that new finding is assigned to a separate follow-up and is not claimed fixed here. Default CI skips this optional native-provider fixture. Managed-Codex credential handoff and the queue-agent integration use automated regression evidence. Cold custom provider-pack uploads still add startup latency. ## Risks - Automatic replacement remains deliberately narrow: local Codex, verified stopped identities, unchanged retained state, and a complete text/completion-only turn. Unknown actions, partial history, or changed ownership remain blocked. - Claude completion permission handling changes an upstream package patch. The exact isolated task bridge must remain pinned; unrelated tools keep their existing permissions. - New task failure projection changes user-visible status. No historical status backfill or database migration is included. - This is a broad lifecycle fix across server and UI. Live proof covers graceful local Claude restart during startup and idle Claude Daytona session recovery across controller shutdown. Live abrupt SIGKILL during local Claude startup also recovered the same run. Unknown ownership or missing action evidence still blocks reuse. Cold custom provider-pack uploads still add startup latency; this change avoids unnecessary repeat uploads. ## Model Used OpenAI GPT-6 (Codex), with reasoning, code execution, browser automation, and tool use. The exact hosted model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
13bae6fa21 |
fix: reject unsupported REST tool connections without stdio validation (#13346)
## Thinking Path > - Paperclip manages AI agents and their connections. > - Connection checks must use the configured transport. > - The tool service treated every remaining transport as local stdio. > - Anthropic's old REST method therefore failed with a templateId error. A REST connection with a valid stdio template could incorrectly pass. > - Anthropic now has a supported AI-account flow. This pull request removes its obsolete REST setup option and limits stdio checks to stdio connections. > - Users can connect an AI account, and existing unsupported connections receive an accurate error. ## Linked Issues or Issue Description Related: #13248 added the supported AI-account flow. Searches for related REST health and templateId bugs found no duplicate fix. **What happened?** The Anthropic REST API-key connection showed `Local stdio MCP connections must use an approved templateId`. Health checks and catalog discovery both fell through to the local stdio path. A REST connection with an approved template could report success and expose the template's catalog without a REST integration. **Expected behavior** Only local stdio connections use command templates. Unsupported transports return an accurate HTTP 422 error. New Anthropic accounts use the supported runtime authentication flow. **Steps to reproduce** 1. Check out the test-only commit `924e6e85a` in a separate worktree and install dependencies. 2. Run `pnpm exec vitest run packages/shared/src/app-definitions.test.ts server/src/__tests__/tool-access-service.test.ts -t 'unsupported REST|obsolete Anthropic'`. 3. The tests exercise saved Anthropic REST configuration and an unsupported REST connection containing an approved stdio template. They cover health checks and catalog discovery separately. 4. Run the same tests on the fix commit. They pass. The full affected files also pass. **Paperclip version or commit** Reproduced against master `6cef9743c`. **Deployment mode** Server transport handling. Reproduced with an isolated embedded PostgreSQL test database. No provider account or live credentials are required. ## What Changed - Restrict stdio health checks and tool discovery to `local_stdio`. - Return and audit `tool_connection_transport_unsupported` with HTTP 422 for unsupported tool transports. - Remove Anthropic's obsolete REST method from the generated catalog and its durable ingestion source. Keep its subscription and API-key AI methods. - Cover the reported error, false-success case, rejected obsolete setup, connection removal, and the UI's AI-account submission path. - Replace impossible reconnect forms for removed methods with supported setup, while preserving connection removal. - Preserve AI-versus-tool intent isolation for legacy requests and reject new unsupported Anthropic tool requests. - Document recovery for existing unsupported connections. ## Verification - Clean-worktree red/green: the same command failed all six regression cases at `924e6e85a` and passed all six at `4d3de9de0`. The failing run includes the reported templateId error. - Green: all 555 tests across the six affected test files passed. - Recovery UI red/green: three added cases failed before the recovery fix and passed afterward; all 200 tests across setup, detail, and advanced controls passed. - After the recovery UI update, UI typecheck/build and token gates passed again. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm check:token-gates` — passed. - Catalog regeneration — passed with the documented `PAPERCLIP_CONTENT_TEMPLATES` override for the local capture corpus. - Full CI on `69fb31fd4` — passed all general and serialized test shards, browser shards, typecheck, build, runner verification, and canary dry run: https://github.com/paperclipai/paperclip/actions/runs/34726975425. - The local serial `pnpm test:run` was stopped after the fixture correction superseded that run; full-suite verification above comes from CI. All 555 affected tests passed locally, including all 17 connection-intent tests after the correction. - Greptile — 5/5, successful check on final commit `69fb31fd4`, no unresolved findings. - No live Anthropic validation was performed. The UI regression uses a fake key and a mocked AI-account response. ## Risks Existing obsolete REST connections remain in needs-attention state. Users must add an account through the supported flow and remove the old connection. Credentials and grants are not transferred automatically. Removal remains covered. The specialized AgentMail and Composio paths keep their existing behavior. There are no schema or permission changes. ## Model Used OpenAI GPT-6 through Codex. The exact serving model ID and context-window capacity are not exposed in this session. Used reasoning, code editing, shell tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e704e1c9ae |
fix: keep AI subscription locks safe through transaction pooling (#13347)
## Thinking Path > - Paperclip manages AI agents and their credentials. > - AI Connections must preserve account ownership from login through execution. > - A subscription environment check can pass and leave a session advisory lock on a pooled database backend. > - The next execution then reports that the subscription is busy. > - Onboarding can also lose its completion callback when the connection list refreshes before the login poll. > - This change keeps each credential lease in one transaction and keeps the login controller mounted until completion. ## Linked Issues or Issue Description Refs: #13247, #13248, #13344. **What happened?** A real hosted subscription login saved successfully and passed its environment check. The first task then failed with an AI connection busy error. A connection-list update could also leave onboarding on Connecting. **Expected behavior** A completed environment check releases its credential lease. A saved login completes onboarding without another sign-in. **Steps to reproduce** 1. Run Paperclip with a transaction-pooling PostgreSQL proxy. 2. Connect a Claude subscription through onboarding and reuse it for an agent. 3. Start a task after its environment check succeeds. 4. For the UI race, refresh the managed connection list before the login completion poll. ## What Changed - Hold the grant lock inside one transaction. Roll back on cleanup or failed acquisition. - Disable the lease transaction's idle timeout so long provider executions retain the lock. - Keep the onboarding login controller mounted while its authorization URL is active; restore saved-account retry if later agent creation fails. - Add lock-lifetime and Claude/Codex completion-order regressions. - Document the pooling requirement. ## Verification - 104 focused connection and onboarding tests passed. - Workspace typecheck, production build, Storybook build, and token gates passed. - All 34 latest-head checks are terminal: 32 passed and two Storybook workflows skipped by their normal filters. CI covers all server/workspace/serialized test shards, browser tests, typecheck, build, runner verification, and canary packaging. The additional local full-suite run is still in progress. - Real browser sign-ins saved Claude and OpenAI subscriptions on both new and existing staging stacks. On the patched new stack, both providers completed tasks and immediate repeat executions with the same saved accounts. Read-only database inspection confirmed live transaction-scoped locks and no remaining locks after completion. On the final deployed head, both shared accounts on the existing stack also completed actual tasks. Both personal accounts passed a fresh environment check followed immediately by execution after a deployment restart. ## Risks - Each active subscription reserves one database connection and holds an otherwise idle transaction until cleanup. This is required to pin the backend through a transaction pool. - Old session locks from earlier versions may need operator cleanup after active runs drain. This change does not bypass existing locks. - No database migration, credential routing, or legacy authentication change. ## Model Used OpenAI GPT-6 in Codex, with code execution and browser tools. The exact deployment model ID and context-window size are not exposed to this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.913.0-canary.0 |
||
|
|
6cef9743c0 |
fix: deliver saved user messages after recovery stops (#13327)
Deliver saved user messages after legacy recovery stops. Validate undelivered comments and queue ownership under the task lock, preserve the operator identity checks from #13315, and prevent duplicate successors. Add a recovery notice with Retry and inline errors, plus service, route, component, and browser coverage. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b2acc674be |
fix: enable isolated subscription login on authenticated self-hosted instances (#13344)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - AI Connections store reusable provider credentials with company and owner boundaries. > - Self-hosted instances can require board authentication while running agents on the local server. > - The login UI treated those instances as unsupported and displayed instructions without a command. > - Removing that restriction must not expose the server operator's existing CLI account. > - This pull request enables owner-scoped login attempts and reuses the existing login UI. > - Users can connect Codex and Claude subscriptions on an authenticated self-hosted instance. ## Linked Issues or Issue Description **What happened?** On an authenticated self-hosted instance, OpenAI subscription setup displayed terminal instructions with no command and a disabled Connect button. Claude also could not complete the local connection flow. **Expected behavior** An authorized board user can prepare an isolated sign-in attempt, sign in on the server, and save the verified account as a Connection. This must not import another user's or the operator's ambient credentials. **Steps to reproduce** Run Paperclip in authenticated mode with a local environment. Open Connections, choose OpenAI or Anthropic, and select Subscription. The previous UI never enabled local login preparation. Related: #13247, #13248, and #10751. This fix preserves the restriction on remote access to the operator's ambient Claude login. ## What Changed - Allow company-authorized users to create, check, cancel, and complete their own isolated local login attempts. - Keep ambient Claude credential import restricted to the local operator. - Support isolated Claude credential files without falling back to the host account or mutating process-wide environment variables. - Use Codex device authorization so sign-in does not depend on a browser callback to the remote server's localhost. - Gate server-host login on authenticated public deployments unless a trusted runtime host is configured. Publish the capability through health so setup shows supported alternatives. - Read isolated Claude credential files through bounded, descriptor-bound opens with ownership, permission, and symlink checks. Try the alternate filename after malformed JSON. - Reuse shared login instructions and lifecycle hooks in onboarding, agent setup, and Connections. Show health-query failures explicitly. - Document authenticated self-hosted behavior and add authorization, isolation, lifecycle, and UI regression tests. ## Verification - Passed 71 focused tests across connection routes, credential isolation, legacy compatibility, the shared login hook, and agent setup. - Passed 89 onboarding regression tests. - Passed `pnpm -r typecheck`, `pnpm build`, Storybook build, and `pnpm check:token-gates`. - Completed real Codex device authorization and Claude browser authorization on an authenticated Linux self-hosted instance. Both accounts were detected automatically and saved as Connected. Both completed attempt directories were removed. - These live checks cover login, credential validation, and connection creation. They do not establish a new model execution or long-running refresh result. - Review follow-up: 68 focused checks passed after rerunning one route socket error; the full route/health rerun passed all 49 tests. The 70-test onboarding suite also passed. Final workspace typecheck, production build, and Storybook build passed again. - The broad local run exposed an instance-name assumption in two new assertions. The fixture now uses an explicit non-default instance, and all 32 connection tests passed with a different inherited instance name. The superseded broad run was stopped; this is not a claim that the full local suite completed. Full CI results will be recorded before merge. ## Risks - The server must have the provider CLI installed. Users still run the displayed command on the server that hosts Paperclip. - Authorization checks must keep login attempts scoped to the company, owner, provider, and reconnect target. Regression tests cover cross-user and cross-company access. - Existing local-trusted Claude behavior stays available. Authenticated remote users cannot use its ambient import path. - No database migration, dependency change, agent binding change, or provider routing change is included. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, and browser tools. The runtime does not expose a more specific model version or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.912.0-canary.12 |
||
|
|
8b84ae5b35 |
fix(ui): improve mobile entity picker sheets (#13343)
Co-Authored-By: Paperclip CodexRunner <codexrunner@paperclip.local> Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
df984cbc2c |
fix: dispatch queued legacy messages with operator identity and task permissions (#13315)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task conversations save messages that arrive during an active turn. > - Legacy adapters deliver these messages in a later turn. > - A run can stop before the saved queue is delivered. > - The Interrupt button previously required an active run, so it could not release this queue. > - This pull request lets a board operator send the saved queue after the run stops and retries queues missed during finalization. > - Manual dispatch must use the clicking operator and must not require permission to create agents. > - The task can continue without a duplicate message or a second execution owner. ## Linked Issues or Issue Description **What happened?** A legacy task retained a queued message after its run stopped. Interrupt was disabled because the queue had no active target. Finalization and deferred message admission can also leave a queue without a successor. **Expected behavior** Interrupt sends the saved messages when no runner is active. Messages that arrive during normal completion are delivered automatically. An uncertain previous execution still requires proof that its process or sandbox stopped. **Steps to reproduce** 1. Queue a user message during a legacy conversation turn. 2. Let the turn stop or simulate a server restart before queue promotion. 3. Open the task with a deferred queue and no active run. 4. Try Interrupt. Before this change, the button is disabled. **Paperclip version or commit** Reproduced against `8f40b4ad4`. **Deployment mode** Legacy conversation adapter. The same persisted queue state is covered with an isolated PostgreSQL fixture. Related public work: #13275 adds active legacy interruption. #13291 addresses automatic sandbox conversation recovery. This change handles explicit saved-queue delivery and late queue promotion. ## What Changed - Accept a null Interrupt target while retaining queue identity, revision, company, and assignee checks. - Save the operator's request on the existing queue. Reuse normal admission after verified stop, including older messages, different authors, and queues whose original wake came from the system. - Strip interruption authority from caller-supplied wake payloads. Only the board queue route can persist that authority. - Retry durable interruption requests after restart and deferred queues after legacy cleanup. - Let an explicit Interrupt retry cleanup for its stopped run, including old ephemeral leases that recorded success without a provider stop receipt. Preserve retained resources, other lease owners, and the automatic retry limit. - Preserve the server's waiting explanation when normalizing and combining queue entries. - Revalidate the consumed board queue receipt at dispatch so a different message author does not cause setup failure. - Use the Interrupt user's execution identity for the new run. Preserve original message authors. Validate the receipt independently at startup and inherit the resulting identity on retry. - Persist authenticated board authority for ordinary manual wakes too. Adopting someone else's queued messages cannot switch a manual run to that author's permissions. Strip caller-supplied authority markers and retain private conversation ownership checks. - Keep the clicking user when a manual wake is merged into an older deferred receipt. Update its requester and payload in the same transaction. - Use the same current-queue/revision API on task details and pipeline conversations; show Interrupt after a legacy target stops. - Keep manual wakes out of active runs, including unscoped agent wakes. They receive their own execution identity; a matching receipt requester is not sufficient because an exact retry can retain a different originating identity. - Authorize both existing-agent wake endpoints with `agent:wake`, available to active non-viewer company members. Keep `agents:create` for hiring. Validate the stored task and current assignee before an exact task retry. - Reject viewer Interrupt requests before saving intent or stopping execution. Keep external chat retry authorization and per-action agent/user permission checks. - Preserve edits and discards until dispatch. Prevent another queue promotion when the same agent already has a successor. Keep independent reviewer recovery available. - Suppress cancelled/failed run toasts for intentional operator interruption. Keep ordinary runtime error notices. - Add UI, route, admission, restart, successor ownership, and toast regression tests. Document the behavior. - Reuse the existing socket reservation helper for both credential-quorum test cases after CI exposed an ambient-port collision. This changes test preparation only; production credential staging is still called exactly once. ## Verification - Failing regression tests reproduced the message-author identity bug and an operator's `agents:create` rejection before the fixes. - All 316 focused tests pass across eight route, queue, identity, authorization, continuation, and responsible-user suites, including the 44-test rerun of queue admission and actual startup after the final manual-wake restriction. Regressions reproduce cross-user merging both with and without a task, and same-requester receipt ambiguity. The cross-company existence guard also passes both tests. - Startup integration tests reach adapter execution under the clicking operator and retain that identity through follow-up. Coverage includes mixed authors, adopted queues, system-origin queues, restarts, forged or stale receipts, viewers, suspended memberships, changed assignees, private conversations, and caller-supplied authority markers. - The earlier queue/cleanup/UI regression suite passed 402 tests. The final review corrections pass another 180 tests across queue admission/persistence, real heartbeat startup, UI API, conversation rendering, and pipeline suites. Regression tests reproduced both review findings before correction. The final head has a 5/5 review with no unresolved threads. Full CI passes on `c2002979c`, including every general and serialized server shard, all browser shards, Paperclip Runner verification, typecheck, build, canary dry run, and the aggregate gates. - Full `pnpm -r typecheck`, `pnpm build`, and UI token gates pass after the final application changes. CI identified an outdated task-page API mock after the shared helper extraction; the fixture now exercises the real helper, and all 131 task-page/API tests pass. The final application build passes with the additional manual-wake restriction. - CI exposed a pre-existing port collision in the Codex credential-quorum fixture. It reproduced locally; both listener cases now use the existing bounded reservation helper. All 41 credential tests pass on rerun. One intervening local run hit a separate ambient bind collision in the two-occupied-port case. - The full local `pnpm test:run` attempt was stopped after host contention caused focused-suite timeouts. The affected focused tests passed on rerun. An expiring trace fixture and a missing private-conversation state were corrected. The successful full CI run is the complete-suite verification. - Hosted Interrupt previously cleared the original queue and produced exactly one successor with neutral interruption feedback. It exposed the dispatch authorization defect. Retry on that earlier build was rejected for missing `agents:create` before creating another run. - Deployed the final application build (`38257f391`) to the scoped hosted instance and verified readiness. The latest PR commit changes only the credential test fixture; application code matches that deployment. A live Retry by the same operator without `agents:create` created one successor attributed to that operator, passing the former dispatch permission gate. Startup then stopped at `configuration_incomplete` because that operator has not configured their required personal Claude Code OAuth secret; the post-deployment run page confirms the operator identity and no provider work started, and the My secrets UI still shows the token as not set. Provider execution remains unverified pending that credential. No permission grants or credentials were changed. ## Risks Queue admission and finalization can race. The task lock, durable queue receipt, current comment IDs, and successor guard prevent duplicate dispatch. Process and lease stop checks, task pauses, approvals, ownership, and budgets remain in force. The API change only allows null on legacy Interrupt; native steering still requires an active run. No schema migration is required. Active non-viewer board members can now invoke existing agents without agent-creation permission. Agent self-invocation rules, raw provider-trace admin access, task retry scope, external chat authorization, and action-specific user/agent permissions remain enforced. ## Model Used OpenAI GPT-6 through Codex. The session does not expose an exact backend model ID or context-window size. Used reasoning, repository search, code execution, tests, and browser tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8d6232e7b0 |
feat: reuse provider sign-in across AI connection workflows (#13248)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users connect provider accounts during onboarding and agent setup. > - They should reuse and manage those accounts through the existing Connectors interface. > - A second login wizard would diverge from the established provider workflows. > - This pull request composes the existing sign-in components into Connections and agent configuration. > - Users can select accounts without changing their agent's harness or model. ## Linked Issues or Issue Description **Problem or motivation** AI credentials are configured separately from Connections. Agents cannot consistently reuse a responsible user's account or a permitted shared account. **Proposed solution** Manage AI accounts with the existing Connections grants and permissions. Keep model and harness selection independent from credential selection. Preserve legacy authentication until validated adoption. **Alternatives considered** A separate credential registry would duplicate ownership and access policy. Automatic fallback would risk using the wrong account. **Roadmap alignment** This extends the shipped Apps, multi-user, secrets, and agent-runtime capabilities. The maintainer requested the feature and reviewed the UI. Related groundwork: #11899 (connection permissions), #10910 (connection wizard), #11692 (Claude subscription profiles), and #11854 (Codex account rotation). ## What Changed - Add compact AI-account management to the existing Connectors pages. - Reuse AgentProviderConnection, AdapterLoginPanel, AdapterLoginChrome, and authentication controllers. - Add the shared connection picker to agent setup/settings and task requests. - Preserve onboarding's sequence and reuse existing accounts. - Add local-login recovery, retry, cancellation, and React StrictMode handling. - Add interactive Storybook scenarios, design-guide examples, and app acceptance checks. This is part 2 of the AI Connections change. The runtime foundation in #13247 is merged. This PR now targets master. ## Verification - Updated against master `47ded8bf9`, including the landed runtime foundation and upstream task-search changes. - Full workspace typecheck, production build, Storybook build, and token gates passed on the integrated branch. Final local-login changes passed 59 focused tests; new-agent and inbox regression suites passed 63 tests. - Browser checks verified automatic local Claude account detection, resumable Codex login commands, retry, focus restoration, and desktop/phone layouts. Commands create their isolated directory before invoking the CLI. - All CI test, browser, build, packaging, and runner jobs passed on final head `dd17d3211931dd70aaa6ea619d83a7f9966dd18e`. The fresh Greptile review is 5/5, the security scan passed, and there are no unresolved review threads. The final CI aggregate gates passed. - Local general-server coverage passed 11,804 tests; three port-collision failures passed in an isolated 25-test rerun. All 6,111 UI tests passed. CLI coverage passed 484 tests; its remaining doctor test requires port 3199, which is occupied by an unrelated report server on this Mac. The complete CLI suite passed in CI. - Live browser testing verified Codex API-key reconnect inside a task card on desktop and phone. Real provider runs resumed and completed with unchanged connection/grant identity and agent routing. - Tested opening, cancelling, reopening, and completing connection creation. A regression confirms Connect another account cannot submit the new-agent form or copy provider keys into agent settings. - Added shared inline repair, automatic local sign-in checks, and responsive connection dialogs. Standalone Daytona installation ignores workspace configuration and suppresses dependency scripts. Its standalone build also passed with CI's exact pnpm 9.15.4. - Destructive live tests are excluded by default. Explicit opt-in, local deployment checks, and matching disposable fixture identities are required before any mutation. ## Risks - Local Codex/Grok creation requires the connection-specific terminal login command. - Browser sign-in uses the existing supported-environment controllers. - This update verifies live local Claude detection and Codex API-key task repair. New subscription authorization/refresh and independent-human/native-runner isolation were not reverified in this update. - No agent automatically adopts managed Connections. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
47ded8bf97 |
feat: manage AI runtime credentials through Connections (#13247)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs need credentials for a specific provider and sign-in method. > - Connections already owns accounts, grants, and access permissions. > - AI authentication should use those same boundaries. > - This pull request adds the storage, API, adoption, and runtime foundation. > - Legacy agents keep their authentication until they explicitly adopt a managed connection. ## Linked Issues or Issue Description **Problem or motivation** AI credentials are configured separately from Connections. Agents cannot consistently reuse a responsible user's account or a permitted shared account. **Proposed solution** Manage AI accounts with the existing Connections grants and permissions. Keep model and harness selection independent from credential selection. Preserve legacy authentication until validated adoption. **Alternatives considered** A separate credential registry would duplicate ownership and access policy. Automatic fallback would risk using the wrong account. **Roadmap alignment** This extends the shipped Apps, multi-user, secrets, and agent-runtime capabilities. The maintainer requested the feature and reviewed the UI. Related groundwork: #11899 (connection permissions), #10910 (connection wizard), #11692 (Claude subscription profiles), and #11854 (Codex account rotation). ## What Changed - Add AI-purpose/runtime-auth contracts and an additive, idempotent migration. - Add Claude, OpenAI, OpenRouter, and Grok provider capabilities and catalog entries. - Store credentials on grants. Resolve responsible-user defaults or explicit permitted grants. - Isolate managed credentials and provider sessions across accounts. Block missing credentials without ambient fallback. - Keep imported legacy secrets unchanged during reconnect. Use independent local Codex/Grok sign-in attempts for rotating credentials. - Add authorization, migration, concurrent refresh, retry, cancellation, and legacy-compatibility tests. This is part 1 of a two-PR stack. The app UI follows in #13248. Merge the foundation first. ## Verification - Updated against master `04e364236`, preserving upstream provider login and connector workflows. - Full workspace typecheck, production build, Storybook build, and token gates passed on the integrated branch. Final local-login changes passed 59 focused tests; new-agent and inbox regression suites passed 63 tests. - Browser checks verified automatic local Claude account detection, resumable Codex login commands, retry, focus restoration, and desktop/phone layouts. Commands create their isolated directory before invoking the CLI. - All current-head CI checks passed on `2a996560a`, including all server/workspace tests, browser shards, runner verification, typecheck, build, and canary dry run. Greptile reviewed that commit at 5/5 with no unresolved threads. Earlier local full-suite attempts hit the Mac PostgreSQL shared-memory limit; the complete suites passed in CI. - Renumbered the additive AI migration to `0276` after upstream migrations and regenerated its snapshot. Existing legacy agents retain their configuration. - Added local login status checks, owner-scoped retry, managed OpenCode remote homes, credential-aware model discovery, and task connection-repair delivery. ## Risks - Managed credential failures intentionally block execution. They do not restore legacy fallback. - Preview-era copied Codex/Grok subscriptions require independent reconnect. - The integrated branch has live provider acceptance coverage. This update verifies local Claude detection and Codex API-key task repair; it does not add a new subscription authorization/refresh or Daytona stress pass. - Runtime-auth connections must stay excluded from tool and channel handling. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.912.0-canary.11 |
||
|
|
1652a5c7f2 |
fix: improve task search relevance with a PostgreSQL rubric (#13335)
Unify full and quick task search around PostgreSQL term coverage, explicit relevance bands, and conservative typo recovery. Preserve matching evidence and navigation, and add a judged corpus, regression tests, and documented performance measurements. Validation: local typecheck, build, focused PostgreSQL tests, and task-list tests pass. All final-head CI gates pass and Greptile is 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
04e364236b |
chore(lockfile): refresh pnpm-lock.yaml (#13336)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com> |
||
|
|
6b45db0c35 |
fix(deps): force one @codemirror/state resolution via pnpm overrides (#13324)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The web UI embeds CodeMirror editors, and CodeMirror validates
extensions with `instanceof`
> - Different CodeMirror packages pin different transitive minors of
`@codemirror/state` (6.7.1 / 6.7.2) and `@codemirror/view` (6.43.9 /
6.43.11), so the bundle ships two module instances
> - The second instance makes valid extensions fail the `instanceof`
check and crashes the editor
> - This pull request adds a shared caret override to `pnpm.overrides`,
the same mechanism the existing `react` and `rollup` overrides use, so
every consumer resolves one copy of each package
> - The lockfile is not edited by hand; the lockfile automation
regenerates it from the manifests
> - The benefit is that the editor stops crashing with "Unrecognized
extension value in extension set"
## Linked Issues or Issue Description
**What happened?**
The production UI throws `Error: Unrecognized extension value in
extension set ([object Object]). This sometimes happens because multiple
instances of @codemirror/state are loaded, breaking instanceof checks.`
Observed 43 times in one week.
**Expected behavior**
The editor loads its extension set without errors. One instance of
`@codemirror/state` and `@codemirror/view` serves every CodeMirror
package.
**Steps to reproduce**
1. Run `grep "'@codemirror/state@" pnpm-lock.yaml` on master: two
versions resolve (6.7.1 and 6.7.2).
2. Build `ui/` and search the output for `Unrecognized extension value`,
a string unique to `@codemirror/state`: two chunks each carry a full
copy, one with a 6.7.1-only code pattern and one without it.
3. Open a view that composes extensions from packages on different
copies: the extension set rejects the foreign-instance extension.
**Paperclip version or commit**
master (
|
||
|
|
4d317274ce |
feat(channels): add experimental iMessage Photon (#13299)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Channels connect external conversations to company tasks and agent execution. > - Slack, Discord, and AgentMail already provide durable delivery and access controls. > - People also need to reach an agent from Apple Messages and send photos. > - Photon provides shared Pro DMs, dedicated numbers, and authenticated event recovery. > - This pull request connects Photon to the existing channel services. > - People can message an agent while Paperclip retains task ownership and approval authority. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: channel services, shared contracts, database constraints, Apps, and agent Channels UI. **Problem or motivation** Paperclip has no iMessage channel. A person cannot use Apple Messages to start a task, send a photo, or answer an agent's pending question. **Proposed solution** Add experimental **iMessage Photon** with Pro-compatible shared DMs or a dedicated Photon Cloud number per agent channel. Reuse channel admission, identity links, task generations, publication, and interaction continuation. Keep groups disabled for shared allocation. Dedicated lines support groups that an operator explicitly enables. Require a fresh linked message and a published agent response before setup completes. **Alternatives considered** Shared allocation has no owned phone number, so it reserves one project and allows DMs only. Dedicated allocation reserves one stable number. Local Mac access needs a separate deployment model. The upstream Photon Chat SDK adapter does not persist the poll mappings and send receipts required here. This change uses the lower-level SDK without adding another agent runtime. **Roadmap alignment** This extends Connected Apps and agent communication through the existing channel subsystem. It does not add a parallel tool connection or agent loop. GitHub searches for Photon and iMessage found no matching provider implementation. **Additional context** This ships behind the existing experimental channel gate. Dedicated-line release qualification remains incomplete. Real Photon Pro DMs passed task/reply, native poll, text answers, confirmation rejection, media, restart, pause, reconnect, revocation, and removal tests. An operator-supplied iPhone camera HEIC also passed the full round trip. Dedicated groups remain unqualified. See [the verification record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) and [the implementation plan](doc/plans/2026-09-11-imessage-photon.md). ## What Changed - Add the provider catalog entry, shared setup contracts, and a forward migration. A global partial index reserves the dedicated number or shared project until its endpoint is archived. - Add Cloud project inspection, vaulted project credentials, selected-line token renewal, and a leased receiver. Persist checkpoint updates under the receiver lease. Shared project replay accepts sparse increasing sequences only after a complete recovery barrier. - Connect DMs and enabled groups to existing task generations, sender authorization, ordered delivery, and publication services. Keep each iMessage conversation on its task after completion; only explicit `/new` or `/close` releases the binding. Publish committed inbound comments live and label their human bubbles “Sent from iMessage” in both task-chat renderers. - Persist immutable text/file send identities, upload receipts, poll IDs, option IDs, per-person drafts, and canonical interaction continuation proofs. - Add source-bound file recovery, bounded HEIC/HEIF conversion, JPEG previews, and related Live Photo companion video retention. - Add the three-step setup flow and channel management surfaces with official branding. Preserve the experimental gate and existing pause/disconnect behavior. - Add interactive production-component Storybooks for setup, access, recovery, and ongoing conversations. Add provider, integration, catalog, and browser regression coverage. Document setup, recovery, supported boundaries, and qualification gaps. ## Verification - Live Photon Pro, SDK 2.1.0: linked iPhone messages create a task and receive native Codex replies in Apple Messages. Unlinked senders cannot start work. - Three real follow-ups each reopened the same completed task. Incoming bubbles appeared on its open page without reload and showed “Sent from iMessage.” The third follow-up ran after restarting the server on `4d7222110`; the agent correctly repeated its previous reply from before the restart. - Native polls after restart, sequential text drafts, required-field correction, explicit submission, approval rejection with a required reason, and native continuation passed against Photon. - PNG, text documents, synthetic HEIC, and a real iPhone camera HEIC passed in both directions. The camera photo produced a 3024×4032 JPEG preview. The native agent described it and returned the received HEIC byte-for-byte. - Pause/resume, reconnect, identity revocation, removal, `/status`, `/new`, `/close`, and stale answers after close passed live. Messages suppressed by pause did not become work on resume. Removal stopped intake and removed credential bindings. - All 304 focused tests passed on `4d7222110`. These cover Photon unit/integration behavior, both task-chat renderers, live comment hydration, completed-task continuity after restart, enabled groups, duplicate delivery, and explicit reset/close. The selected Teams completion-boundary regression also passed. Full workspace typecheck/build and token gates passed for the conversation fix; the final UI changes passed their affected typecheck/build and tests. - All 26 new Photon Storybook Playwright cases passed in light and dark themes, including the complete shared-DM setup journey and 390px mobile follow-ups. UI typecheck and the Storybook build passed. These stories use simulated Photon responses and do not replace the live evidence above. - The full chat-adapters browser suite previously passed all 39 cases. Migration checks passed, and migration 0275 applied to the isolated live instance with the earlier Photon migration already applied. - The local full Vitest run was previously interrupted by the host's embedded-Postgres shared-memory limit; it is not a full-suite pass. All 30 applicable CI checks passed on preceding head `7a5419cac`, with two skipped checks and Greptile 5/5. Head `24f8e1aae` adds an explicit required-story discovery guard to the 26 passing Storybook cases. Greptile rates this final head 5/5 with no unresolved review threads. All 30 applicable CI checks passed, with two optional checks skipped. - A repeated live send key suppressed the duplicate but returned gRPC 6 / SDK `internalError` without an original receipt. Paperclip keeps unknown delivery unresolved. This provider behavior is covered by a regression test. - See [the verification record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) for package versions, redacted live evidence, deterministic coverage, and remaining qualification gaps. ## Risks - Dedicated group qualification remains unrun; groups are disabled for the approved Pro scope. Real iPhone camera HEIC passed transport, preview generation, agent inspection, and return. Keep the channel experimental; the dedicated-line release matrix remains incomplete. - Shared recovery and attachment aliases were verified against the live gateway. Duplicate writes currently return an error without the original receipt; unresolved sends require operator resolution. The implementation fails visibly on invalid replay ordering, a reset cursor, or changed identity. - The HEIF converter passed on macOS arm64 and in Linux CI. Windows HEIF binaries have not been executed in this work. Linux musl has no packaged converter. Unsupported conversion retains the original and reports the missing preview. - The migration adds a global reservation across companies for Photon numbers and shared projects. Paused and revoked endpoints keep that reservation until removal. - Integration touches shared channel services. Existing provider browser coverage passes; broad repository verification is recorded above. - `pnpm-lock.yaml` is intentionally excluded under repository policy. The repository bot owns lockfile updates. The additional Superagent supply-chain scan is neutral/inconclusive because these new dependencies are not yet in the committed lockfile. Its security scan passed; all required CI checks pass. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository inspection, code execution, browser testing, and tool use. The exact served model identifier and context-window size are not exposed in this session. No sub-agents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
44f6312cd8 |
fix(ci): reuse one available Cloud registry cache (#13334)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud needs a verified image for each merged source commit. > - Fresh builders restore compiled native dependencies from registry caches. > - The current workflow imports up to eleven historical cache manifests at once. > - Live builds missed native layers that a fresh builder reused from one manifest. > - This PR selects the nearest available cache and tests reuse across fresh builders. ## Linked Issues or Issue Description Refs #13329 and #13330. A search of open cache PRs found no duplicate of this change. **What existing behavior does this improve?** Remote Docker cache reuse on fresh Cloud image builders. **Current behavior** [Cloud run 34714483272](https://github.com/paperclipai/paperclip/actions/runs/34714483272/job/103609096836) imported the previous cache manifest successfully but rebuilt `cargo-chef` and Rust dependencies. The dependency compile took 3m43s. The preceding image build had already exported those layers. A controlled [fresh-builder diagnostic](https://github.com/paperclipai/paperclip/actions/runs/34715336530) used the same source and registry cache. The single-manifest job reused both layers immediately. The multiple-manifest job rebuilt them and failed the cache assertion. Both jobs used GitHub-hosted runners with read-only access. **Proposed behavior** Inspect cache manifests in first-parent order and import only the nearest available one. Keep full-SHA cache exports, the ten-commit search bound, and the legacy fallback. If caches cannot be read, permit a cold build. **Reason and benefit** Avoid the observed cache misses without changing image contents or builder sizes. Expected savings include about four minutes of native tool/dependency compilation when those inputs are unchanged. The final merge-to-deployable gain still needs a post-merge measurement. **Breaking changes** No image, artifact, deployment, or runner-routing contract changes. ## What Changed - Select one available ancestor cache after Docker login and Buildx setup. - Preserve separate writable cache tags for each full source SHA. - Test cache ordering, missing caches, registry errors, and workflow integration. - Add the selector tests to the existing release-registry suite. - Export a local test cache, remove the first builder, and verify a source rebuild on a fresh builder. - Document cache selection and the stronger Docker check. ## Verification - Passed 456 focused workflow, routing, readiness, preview-artifact, and cache-selector tests. - Passed shell syntax, ShellCheck for the changed probe, actionlint workflow validation, and `git diff --check`. actionlint's shell checks were disabled for the workflow validation because unchanged migration-label commands trigger existing SC2012 notes. - The fresh-builder registry diagnostic proves the single-cache behavior. The [permanent two-builder probe passed](https://github.com/paperclipai/paperclip/actions/runs/34715771048/job/103612624090), including a changed real binary and dependency-declaration invalidation. - Passed all 35 latest-head checks (green or intentionally skipped), including full typecheck, test, build, and browser suites in [PR CI run 34715771217](https://github.com/paperclipai/paperclip/actions/runs/34715771217). - The real selector CLI inspected registry metadata and chose the nearest available ancestor cache. - Fresh Greptile review is 5/5 with no open findings. The PR title was corrected to meet the source-change naming rule; the review check passed after that correction. - Local full-suite runs and Docker builds are unavailable because the local Docker daemon is unresponsive after disk exhaustion. CI provides the Linux verification. ## Risks - Missing or unreadable caches cause a slower cold build. The selector logs that condition and preserves image publication. - Inspecting several missing ancestors adds lookup time. Each lookup has a ten-second timeout and the search is bounded. - The Docker test now exports a local cache. It removes the first builder before starting the second to release disk space, then cleans up its builders and files. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0ce7df2648 |
ci: keep Cloud readiness markers out of the builder queue (#13330)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud consumes versioned readiness markers for each merged source commit. > - Those markers can be published only after the required checks and artifacts pass. > - The marker jobs currently wait for the same AWS runner capacity as builds and tests. > - A busy builder pool can delay readiness after all required work has finished. > - This PR moves the small readiness jobs to GitHub-hosted runners while retaining every dependency gate. ## Linked Issues or Issue Description Refs #13326 and #13328. **What existing behavior does this improve?** Time from completed Cloud verification to a deployable marker, and AWS capacity occupied by artifact polling. **Current behavior** In [Cloud readiness run 34711557083](https://github.com/paperclipai/paperclip/actions/runs/34711557083), all builds and tests finished at 18:43:05 UTC. The source marker did not start until 18:44:21, and the deployable marker did not start until 18:44:37. Merge-to-deployable was 11m33s, although the prerequisite work finished in 9m44s. **Proposed behavior** Run the artifact wait and both versioned marker jobs on `ubuntu-latest`. Keep the compute jobs on the approved post-merge AWS fleet. **Reason and benefit** Avoid builder-pool queue delays after verification finishes. This also removes the long artifact-wait job from AWS capacity. Expected savings depend on queue depth: the observed run had over 90 seconds of avoidable marker waiting. The marker commands themselves take only seconds. **Breaking changes** Runner placement changes for three bookkeeping jobs. Marker names, exact-source artifact checks, required verification, and image verification stay the same. **Additional context** Searched the related runner and Cloud readiness work. This addresses queue time observed after the parallel verification change. ## What Changed - Place the artifact wait, source-verification marker, and deployable marker on GitHub-hosted runners. - Keep all existing job dependencies, source guards, permissions, and commands. - Extend routing regressions to enforce this placement and retain fail-closed readiness gates. - Document why readiness bookkeeping uses separate runner capacity. ## Verification - Passed 433 workflow, routing, source-verification, and Cloud readiness tests with `node --test .github/scripts/tests/*.test.mjs scripts/cloud-source-verification.test.mjs scripts/cloud-readiness.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`. - Passed `actionlint` and `git diff --check`. - Passed all latest-head CI gates in [run 34712624340, attempt 2](https://github.com/paperclipai/paperclip/actions/runs/34712624340), including typecheck, build, browser, Runner, and all general/serialized tests. - Attempt 1 had one localhost readiness timeout in an unchanged test. A single targeted retry passed all 166 files (3,225 tests passed, one existing skip), including all seven tests in that file. No timeout, assertion, or application source was changed; the retry is documented in the PR comment. - Local full-suite verification is limited by local disk exhaustion; the focused checks above pass. - Fresh Greptile review is 5/5 with no unresolved findings. ## Risks - GitHub-hosted capacity can also queue, but these jobs no longer compete with AWS build/test demand. The change does not reserve instances or change box sizes. - Readiness must still fail if any prerequisite fails. The existing `needs` relationships and success-only execution are preserved and tested. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.912.0-canary.10 |
||
|
|
7435b2ee9c |
ci: cache compiled Docker Rust dependencies separately from source (#13329)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud deploys images that contain the native Rust Runner. > - The image already builds that Runner before copying ordinary app source. > - A Rust source change still invalidates its entire compiled dependency layer. > - Compiled dependencies can survive source changes when their recipe is unchanged. > - This PR adds a separate locked dependency build before compiling the real workspace. ## Linked Issues or Issue Description Refs #13195. A search of related Docker and Cargo cache PRs found no duplicate dependency-recipe change. **What existing behavior does this improve?** Docker image build time after Rust source or embedded protocol changes. **Current behavior** The `runner-build` stage compiles dependencies and workspace code in one layer. In Cloud readiness run 34698143548, that stage took about 3m48s when its cache was unavailable. **Proposed behavior** Generate a recipe with pinned cargo-chef 0.1.73. Build locked release dependencies in `runner-deps`, then copy and compile real Rust source and embedded protocol inputs in `runner-build`. Source edits can reuse the dependency layer from the existing registry cache. **Reason and benefit** Reduce dependency recompilation during source changes and merge bursts. Expected savings are roughly 2–4 minutes when the old native layer would miss but dependency layers are available. Full cold builds also pay for the recipe tool installation. Ordinary app-only cache hits gain little from this change. **Breaking changes** None to the shipped application or image tags. The recipe tool and compiled dependencies remain in build stages. ## What Changed - Install a pinned recipe generator with its locked dependencies and the existing package-owned compiler. - Add recipe planning and compiled dependency stages. Use the same release profile, package, binary, and lockfile enforcement as the real native build. - Remove generated source stubs before copying actual source. Preserve protocol inputs, timestamp normalization, binary staging, and application checks. - Add Docker cache wiring regressions and update the Docker cache documentation. - Run a two-build probe in Docker Runner check. It requires dependency reuse, changed real binary metadata after a source edit, and a changed recipe after a dependency declaration edit. It uses a disposable tracked-source context and exports only small metadata files. ## Verification - Passed all five Docker build-stamp and dependency-cache tests with `pnpm exec vitest run server/src/__tests__/docker-build-stamp.test.ts`. - Passed the local ARM64 `docker buildx build --target runner-build --progress plain`. Local Docker then hit storage errors during a runtime probe; cache invalidation verification continues on GitHub-hosted Linux. - Passed `bash -n scripts/check-docker-runner-cache.sh`, `actionlint`, and `git diff --check`. - Passed a [Linux AMD64 cache probe](https://github.com/paperclipai/paperclip/actions/runs/34711042199) against the PR source: dependencies compiled in 3m49s for the baseline and were `CACHED` after a source edit; real source compilation took about 37 seconds. Binary metadata changed and dependency declaration changes altered the recipe. The permanent probe is also running in latest-head Docker Runner check. - Passed latest-head [Docker Runner check](https://github.com/paperclipai/paperclip/actions/runs/34711145160), including the permanent source/dependency invalidation probe. - Passed full [PR verification](https://github.com/paperclipai/paperclip/actions/runs/34711145352/attempts/2): typecheck, all grouped tests, native verification, build, release dry run, and browser checks. One unrelated signoff-policy browser test failed waiting for a heartbeat run on attempt 1; only that failed shard and dependent checks were retried, and passed. - Latest-head Greptile is 5/5 with no unresolved findings. Full local tests/build were limited by local disk exhaustion; Linux CI completed those checks. ## Risks - The two-build CI probe has a 20-minute job limit to cover the cold build and source rebuild. It adds no AWS routing. - A fully cold build must install cargo-chef and populate the dependency layer. Both become reusable registry layers; no Actions cache is added. - The recipe and final build must keep the same compiler, build profile, package, binary, and directory layout. A source-change rebuild probe checks real cache reuse and binary invalidation. - Dependency or compiler changes still require rebuilding dependencies. Existing image verification and full-SHA publication gates remain unchanged. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.912.0-canary.9 |
||
|
|
8df3ee2cf5 |
ci: rebalance serialized tests with current suite durations (#13328)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud deployments wait for verified source commits. > - Verification splits serialized server tests across independent runners. > - The shard duration estimates came from August and no longer match current tests. > - Stale estimates put much more work on one runner than the others. > - This PR refreshes the estimates from a complete successful run to balance the existing runners. ## Linked Issues or Issue Description **What existing behavior does this improve?** The time spent waiting for the slowest serialized server-test shard in PR and release verification. **Current behavior** In [Cloud readiness run 34705914878](https://github.com/paperclipai/paperclip/actions/runs/34705914878), the five serialized shards spent 479, 371, 322, 390, and 365 seconds running tests. The recovery suite had a 55-second estimate but now takes about 156 seconds including process overhead. **Proposed behavior** Use fresh per-suite measurements with the existing deterministic duration balancer. Applying the same measured costs to the new assignment gives 385, 385, 386, 385, and 385 seconds. This predicts about 93 seconds less waiting for the slowest shard, before runner/setup overhead. Live CI will confirm the result. **Reason and benefit** Use the existing runners more evenly. No extra runner, test parallelism, cache, timeout, or routing change is needed. **Breaking changes** Suite-to-shard assignments change. The full suite set, assertions, and per-suite process isolation stay the same. **Additional context** Searched related CI and shard PRs. This updates the existing duration manifest, without duplicating a pending sharding implementation. ## What Changed - Refresh all 145 serialized suite weights from the same successful release verification run. - Record source job IDs and the measurement method in the manifest. Durations include process startup, imports, collection, tests, and shutdown. ## Verification - Passed all 30 shard and release-workflow tests: `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`. - Confirmed every measured suite appears exactly once across all five source logs. - Compared old and new assignments using the same measured weights. The maximum fell from 478862ms to 385551ms. - In [PR CI run 34710696242](https://github.com/paperclipai/paperclip/actions/runs/34710696242), all five serialized jobs passed in 7m05s–7m27s including setup. The measured assignment is now balanced in a live run. - The same run passed full typecheck, all grouped tests, native verification, build, release dry run, and browser checks. Local full-suite verification on this base was limited by disk exhaustion; local typecheck and targeted shard tests passed. - Latest-head Greptile is 5/5 with no open findings. Every current-head CI check must be green or intentionally skipped before merge. ## Risks - Individual durations vary with load and future test changes. These estimates affect assignment only; missing or renamed suites receive the existing median weight. - Both PR and release verification read this manifest, so both receive the new assignments. Each suite still runs in its own serialized Vitest process. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ed50a39c3f |
fix: preserve NUL characters in run-event payloads (#13325)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
d2e940f4c1 |
ci: run release Runner protocol and Rust checks in parallel (#13326)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud deploys verified images from merged source commits. > - Cloud readiness waits for every release verification check. > - Runner verification currently runs long TypeScript tests before Rust checks. > - These checks can run on independent runners with their own build directories. > - This PR runs them in parallel while preserving all checks and the shared dependency cache. ## Linked Issues or Issue Description Refs #13194. Related prior work: #13142 and #13259. A search found no duplicate parallel release-check change. **What existing behavior does this improve?** Time from merge to Cloud source verification and deployment readiness. **Current behavior** Recent successful runs take roughly 13 minutes from merge to deployable. In run 34705914878, Runner verification took 11m23s. Protocol tests finished before Rust tests and API authority checks started. **Proposed behavior** Run protocol and Rust verification in two matrix jobs. Cloud readiness still requires both jobs to pass. **Reason and benefit** Remove the serial dependency between independent checks. Expected improvement is about 2–3 minutes on a typical cached run, until the image build or server tests become the longest job. This is an estimate; post-merge timing will confirm it. **Breaking changes** Individual release Runner job names gain a lane suffix. Cloud source and readiness marker names stay the same. PR runner routing is unchanged. ## What Changed - Split release Runner checks into protocol and Rust lanes. Keep every constituent of `check:all` exactly once. - Restore the existing Rust dependency cache in both lanes. Allow only the Rust lane to save it after warming both build profiles. - Add coverage and cache authorization regressions. Document the parallel verification and single cache writer. ## Verification - Passed 477 workflow and source-verification tests with `node --test .github/scripts/tests/*.test.mjs scripts/cloud-source-verification.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`. - Passed `actionlint`, `git diff --check`, and the private AWS routing regression suite. - Passed local `pnpm -r typecheck` and the standalone `check:runner && check:api-authority` lane, including all 1,671 API tests before the protocol lane had built TypeScript output. - The broad local protocol run under Node 25 had four failures. The two affected files passed under CI's Node 24.19.0: 67 passed, 6 platform skips. - Local `pnpm test:run` aborted when disk space ran out; local `pnpm build` could not run afterward. These are local verification limits. [Linux CI run 34710421424](https://github.com/paperclipai/paperclip/actions/runs/34710421424) passed full typecheck, all grouped tests, native verification, build, release dry run, and browser checks. Native protocol CI passed 1,986 tests, plus 1,671 API tests and the Rust suites. - Latest-head Greptile is 5/5 with no open findings. All 33 current-head checks are successful or intentionally skipped. ## Risks - Uses one additional short-lived verification runner per release verification. The existing AWS exact-master restriction remains in place. - The Rust lane warms debug dependencies so its cache save also serves protocol tests. Both lanes always rebuild workspace code. - A workflow regression could omit a check. The new coverage test compares the matrix checks directly with `check:all`; Cloud readiness depends on the complete reusable workflow. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a59f5a8adc |
fix(server): stop reporting expected managed-cloud transients to Sentry (#13323)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The server reports crashes to Sentry so operators can find real
faults
> - Three expected conditions report as crashes: a client that closes
the connection mid-request, one stale pooled database socket after a
pooled endpoint recycles, and the short boot window where a supervised
cloud stack runs a new app image before its migration runner has caught
up
> - These events arrive in the hundreds and bury real errors
> - This pull request classifies each condition as expected and stops
the Sentry capture for exactly that condition, with behavior unchanged
everywhere else
> - The benefit is a Sentry feed where each event is a real fault
## Linked Issues or Issue Description
**What happened?**
Three noise classes fill the backend Sentry project on managed cloud
fleets:
1. `Error: aborted` (ECONNRESET) reports as a 500 crash when a client
closes the tab or loses its network mid-request. Observed 18 times in
one week from routine client disconnects.
2. `Error: write CONNECTION_CLOSED ...` reports from many query paths
after a pooled Postgres endpoint suspends. The existing single retry in
cloud actor resolution still fails, because a suspended endpoint kills
every pooled socket at once and the one replay draws another dead
socket.
3. `Error: PostgreSQL has pending migrations (...). Refusing to start`
reports from every supervised stack during a fleet upgrade. The
supervisor delivers the new app image before it runs the migration
runner, so each stack crash-loops briefly by design. One fleet roll
produced 329 events (11 per container).
**Expected behavior**
A client disconnect ends the request quietly. A transient dead socket is
replayed until a live socket answers. A supervised mid-upgrade boot
refusal logs and exits nonzero without a Sentry capture, while the same
refusal on a self-hosted deployment keeps reporting.
**Steps to reproduce**
1. Abort an HTTP request mid-flight: the error handler reports a crash
to Sentry.
2. Suspend a pooled Postgres endpoint under an idle server, then issue
two quick requests: the first replay can draw a second dead socket and
surface `CONNECTION_CLOSED`.
3. On a deployment with `PAPERCLIP_CLOUD_API_ORIGIN` set, add a
migration file without running the migration runner and boot: the
refusal reports to Sentry.
**Paperclip version or commit**
master (
|
||
|
|
24007a9804 |
fix: promote review tasks when continuations start (#13318)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat service starts and tracks agent work on issues. > - An issue can be in review before a user comment starts more agent work. > - The previous checkout rule did not claim an issue from review for this continuation. > - The issue could stay in review while an agent actively worked on it. > - This pull request lets a resolved interaction continuation claim an issue from review. > - The existing checkout update then sets the issue to in progress. > - The benefit is that issue status now shows active agent work correctly. ## Linked Issues or Issue Description **What happened?** An issue stayed in review when a new agent continuation started work on it. **Expected behavior** The issue must move to in progress when an agent starts work. An idle issue must stay in review. **Steps to reproduce** 1. Put an assigned issue in review. 2. Resolve an interaction that starts a continuation. 3. Start the heartbeat run. 4. Observe that the issue remains in review while the run works. **Paperclip version or commit** The problem reproduces on the master branch before this change. ## What Changed - Allow a resolved interaction continuation to claim an assigned issue from review. - Keep idle review issues unchanged. - Add regression tests for both behaviors. ## Verification - Run `node_modules/.bin/vitest run server/src/__tests__/heartbeat-auto-checkout.test.ts`. - Confirm that all four tests pass. ## Risks - Low risk. The change only expands checkout eligibility for resolved interaction continuations. - The existing guarded checkout update still controls ownership and the status update. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5.5. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c9021c6721 |
fix: require explicit native completion reviews (#13314)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runs report their outcome through paperclip_finish. > - The server previously turned incomplete reports into human approval requests. > - Those requests could block a later successful run, even when no person had requested review. > - This pull request creates review cards only for explicit attention requests and withdraws proven old fallback cards. > - Agents receive useful completion feedback, while explicit approval gates and task state protections remain in force. ## Linked Issues or Issue Description Related: #13266 removed reviews caused by policy upgrades. This change removes the separate completion fallback. **What happened?** An agent reported needs_review while waiting for checks without requesting a human decision. Paperclip created a generic Native completion review. A later successful report could not complete the task because that old card remained pending. **Expected behavior** Ordinary low-risk work completes after a valid done report, a successful run, and workspace finalization. Incomplete work stays with the agent. Explicit approval requests remain visible and must be resolved. **Steps to reproduce** 1. Complete a native run with needs_review and no attention requests. 2. Continue the task and submit a successful done report. 3. Observe that the old implementation leaves the task in review behind a generic confirmation card. ## What Changed - Require explicit attention requests to create native review cards. Route each request independently and preserve pending or declined decisions. - Withdraw only pending system cards with matching old decision, assessment, effect, contract, and prompt provenance. Preserve history and explicit or answered requests. - Reassess an affected current result without overwriting later task edits, runs, contracts, or workspace failures. - Return pending approval links and required actions through the completion tool. Reject contradictory done reports and empty review requests before accepting a result. - Allow one corrective continuation for incomplete results, then expose a recovery action. - Update status fixtures, database regressions, runner tests, and the completion contract documentation. ## Verification - `pnpm -r typecheck` passed after merging current master. - `pnpm build` passed after merging current master. - The combined branch passed 89 completion and Agent Chat tests. Other targeted tests passed: 170 external-chat and reconciliation tests; 50 runner-resume and control-plane tests; 13 arbiter tests; 7 chat delivery tests; 21 runner completion and runtime-context tests. - The full local test attempt exposed old review fixtures and a missing fake-provider binary. The fixtures are fixed and the helper is built. All affected suites pass in fresh reruns. The timing-sensitive Discord test also passed on rerun. - All latest-head CI checks passed, including build, typecheck, general and serialized tests, runner verification, browser tests, and canary dry run. Greptile is 5/5 with zero unresolved comments. ## Risks - Cleanup changes existing pending cards. It requires exact system provenance and only applies to low-risk agent-claim contracts. It does not delete history or dismiss explicit requests. - Status still commits after the turn and workspace finalization. Completion feedback reports current constraints and does not claim an early status commit. - Incomplete reports now request a bounded corrective run instead of an automatic approval. Repeated failures expose recovery. ## Model Used OpenAI Codex, GPT-6, with repository inspection, code execution, and test tools. The exact deployment identifier and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7e6d512597 |
fix(onboarding): make chief-of-staff hiring reliable (#13317)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The first agent helps the board define work and hire other agents. > - That agent can have the general role while its instructions require hiring skills. > - Missing skills and blocked schema discovery make valid requests fail. > - Repeated confirmation and invalid waiting guidance can turn these failures into extra runs. > - This PR supplies the required skills, opens read-only schema discovery, and corrects the guidance. > - The agent can complete an authorized hire while company approval and duplicate checks still apply. ## Linked Issues or Issue Description Refs #13068 — the first-task onboarding flow that this change repairs. Refs #12029 — related drift between the sandbox allowlist and bundled hiring guidance. This PR adds schema access; it does not replace the earlier hiring-route fix. **What happened?** A general-role onboarding chief received hiring instructions without the core hiring skills. Sandbox requests to the documented OpenAPI endpoint failed. The agent then guessed question and hire payloads. The persona required new confirmation after validation errors and described waiting states that agents cannot set. **Expected behavior** A direct request authorizes the requested hire. The chief asks only for material missing details, uses valid API payloads, and completes the task. Formal company approval gates still apply. A saved human-input card gives the task a valid waiting state. **Steps to reproduce** 1. Create an onboarding chief with role `general` through the board. 2. Ask it to hire a friendly robot with a supplied name and responsibilities. 3. Check its assigned skills, schema requests, question cards, hire requests, and final task state. **Paperclip version or commit** Reproduced on the first-task onboarding implementation after #13068. The live local verification used this branch at `112f44610`. **Deployment mode** The original failure used a hosted sandbox with legacy Codex ACP. Live verification used an isolated local instance and real `codex_local` execution. Queue and HTTP/2 transport access is covered by automated tests. ## What Changed - Give board-created onboarding chiefs the existing core skills regardless of role. Preserve explicit skill version pins, including aliases. Keep ordinary general-agent defaults and authorization checks. - Allow exactly `GET /api/openapi.json` through both sandbox bridge transports. - Publish validator-tested question, free-text, hire, and waiting examples. Regenerate the runner API reference and capability inventory. - Clarify direct authorization, material ambiguity, and correction of confirmed pre-creation validation failures. Preserve uncertain-outcome reconciliation, duplicate protection, and company approval gates. - Align disposition instructions with agent permissions and the saved human-input waiting path. ## Verification - After rebasing onto current `master`: 69 targeted server tests, 110 queue/HTTP2 bridge tests, and 4 capability inventory tests passed. These cover core skill defaults, version pins, actor restrictions, schema access, published examples, hire validation, idempotency, and approval gates. Waiting recovery tests and live question flows also passed before the rebase. - `pnpm -r typecheck` and `pnpm build` passed again after the rebase. Frozen dependency installation and both generated capability checks passed. - Ran the full `pnpm test:run` suite. The initial run had 14 failed server files due to local database resource limits, a missing built test fixture, and socket failures. All 14 files passed after fixture repair and isolated retries. UI, CLI, workspace packages, database tests, and all 145 serialized server files passed. - Real one-request hiring replay: one hire, one successful run, task done in 2m16s. No repeated approval or recovery escalation. - Real two-turn browser conversation: start with an unspecified hire, then supply a name and friendly robot responsibilities. One clarification card, one hire, two successful runs, task done in 3m27s of execution. No failed writes, confirmation cards, or recovery actions. - Assigned the hired robot a welcome-message task through the browser. It produced a warm message under 100 words and finished in one successful 66-second run, with no questions or recovery actions. - The two-turn flow still asked an optional preferences question and gave a technical final reply. These are remaining presentation limits. - Greptile: 5/5 on `b71f83ba2`, with zero unresolved review threads. Fixed its generator finding and passed 1,655 published-example/runtime API tests plus server typecheck. All latest-head CI checks are green (32 passed; 2 unrelated Storybook checks skipped). The signoff-policy browser test initially timed out while waiting for an approver run. Its shard passed on one rerun without code changes. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34698211049). ## Risks - Onboarding chiefs receive more default skills. Ordinary general agents retain existing defaults, and explicit versions take precedence. - Prompt guidance can affect model behavior. The live replays are examples, not a guarantee that every model follows the guidance. - Retry guidance applies only when validation confirms that nothing was created. Uncertain outcomes still require checking existing agents. - No database migration or new public endpoint. Existing company boundaries, approval gates, and bounded recovery remain in force. ## Model Used OpenAI Codex, model `gpt-6-astra`, with reasoning, tool use, code editing, and live browser verification. The exact context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.912.0-canary.8 |