mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-11 05:31:46 +02:00
120ae5428fa29bee300bcf806491cd4d965fbb7c
1066
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
120ae5428f |
feat(server): add a one-click relink action for detached custom-image templates (#11641)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip environments can use captured custom images for agent runs > - A configuration fingerprint change can detach a valid custom-image template > - Operators need a safe way to confirm that the image still matches the boot source > - This pull request adds a guarded relink action with drift classification and audit logging > - The benefit is a deliberate relink without a new sandbox boot or provider snapshot ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting environment, server, and UI behavior. **Problem or motivation** A custom-image template detaches when the environment configuration fingerprint changes. The runtime then uses the base image, even when the boot source did not change. The only prior remedy required a full re-capture. **Proposed solution** Add an operator-triggered relink action. Classify configuration drift from a server-owned boot-relevant snapshot. Relink knob-only drift without confirmation. Require explicit confirmation for boot-source or unclassified drift. Guard the route for instance administrators and record a safe activity event. **Alternatives considered** Keep requiring a full re-capture. This adds a sandbox boot and provider snapshot for cases where the image remains correct. **Roadmap alignment** The roadmap has no matching custom-image relink item. This change addresses an environment operation gap. **Additional context** The relink response exposes raw drift values only in the transient 409 response to the instance administrator. The service never persists or logs fingerprints or configuration values. Reserved identity-path segments fail closed. ## What Changed - Add `relinkActiveTemplate` with drift classification and conditional fingerprint update. - Persist a server-owned boot-relevant configuration snapshot during capture. - Add the guarded relink route with strict request validation and activity logging. - Add the relink action and confirmation flow to the environment page. - Add service, route, UI, and OpenAPI coverage. ## Verification - Run the focused service suite: `pnpm vitest run server/src/services/environment-custom-images-service.test.ts`. - Run the focused route suite: `pnpm vitest run server/src/routes/environment-custom-image-routes.test.ts`. - Run the focused UI suite: `pnpm vitest run ui/src/pages/CompanyEnvironments.test.tsx`. - Run server and UI TypeScript checks. - Confirm the OpenAPI snapshot matches the new route. - Confirm all required GitHub checks pass on commit `e46fdcfe94a719be854adf8849d30714e5b70b93`. - Confirm Greptile reports 5/5 with no unresolved review threads. ## Risks The relink action can keep an image after configuration drift. The service requires explicit confirmation for boot-source or unclassified drift. Reserved path segments produce a safe unresolved marker and never enter stored values. ## Model Used OpenAI GPT-5 Codex. The model used repository inspection, GitHub operations, and PR preparation with tool use and code execution. The runtime did not expose a context-window value or a separate reasoning-mode value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
393da0f67c |
fix(adapter-utils): graft unrelated imported histories instead of failing the run (#11638)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - At run finalize, the host imports the sandbox git history and reconciles it with the local worktree in `integrateImportedGitHead` > - Transported workspaces are depth-1 shallow clones, so the boundary commit reads as parentless inside the sandbox > - A `git commit --amend` there rewrites the boundary commit into a root commit, and the re-imported history no longer connects to the host history > - `git merge-tree` has no common base to merge against, so the sync throws "Failed to merge concurrent remote git histories" and the run fails with its work stranded in the sandbox > - This pull request grafts the imported tree onto the current head as a single commit instead of failing > - The benefit is that a history rewrite inside the sandbox can no longer lose a run's work ## Linked Issues or Issue Description No existing issue found. I searched issues and PRs for "unrelated histories", "Failed to merge concurrent", and "shallow". Depends on #11637 (merged; the graft commit reuses its identity constant). This PR is now rebased onto `master`. **What happened?** An agent run amended a commit inside its sandbox workspace to address review feedback. The sandbox clone is depth-1 shallow, so git treated the boundary commit as parentless and the amend produced a root commit. At finalize, the host-side sync failed with `Failed to merge concurrent remote git histories for <sha>` and the run was marked failed. A follow-up run had to repair the branch by hand: fetch the true parent from origin and rebuild the commit with `git commit-tree`. **Expected behavior** The sync must never strand completed work. When the imported history shares no ancestor with the local one, the imported tree should still land on the current head, with the imported message preserved and the graft recorded. **Steps to reproduce** 1. Start a run whose workspace transport uses the shallow clone path (`withShallowGitWorkspaceClone`, depth 1). 2. Inside the sandbox workspace, run `git commit --amend` on the boundary commit. The result is a parentless root commit. 3. Finish the run. The host-side `integrateImportedGitHead` finds no merge base, `merge-tree` fails, and the run fails. ## What Changed - `git-workspace-sync.ts`: new exported `createUnrelatedHistoryGraftCommit` helper. It reads the imported head's tree and message, and creates one commit on top of the current head with the deterministic sync identity and a trailer that records the graft and both shas. - `integrateImportedGitHead` (both the remote-git-sync version and the SSH copy in `ssh.ts`): when `merge-base` reports no common ancestor, graft instead of throwing. The ref update keeps the same compare-and-swap and concurrent-retry semantics as the merge path. - The graft is gated on `git merge-base` exiting with status 1 — the no-ancestor signal. Operational failures (timeout, missing object, repository error) keep the loud merge failure instead of rewriting the tip. - New regression tests: one builds the exact shallow-amend shape (a root commit rebuilt from the base tree) and asserts the graft lands on the current head with the imported tree, subject, and graft trailer; one integrates a well-formed sha the repository does not hold and asserts the integration still throws with the branch tip unchanged. ## Verification - `pnpm vitest run packages/adapter-utils/src/git-workspace-sync.test.ts` — 19/19 pass (includes the new graft test and the merge-base failure-discrimination test). - `pnpm --filter @paperclipai/adapter-utils typecheck` — clean. - Full `pnpm vitest run packages/adapter-utils`: every file passes except `local-process-sandbox.test.ts`, which fails identically on an untouched `master` checkout on macOS (bubblewrap-dependent, pre-existing, unrelated). ## Risks - Behavioral shift: unrelated imported histories previously failed the integration; now they land as a squash-graft. In this degenerate case there is no base to merge against, so the imported tree is taken wholesale and concurrent local-only tree changes are superseded at the tip. The local commits keep their place in the graft's ancestry, and the trailer records both shas, so nothing is unrecoverable. The old behavior lost the imported work instead, which is the worse failure for an autonomous run. - The graft reuses the imported head's commit message, so branch history still reads naturally after a sandbox rewrite. ## Model Used - Claude Fable 5 (`claude-fable-5`), extended thinking, via Claude Code CLI (tool use for code exploration, test runs, and verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no doc surface describes this internal sync path) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
9ea8143c87 |
fix(adapter-utils): give sync-created merge commits a deterministic git identity (#11637)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Runs execute in transported workspaces; at finalize, the host syncs the sandbox git history back into the local worktree > - When both sides advanced, `integrateImportedGitHead` reconciles them with `git merge-tree` plus `git commit-tree` on the host > - Execution hosts are often containers with no git config and no resolvable hostname, so `commit-tree` fails with "Author identity unknown" > - That one local command failure marks the whole run as failed, even though the run's work succeeded > - This pull request gives sync-created merge commits an explicit, deterministic identity at the call site > - The benefit is that workspace finalize no longer depends on ambient host git configuration ## Linked Issues or Issue Description No existing issue found. I searched issues and PRs for "Author identity unknown", "unrelated histories", and "commit-tree identity". **What happened?** A run finished its work, but workspace finalize failed. The host-side sync ran `git commit-tree <tree> -p <localHead> -p <importedHead> -m "Paperclip remote git sync merge <sha>"`. Git exited with `Author identity unknown ... fatal: unable to auto-detect email address (got 'node@<container-id>.(none)')`. The adapter recorded the whole run as failed, and the host worktree kept the stale head. Any container deployment without a global gitconfig reproduces this; I observed it on a Paperclip Cloud stack. **Expected behavior** Commits that the sync machinery itself creates must not depend on ambient host git configuration. The merge commit is machine-authored, so it should carry a deterministic Paperclip identity. **Steps to reproduce** 1. Run the Paperclip server in a container with no `user.name`/`user.email` git config and a hostname git cannot turn into an email. 2. Let a run's sandbox branch diverge from the host worktree, so both sides advance. 3. Workspace finalize calls `integrateImportedGitHead`. The `git commit-tree` step fails with "Author identity unknown" and the run fails. ## What Changed - `git-workspace-sync.ts`: new exported `GIT_SYNC_COMMIT_IDENTITY_ARGS` (`-c user.name=Paperclip -c user.email=noreply@paperclip.ing`), applied to the `commit-tree` call in `integrateImportedGitHead`. - `ssh.ts`: the SSH-sync copy of `integrateImportedGitHead` applies the same identity args to its `commit-tree` call. - New regression test: builds divergent histories in a repo with no configured identity and asserts the sync merge commit is created with the deterministic identity, correct parents, and merged tree. ## Verification - `pnpm vitest run packages/adapter-utils/src/git-workspace-sync.test.ts` — 18/18 pass. - `pnpm --filter @paperclipai/adapter-utils typecheck` — clean. - Negative proof: with the source fix stashed, the new test fails on the identity assertion. - Full `pnpm vitest run packages/adapter-utils`: every file passes except `local-process-sandbox.test.ts`, which fails identically on an untouched `master` checkout on macOS (bubblewrap-dependent, pre-existing, unrelated). ## Risks Low risk. The change only adds `-c` identity flags to two machine-generated commit invocations. `GIT_AUTHOR_*` / `GIT_COMMITTER_*` environment variables still take precedence over `-c` when an operator sets them, so existing deployments that configure an identity keep their behavior. ## Model Used - Claude Fable 5 (`claude-fable-5`), extended thinking, via Claude Code CLI (tool use for code exploration, test runs, and verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no doc surface describes this internal sync path) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
6b8e42168e |
Add governed secret alias confirmation cards (#11486)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need scoped secret bindings to use external services safely. > - Agents could not request an existing secret under a new config name without an internal secret identifier. > - Existing binding proposals were only visible in Settings and did not create an issue-thread approval path. > - A confirmation card could record acceptance without proving that the binding was created. > - This pull request extends the existing secret proposal system with safe source references and governed issue-thread confirmation cards. > - The benefit is a one-click flow that creates the binding or shows a clear failure without exposing secret material. ## Linked Issues or Issue Description Related prerequisite: #11482. **Subsystem affected** Cross-cutting: server REST APIs, shared interaction contracts, database proposal schema, and issue-thread UI. **Problem or motivation** An agent can need an existing bound secret under a second config name. The agent cannot safely discover the internal secret identifier. The existing proposal is also easy for the operator to miss because it only appears in Settings. A generic confirmation can record acceptance without executing the binding. **Proposed solution** Let an agent create a binding proposal from one of its existing config paths. Mint a server-owned, human-only confirmation card on the checked-out issue. Recheck the operator's target-agent permission under the proposal row lock. Execute the existing proposal transaction after card acceptance. Store an `executed` or `failed` result on the card. Render the complete lifecycle in the issue thread and attention resolver. **Alternatives considered** A new alias subsystem would duplicate proposal quotas, expiry, authorization, and binding synchronization. A text-only issue comment would not provide a governed action or an execution result. An agent-supplied card payload would permit metadata smuggling. This change uses the existing proposal transaction and a server-owned payload instead. **Roadmap alignment** This change extends the completed "Secrets Manager with per-agent access" roadmap item. It preserves scoped bindings and audited resolution. The required GitHub search found no other open duplicate issue or pull request. ## What Changed - Added safe source-config-path binding proposals and preserved user-secret ownership checks. - Added a proposal-to-interaction link and an idempotent database migration. - Minted human-only `request_confirmation` cards with server-owned `secretProposal` metadata. - Rejected agent-supplied governed metadata and agent addressees. - Rechecked `agent_config:update` authority under the proposal lock before execution. - Recorded `executed` or `failed` results and posted a failure comment when no binding was created. - Settled failed accepted proposals atomically and mirrored rejection, withdrawal, and expiry in both directions. - Emitted `secret.binding.created` for new agent binding writes. - Added a dedicated issue-thread card for pending, executed, failed, rejected, withdrawn, and expired states. - Showed only the source label, target agent, config path, skeptical justification, expiry, and safe failure code. - Replaced resolved attention-query entries immediately with the stitched server result. - Added focused server, database, UI, and state-transition tests. - Added Storybook fixtures for every review state and documented the API and agent behavior. ## Verification - `pnpm exec vitest run ui/src/components/IssueThreadInteractionCard.test.tsx ui/src/components/AttentionInteractionResolver.test.ts` — 58 passed. - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - `pnpm build-storybook` - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/db typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/db check:migrations` - `NODE_ENV=test pnpm exec vitest run server/src/__tests__/issue-thread-interaction-routes.test.ts server/src/__tests__/secret-proposals-routes.test.ts server/src/__tests__/secrets-routes.test.ts server/src/__tests__/agents-service-secret-bindings.test.ts` — 142 passed. - `NODE_ENV=test pnpm --filter @paperclipai/db exec vitest run src/company-secret-proposals-migration.test.ts --silent` — 1 passed. - `pnpm -r typecheck` - `pnpm test:run` — server 4,175 passed, UI 4,109 passed; the CLI AWS-doctor case passes 8/8 with runtime-injected static AWS credential variables unset. - `pnpm build` - `git diff --check origin/master...HEAD` ## Risks - Migration `0221` adds one nullable foreign key and one index. It uses idempotent guards. - The accept route performs a governed write after it records card acceptance. A failed write is visible and settles the proposal as rejected. - Concurrent proposal and card resolution must use proposal-before-interaction lock order. A race test covers direct approval against card rejection. - The new audit event increases activity rows for newly added agent bindings. It does not include secret values or fingerprints. - The card includes only safe proposal metadata. It does not include secret value, fingerprint, version, or internal secret identifiers. - The UI uses the stitched resolution result. Focused tests cover immediate cache replacement and every terminal state. > This work extends an existing completed roadmap capability. The GitHub duplicate search returned no other open related work. ## Model Used - OpenAI Codex with model ID `gpt-5`. The runtime did not expose its context-window size. Reasoning, repository tools, code execution, database integration tests, UI rendering, and GitHub tools were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b446ff59bf |
refactor(acpx-engine): coordinator-owned ACP run lifecycle with a typed resource ledger (#11576)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters run agent sessions through the ACPX engine > - The ACPX engine handled one run attempt as a long implicit procedure > - That shape made resource ownership, cleanup order, and failure behavior hard to verify > - This pull request gives the attempt a coordinator, a typed resource ledger, separate run sites, and explicit turn and settlement sequences > - The benefit is clear ownership, one cleanup path, safer session reuse, and testable failure behavior ## Linked Issues or Issue Description **What existing behavior does this improve?** The ACPX engine manages startup, turn execution, session reuse, and cleanup inside one large run procedure. **Current behavior** The run procedure owns several resources through implicit control flow. Cleanup and session reuse behavior depend on lane-specific branches and error paths. **Proposed behavior** The coordinator owns the run attempt. A typed ledger records six resources and their states. Host and sandbox run sites own lane-specific acquisition. Turn and settlement sequences expose typed outcomes. The engine emits allowlisted phase telemetry. **Reason and benefit** Explicit ownership makes cleanup and failure behavior easier to inspect. The fault matrix and characterization tests protect the external result while the refactor reduces hidden control flow. **Breaking changes** None to the public adapter contract. The host warm-save path now closes and relaunches the runtime because a transferred runtime could retain a run-scoped credential. A cold session-handshake failure now closes the created runtime. **Additional context** This pull request contains the ACPX engine lifecycle refactor, its tests, and the lifecycle document. ## What Changed - Add a run coordinator for startup, turn execution, settlement, and result reproduction. - Add a typed resource ledger with open, sealed, and consumed states. - Add host and sandbox run sites for lane-specific resource acquisition. - Replace separate runtime maps with a generic session reuse store. - Split session fingerprint identity from the outer session key. - Add typed turn and settlement sequences with one cleanup owner. - Add a closed allowlist for phase telemetry. - Add characterization tests and a 17-case fault matrix. - Add `doc/acp-run-lifecycle.md`. ## Verification - `npx vitest run packages/adapter-utils/src/acpx-engine/` passes 18 files and 286 tests at the submitted commit. - `pnpm --filter @paperclipai/adapter-utils typecheck` reports 0 errors at the submitted commit. - Run the full pull request checks after GitHub starts CI. - Run Greptile review after the pull request opens. ## Risks - The refactor changes internal control flow across the ACPX engine. - Host warm-save behavior now closes and relaunches the runtime. - Settlement changes the handling of a cold session-handshake failure from a leak to a close. - The characterization baselines and fault matrix reduce the risk of an external behavior change. > Paperclip is the open source app people use to manage AI agents for work > The adapter layer runs agent sessions through the ACPX engine > The engine needs explicit lifecycle ownership for reliable cleanup > This pull request adds coordinator-owned phases and a typed resource ledger > The result makes lifecycle behavior easier to test and review ## Model Used OpenAI GPT-5 Codex. Exact model ID: GPT-5. The model used tool execution, repository inspection, and code review support. The implementation author supplied the submitted code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1c366a9059 |
fix(server): reject invalid agent credentials instead of downgrading to the local user actor (#11589)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server authenticates each agent request in `actorMiddleware` before it attributes chat comments > - When an agent bearer token failed verification, the middleware called `next()` with no error and the request continued without an agent actor > - The request then fell back to the local user actor, so the server stored agent replies as user comments > - The task chat UI renders user comments in blue bubbles, so agent messages appeared as blue user bubbles > - This pull request rejects invalid agent credentials with 401 instead of a silent downgrade > - The benefit is that agent messages keep agent attribution, and broken credentials fail loudly with a clear retry message ## Linked Issues or Issue Description **What happened?** A user cancelled an onboarding question card. The agent posted a follow-up reply. The reply appeared in a blue bubble, which the UI reserves for human messages. The agent run held an expired local agent JWT. The auth middleware could not verify the token, called `next()` without an actor, and the request fell back to the local user identity. The server stored the agent comment as a user comment. **Expected behavior** Agent messages always render as agent bubbles. A request with invalid agent credentials must fail with 401 so the adapter can refresh credentials and retry. It must not post content under a human identity. **Steps to reproduce** 1. Start a local Paperclip instance. 2. Give an agent run an expired or malformed agent JWT. 3. Let the agent post an issue comment through the API bridge. 4. Before this change: the comment is stored with the local user identity and renders as a blue bubble. After this change: the request fails with 401 and a message that tells the caller to obtain fresh credentials. ## What Changed - `server/src/middleware/auth.ts`: a bearer token that fails verification now produces a 401 `unauthorized` error instead of a silent fall-through to the anonymous/local-user actor. - The 401 message states the cause: expired token, unverifiable token, empty bearer token, missing agent record, agent record in another company, terminated agent, or agent pending approval. - The API-key path now also rejects an agent record whose company does not match the key. - `packages/adapter-utils/src/execution-target.ts`: the bridge proxy now writes a `comment id: <id>` marker to the run log for each posted issue comment, so misattributed comments can be traced to a run. - `ui/src/components/task-chat/task-chat-adapter.test.ts`: a regression test asserts that a recovered `local-board` comment with a derived agent author renders as an agent bubble, not a user bubble. - `server/src/__tests__/agent-auth-middleware.test.ts` and `packages/adapter-utils/src/execution-target-sandbox.test.ts`: new tests cover each rejection path and the log marker. ## Verification - Run `pnpm vitest run src/__tests__/agent-auth-middleware.test.ts` in `server/` — 14 tests pass. - Run `pnpm vitest run execution-target-sandbox` at the repo root — 44 tests pass. - Run `pnpm vitest run src/components/task-chat/task-chat-adapter.test.ts` in `ui/` — 4 tests pass. - Manual check: post an issue comment with an expired agent JWT; the API returns 401 with a retry message and no comment is stored. ## Risks - Behavioral shift: requests that previously continued as anonymous or local-user actors after a failed agent-token verification now receive 401. Any caller that relied on the silent downgrade must refresh its credentials. This is the intended fix, and the adapters already handle 401 with a credential refresh. - No schema or migration changes. Low risk otherwise. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude (Anthropic), model ID `claude-fable-5`, via Claude Code with extended thinking and tool use (agent harness with shell, file, and git tools). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8087661bb8 |
fix: bound workspace Git scans (#11572)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Workspaces let users and agents inspect files that belong to an issue > - Changed-file views use full-tree Git status scans > - Many issue views could start those scans at the same time and make the server unresponsive > - Route-level limits did not protect the process or coalesce work for one repository > - This pull request adds one bounded scheduler for every expensive workspace Git scan > - It also starts browser scans only when the file panel is open and visible > - The benefit is bounded child-process use and responsive health checks during request storms ## Linked Issues or Issue Description **What happened?** Many changed-file requests could start full `git status --porcelain=v1 -z --untracked-files=all` scans at the same time. One production incident produced about 270 direct Git child processes. The Node process stayed alive but stopped answering health requests in time. **Expected behavior** Paperclip must bound expensive Git work across all companies, actors, issues, repositories, and browser tabs. Duplicate requests for one worktree must share work. Excess requests must fail fast with a retryable response. Hidden or closed file panels must not start scans. **Steps to reproduce** 1. Open changed-file views for many issue and actor keys. 2. Send requests for two large workspace roots at the same time. 3. Observe that route-level limiter keys allow many full Git scans to run together. 4. Observe delayed health responses and accumulated Git children. **Paperclip version or commit** Reproduced on master before commit `43ab441f0f`. **Deployment mode** Self-hosted server with local workspace repositories. ## What Changed - Add a process-wide scheduler with configurable concurrency, queue capacity, timeout, and cache TTL. - Add fair admission, a bounded queue, canonical worktree keys, single-flight joins, and bounded result caching. - Add subprocess timeouts, TERM-to-KILL escalation, bounded output, waiter cancellation, and slot cleanup. - Route full-tree status work from file resources, workspace runtime, execution workspaces, and adapter overlay sync through the scheduler. - Return stable retryable `503` and `504` error codes for saturation and timeout. - Add structured logs with safe workspace hashes, durations, queue state, cache use, joins, and terminal outcomes. - Gate UI queries on panel and document visibility. Cancel queries on close, hide, unmount, and workspace change. - Disable focus and reconnect bursts. Keep one explicit refresh action and a retryable unavailable state. - Document the 10-second default freshness tradeoff and all configuration variables. - Add unit, route, UI, adapter, and deterministic 500-request load coverage. ## Verification - `pnpm -r typecheck` - `pnpm build` - `pnpm check:token-gates` - `pnpm --filter @paperclipai/server exec vitest run src/services/workspace-git-operation-scheduler.test.ts src/__tests__/file-resources-git-scan-load.test.ts --reporter=dot` — 16 tests passed. - `pnpm --filter @paperclipai/ui exec vitest run src/components/WorkspaceFileBrowser.test.tsx src/lib/page-visibility.test.ts --reporter=dot` — 38 tests passed. - `pnpm --filter @paperclipai/adapter-utils exec vitest run src/git-workspace-sync.test.ts --reporter=dot` — 16 tests passed. - Existing file-resource, workspace-runtime, and execution-workspace regression selections passed. - Two cleanup safety regressions prove failed scans preserve the worktree before archive and at the final deletion fence. - Before: the incident produced about 270 Git children and health requests timed out. - After: 500 concurrent requests across 500 issue keys, 73 actors, and two roots started two underlying scans. Peak scan concurrency was 2. All 500 requests succeeded. Health p99 was 4.94 ms. The harness found zero unreaped children. - The full local Vitest run passed 4,267 tests. Ten existing fixed-port HTTPS exposure tests could not run because this host already owns Tailnet listeners on ports 42000 and 52000. Clean GitHub CI is the final full-suite result. - Latest-head GitHub CI passed all required test, typecheck, build, canary, e2e, policy, and security gates. - Greptile completed at 5/5 with zero unresolved comments, recommendations, or follow-ups. ## Risks - Changed-file results can be up to 10 seconds old by default. Explicit refresh remains available. - A full queue returns a retryable `503` instead of waiting without a bound. - A scan that exceeds the default 8-second deadline returns a retryable `504` and terminates its process group. - Operators can tune all limits with documented environment variables. Safe defaults protect local and shared servers. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5 family. The runtime does not expose the exact deployment ID or context-window size. High reasoning, tool use, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
48f4ae16ac |
fix(codex-local): keep a promoted device-login credential when re-seeding the managed home (#11578)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The codex_local adapter supports a device login that runs in a trusted sandbox and promotes the credential into the per-company managed Codex home > - The managed-home re-seeding step treats every regular-file `auth.json` as apikey-mode residue and removes it so the shared-home symlink can be restored > - The promotion writes the company credential as a regular file, so the first environment Test or run after a successful login deletes it > - On a server with no shared Codex login — any containerized deployment — nothing replaces the file, and the UI reports that the sandbox has no ready authentication right after it reported a successful login > - This pull request makes the cleanup identity-anchored: a subscription credential whose identity the shared source does not hold survives re-seeding > - The benefit is that a device login stays usable after Test and runs, on hosts with and without a shared Codex login ## Linked Issues or Issue Description No public GitHub issue covers this. The problem is described in-PR following the bug template. Related public PRs: [#11237](https://github.com/paperclipai/paperclip/pull/11237) added the sandbox device login and the credential promotion, [#11097](https://github.com/paperclipai/paperclip/pull/11097) added its building blocks, and the `ensureSymlink` heal for stale copies came from the fix for #5028. [#9621](https://github.com/paperclipai/paperclip/pull/9621) touches the adjacent sandbox auth sync-back lane but not this defect. **Subsystem affected** packages/adapters/codex-local — managed `CODEX_HOME` seeding (`codex-home.ts`). **Current behavior** A successful device login promotes the subscription `auth.json` into the company Codex home as a regular file, and the UI reports the login as authenticated. The next `seedManagedCodexHome` call — the environment Test probe and every execute both run it — removes any regular-file `auth.json` when no API key is configured, because the cleanup assumes such a file is apikey-mode residue left by a previous run. It then symlinks `auth.json` from the shared source home. On a server whose shared home has no Codex login (a container image, for example), there is no source to symlink, so the home ends with no credential at all. The Test probe then reports "The sandbox has no ready authentication for this adapter" immediately after a successful login, and a fresh login repeats the same cycle. On a server whose shared home does hold a login, the symlink silently replaces the promoted account with the host account. **Expected behavior** The credential a device login promoted stays in the company home across Test probes and runs. The #5028 heal (a stale regular-file copy of the shared credential becomes a symlink to the live source) and the apikey-residue cleanup keep working. **Steps to reproduce** 1. Run the server in an environment whose shared Codex home (`$CODEX_HOME` or `~/.codex`) has no `auth.json`. 2. Complete a Codex device login for a company; the promotion writes the company home `auth.json` and the UI reports authenticated. 3. Click Test on a codex_local agent (or start a run). The probe reports no ready authentication, and the promoted `auth.json` is gone from the company home. **Proposed solution** Make the cleanup identity-anchored, the same rule the promotion and the cache vend already use. A regular-file `auth.json` survives re-seeding when it holds a usable subscription identity that the shared source does not also hold, and the shared symlink does not replace it. A same-identity regular file is still the #5028 stale copy and is still healed into the symlink, because the symlink serves the same account with live, rotating tokens. An apikey-mode or unreadable file is still removed. ## What Changed - `seedManagedCodexHome` reads the target `auth.json` before the cleanup and keeps it when `readSubscriptionAccountId` yields an identity the shared source `auth.json` does not hold. The kept file is excluded from the shared symlink pass, and the function logs a fixed line when it keeps the file. - The function doc comment states the kept-promoted-credential rule. - Four new `seedManagedCodexHome` test cases: a promoted credential with no shared auth, a promoted credential with a different shared identity, the same-identity #5028 heal, and apikey-mode residue removal. ## Verification ```sh cd packages/adapters/codex-local npx tsc --noEmit # clean npx vitest run # 28 files, 321 passed, 1 skipped ``` The four new cases fail on the previous code: the first two observed the promoted file deleted (and, with a shared login present, replaced by the shared symlink). ## Risks Low risk. The change narrows one deletion path. Deployments that never use the device login see no difference: without a promoted subscription file, the cleanup and the symlink behave exactly as before, and the #5028 heal is pinned by an existing test plus a new same-identity test. The one deliberate behavioral shift: after a device login, the promoted company credential now stays authoritative over the shared host login for that company — which is the promotion's documented contract ("the company credential slot"). ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, with tool use and code execution — investigation, implementation, and tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
7ef75f5636 |
fix(opencode-local): retry models preflight during transient contention (#9225)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Local CLI adapters are responsible for starting agent runtimes and validating that their configured models are usable before a run starts. > - The OpenCode local adapter checks `opencode models` during model discovery and preflight validation. > - On hosts with a shared Ollama daemon, that lightweight metadata call can transiently queue behind an active generation and time out or return a short failure. > - Treating that transient contention as a hard adapter failure prevents otherwise valid local OpenCode runs from starting. > - This pull request adds a small bounded retry/backoff around OpenCode model discovery while keeping the existing per-attempt timeout and surfacing a final failure when retries are exhausted. > - The benefit is fewer false adapter failures during local Ollama contention without changing shared Ollama configuration or hiding genuinely stuck model discovery. ## Linked Issues or Issue Description No public GitHub issue exists for this adapter reliability bug. Bug description: - What happened: `opencode models` can transiently time out or fail while a shared local Ollama daemon is busy serving another OpenCode generation, causing the adapter preflight to fail before the actual run starts. - Expected behavior: transient model-list contention should be retried briefly before declaring the adapter unavailable. - Steps to reproduce: run an OpenCode local adapter using an Ollama-backed model while another `opencode run` is actively generating against the same daemon, then trigger model discovery/preflight during that contention window. - Paperclip version/commit: observed on the current Paperclip master-line OpenCode local adapter before this change. - Deployment mode: local trusted / local CLI adapter execution with a shared local Ollama daemon. Related search: - Searched public GitHub issues for `opencode models preflight retry`; no matching issue found. - Searched public GitHub PRs for `opencode models preflight retry`; no matching PR found. The only search hit was unrelated OpenClaw gateway authentication work (#6121). ## What Changed - Added bounded retry/backoff to OpenCode model discovery: three total attempts with 2s and 4s waits between failures. - Preserved the existing 20s per-attempt `opencode models` timeout. - Retry covers timeout and non-zero process exits, while spawn-level failures still surface immediately. - Added unit coverage for transient fail -> timeout -> success behavior and exhausted retry behavior. - Updated existing OpenCode environment diagnostic tests with explicit timeouts for the intentional retry/backoff path. ## Verification - `pnpm --filter @paperclipai/adapter-opencode-local exec vitest run src/server/models.test.ts src/server/execute.test.ts` -> 2 files passed, 13 tests passed. - `pnpm --filter @paperclipai/adapter-opencode-local typecheck` -> passed. - `pnpm vitest run server/src/__tests__/opencode-local-adapter-environment.test.ts` -> 1 file passed, 3 tests passed. - Branch diff against current `upstream/master` is limited to `packages/adapters/opencode-local/src/server/models.ts`, `packages/adapters/opencode-local/src/server/models.test.ts`, and `server/src/__tests__/opencode-local-adapter-environment.test.ts`. ## Risks Low risk. This only changes OpenCode model discovery behavior and keeps the preflight bounded. A genuinely unavailable `opencode models` call still fails after three attempts, and command spawn failures are not masked. ## Model Used OpenAI Codex, GPT-5.5 coding agent, tool-enabled repository editing and shell verification in a local Paperclip workspace. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Test <test@paperclip.ing> |
||
|
|
d77eeb8914 |
fix(sandbox-bridge): allow the agent-hire skill's routes through the callback bridge (#8978)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Managed agents run inside a sandbox and reach the Paperclip server only through the sandbox callback bridge, which forwards a fixed route allowlist (`DEFAULT_SANDBOX_CALLBACK_BRIDGE_ROUTE_ALLOWLIST`) > - The `paperclip-create-agent` skill instructs an agent to call adapter/icon discovery endpoints, compare existing agent configurations, submit a hire request, and link the resulting approval to its source issue > - None of those routes were on the bridge allowlist, so a sandboxed agent following the skill correctly hit `Route not allowed` on every call — including the hire `POST` itself — making hiring impossible from inside a sandbox > - This pull request adds the six routes the skill uses to the bridge allowlist, while keeping direct agent creation (`POST /api/companies/:id/agents`) denied > - The benefit is that hiring works end-to-end for sandboxed agents through the approval-gated `agent-hires` path, without widening the bridge beyond what the skill needs ## Linked Issues or Issue Description No public issue exists; describing the bug in-PR (bug template fields): - **What happened:** A managed agent running in a sandbox followed the `paperclip-create-agent` skill and got `Route not allowed` from the callback bridge on every endpoint the skill documents — adapter discovery (`/llms/agent-configuration.txt`, `/llms/agent-configuration/:adapterType.txt`, `/llms/agent-icons.txt`), config comparison (`GET /api/companies/:id/agent-configurations`), the hire submission (`POST /api/companies/:id/agent-hires`), and approval linking (`POST /api/issues/:id/approvals`). - **Expected behavior:** An agent with hiring permission can complete the hire flow from inside a sandbox; the bridge forwards the skill's routes and the server enforces authorization (`canCreateAgents`). - **Impact:** Hiring by sandboxed agents was fully broken — the failure is in the transport allowlist, not permissions, so no configuration could work around it. Related: #8981 (companion fix making the `paperclip-create-agent` skill available to agents that can hire; supersedes #8823). The two changes serve the same end-to-end hire flow but are independently mergeable — this PR is purely the bridge transport allowlist. Supersedes #8853. ## What Changed - `packages/adapter-utils/src/sandbox-callback-bridge.ts`: add six routes used by the `paperclip-create-agent` skill to `DEFAULT_SANDBOX_CALLBACK_BRIDGE_ROUTE_ALLOWLIST` (three `GET /llms/...` discovery routes, `GET .../agent-configurations`, `POST .../agent-hires`, `POST /api/issues/:id/approvals`), with a comment documenting why direct agent creation stays denied - `packages/adapter-utils/src/sandbox-callback-bridge.test.ts`: assert the six routes are allowed, and add negative cases proving the regexes do not over-match (no `POST .../agents`, no non-`.txt` or arbitrary `/llms` files, no `agent-hires` sub-resources) ## Verification - `npx vitest run packages/adapter-utils/src/sandbox-callback-bridge.test.ts` — 13/13 tests pass locally - `npx tsc --noEmit -p packages/adapter-utils` — clean - Manual: run a managed agent in a sandbox, invoke the `paperclip-create-agent` skill, and confirm the discovery calls, hire `POST`, and approval linking all pass through the bridge; `POST /api/companies/:id/agents` still returns `Route not allowed` ## Risks - Low risk: additive allowlist entries only; anchored regexes with `[^/]+` segments prevent over-matching (covered by tests) - The bridge allowlist bounds surface area but does not replace server-side authorization — the hire `POST` remains approval-gated and permission-checked (`canCreateAgents`) on the server > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude (Anthropic) — Fable 5 (`claude-fable-5`), extended thinking, agentic tool use via Claude Code; original diff authored with Claude Opus 4.8 (1M context) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
93ff6a8771 |
fix(cursor-cloud): drop unreachable Paperclip API callback for remote… (#8546)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents run via per-adapter execute paths; the `cursor_cloud` adapter runs the agent in Cursor's cloud (remote), orchestrated server-side via the Cursor Agent SDK > - Local adapters receive a run-scoped Paperclip JWT (`supportsLocalAgentJwt=true`) injected as `PAPERCLIP_API_KEY` so the agent can call the Paperclip API; `cursor_cloud` is intentionally `supportsLocalAgentJwt=false` (no JWT minted for a remote worker) > - But `buildPaperclipEnv` always sets `PAPERCLIP_API_URL` (defaulting to the local runtime host), so the remote cloud worker is handed a callback URL it can neither reach nor authenticate against > - Any agent-initiated Paperclip API call from the cloud worker therefore fails with a 401 (or is unreachable), producing log noise and confusing failures > - This pull request drops the callback wiring when there is no usable key, so cloud-side Paperclip tools degrade to a clean no-op > - The benefit is no spurious 401s from remote cloud runs, with run results unaffected (delivered server-side via the Cursor Agent SDK) ## Linked Issues or Issue Description No existing public issue — describing the bug inline (per `.github/ISSUE_TEMPLATE/bug_report.yml`): **What happened** `cursor_cloud` runs emit 401s when the remote cloud agent attempts Paperclip API calls. Root cause: `buildPaperclipEnv` (`packages/adapter-utils/src/server-utils.ts`) always sets `PAPERCLIP_API_URL` (local runtime default), while `cursor_cloud` has `supportsLocalAgentJwt=false`, so no `PAPERCLIP_API_KEY` is minted — URL present, key absent → 401 / unreachable from `buildWakeEnv` in `packages/adapters/cursor-cloud/src/server/execute.ts`. **Expected behavior** A remote cloud worker that is not issued a run JWT should not attempt (and fail) Paperclip API callbacks. **Steps to reproduce** 1. Configure a `cursor_cloud` agent (runs in Cursor's cloud; `supportsLocalAgentJwt=false`). 2. Trigger a run that causes the cloud agent to make a Paperclip API call. 3. Observe a 401 (or unreachable) because `PAPERCLIP_API_URL` points at an unreachable local runtime and no key is present. **Paperclip version** Reproduced on current `master` (cutover base `e68188c43`). **Deployment mode** Self-hosted control plane, `cursor_cloud` adapter (remote execution in Cursor's cloud). **Related PRs (searched; none duplicate this fix):** - #8197 — `claude_local` opt-out of the sandbox *bridge* for direct-reachable remote SSH targets. Related family, but the opposite situation: that path keeps the callback because the remote is reachable **and** has a run token. `cursor_cloud` has neither, so here the callback is removed. - #8130, #4794, #8025 — `PAPERCLIP_API_URL`/loopback injection for **local** agents (distinct from the remote cloud worker case). - #401 — alternative agent-auth scheme (run-ID header when no bearer token); different approach, not overlapping with this targeted fix. ## What Changed - `packages/adapters/cursor-cloud/src/server/execute.ts`: in `buildWakeEnv`, when there is no usable `PAPERCLIP_API_KEY`, delete `PAPERCLIP_API_URL` and `PAPERCLIP_API_BRIDGE_MODE` so the remote worker performs no Paperclip API callbacks. Informational `PAPERCLIP_*` vars (run id, agent id, company id, task, wake reason) still flow. When a key *is* present (operator-provided), the URL is retained. - `packages/adapters/cursor-cloud/src/server/execute.test.ts`: new test asserting no callback vars are injected when no run JWT is present; positive assertion that the URL is retained when a key is present. ## Verification - `pnpm exec vitest run packages/adapters/cursor-cloud/src/server/execute.test.ts` → **5/5 pass**. - `pnpm --filter @paperclipai/adapter-cursor-cloud typecheck` → **green**. - Confirmed result delivery does not depend on this callback: `execute()` reads results server-side via `Agent.getRun()` and `run.wait()`. ## Risks - **Low risk.** Only affects the env handed to remote `cursor_cloud` workers. No schema/migration/behavioral change to result delivery (which is server-side). When an operator explicitly provides `PAPERCLIP_API_KEY`, the callback URL is retained, preserving intentional callback setups. ## Model Used - **Claude Opus 4.8** (Anthropic), extended/high reasoning mode, via the Cursor agent with tool use + code execution. Diagnosis grounded in the adapter/runtime code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (none found) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change (`fix/cursor-cloud-skip-unreachable-callback`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes (N/A — no documented behavior changes) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending CI run) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending review) - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Sebastian Heyneman <sebastian@joinnova.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
14fd8aee36 |
fix(server): flag truncated issue descriptions (#4771)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies. > - The issue list API is one of the surfaces API consumers use to synchronize issue metadata. > - The list endpoint intentionally returns a bounded `description` preview so large descriptions do not bloat list responses. > - Before this change, that preview looked like a complete field value because the response did not say whether it had been shortened. > - That made round-trip clients vulnerable to accidentally PATCHing a preview back over the full description. > - This pull request keeps the existing preview behavior but adds an explicit `descriptionTruncated` flag. > - The benefit is backwards-compatible visibility into truncated issue descriptions, so clients can avoid data-loss workflows. ## Linked Issues or Issue Description Fixes #4758. Related PR: #4792 also targets #4758, but it includes unrelated logger changes and currently has separate review/security concerns. This PR keeps the fix scoped to the issue-list description truncation API behavior. ## What Changed - Added `descriptionTruncated` to the issue list projection when `description` exceeds the existing 1200-character preview limit. - Exposed `descriptionTruncated?: boolean` on the shared `Issue` type. - Added service tests for truncated descriptions, exact-limit descriptions, null descriptions, and multibyte-safe preview truncation. ## Verification June 18, 2026 refresh after rebasing onto current `origin/master`: - `pnpm install --frozen-lockfile` - `pnpm exec vitest run server/src/__tests__/issues-service.test.ts` - `pnpm --filter @paperclipai/server typecheck` - `pnpm typecheck` - `git diff --check origin/master...HEAD` - GitHub PR checks are green on head `12e828e6`. Earlier pre-review verification also included `pnpm test`. ## Risks - Low risk. This is an additive API response field; existing clients can ignore it. - The list endpoint still returns the same bounded `description` preview. Clients that need full text should continue fetching the issue detail, but can now detect when that is necessary. - No database migration or UI behavior change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex based on GPT-5, via Codex desktop on April 29, June 15, and June 18, 2026. Used tool-assisted repository inspection, code editing, local test execution, GitHub CLI workflows, and PR review follow-up. Exact context window size is not surfaced by the tool. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots (N/A: no UI change) - [x] I have updated relevant documentation to reflect my changes (N/A: additive API field covered by tests) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Sami Rusani <sr@samirusani> |
||
|
|
c1c46f1e4e |
feat: Claude login on the new-agent page before agent creation (#11347)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Claude local adapter supports subscription login through a sandbox > - The new-agent page must show login before the user creates an agent > - Test results must not expose raw sandbox diagnostics or secret values > - This pull request adds the login UI to both Test lanes and closes the diagnostic boundary > - The branch also adds durable cleanup recovery for failed sandbox teardown > - Reusable sandboxes must retain both their recorded teardown configuration and a valid lifecycle path until destruction succeeds > - The benefit is a usable login flow with fixed public checks, redacted server logs, and recoverable sandbox cleanup ## Linked Issues or Issue Description Related public work: [#9488](https://github.com/paperclipai/paperclip/pull/9488) adds first-class recognition for `CLAUDE_CODE_OAUTH_TOKEN` in headless and remote runs. Related public issue: [#2681](https://github.com/paperclipai/paperclip/issues/2681) requests Claude Code subscription support. This pull request adds the login transport and new-agent UI flow that those changes do not provide. **Subsystem affected:** Claude local adapter, server login probes, sandbox provider setup, cleanup recovery, and the new-agent UI. **Problem or motivation:** The Test lanes did not show the sandbox login panel in all supported cases. Test results also exposed raw probe diagnostics, and JSON escapes could end secret redaction early. **Proposed solution:** Surface the login capability through the bundled provider manifest. Prepare the same probe runtime in the ACP lane. Send diagnostics only to redacted server logs. Keep Test checks on fixed public messages. Normalize login URL hints to allowlisted HTTPS Claude and Anthropic hosts. Consume JSON escapes during redaction. Preserve failed sandbox cleanup state across retries and restarts, and prevent deletion from severing the lifecycle context of a live reusable sandbox. **Alternatives considered:** Keep raw diagnostics in Test checks or trust login URL text from the sandbox. Both choices increase information exposure. Keep separate probe behavior in the ACP lane. That choice would leave the two Test lanes inconsistent. ## What Changed - Surface the sandbox login panel on both Test lanes. - Reconcile the bundled Daytona plugin manifest so `supportsSetupTokenLogin` reaches the UI capability gate. - Prepare the ACP Test lane with the same probe runtime as the CLI Test lane. - Add the `claude_acp_login_probe_unavailable` warning when the ACP probe cannot run. - Send raw sandbox diagnostics only to redacted server logs. - Keep Test checks on fixed public messages in the ACP, managed-config, and CLI paths. - Normalize login URL hints to allowlisted HTTPS Claude and Anthropic hosts. - Redact JSON and escaped-JSON secret values, including escaped quotes and backslashes. - Preserve orphan cleanup records across provider failures, restarts, and unavailable plugins. - Atomically block environment deletion while a live reusable sandbox lease still depends on it. - Verify pending cleanup destroys plugin sandboxes with the provider configuration recorded on the lease, even after the current environment configuration changes. ## Verification - Head under review: `506b7fa2d83c36bfa5fd722ee9d95b0c7431c241`. - Focused environment route/service/runtime coverage passes: 196 tests across 3 files. - `pnpm -r typecheck` passes. - `pnpm build` passes. - The full Vitest run completed with 4,754 passing and 28 failing tests. All 23 source-test failures reproduce unchanged on parent head `58cfe61a33191ce03d965d65085d26064b4888ba`; the other 5 are duplicate executions from stale `server/dist` output. The failures are unrelated macOS path/listener and scheduler-fixture failures, so there is no new bad commit for bisect to localize. - All required CI checks pass for the current head, including build, typecheck/release registry, all server and workspace shards, serialized server suites, canary, and e2e. - A fresh Greptile review for `506b7fa2d83c36bfa5fd722ee9d95b0c7431c241` reports 5/5, “safe to merge,” with no blocking failure remaining. ## Risks - A probe or redaction change could hide useful server diagnostics. - An allowlist change could reject a valid Claude login URL. - Cleanup recovery changes could affect provider teardown ordering. - An environment with a live reusable sandbox can no longer be deleted until the owning issue or execution workspace completes teardown. - The implementation keeps public Test messages fixed and sends detail to redacted server logs. ## Model Used OpenAI GPT-5 via Codex — exact model ID: GPT-5; tool use and code execution enabled; extended reasoning enabled. The implementation author used AI-assisted development. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and documented the result - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation or confirmed no separate documentation change is needed - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3061ce6901 |
feat(sandbox): stream session output by capability, drop three operator flags (#11557)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandboxed agents use provider capabilities to select safe execution paths > - Session output still depends on three operator flags that duplicate capability data > - Duplicate flags can drift from the verified sandbox capability snapshot > - This pull request makes the capability snapshot the only streaming decision and removes the obsolete flags > - The benefit is default streaming with a poll fallback when a capability or stream fails ## Linked Issues or Issue Description **What existing behavior does this improve?** ACP sandbox session-output streaming and sandbox execution configuration. **Subsystem affected** Cross-cutting (multiple of the above): server/, packages/shared/, packages/adapter-utils/, and packages/plugins/. **Current behavior** Session-output streaming requires operator flags in the server and Daytona plugin configuration. Saved configurations can retain a removed key. **Proposed behavior** The verified capability snapshot selects streaming. The Daytona plugin uses persistent sessions by default, keeps bypass commands one-shot, and falls back from the log stream to polling. Removed configuration keys become inert. **Reason and benefit** One capability source prevents configuration drift. The fallback keeps output available when capability resolution or log streaming fails. **Breaking changes** The three operator flags no longer control session-output streaming. Existing saved keys load but have no effect. ## What Changed - Remove `useSessions` and `useLogStream` from the Daytona plugin configuration and manifest. - Remove `streamAgentSessionOutput` from server configuration, shared types, and execution-target plumbing. - Select streaming from `persistentProcessSessions` and `independentControlCommands`. - Keep poll fallback on capability resolution failure and stream failure. - Strip removed keys from strict fake-sandbox and catchall plugin configuration. - Update the sandbox capability documentation and focused tests. ## Verification - `tsc --noEmit` passed in `packages/shared`, `packages/adapter-utils`, `server`, and the Daytona plugin. - Daytona `plugin.test.ts` passed 139 tests. - Server capability, configuration, route, and runtime suites passed 160 tests. - `packages/adapter-utils` `execution-target-sandbox.test.ts` passed 44 tests. - The capability matrix covers stream, poll, and resolution-failure paths. - Removed-key tests cover strict fake-sandbox and catchall plugin schemas. ## Risks - A capability snapshot that lacks either required session capability uses polling. - A log stream failure uses polling and can increase request count. - Existing removed configuration keys no longer change behavior. - The isolated-worktree Daytona Vitest run has a pre-existing missing `packages/adapters/droid-local` reference. CI and standard checkouts use the committed configuration. ## Model Used OpenAI Codex, GPT-5, tool use and code review assistance. The exact runtime context window is managed by the Codex platform. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e71ce9a9d3 |
feat: sandbox provider capability contract with fail-closed effective resolution (#11463)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs work through adapters and sandbox providers > - Providers need a clear contract so the server can use only verified capabilities > - A declared capability must not grant a method that the live worker did not verify > - This pull request adds manifest declarations and fail-closed effective capability resolution > - The benefit is safe provider reuse across execution targets and run lifecycles ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** Sandbox providers expose different runtime methods. The server needs one safe capability contract that accounts for provider declarations, worker verification, and narrowing configuration. **Proposed solution** Add strict manifest validation for five sandbox capabilities. Resolve effective capabilities as the subset of verified, declared, and narrowed values. Store the result as a frozen execution-target snapshot. **Alternatives considered** Trusting the manifest alone could grant methods that the worker does not support. Trusting only a fixed built-in list would reject valid third-party providers. The intersection rule keeps the verified runtime ceiling and supports both provider types. **Roadmap alignment** This change supports the ACP run lifecycle track and the sandbox provider contract work in the current roadmap. **Additional context** The legacy `supportsReusableLeases` field remains supported. The nested capability validator rejects unknown keys. Missing or unavailable verification resolves all capabilities to `false`. ## What Changed - Add strict `sandboxCapabilities` manifest validation with legacy reusable-lease compatibility. - Carry declarations through the ready-driver projection. - Add fail-closed effective resolution from verified, declared, and narrowed capabilities. - Add narrowing for provider configuration, Kubernetes Job leases, and Daytona sessions. - Add a frozen read-only capability snapshot to execution targets. - Add focused tests and keep existing characterization baselines covered. - Add and update sandbox provider capability documentation. ## Verification - `npx vitest run packages/shared/src/validators/plugin.test.ts` - `npx vitest run server/src/__tests__/plugin-environment-driver-sandbox-capabilities.test.ts` - `npx vitest run server/src/__tests__/sandbox-capability-contract.test.ts` - `npx vitest run server/src/__tests__/environment-execution-target-capabilities.test.ts` - `npx vitest run packages/adapter-utils/src/acpx-engine/startup-characterization.test.ts packages/adapter-utils/src/acpx-engine/turn-characterization.test.ts packages/adapter-utils/src/acpx-engine/settlement-characterization.test.ts packages/adapter-utils/src/acpx-engine/composed-run-characterization.test.ts` - Package typechecks for shared, server, and adapter-utils pass. - Stage-2 security review suites pass with 28 tests. ## Risks The resolver fails closed when verification is absent or unavailable. Providers that rely on undeclared capabilities may see narrower behavior until they expose verified worker methods. The change does not alter the existing native-sync guard. ## Model Used OpenAI Codex, GPT-5, exact runtime model ID `gpt-5`, tool use and code execution. The implementation author used this model to assist with the change. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f4802b1bbc |
feat(runtime-exposure): least-privilege Tailscale HTTPS broker, shared contract, and persisted exposure state (#11524)
<!-- Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip starts and supervises managed runtime services for a project's execution workspaces, so an agent's branch can be previewed while it works > - Those services only listen on plain loopback HTTP. A person on another device, or on a phone, cannot open the preview > - A Tailscale HTTPS mapping solves this, but `tailscale serve` needs host privileges that the Paperclip server process must not hold > - This pull request adds the foundation only: a separate least-privilege host broker, the shared exposure contract, and the database columns that hold exposure state > - Nothing calls the broker yet, so there is no behavior change. The benefit is that the privileged surface is small, reviewable, and isolated before any lifecycle code depends on it ## Linked Issues or Issue Description No public GitHub issue exists. The change follows the feature request template. **Subsystem affected** Managed workspace runtime services, the shared type and validator package, and the database schema. **Problem or motivation** A managed runtime service binds to loopback only. There is no supported way to reach that preview from another device. Adding HTTPS directly to the server would mean the server process runs `tailscale serve`, which needs privileges far wider than the task requires. A compromised or buggy server could then map any port to the tailnet. **Proposed solution** Split the privileged work into a separate broker process with a narrow protocol, and define one shared contract that the server, the UI, the runtime, and the broker all read. Land this foundation first, with no caller, so the privileged code can be reviewed on its own. **Alternatives considered** - Call `tailscale serve` from the server process. This was rejected because it gives the server unrestricted mapping authority. - Use `sudo` for single `tailscale` commands. This was rejected because the argument list is the only guard, and it is easy to widen by accident. - Use a generic reverse proxy. This was rejected because it does not remove the need for a privileged Tailscale mapping step. **Roadmap alignment** This supports the existing managed workspace runtime capability. It adds no new product surface on its own. **Additional context** The broker is the security boundary of the feature, so it is deliberately the first slice. Three later pull requests build on it: the server exposure lifecycle, the runtime lease and recovery integration, and the leased-port mediator. ## What Changed - Add the `@paperclipai/tailscale-https-broker` workspace package. The broker listens on a unix socket, authorizes each peer with `SO_PEERCRED`, and answers a small request protocol. - Restrict what the broker will map. It accepts only same-number HTTPS-to-loopback pairs inside the Paperclip port range, refuses protected ports, and confirms that the loopback port belongs to a Paperclip-owned listener. - Parse every request with a strict JSON reader that rejects duplicate keys, prototype keys, and unknown fields. - Write an append-only audit record for each broker decision. - Add the shared exposure contract in `@paperclipai/shared`: the `RuntimeExposureConfig`, `RuntimeExposureState`, and `RuntimeExposureStatus` types, their zod validators, the app and HMR port rules, and the loopback-bind helpers. - Persist exposure state on `workspace_runtime_services` with the new `exposure` column, plus the server-private `exposure_handle` and `backend_url` columns that are never serialized to API clients. - Add the `execution_workspace_runtime_leases` table that the later lease slice uses. - Extend the runtime read-model test fixture for the three new columns. ## Verification Focused checks, all run on this branch: - `pnpm --filter @paperclipai/tailscale-https-broker test` — 12 files, 82 tests pass. This covers peer credentials, port policy, protected ports, the serve config writer, the strict JSON reader, argv parsing, and the socket server. - `pnpm --filter @paperclipai/tailscale-https-broker typecheck` — clean. - `npx vitest run --root packages/shared src/runtime-exposure src/validators/runtime-exposure.test.ts` — 3 files, 40 tests pass. - `pnpm --filter @paperclipai/db typecheck` — runs `check:migrations` first. Migration numbering and migration safety both pass. - `pnpm --filter @paperclipai/shared typecheck` — clean. - `pnpm --filter @paperclipai/ui typecheck` — clean. - `npx vitest run --root server src/services/workspace-runtime-read-model.test.ts` — 3 tests pass. - `npx tsc --noEmit -p server/tsconfig.json` — 139 errors, which is exactly the count on `master` before this branch. All 139 come from the unbuilt `@paperclipai/plugin-sdk` package. To confirm the exposure state is inert, start a managed runtime service as usual. The new columns stay null and the service behaves as it does today. ## Risks - Migration risk is low. Both migrations only add a table and three nullable columns. No column is backfilled and no existing column changes. The migration safety check passes. - Behavior risk is low. No code path calls the broker in this pull request, and the shared exposure fields are optional. - The broker is privileged, so it is the real risk surface. It is mitigated by peer-credential authorization, a fixed port range, a protected-port deny list, same-number pair enforcement, listener-ownership checks, strict JSON parsing, and an audit trail. Reviewers should read `packages/tailscale-https-broker/src/authorization.ts` and `src/port-policy.ts` closely. - The broker requires a `tailscale` version floor, which its README records. An older host CLI makes the broker refuse to start rather than map incorrectly. - `pnpm-lock.yaml` changes because a new workspace package is added. The diff is the new importer block, plus one duplicate `tinyexec` entry that pnpm removed. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. ## Model Used Claude Opus 5 (`claude-opus-5`), 1M context window, extended thinking, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
10d0555189 |
fix(interactions): authorize resolvers consistently (#11376)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Issue interactions give agents and people a structured decision record. > - Resolver routes used different authorization rules. > - Some routes blocked valid agents, including task watchdogs with normal issue access. > - The API did not show who could resolve a pending interaction. > - This pull request gives every interaction kind one resolver policy evaluator. > - The benefit is a clear decision path with consistent governance and company isolation. ## Linked Issues or Issue Description Fixes: #8087 Refs: #7403 Related PR: #11082 proposes board-only confirmation rules. This change keeps human-only review as an explicit policy. **What happened?** Agents could create issue interactions. Some resolver routes still required board access. This left valid agent confirmations pending. Task watchdogs could see the same problem without board identity. **Expected behavior** Every interaction kind must use one resolver policy contract. The contract must support `anyone`, `not_creator`, and `human_only`. It must also apply all normal governance controls. **Steps to reproduce** 1. Create a `request_confirmation` interaction as an agent. 2. Resolve it with another authorized agent. 3. Observe the board-only denial. **Paperclip version or commit** The problem exists on `master` before this change. **Deployment mode** Local development with `pnpm dev`. ## What Changed - Add canonical policies for `anyone`, `not_creator`, and `human_only`. - Use one server evaluator for every interaction kind. - Apply named addressees, company limits, review rules, and task watchdog scope. - Charge cross-issue resolutions to the existing per-run action limit. - Return the effective resolver audience in attention and interaction data. - Show the audience, governance choices, and denial reasons in the board UI. - Add telemetry, API documents, product documents, and regression fixtures. - Add migration provenance for safe legacy behavior. - Make migration `0218` safe for complete replays and partial prior runs. ## Product Rules - An interaction records a response. It does not grant authority for the next action. - `anyone` lets any authorized issue participant respond. - `not_creator` requires a responder other than the interaction creator. - `human_only` requires an authorized person. - A named addressee, company policy, or governed action can narrow the audience. - These controls cannot widen the audience. - A task watchdog uses the same rules as an ordinary agent. - A task watchdog does not receive board authority. - An agent resolution on another issue uses the shared cross-issue action limit. - Legacy pending interactions keep their earlier restrictions. - The UI shows the effective audience and a permanent denial reason. ## Verification - `pnpm --filter @paperclipai/db check:migrations` - `pnpm --filter @paperclipai/db typecheck` - `pnpm exec vitest run packages/db/src/issue-thread-interaction-resolver-policy-migration.test.ts` - The focused PostgreSQL test applies migration `0218` twice. - The test also completes a partial prior run and preserves existing provenance. - The latest GitHub head has 29 successful checks. - The opt-in Storybook visual check skipped as expected. - Greptile reports 5/5 with no open comments. ## Risks - New interaction writes use `anyone` by default. - Callers must select `not_creator` or `human_only` when they need stricter review. - Legacy pending interactions keep the old creator and human restrictions. - Migration `0218` fills only missing provenance fields during recovery. - Cross-issue resolutions can reach the existing action limit. - The shared evaluator affects every interaction kind. - Route, service, database, shared contract, and UI tests cover these rules. > This work matches the Agent Reviews and Approvals direction in `ROADMAP.md`. It does not duplicate a planned item. ## Model Used OpenAI Codex, GPT-5. The runtime does not expose the exact deployment ID or context window. The agent used reasoning, repository tools, shell commands, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked public issues or described the issue with the required labels - [x] I have not referenced internal Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation - [x] I have considered and documented the risks - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open comments - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
92047cac46 |
test(adapter-utils): make the next sandbox flake diagnosable (#11483)
`execution-target-sandbox` has failed twice in CI and not once in several hundred local runs. This does not fix it. It makes the next occurrence carry its own evidence, because a third unreproducible failure would teach nothing. The observed signature was an empty stdout with exit code 0 - the child exited cleanly having produced nothing, which is what a lost stdin frame looks like from the test's side. Three mechanisms were checked and ruled out rather than assumed: the helper resolving on `exit` rather than `close` (a 200-iteration probe produced no truncations, and the failure was empty rather than partial); the wrapper reporting exit before stdout drains (it already listens on `close`); and frame writes racing (the stream wrapper's `writeEvent` is synchronous and sequence-numbered). Two candidates remain and the runtime tree separates them. A stdin queue frame still present means the host wrote it and the wrapper never consumed it; a drained queue with no output means it was consumed and the reply was lost on the way back. The report prints that tree, both proxy streams, the exit code, and the elapsed time - the last because the bridge and proxy run on 5s budgets that are generous locally and tight on a runner sharing a box with 19 other lanes. Timeouts are deliberately unchanged. Raising them would probably make the symptom go away, which is the reason not to do it blind. The first revision capped the tree walk one level above the queue frames, so "the queue is empty" and "the walk never looked" printed identically - the distinction the report exists to make. Caught in review. Verifying that the reporter printed something was not enough; it had to print the thing that discriminates, which is now checked by planting a frame and forcing the assertion. adapter-utils typecheck clean; 44 pass, stable across repeated runs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9e9f744f58 |
Show blocker links in the task chat (#11456)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip helps operators supervise agent work through tasks and task threads. > - The redesigned task thread shows the current work and its state. > - A blocked task did not show the dependency that prevented progress. > - Operators had to leave the thread to find the direct and final blockers. > - This pull request adds compact blocker links at the top and bottom of the task thread. > - The benefit is that operators can identify and open the relevant tasks without adding a large notice to the thread. ## Linked Issues or Issue Description **What existing behavior does this improve?** The redesigned task thread did not show which task directly blocked the current task or which task ultimately blocked its dependency chain. **Subsystem affected** `server/`, `packages/shared/`, and `ui/` task-blocker presentation. **Current behavior** A blocked task can open in the redesigned thread without a visible dependency link at the top or bottom of the conversation. **Proposed behavior** Show one compact amber row for the direct blocker. Show a second row for the selected final blocker when one exists. Render the rows at both ends of the thread. **Reason and benefit** Operators can see the reason for the blocked state and open the relevant task from the conversation. The compact rows preserve thread density. **Breaking changes** None. The new blocker-attention fields are optional. Existing clients remain compatible. ## What Changed - Added a compact task-chat component for direct and selected final blocker links. - Added the blocker rows to the top and bottom of populated and empty task threads. - Added link-ready blocker-attention details so an intermediate selected task stays on its correct direct chain. - Included blocker-link changes in the thread content key so pinned threads follow a newly added bottom row. - Added component, scrolling, server contract, and Storybook coverage for the new states. ## Verification - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx server/src/__tests__/issue-blocker-attention.test.ts` (38 tests passed) - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` ## Risks - Low risk. The rows only render while the task status is `blocked` and an unresolved blocker is available. - Long titles are truncated to keep each blocker on one line. The full task label remains available in the link title. - Older server payloads keep the original leaf-selection behavior because the new sampled details are optional. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The run used tool-enabled reasoning and code execution. The context-window size was not exposed to the run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cd501499a2 |
test: add ACPX run lifecycle characterization baselines (#11461)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The adapter runtime starts, turns, settles, and composes ACPX runs > - Recent lifecycle corrections changed several order and cleanup rules > - Those rules need regression coverage before the planned engine refactor > - This pull request adds characterization suites for the corrected behavior > - The benefit is a clear test baseline for the next refactor ## Linked Issues or Issue Description **What existing behavior does this improve?** The ACPX adapter runtime and server heartbeat lifecycle need stable regression coverage for their current corrected behavior. **Subsystem affected** Cross-cutting (multiple of the above): `packages/adapter-utils` and `server` test suites. **Current behavior** The runtime has corrected rules for startup, turns, settlement, composed results, and heartbeat terminalization. The repository lacks a single characterization baseline for these rules. **Proposed behavior** Keep the current lifecycle rules pinned by five test suites. Let the later engine refactor change behavior only when it updates these tests with a clear reason. **Reason and benefit** The suites expose order, cleanup, transport, timeout, retry, result, and lease-release changes during the refactor. They also record one known latent defect as current behavior. **Breaking changes** None. This pull request adds tests only. ## What Changed - Add startup characterization coverage for commands, launch values, session fingerprints, sync order, bridge overlap, and cleanup paths. - Add turn characterization coverage for inputs, events, transports, timeout and cancel behavior, retry rules, errors, and usage. - Add settlement characterization coverage for teardown, adapter sync-back, workspace restore order, native sync, and error policy. - Add composed-run characterization coverage for result forms, finalization sets, and host-lane warm save and warm hit behavior. - Add server coverage that checks run terminalization before environment lease release. ## Verification - Run `npx vitest run packages/adapter-utils/src/acpx-engine/startup-characterization.test.ts packages/adapter-utils/src/acpx-engine/turn-characterization.test.ts packages/adapter-utils/src/acpx-engine/settlement-characterization.test.ts packages/adapter-utils/src/acpx-engine/composed-run-characterization.test.ts packages/adapter-utils/src/acpx-engine/execute.test.ts`. - Run `npx vitest run server/src/__tests__/heartbeat-run-terminalize-before-release.test.ts`. - The adapter-utils run passes 178 tests, and the server run passes 4 tests. - Check `pnpm --filter @paperclipai/adapter-utils typecheck`. - Check `pnpm --filter @paperclipai/server typecheck`. ## Risks Low risk. The change adds test files and does not change production code. One known cold ensure-session cleanup defect remains pinned as current behavior. ## Model Used OpenAI Codex, GPT-5, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e52b8a343f |
fix: ACP run lifecycle corrections — failure settlement, workspace sync-back, lease cleanup (#11454)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters run ACP sessions and manage runtime, workspace, and lease resources. > - Several failure paths left runtime bridges, staged workspaces, or environment leases active after an error. > - These leaks reduce run reliability and can leave later runs without clean resources. > - This pull request closes the failure paths, applies one teardown policy, and adds regression tests. > - The benefit is consistent failure settlement and safer reuse of agent workspaces and leases. ## Linked Issues or Issue Description **What happened?** ACP runs could leave runtime bridges, staged workspaces, or environment leases active after failures. Claude and Gemini ACP runs did not restore the sandbox workspace on teardown. Lease release stopped when one lease returned an error. **Expected behavior** Each ACP failure must return an error result and settle its resources. Teardown must run each step, release leases independently, and restore the host workspace when the sandbox ends. Pending cleanup leases must receive bounded retry attempts. **Steps to reproduce** 1. Run an ACP session that fails after runtime creation or during turn preparation. 2. Run an ACP session that fails during a warm hit or staged runtime handoff. 3. Run lease cleanup with more than one lease when the first release returns an error. 4. Inspect the result phase, teardown calls, workspace state, and lease metadata. 5. Run the regression suites listed in the Verification section. ## What Changed - Settle every ACP failure after runtime creation with an error result and one sandbox.startup span closure. - Close the ACP runtime and remove warm entries after every pre-turn failure. - Run all teardown steps, record teardown errors, release staging leases in finally, and prevent duplicate teardown. - Dispose staged runtimes after seam failures and remove borrowed staged entries with identity guards. - Add fail-open workspace sync-back teardown for Claude and Gemini ACP adapters. - Isolate lease release errors and add bounded retry sweeps for stranded pending_cleanup leases. - Atomically claim pending_cleanup retries and clamp attempt readers to keep the five-attempt bound. - Default absent provider reusableLeases values to false and align the fake provider with its runtime declaration. - Add regression tests for engine, adapter, server, and shared environment behavior. ## Verification - [x] `npx vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts` — 124 tests passed. - [x] `npx vitest run packages/adapters/codex-local/src/server/acp.test.ts packages/adapters/claude-local/src/server/acp.test.ts packages/adapters/gemini-local/src/server/acp.test.ts` — 61 tests passed. - [x] `npx vitest run server/src/__tests__/environment-runtime.test.ts server/src/__tests__/heartbeat-pending-cleanup-sweep.test.ts server/src/__tests__/reusable-leases-default.test.ts server/src/__tests__/environment-routes.test.ts packages/shared/src/environment-support.test.ts` — passed. - [x] All listed suites ran from the repository root. - [x] GitHub CI completed successfully for `cfc349c9f232711433897915112a1c52c0e462ca`. - [x] Greptile completed with a 5/5 confidence score and no blocking finding. ## Risks The engine changes affect failure settlement and teardown order across ACP runs. The server changes add retry state to existing lease metadata without a schema migration. The adapter changes restore workspaces after sandbox execution. Regression tests cover the changed paths. GitHub CI and Greptile passed for the current head. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. This change fixes runtime reliability and does not duplicate a roadmap feature. ## Model Used OpenAI GPT-5 Codex. The model used tool-based repository inspection, GitHub operations, and code review support. The runtime does not expose a context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (for example, `docs/...` or `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bc9f70f54c |
fix(plugin-daytona): bound the sandbox liveness calls with a per-call timeout (#11408)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox provider plugins run agent work in remote execution environments > - The Daytona sandbox liveness read can stay pending when the connection stops responding > - A pending read blocks the plugin until a broad host-to-worker limit expires > - This pull request adds bounded deadlines to Daytona liveness calls and clears stale handles > - The benefit is a fast and clear error when a Daytona connection stops responding ## Linked Issues or Issue Description Refs #11341 **What happened?** The Daytona sandbox liveness read had no per-call timeout. A silent connection failure left the read pending until the broad host-to-worker RPC limit expired. **Expected behavior** The plugin should stop a liveness call within a defined limit and report a clear timeout error. **Steps to reproduce** 1. Create a Daytona sandbox handle. 2. Make the cached handle freshness read never resolve. 3. Run the next sandbox operation. 4. Observe that the operation waits for the outer RPC limit without a liveness timeout. **Paperclip version or commit** `master` before this change. **Deployment mode** Any deployment mode that uses the Daytona sandbox provider. ## What Changed - Add `withLivenessTimeout` with timer cleanup and `SandboxLivenessTimeoutError`. - Bound `refreshData` with configurable `livenessTimeoutMs`, which defaults to 30000 milliseconds. - Bound sandbox start and recovery calls with the SDK timeout plus a 5000 millisecond margin. - Reject `livenessTimeoutMs` values above 86400000 milliseconds and document the setting. - Evict a cached handle after a failed freshness refresh so the next operation fetches a new handle. - Add a test for a never-resolving freshness refresh and the cached-handle eviction. ## Verification - Run the Daytona plugin test suite with its package Vitest configuration. - Confirm that 150 of 150 tests pass. - Confirm that the new test reports a bounded timeout and a fresh handle on the next operation. - Confirm that GitHub Actions reports green status checks after the pull request starts. ## Risks This change adds an early timeout only to Daytona liveness calls. A value of 0 or less disables the extra bound. The default leaves normal SDK calls within their expected time limit. The main risk is a timeout value that is too short for a slow but healthy connection. ## Model Used OpenAI Codex, GPT-5. The model used tool calls and code execution. The model supplied the PR handoff and did not author the code in this pull request. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fdb9a4880d |
fix(security): route paperclipai CLI guidance through safe npx form (CWE-78) (#11400)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip provides CLI commands and guidance for operators and agents > - The `pnpm paperclipai` script can pass argument values through a shell > - Shell re-parsing can execute command substitutions inside quoted values > - This pull request routes guidance through inert-argv `npx paperclipai` commands and adds regression coverage > - The benefit is safer operator guidance across documentation and runtime hints ## Linked Issues or Issue Description This pull request fixes a command-injection-class defect in Paperclip CLI guidance. **What happened?** The `pnpm paperclipai <sub> --flag "$VALUE"` form can re-parse argument values through a shell. A command substitution inside a quoted value can execute on the host. **Expected behavior** Paperclip guidance must pass CLI values as inert argument values. Host-derived values must not appear in copyable commands. **Steps to reproduce** 1. Run a Paperclip guidance command that uses the `pnpm paperclipai` script. 2. Provide a quoted value that contains a command substitution. 3. Observe that the shell can evaluate the substitution before the CLI starts. 4. Compare the result with the `npx paperclipai` form. **Paperclip version or commit** `5670984b75d109950c968542a0111ebb6967f4da` **Deployment mode** All deployment modes that show or use the affected CLI guidance. **Installation method** Built from source and installed CLI guidance. **Agent adapter(s) involved** Not adapter-specific (core bug). **Database mode** Not database-related. **Access context** Both. **Additional context** The earlier merged PR [#11343](https://github.com/paperclipai/paperclip/pull/11343) used the unsafe `pnpm exec paperclipai` form. This fresh PR replaces that guidance with the safe `npx paperclipai` form. ## What Changed - Standardize documentation and runtime hints on `npx paperclipai`. - Remove the broken `pnpm exec paperclipai` guidance. - Use a static `<host>` placeholder in private-hostname guidance. - Add regression tests for unsafe forms, continued lines, static hosts, and offline guidance. ## Verification - `git diff --check origin/master...origin/fix/paperclipai-cli-npx-safe-invocation` passes. - The branch adds `server/src/__tests__/cli-invocation-safety.test.ts` and updates private-hostname tests. - CI must run the new tests, typecheck, lint, and build checks. - Local Vitest execution was not available because this worktree has no installed Vitest binary. ## Risks - The change affects operator and agent documentation text. - The runtime hints now show `<host>` instead of a request-derived host value. - No database schema or migration changes exist. - CI will detect any missed unsafe invocation or type error. ## Model Used OpenAI GPT-5, exact model ID `gpt-5`, with tool use and code-review assistance. The model used repository inspection, Git operations, and PR preparation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] CI ran the test suites and they pass; local test execution was unavailable in this worktree - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I addressed all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ed8075b535 |
fix(adapter-utils): order stdin file writes in the sandbox process-session bridge (#11406)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters can run through a sandbox process-session bridge > - The bridge writes streamed standard input to files before the remote process reads them > - Concurrent file writes can make a later chunk visible before an earlier chunk > - The remote process can then parse a tail fragment and wait forever for the missing head > - This pull request serializes host writes and makes an unexpected file gap a loud error > - The benefit is ordered input with a bounded failure path for sandbox ACP sessions ## Linked Issues or Issue Description No public GitHub issue exists for this change. The description below follows `.github/ISSUE_TEMPLATE/bug_report.yml`. **What happened?** A sandbox ACP process-session bridge could stall after its handshake when a run sent a large prompt. The host sent one un-awaited file write for each standard input chunk. A small later chunk could finish before a large earlier chunk. The remote poller then sent the tail bytes first. The agent parser raised an error on the tail fragment, and the head bytes stayed buffered without a newline. **Expected behavior** The bridge must expose standard input files in sequence. The remote poller must report a clear error when an earlier file remains missing beyond the retry budget. **Steps to reproduce** 1. Start an ACP session through a sandbox process-session bridge. 2. Send a prompt that produces multiple standard input file chunks. 3. Delay finalization of an earlier chunk while a later chunk completes. 4. Observe that the remote parser can receive the later chunk first and the session can stop without a clear error. **Paperclip version or commit** The change targets the current `master` branch at the submitted commit. **Deployment mode** The bug affects sandbox execution. **Agent adapter(s) involved** The failure affects the ACP process-session bridge. **Database mode** Not database-related. **Additional context** The fix keeps the existing per-file atomic write behavior. It adds ordering at the host write boundary and a bounded ordering check in the shared wrapper poll tail. ## What Changed - Add a per-session promise chain for host standard input file writes. - Keep a failed write from blocking later chain entries. - Track the next expected sequence number in the shared wrapper poll tail. - Hold later files while an earlier file is missing within the existing retry budget. - Emit a loud error and advance after the retry budget expires. - Add regression tests for host ordering, gap holding, and the loud error path. ## Verification - `npx vitest run packages/adapter-utils/src/execution-target-stdin-race.test.ts` — 9 tests passed. - `npx vitest run packages/adapter-utils/src/execution-target-sandbox.test.ts` — 43 tests passed. - `pnpm --filter @paperclipai/adapter-utils typecheck` — clean. - With the source fix reverted, the 3 new tests fail and the 6 original tests pass. - CI must pass on the pull request before merge. ## Risks Low risk. - The host now serializes writes for each session, which can reduce write parallelism. - A failed write still emits one error and destroys the socket, as before. - The wrapper can emit a loud error after the existing retry budget when a file gap persists. - The change does not alter the atomic per-file write behavior. ## Model Used OpenAI GPT-5, model ID `gpt-5`, with tool use and code execution. The context window and internal reasoning details are not disclosed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6f26f2a450 |
fix(adapter-utils): add per-iteration timeout and watchdog to the sandbox callback bridge poll loop (#11341)
## Thinking Path > - Paperclip runs AI agents through adapters and sandboxed execution paths > - The sandbox callback bridge carries file requests between the host and a sandbox > - The poll loop waited forever when a sandbox call stopped responding > - A permanent wait stranded queued requests and hid the run failure > - This pull request adds bounded timeouts, abort handling, recovery backstops, and trace reporting > - The benefit is prompt request failure, safe mutation outcomes, run-level error reporting, and trace visibility ## Linked Issues or Issue Description **What happened?** The sandbox callback bridge could wait forever when a client call stopped responding without a rejection. **Expected behavior** The bridge should fail queued requests and report a run-level error when the sandbox channel stops responding. **Steps to reproduce** 1. Start a sandbox callback bridge. 2. Queue a request. 3. Make the sandbox call stop responding. 4. Observe that the request does not receive a failure response. **Paperclip version or commit** Commit `edc4f71b460c600f97cf44cb486d5cac72ca2db9`. **Deployment mode** Built from source. **Installation method** Built from source with pnpm. **Agent adapter(s) involved** Custom or external sandbox callback bridge. **Database mode** Not database-related. ## What Changed - Add a per-iteration timeout for `listJsonFiles` and `processRequestFile`. - Add a watchdog that fails pending requests when the loop makes no progress. - Abort a hung handler and use a non-retryable 504 backstop when its outcome can be indeterminate. - Retry recovery writes and keep queued requests when a recovery write fails. - Forward the indeterminate-outcome header through the execution target. - Record worker failures through the `sandbox.callbackBridge.workerFailed` trace span. - Add tests for timeout, watchdog, recovery, mutation safety, header forwarding, and fast-request behavior. ## Verification - Run `pnpm exec vitest run packages/adapter-utils/src/sandbox-callback-bridge.test.ts packages/adapter-utils/src/execution-target-sandbox.test.ts`. - Confirm that the PR test, typecheck, build, end-to-end, serialized test, and security checks pass. - Confirm that the current PR head is `edc4f71b460c600f97cf44cb486d5cac72ca2db9`. - Confirm that the PR changes four files: the callback bridge, its tests, the execution target, and its tests. ## Risks The default timeout can fail a slow but valid sandbox call. The defaults remain configurable, and the iteration timeout stays below the sandbox response deadline. A mutation that may have committed returns a non-retryable 504 outcome so the caller does not apply it twice. ## Model Used OpenAI Codex, GPT-5 current runtime, with extended reasoning and tool use. The exact context window is not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a53cc8819b |
fix(claude-local): pipe print prompt via stdin (#9500)
Fixes #2444. Refs #4947. The `claude_local` adapter launched Claude Code as `claude --print - --output-format stream-json --verbose`. Paperclip writes the rendered task prompt to Claude's stdin, but current Claude Code releases can treat the stale `-` positional marker as the prompt itself, so Claude received the literal string `"-"` instead of the issue body. The customer's task ran against no content at all. The fix keeps `--print` mode and stdin delivery, and removes the stale `-`. Adds regression coverage on both sides of the delivery path: a `claude_local` assertion that `--print` is present, `"-"` is absent and the prompt still reaches stdin, and an adapter-utils case proving the sandbox run-log command wrapper preserves stdin while streaming logs. Authored by @elJayAdvisor, whose commit is included unchanged with their authorship. The branch had gone stale and was showing CONFLICTING; the conflict was in `execution-target-sandbox.test.ts`, where their new test was added at the same point as master's `creates the process session directories only in the launch exec` case and git interleaved the two into one hunk. Resolved by taking master's file and re-inserting their test whole, after checking every helper it needs still exists there. Verified: the bug was still live on master at `execute.ts:838`; the regression test genuinely catches it — restoring the stale `-` fails `expect(captured.argv).not.toContain("-")`; `@paperclipai/adapter-claude-local` and `@paperclipai/adapter-utils` typecheck clean; 67 pass across the two test files. All CI gates green; Greptile 5/5. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
eabecc6f77 |
feat(annotations): include issue document annotations in agent review context (#11332)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Reviewers annotate plans and issue documents with inline comments, and assigned agents act on that feedback > - The server already builds a bounded review context from open plan annotations and includes it in agent wake payloads > - Non-plan issue documents did not get the same treatment: their open annotation threads never reached the agent, and the properties pane did not surface their annotations > - This pull request extends the review-context path and the properties-pane UI to issue documents, at parity with plans > - The benefit is that agent feedback on any issue document reaches the assigned agent, not only feedback on the plan ## Linked Issues or Issue Description **What existing behavior does this improve?** The review-context pipeline that delivers inline annotation feedback to assigned agents, and the properties pane that surfaces those annotations to reviewers. **Subsystem affected** The server review-context path (`server/src/services/plan-review-context.ts`, wake payload assembly in `server/src/services/heartbeat.ts`, `server/src/routes/issues.ts`), shared wake-payload types (`packages/shared`, `packages/adapter-utils`), and the issue properties pane (`ui/src/components/issue-properties/`). **Current behavior** A reviewer can annotate any issue document, not only the plan. The agent wake payload includes open annotation threads for the plan document only. Feedback left on other issue documents is invisible to the assigned agent. In the properties pane, the Artifacts tab also gives no way to see or open a document's annotations. **Proposed behavior** Add `buildDocumentReviewContext` beside the existing plan builder. It collects open annotation threads for all non-plan issue documents, applies the same thread, comment, and character budgets across documents, and reports truncation. Include the result as a new `documentReviewContext` field in agent wake payloads and in the issue wake-context route. Keep the plan context on its legacy builder and field so plan-only wakes stay byte-for-byte compatible. Render the new context in the adapter wake-payload text, and surface annotation counts and the annotation panel for documents in the properties pane's Plans and Artifacts tabs. **Reason and benefit** The floating annotation popover and persistent highlight UI landed earlier; this change completes the loop so agent feedback on any issue document reaches the assigned agent, not only feedback on the plan. **Breaking changes** None. The wake payload gains a new optional `documentReviewContext` field; the existing plan context field and its legacy builder are unchanged, so plan-only wakes stay byte-for-byte compatible. ## What Changed - Add `buildDocumentReviewContext` in `server/src/services/plan-review-context.ts`: bounded review context (shared thread/comment/character budgets, per-document legacy limits) over all non-plan issue documents - Include `documentReviewContext` in agent wake payloads (`server/src/services/heartbeat.ts`) and in the issue wake-context response (`server/src/routes/issues.ts`) - Add shared `DocumentReviewContext` / `DocumentReviewContextDocument` types in `packages/shared` - Normalize and render the new context in adapter wake-payload text (`packages/adapter-utils/src/server-utils.ts`), with tests - Show a `DocumentAnnotationsCountChip` and the annotation panel for documents in the properties pane Plans and Artifacts tabs, with tests - Extend server document-annotations service tests to cover the new context builder ## Verification - Run `npx vitest run packages/adapter-utils/src/server-utils.test.ts server/src/__tests__/document-annotations-service.test.ts` from the repo root — 104 tests pass - Run `TZ=UTC npx vitest run ui/src/components/issue-properties/IssuePropertiesDocumentAnnotations.test.tsx ui/src/components/IssueProperties.test.tsx ui/src/components/IssueDocumentAnnotations.test.tsx ui/src/components/DocumentAnnotationPopover.test.tsx` from the repo root — 75 tests pass (one pre-existing monitor-row case asserts UTC timestamps, so use `TZ=UTC` locally; CI runs in UTC) - `pnpm run typecheck` in `server/` passes - Manual: annotate a non-plan issue document, then wake the assigned agent with a comment — the wake payload lists the open document annotation threads; the Artifacts tab shows the annotation count chip and opens the panel ## Risks - The wake payload gains a new optional `documentReviewContext` field; consumers that ignore unknown fields are unaffected, and the plan context field is unchanged - The context is new input to agent wakes; shared budgets (same limits as the plan context) bound token cost across all documents - Low UI risk: the properties-pane changes reuse the existing annotation components > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude (Anthropic), model ID `claude-fable-5` (Claude Fable 5), with extended thinking and agentic tool use (Claude Code harness) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9b1fd42ac1 |
test(grok-local): isolate billing env in usage cost test (#11285)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Local adapters report run output, token use, and cost data. > - The Grok local adapter now reports real token use and cost data. > - Its new billing test must prove the no-key path and the API-key path. > - The no-key assertion used the caller environment without isolation. > - This made the test fail when `XAI_API_KEY` was already set. > - This pull request isolates that environment state in the test. > - The benefit is stable coverage for the cost gate from #10433. ## Linked Issues or Issue Description Refs #10433 **What happened?** The Grok local usage and cost test asserted subscription billing while it still used the ambient process environment. If `XAI_API_KEY` was set before the test ran, the adapter selected API billing instead. The subscription assertion could then fail on a developer machine or a CI runner with provider credentials. **Expected behavior** The test should prove the subscription path with no `XAI_API_KEY`. It should also prove the API billing path with a test key. **Steps to reproduce** 1. Start from `master` after #10433. 2. Set `XAI_API_KEY` in the shell environment. 3. Run `vitest` for `packages/adapters/grok-local/src/server/execute.test.ts`. 4. Observe that the subscription half can take the API billing branch without test isolation. **Paperclip version or commit** `master` after #10433. **Deployment mode** Built from source. ## What Changed - Isolated `XAI_API_KEY` with save, delete, set, and restore logic around both billing assertions. - Gave the subscription and API billing checks separate run ids and temp roots. ## Verification - `XAI_API_KEY=ambient-test-key corepack pnpm exec vitest run packages/adapters/grok-local/src/server/execute.test.ts` - `corepack pnpm --filter @paperclipai/adapter-grok-local typecheck` ## Risks Low risk. This changes test setup only. It does not change Grok local adapter runtime behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex local coding agent. The agent used shell tools, GitHub CLI, and local test execution. The context window size was not exposed in this run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Claude <noreply@paperclip.ing> |
||
|
|
44694328a3 |
fix(issues): make DELETE /api/issues/:id succeed for issues with dependents (#11331)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server provides issue APIs and the database stores issue child rows > - The issue delete endpoint removes the parent issue before dependent rows > - Several issue foreign keys had no delete policy, so PostgreSQL returned a foreign-key error > - This pull request adds safe cascade and set-null policies and a clear conflict response > - The benefit is reliable issue deletion with a useful error when a restricted audit row still blocks deletion ## Linked Issues or Issue Description Fixes #7728 Fixes #4660 Fixes #7991 Fixes #4627 Fixes #5086 **What happened?** `DELETE /api/issues/:id` returned HTTP 500 when dependent comments, thread interactions, read states, inbox archives, feedback votes, or ledger rows referenced the issue. The database raised SQLSTATE 23503 because several foreign keys had no delete policy. **Expected behavior** The endpoint must remove dependent rows that have no meaning without the issue. It must keep ledger rows with a null issue reference. It must return HTTP 409 when a restricted decision audit row still references the issue. **Steps to reproduce** 1. Create an issue. 2. Add a comment or thread interaction that references the issue. 3. Send `DELETE /api/issues/:id`. 4. Observe the HTTP 500 response. **Paperclip version or commit** Commit `1f8f456f8340823fe2bd891ae8933d942f190b7b`. **Deployment mode** Local dev with embedded PGlite or external PostgreSQL. ## What Changed - Add `CASCADE` to five issue child foreign keys. - Add `SET NULL` to the finance and cost event issue foreign keys. - Keep decision audit references restricted. - Map SQLSTATE 23503 from the issue delete service to HTTP 409. - Add migration 0217 for the seven changed tables. - Add regression tests for cascade deletion and restricted decision references. ## Verification - Run `pnpm --filter @paperclipai/db typecheck`. - Run `pnpm --filter @paperclipai/server typecheck`. - Run `npx vitest run src/__tests__/issue-remove-cascade.test.ts` from `server/`. - The regression test applies migration 0217 to a fresh embedded PostgreSQL database. ## Risks - Migration 0217 changes only seven foreign keys that reference `issues.id`. - Cascade deletion removes child rows that cannot exist without the parent issue. - Set-null preserves finance and cost ledger rows. - Decision audit rows remain protected, so the endpoint can return HTTP 409. ## Model Used Codex, based on GPT-5, with tool use and code-review support. The implementation author used an AI coding agent. This PR handoff uses the same model family to validate the commit and manage the pull request. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3cd596725d |
build(deps): bump @modelcontextprotocol/sdk from 1.29.0 to 1.30.0 (#10729)
Bumps [@modelcontextprotocol/sdk](https://github.com/modelcontextprotocol/typescript-sdk) from 1.29.0 to 1.30.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/modelcontextprotocol/typescript-sdk/releases">@modelcontextprotocol/sdk's releases</a>.</em></p> <blockquote> <h2>1.30.0</h2> <h2>What's Changed</h2> <ul> <li>fix(server): prioritize zod issues and format them by <a href="https://github.com/mozmo15"><code>@mozmo15</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/1503">modelcontextprotocol/typescript-sdk#1503</a></li> <li>chore(ci): switch publish to OIDC trusted publishing by <a href="https://github.com/felixweinberger"><code>@felixweinberger</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/1839">modelcontextprotocol/typescript-sdk#1839</a></li> <li>Add end-to-end test suite by <a href="https://github.com/felixweinberger"><code>@felixweinberger</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2167">modelcontextprotocol/typescript-sdk#2167</a></li> <li>v1 stdio buffer limit by <a href="https://github.com/KKonstantinov"><code>@KKonstantinov</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2239">modelcontextprotocol/typescript-sdk#2239</a></li> <li>fix: support Zod 3.25 method literals by <a href="https://github.com/mattzcarey"><code>@mattzcarey</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2368">modelcontextprotocol/typescript-sdk#2368</a></li> <li>Validate Content-Type by parsed media type instead of substring match (v1.x) by <a href="https://github.com/felixweinberger"><code>@felixweinberger</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2444">modelcontextprotocol/typescript-sdk#2444</a></li> <li>fix: send SSE keep-alive comment frames from Streamable HTTP server transport (v1.x) by <a href="https://github.com/mattzcarey"><code>@mattzcarey</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2538">modelcontextprotocol/typescript-sdk#2538</a></li> <li>fix(deps): widen <code>@hono/node-server</code> past GHSA-frvp-7c67-39w9 by <a href="https://github.com/arimu1"><code>@arimu1</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2549">modelcontextprotocol/typescript-sdk#2549</a></li> <li>Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x) by <a href="https://github.com/felixweinberger"><code>@felixweinberger</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2547">modelcontextprotocol/typescript-sdk#2547</a></li> <li>chore: bump version to 1.30.0 by <a href="https://github.com/felixweinberger"><code>@felixweinberger</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2563">modelcontextprotocol/typescript-sdk#2563</a></li> </ul> <h2>New Contributors</h2> <ul> <li><a href="https://github.com/mozmo15"><code>@mozmo15</code></a> made their first contribution in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/1503">modelcontextprotocol/typescript-sdk#1503</a></li> <li><a href="https://github.com/arimu1"><code>@arimu1</code></a> made their first contribution in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2549">modelcontextprotocol/typescript-sdk#2549</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/modelcontextprotocol/typescript-sdk/compare/v1.29.0...1.30.0">https://github.com/modelcontextprotocol/typescript-sdk/compare/v1.29.0...1.30.0</a></p> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/2d889f2b329e46680ec9bdd565de4616c497825a"><code>2d889f2</code></a> chore: bump version to 1.30.0 (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2563">#2563</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/e3f3daa12cc2603919939b72136ce9d9e800b868"><code>e3f3daa</code></a> Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x)...</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/bb5a718cbf90796bacbf62218b359196d210426b"><code>bb5a718</code></a> fix(deps): widen <code>@hono/node-server</code> past GHSA-frvp-7c67-39w9 (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2549">#2549</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/1dad2634ce5799fb386283d14291d1b4935a9a52"><code>1dad263</code></a> fix: send SSE keep-alive comment frames from Streamable HTTP server transport...</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/69749aa5081ddfe675d36da8d96c7e27d83742b8"><code>69749aa</code></a> Validate Content-Type by parsed media type instead of substring match (v1.x) ...</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/369513df7b0e9d8a979c86f68ba1930e0d5f27f0"><code>369513d</code></a> fix: support Zod 3.25 method literals (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2368">#2368</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/e7ee57c2f33b8290a78a3cefa27ab635fe67fbff"><code>e7ee57c</code></a> v1 stdio buffer limit (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2239">#2239</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/c36e1ef5bb3b07b0c23fc6d28d4a6b56ebdd9512"><code>c36e1ef</code></a> Add end-to-end test suite (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2167">#2167</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/bf1e022bd219f678b3865093d58595c6c8a67f1a"><code>bf1e022</code></a> chore(ci): switch publish to OIDC trusted publishing (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/1839">#1839</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/9edbab7a09f31a288a27df3220edbebff45dbb6c"><code>9edbab7</code></a> fix(server): prioritize zod issues and format them (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/1503">#1503</a>)</li> <li>See full diff in <a href="https://github.com/modelcontextprotocol/typescript-sdk/compare/v1.29.0...1.30.0">compare view</a></li> </ul> </details> <details> <summary>Maintainer changes</summary> <p>This version was pushed to npm by <a href="https://www.npmjs.com/~GitHub%20Actions">GitHub Actions</a>, a new releaser for <code>@modelcontextprotocol/sdk</code> since your current version.</p> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
817225415c |
build(deps-dev): bump rollup from 4.62.2 to 4.62.4 (#11319)
Bumps [rollup](https://github.com/rollup/rollup) from 4.62.2 to 4.62.4. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/rollup/rollup/releases">rollup's releases</a>.</em></p> <blockquote> <h2>v4.62.4</h2> <h2>4.62.4</h2> <p><em>2026-08-01</em></p> <h3>Bug Fixes</h3> <ul> <li>Resolve a regression when using Rollup on older Linux distributions (<a href="https://redirect.github.com/rollup/rollup/issues/6467">#6467</a>)</li> </ul> <h3>Pull Requests</h3> <ul> <li><a href="https://redirect.github.com/rollup/rollup/pull/6463">#6463</a>: docs: add llms.txt documentation index for LLMs and agents (<a href="https://github.com/abyworkings-coder"><code>@abyworkings-coder</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6464">#6464</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6465">#6465</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6466">#6466</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6467">#6467</a>: ci: fix linux-gnu glibc regression and enforce glibc ≤ 2.28 compatibility (<a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> </ul> <h2>v4.62.3</h2> <h2>4.62.3</h2> <p><em>2026-07-26</em></p> <h3>Bug Fixes</h3> <ul> <li>Sanitize illegal characters preserved modules input base (<a href="https://redirect.github.com/rollup/rollup/issues/6439">#6439</a>)</li> </ul> <h3>Pull Requests</h3> <ul> <li><a href="https://redirect.github.com/rollup/rollup/pull/6421">#6421</a>: docs: update x_google_ignoreList link to canonical URL (<a href="https://github.com/DucMinhNe"><code>@DucMinhNe</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6422">#6422</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6423">#6423</a>: chore(deps): update actions/checkout action to v7 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6424">#6424</a>: chore(deps): update dependency eslint-plugin-unicorn to v68 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6425">#6425</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6426">#6426</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6432">#6432</a>: fix: make isLegal idempotent by not using a global-flag regex (<a href="https://github.com/spokodev"><code>@spokodev</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6433">#6433</a>: docs: clarify sideEffects and moduleSideEffects (<a href="https://github.com/ishaanlabs-gg"><code>@ishaanlabs-gg</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6434">#6434</a>: chore(deps): update dtolnay/rust-toolchain digest to 4be7066 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6435">#6435</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6436">#6436</a>: chore(deps): update actions/cache action to v6 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6438">#6438</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6439">#6439</a>: Sanitize input base before computing preserved module chunk names (<a href="https://github.com/MahinAnowar"><code>@MahinAnowar</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6443">#6443</a>: chore(deps): update dependency eslint-plugin-unicorn to v71 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6444">#6444</a>: fix(deps): update rust crate swc_compiler_base to v60 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6446">#6446</a>: chore(deps): update dtolnay/rust-toolchain digest to 4cda84d (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6447">#6447</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6448">#6448</a>: chore(deps): update actions/setup-node action to v7 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6449">#6449</a>: chore(deps): update dependency eslint-plugin-unicorn to v72 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6450">#6450</a>: chore(deps): update dependency pinia to v4 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6451">#6451</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6455">#6455</a>: docs: fix broken commonjs namedExports link in troubleshooting (<a href="https://github.com/Hashim1999164"><code>@Hashim1999164</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/rollup/rollup/blob/master/CHANGELOG.md">rollup's changelog</a>.</em></p> <blockquote> <h2>4.62.4</h2> <p><em>2026-08-01</em></p> <h3>Bug Fixes</h3> <ul> <li>Resolve a regression when using Rollup on older Linux distributions (<a href="https://redirect.github.com/rollup/rollup/issues/6467">#6467</a>)</li> </ul> <h3>Pull Requests</h3> <ul> <li><a href="https://redirect.github.com/rollup/rollup/pull/6463">#6463</a>: docs: add llms.txt documentation index for LLMs and agents (<a href="https://github.com/abyworkings-coder"><code>@abyworkings-coder</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6464">#6464</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6465">#6465</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6466">#6466</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6467">#6467</a>: ci: fix linux-gnu glibc regression and enforce glibc ≤ 2.28 compatibility (<a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> </ul> <h2>4.62.3</h2> <p><em>2026-07-26</em></p> <h3>Bug Fixes</h3> <ul> <li>Sanitize illegal characters preserved modules input base (<a href="https://redirect.github.com/rollup/rollup/issues/6439">#6439</a>)</li> </ul> <h3>Pull Requests</h3> <ul> <li><a href="https://redirect.github.com/rollup/rollup/pull/6421">#6421</a>: docs: update x_google_ignoreList link to canonical URL (<a href="https://github.com/DucMinhNe"><code>@DucMinhNe</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6422">#6422</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6423">#6423</a>: chore(deps): update actions/checkout action to v7 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6424">#6424</a>: chore(deps): update dependency eslint-plugin-unicorn to v68 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6425">#6425</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6426">#6426</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6432">#6432</a>: fix: make isLegal idempotent by not using a global-flag regex (<a href="https://github.com/spokodev"><code>@spokodev</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6433">#6433</a>: docs: clarify sideEffects and moduleSideEffects (<a href="https://github.com/ishaanlabs-gg"><code>@ishaanlabs-gg</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6434">#6434</a>: chore(deps): update dtolnay/rust-toolchain digest to 4be7066 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6435">#6435</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6436">#6436</a>: chore(deps): update actions/cache action to v6 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6438">#6438</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6439">#6439</a>: Sanitize input base before computing preserved module chunk names (<a href="https://github.com/MahinAnowar"><code>@MahinAnowar</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6443">#6443</a>: chore(deps): update dependency eslint-plugin-unicorn to v71 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6444">#6444</a>: fix(deps): update rust crate swc_compiler_base to v60 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6446">#6446</a>: chore(deps): update dtolnay/rust-toolchain digest to 4cda84d (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6447">#6447</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6448">#6448</a>: chore(deps): update actions/setup-node action to v7 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6449">#6449</a>: chore(deps): update dependency eslint-plugin-unicorn to v72 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6450">#6450</a>: chore(deps): update dependency pinia to v4 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6451">#6451</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6455">#6455</a>: docs: fix broken commonjs namedExports link in troubleshooting (<a href="https://github.com/Hashim1999164"><code>@Hashim1999164</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6456">#6456</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6457">#6457</a>: chore(deps): update dependency magic-string to v1 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/rollup/rollup/commit/ddc4ffab628944e45dbb8d66d58aae818015440f"><code>ddc4ffa</code></a> 4.62.4</li> <li><a href="https://github.com/rollup/rollup/commit/86d171076b2855c2036cac1f17eb94eea2b9cd7a"><code>86d1710</code></a> Update audit resolve</li> <li><a href="https://github.com/rollup/rollup/commit/7beedfa94963a32c14116994d1bae48d33673245"><code>7beedfa</code></a> ci: fix linux-gnu glibc regression and enforce glibc ≤ 2.28 compatibility (<a href="https://redirect.github.com/rollup/rollup/issues/6">#6</a>...</li> <li><a href="https://github.com/rollup/rollup/commit/9c2c58d55632d25a56cc72f4ae9d43d66ef54040"><code>9c2c58d</code></a> docs: add llms.txt documentation index for LLMs and agents (<a href="https://redirect.github.com/rollup/rollup/issues/6463">#6463</a>)</li> <li><a href="https://github.com/rollup/rollup/commit/dc692883d8c7575692613b76a0d0d45afcf1a4ef"><code>dc69288</code></a> chore(deps): lock file maintenance (<a href="https://redirect.github.com/rollup/rollup/issues/6466">#6466</a>)</li> <li><a href="https://github.com/rollup/rollup/commit/5ee08215eaafc3ea495a76fb317ec89aa15c0690"><code>5ee0821</code></a> chore(deps): lock file maintenance (<a href="https://redirect.github.com/rollup/rollup/issues/6465">#6465</a>)</li> <li><a href="https://github.com/rollup/rollup/commit/4501389a63d588ce05f141b6452359e7a510a496"><code>4501389</code></a> fix(deps): update minor/patch updates (<a href="https://redirect.github.com/rollup/rollup/issues/6464">#6464</a>)</li> <li><a href="https://github.com/rollup/rollup/commit/a80a1974c584bfa8b694fb5d1a20f3fa75ebaf0a"><code>a80a197</code></a> 4.62.3</li> <li><a href="https://github.com/rollup/rollup/commit/e87e19b31e87a1dd6a749a6c11afe4cb9183cf8d"><code>e87e19b</code></a> Update audit resolve</li> <li><a href="https://github.com/rollup/rollup/commit/72f98e99228ad90a8e13d79709c009a827a4ef34"><code>72f98e9</code></a> Fix build:docs after rollup update (<a href="https://redirect.github.com/rollup/rollup/issues/6460">#6460</a>)</li> <li>Additional commits viewable in <a href="https://github.com/rollup/rollup/compare/v4.62.2...v4.62.4">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
3040db3343 |
build(deps): bump @agentclientprotocol/claude-agent-acp from 0.63.0 to 0.66.0 (#11314)
Bumps [@agentclientprotocol/claude-agent-acp](https://github.com/agentclientprotocol/claude-agent-acp) from 0.63.0 to 0.66.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/claude-agent-acp/releases">@agentclientprotocol/claude-agent-acp's releases</a>.</em></p> <blockquote> <h2>v0.66.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.65.0...v0.66.0">0.66.0</a> (2026-08-07)</h2> <h3>Features</h3> <ul> <li><strong>deps-dev:</strong> Bump globals from 17.8.0 to 17.9.0 in the minor group (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/960">#960</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/7f27c47c5c7c49e65014e9f7dc55cba17352d33b">7f27c47</a>)</li> <li>expose provider-neutral ACP goal extension (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/964">#964</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8b31dea11bed54f86c41217759159c415611346c">8b31dea</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>publish and replace Claude goals reliably (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/967">#967</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/f8fd3ab8224420f8ced570e974cde09612939d6b">f8fd3ab</a>)</li> </ul> <h2>v0.65.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.64.2...v0.65.0">0.65.0</a> (2026-08-05)</h2> <h3>Features</h3> <ul> <li><strong>deps-dev:</strong> Bump nanoid from 3.3.16 to 3.3.17 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/951">#951</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/b965dd21917e822b56f5012c3572902f26c065c9">b965dd2</a>)</li> <li><strong>deps-dev:</strong> Bump tinyexec from 1.2.4 to 1.3.0 in the minor group (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/959">#959</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/15b4eb46f329566837eae58f2ee4b05e3e81bf64">15b4eb4</a>)</li> <li><strong>deps:</strong> Bump <code>@hono/node-server</code> from 1.19.17 to 2.1.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/956">#956</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/f9123f3e18560b580398aabf49e2190f69746976">f9123f3</a>)</li> <li><strong>deps:</strong> Bump fast-uri from 3.1.4 to 3.1.5 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/952">#952</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/098843842895dcb450746bd064dbb1509e3049d1">0988438</a>)</li> <li><strong>steering:</strong> settle a steered turn at idle, not at the interrupt (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/958">#958</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/a84b81080a4127edf40bc448fc8bf2b15503304d">a84b810</a>)</li> </ul> <h2>v0.64.2</h2> <h2>Bug Fixes</h2> <ul> <li>restore the single-tool representation for ExitPlanMode (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/942">#942</a>) (4302a4b)</li> </ul> <h2>v0.64.1</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.64.0...v0.64.1">0.64.1</a> (2026-08-02)</h2> <h3>Bug Fixes</h3> <ul> <li>release 0.65.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/939">#939</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/0936ec281ec730714c605e3da732069ff47d8969">0936ec2</a>)</li> </ul> <h2>v0.64.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.63.0...v0.64.0">0.64.0</a> (2026-07-30)</h2> <h3>Features</h3> <ul> <li><strong>deps:</strong> Bump actions/checkout from 7.0.0 to 7.0.1 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/925">#925</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8e099e844254c3e91508c79a02e3e7dc2239fcbb">8e099e8</a>)</li> <li><strong>deps:</strong> Bump the minor group with 7 updates (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/928">#928</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/3f609219592e63b947539f79c696b3cedb421060">3f60921</a>)</li> </ul> <h3>Bug Fixes</h3> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/claude-agent-acp/blob/main/CHANGELOG.md">@agentclientprotocol/claude-agent-acp's changelog</a>.</em></p> <blockquote> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.65.0...v0.66.0">0.66.0</a> (2026-08-07)</h2> <h3>Features</h3> <ul> <li><strong>deps-dev:</strong> Bump globals from 17.8.0 to 17.9.0 in the minor group (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/960">#960</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/7f27c47c5c7c49e65014e9f7dc55cba17352d33b">7f27c47</a>)</li> <li>expose provider-neutral ACP goal extension (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/964">#964</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8b31dea11bed54f86c41217759159c415611346c">8b31dea</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>publish and replace Claude goals reliably (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/967">#967</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/f8fd3ab8224420f8ced570e974cde09612939d6b">f8fd3ab</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.64.2...v0.65.0">0.65.0</a> (2026-08-05)</h2> <h3>Features</h3> <ul> <li><strong>deps-dev:</strong> Bump nanoid from 3.3.16 to 3.3.17 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/951">#951</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/b965dd21917e822b56f5012c3572902f26c065c9">b965dd2</a>)</li> <li><strong>deps-dev:</strong> Bump tinyexec from 1.2.4 to 1.3.0 in the minor group (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/959">#959</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/15b4eb46f329566837eae58f2ee4b05e3e81bf64">15b4eb4</a>)</li> <li><strong>deps:</strong> Bump <code>@hono/node-server</code> from 1.19.17 to 2.1.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/956">#956</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/f9123f3e18560b580398aabf49e2190f69746976">f9123f3</a>)</li> <li><strong>deps:</strong> Bump fast-uri from 3.1.4 to 3.1.5 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/952">#952</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/098843842895dcb450746bd064dbb1509e3049d1">0988438</a>)</li> <li><strong>steering:</strong> settle a steered turn at idle, not at the interrupt (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/958">#958</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/a84b81080a4127edf40bc448fc8bf2b15503304d">a84b810</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.64.1...v0.64.2">0.64.2</a> (2026-08-02)</h2> <h3>Bug Fixes</h3> <ul> <li>restore the single-tool representation for ExitPlanMode (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/942">#942</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/4302a4b0b6df821b164cbe4857f26cf5b44b532c">4302a4b</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.64.0...v0.64.1">0.64.1</a> (2026-08-02)</h2> <h3>Bug Fixes</h3> <ul> <li>release 0.65.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/939">#939</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/0936ec281ec730714c605e3da732069ff47d8969">0936ec2</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.63.0...v0.64.0">0.64.0</a> (2026-07-30)</h2> <h3>Features</h3> <ul> <li><strong>deps:</strong> Bump actions/checkout from 7.0.0 to 7.0.1 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/925">#925</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8e099e844254c3e91508c79a02e3e7dc2239fcbb">8e099e8</a>)</li> <li><strong>deps:</strong> Bump the minor group with 7 updates (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/928">#928</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/3f609219592e63b947539f79c696b3cedb421060">3f60921</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li><strong>steering:</strong> add opt-in host-owned fallback (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/919">#919</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/43af4ec29ea5396c2614813af05967bfb0b1bac8">43af4ec</a>), closes <a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/903">#903</a></li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/6b405138fc82be947964612fac04e56654827b66"><code>6b40513</code></a> chore(main): release 0.66.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/961">#961</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8aaf608b4e5a9bead4dd3a4abb060c916281dca6"><code>8aaf608</code></a> ci: fix release flow (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/971">#971</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/f8fd3ab8224420f8ced570e974cde09612939d6b"><code>f8fd3ab</code></a> fix: publish and replace Claude goals reliably (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/967">#967</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/133337ffe5f1137305fa18befafa06f483db32b1"><code>133337f</code></a> ci: simplify the release flow and make it agent-friendly (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/965">#965</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8b31dea11bed54f86c41217759159c415611346c"><code>8b31dea</code></a> feat: expose provider-neutral ACP goal extension (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/964">#964</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/bba912728f7acde6a415b70bf4e6d3b6be99947d"><code>bba9127</code></a> ci: validate PR titles against release-please conventions (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/962">#962</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/7f27c47c5c7c49e65014e9f7dc55cba17352d33b"><code>7f27c47</code></a> feat(deps-dev): Bump globals from 17.8.0 to 17.9.0 in the minor group (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/960">#960</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/6d608cb399001329b5f485d750e1114ce7293439"><code>6d608cb</code></a> chore(main): release 0.65.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/957">#957</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/a84b81080a4127edf40bc448fc8bf2b15503304d"><code>a84b810</code></a> feat(steering): settle a steered turn at idle, not at the interrupt (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/958">#958</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/b965dd21917e822b56f5012c3572902f26c065c9"><code>b965dd2</code></a> feat(deps-dev): Bump nanoid from 3.3.16 to 3.3.17 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/951">#951</a>)</li> <li>Additional commits viewable in <a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.63.0...v0.66.0">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
d68cf32ae9 |
build(deps): bump @codemirror/view from 6.43.1 to 6.43.8 (#11321)
Bumps [@codemirror/view](https://github.com/codemirror/view) from 6.43.1 to 6.43.8. <details> <summary>Commits</summary> <ul> <li>See full diff in <a href="https://github.com/codemirror/view/commits">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
f0e6c0f549 |
feat(server): receive and apply the Paperclip Cloud onboarding seed (#11098)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud provisions a dedicated tenant stack for each
customer. During signup it asks for a mission, a name and role for the
first agent, and a first task.
> - Cloud pushes those answers into the new stack at activation, as
`POST /api/companies/:companyId/onboarding-seed`.
> - No route served that path. The tenant answered 404, so Cloud
recorded the push as unacknowledged and retried on every portfolio
fetch.
> - The failure was soft. The answers stayed durable in Cloud and the
stack still activated. But the stack opened on the empty first-run
wizard, and it asked the customer again for what they had already given.
> - This pull request adds the receiving endpoint. It validates the
seed, applies it, and acknowledges it.
> - The benefit is that a seeded stack opens with the mission, the agent
and the first task already in place.
## Linked Issues or Issue Description
No public GitHub issue covers this. The problem is described in-PR,
following the feature template.
**Subsystem affected**
server/ — Express REST API and orchestration services. Also
`packages/db` (one new table) and `packages/shared` (one new validator).
**Problem or motivation**
Paperclip Cloud collects onboarding answers at signup and pushes them to
the tenant stack at activation. The tenant had no route for that
request. It answered 404. Cloud treats a non-2xx as "not yet applied",
so it kept the answers and retried, but the stack itself stayed
unseeded. A customer who had already named their mission, their first
agent and their first task arrived at an empty first-run wizard that
asked for all three again.
**Proposed solution**
Serve `POST /api/companies/:companyId/onboarding-seed`. Validate the
body, apply it to the company, then acknowledge it.
The seed is customer free text, so it is bounded and validated in
`packages/shared` and read from the JSON body only. It is never read
from an `x-paperclip-cloud-*` header. That header set is the trusted
identity envelope: every member is derived server-side from the host
plus verified domain records, and that is exactly what makes it
trustworthy. Mixing user content into it would remove the property. A
test plants a mission on a cloud header and asserts that the body value
wins.
Application reuses the shapes the first-run wizard already produces, so
a seeded stack and a manually onboarded one look the same afterwards:
- The mission becomes the company-level goal. A multi-line mission
splits into a title and a description, as the wizard does.
- The agent becomes the company's first hire. Its free-text role ("Chief
of Staff") lands on `title`. The structural `role` stays `ceo`, which is
what the org chart and the default-instructions lookup read.
- The first task becomes an issue in the Onboarding project, assigned to
that agent.
Cloud retries until it gets a 2xx, and it reads any 2xx as "the tenant
holds this content". So the endpoint is idempotent per `revision`. A new
`company_onboarding_seeds` table records the applied revision together
with the goal, the agent and the issue it produced. A replay of a
revision that already matches is a successful no-op. A later revision —
the customer edited their answers — updates those three rows in place
instead of creating a second agent and a second task. The record is
written last, after every other write has landed, so a partial
application cannot present itself as acknowledged.
Everything is applied before the 200 is sent. This is an ordering
guarantee, not eventual consistency. The tests read the database
immediately after the response, with no waiting and no polling, so a
lazy receiver fails them on a fast machine as well as a slow one. That
matters because the redirect into the tenant dashboard is gated on this
acknowledgement.
**Alternatives considered**
Store the seed and let the tenant UI apply it on first load. Rejected:
the dashboard redirect is gated on the acknowledgement, so a background
apply would let the dashboard open before the agent and the task exist.
The whole point is that it must not.
Reuse `POST /companies/:companyId/agents` and `POST
/companies/:companyId/issues` over HTTP from Cloud. Rejected: it needs
three round trips with no shared idempotency key, and it moves the "did
all of it land?" decision to the caller.
**Roadmap alignment**
This completes an existing Cloud-to-tenant contract. It does not add a
new user-facing surface.
## What Changed
- Add `POST /api/companies/:companyId/onboarding-seed` in
`server/src/routes/onboarding-seed.ts`. It authenticates exactly as
`POST /api/companies/:companyId/logo` does, through
`assertCompanyAccess`.
- Add `server/src/services/onboarding-seed.ts`. It applies the mission,
the agent and the first task, and records the applied revision last.
- Add the `company_onboarding_seeds` table: schema, migration `0216`,
and journal entry. It holds the applied revision and the ids of the
goal, agent and issue the seed produced.
- Add `applyOnboardingSeedSchema` in `packages/shared`. It bounds
mission to 2000, agent name to 80, agent role to 120, task title to 200,
and task details to 2000 — the same limits Cloud enforces before it
sends.
- Mount the router in `server/src/app.ts` and register the path in the
OpenAPI document.
- Add `server/src/__tests__/onboarding-seed-route.test.ts` with 13
tests.
- The seeded agent is created on `claude_local`. This mirrors the
teams-catalog default for agents created server-side, where no human
runs an environment test first. `PAPERCLIP_ONBOARDING_SEED_ADAPTER_TYPE`
overrides it.
## Verification
```sh
pnpm typecheck # whole workspace, passes
npx vitest run \
server/src/__tests__/onboarding-seed-route.test.ts \
server/src/__tests__/openapi-routes.test.ts # 15 passed
```
The suite runs against embedded Postgres with migrations applied, so
migration `0216` is exercised by every test.
The route tests cover:
- the happy path — mission, agent and task all applied, read immediately
after the 200
- replay of the same revision — no second agent, no second task, no
second goal, no second project
- a later revision — the goal, agent and task are updated in place
- a multi-line mission splitting into a goal title and description
- a revision-only seed
- the activity log entry written once, and not again on a replay
- a caller without access to the company — 403, and nothing written
- a body with no revision — 400
- each field bound past its limit — 400
- a mission planted on an `x-paperclip-cloud-*` header — ignored, body
wins
- an existing Onboarding project — reused, not duplicated
Not verified here: the full Cloud-to-tenant walk against a live stack.
That needs a deployed Cloud and a provisioned tenant together, which is
separate staging work.
## Risks
Migration `0216` creates one new table. It adds no column to an existing
table, rewrites nothing, and backfills nothing, so it is safe to apply
online. The migration safety check passes.
The endpoint writes to a company. Access is enforced by
`assertCompanyAccess`, the same gate the company logo write uses, and a
test covers the denial.
Behavioral note for stacks that already hold data. If a company already
has a non-built-in `ceo` agent, a first seed updates that agent's name
and title rather than creating a second lead. Likewise a seed adopts an
existing company-level goal rather than adding a parallel one. This is
deliberate: the seed is the customer's own stated answer from signup,
and two competing missions or two leads would be worse than one updated
in place. In the intended case — a stack that Cloud has just activated —
none of these exist yet.
The seeded agent is created on `claude_local` with an empty adapter
config. It is idle and needs the usual credential setup before it runs.
Seeding it does not start it.
## Update — rebased onto master + review hardening
Master moved on after this PR was cut, so it was **rebased onto
`master`** and
the seed migration was **renumbered from `0212` to `0216`** (the merged
#11101
took `0212_onboarding_first_task_unique`); the drizzle journal was
re-stitched
and `check:migrations` passes.
Two things landed on top of the original receiver:
- **Mission-only walk contract (PAP-67 r17.4).** The tenant now owns the
first
agent and the first task via #11101's server-owned onboarding path,
which
stamps `ONBOARDING_FIRST_TASK_ORIGIN_KIND` and races safely on the
partial
unique index `issues_onboarding_first_task_uq`. A comment in the apply
path
documents why this receiver leaves the first task to that path on the
cloud
walk, and a paperclip-cloud `node:test`
(`src/onboarding/walk-seed.test.ts`)
asserts the walk's seed carries no `agent`/`firstTask`. The receiver
retains
the agent/first-task code for its documented body contract, kept inert
on the
cloud path by the mission-only seed.
- **Three Greptile P1 fixes** (`95622fa37`): concurrent application is
now
serialized under a per-company `pg_advisory_xact_lock` (no duplicate
goal/agent/project/task on overlapping pushes); a revised first task
carries
its resolved `assigneeAgentId`/`goalId`; and the
`company.onboarding_seed_applied`
audit write is best-effort so a logging failure can't leave the entry
permanently absent. Two new regression tests cover the first two.
## Model Used
Claude Opus 5 (`claude-opus-5`), 1M context window, extended thinking,
with tool use and code execution. Used for the original codebase
investigation, the implementation, and the tests. The rebase, migration
renumber, mission-only contract, and the three P1 fixes were done with
Claude Opus 4.8 (`claude-opus-4-8`), extended thinking, with tool use
and code execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
1e07d5b9aa |
fix(db): give the last two embedded-Postgres migration tests a timeout (#11313)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The `@paperclipai/db` package owns the database schema and its migrations > - Some migration tests start an embedded Postgres server and replay a migration against it > - An embedded Postgres server needs 7 to 12 seconds to start on a CI runner > - Vitest stops a test after 5 seconds unless the test sets its own timeout > - Two of these tests do not set a timeout, so they fail on CI before they assert anything > - This pull request gives both tests a 30 second timeout > - The benefit is that unrelated pull requests stop failing on a test they did not change ## Linked Issues or Issue Description No public issue exists for this. The problem follows. **What happened?** The test `packages/db/src/company-secret-proposals-migration.test.ts` fails on CI. The error is `Test timed out in 5000ms`. The test never reaches its assertions. The suite reports `1 failed | 104 passed`. The failure is not caused by the branch under test. It appeared on three different branches in a few hours: | Run | Head | Failing jobs | | --- | --- | --- | | 31630781317 | `95622fa3` | `General tests (workspaces-b)`, `verify`, `e2e shard (2/3)`, `e2e` | | 31652020976 | `feba90c9` | `General tests (workspaces-b)`, `verify`, `e2e shard (3/3)`, `e2e` | | 31651467721 | `f2115207` | `General tests (workspaces-a (1/2))`, `verify` | The `verify` job reads the result of the general tests. One timeout therefore turns into two red checks. A reviewer sees two failures and reads them as a regression. **Expected behavior** The test starts an embedded Postgres server, replays the migration, and asserts the schema. It must pass on a normal CI runner. **Steps to reproduce** 1. Open any pull request against `master`. 2. Wait for the job `General tests (workspaces-b)`. 3. Read the failure. The test times out after 5000 ms. The failure needs a slow runner. A fast development machine starts embedded Postgres in less than 5 seconds, so the test passes there. **Paperclip version or commit** `master` at `a09d7dcc0`. **Deployment mode** CI only. GitHub Actions, `ubuntu24` runner image. ## What Changed - `packages/db/src/company-secret-proposals-migration.test.ts` — the test now uses a 30 second timeout. The migration suites in this package already use 20 to 60 seconds. 30 seconds is the most common value. - `packages/db/src/status-card-migrations.test.ts` — the same change. This test has the same defect. It does not fail yet because it replays fewer statements. A fix to only one test moves the problem instead of removing it. - Both tests get a comment. The comment tells the next author why the 5 second default is too short. These two tests were the only embedded-Postgres migration tests in the package without a timeout. ## Verification - Run `pnpm vitest run src/company-secret-proposals-migration.test.ts src/status-card-migrations.test.ts` in `packages/db`. Both tests pass. - These suites skip themselves when the Postgres binaries are absent. A pass alone therefore proves nothing. Run the command with `--reporter=verbose`. The output contains Postgres `NOTICE` messages, for example `relation "status_cards" already exists, skipping`. These messages prove the tests ran real SQL. - Run the same command with `--testTimeout=1`. Both tests still pass. This proves the per-test timeout overrides the global timeout. Before this change, the same command fails immediately. - All CI jobs on this pull request pass. The job `General tests (workspaces-b)` passes. This job failed on the three runs listed above. Not done: no attempt to reproduce the timeout on a development machine. A fast machine starts embedded Postgres in less than 5 seconds, so the failure does not occur there. ## Risks Low risk. The change adds two timeout arguments to tests. It changes no source code, no schema, and no dependency. A longer timeout cannot hide a regression here. The tests assert the same conditions as before. A migration that truly hangs now fails after 30 seconds. Before, it failed after 5 seconds with a message that pointed at the wrong cause. The `e2e` failures on the runs above have a different cause. The spec `mcp-user-stories.spec.ts › US-9` fails with `502 — fetch failed` and `fetch failed: bad port`. These errors come from MCP tool-connection health checks. The failures hit different shards on different runs. This pull request does not change that behavior. `e2e shard (2/3)` passes here, which supports the view that those failures are unstable infrastructure. To revert, remove the two timeout arguments. ## Model Used Claude Opus 5 (`claude-opus-5`), through Claude Code. Extended thinking enabled. Tool use enabled: file read and edit, shell command execution for the local test runs, and the GitHub CLI to read the failing CI logs. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b7b8fbf688 |
fix(adapter-utils): let explicit PAPERCLIP_API_URL override the derived runtime URL in run env (#10339)
## Thinking Path
> - Paperclip orchestrates AI agents for zero-human companies
> - Every agent run gets a run-scoped bridge into the Paperclip API
through the injected `PAPERCLIP_API_URL` / `PAPERCLIP_API_KEY` env vars,
built by `buildPaperclipEnv` in
`packages/adapter-utils/src/server-utils.ts`
> - `buildPaperclipEnv` resolves that URL as `PAPERCLIP_RUNTIME_API_URL
?? PAPERCLIP_API_URL ?? http://<listen-host>:<port>`, and the server
always exports `PAPERCLIP_RUNTIME_API_URL` derived from
`authPublicBaseUrl` at boot
> - When `authPublicBaseUrl` points at an address that is not reachable
from inside the runtime container (e.g. a VPN/tailnet-only address used
to keep the web UI off the public internet), every local run receives a
dead API URL (`curl` exit 7) and agents only survive by hand-rolling a
localhost fallback
> - An operator-set `PAPERCLIP_API_URL` is the documented escape hatch —
`docs/deploy/environment-variables.md` states the server "preserves the
value" when set externally and that the run-level var "inherits the
server-level value" — but the run env builder inverts the precedence, so
the override never actually reaches runs
> - This pull request swaps the precedence in `buildPaperclipEnv` so an
explicit `PAPERCLIP_API_URL` wins over the derived runtime URL, aligning
the behavior with the documented contract
> - The benefit is that operators with split-horizon topologies (public
auth URL != container-reachable URL) can point agent runs at a reachable
endpoint with one env var, with zero behavior change for deployments
that do not set it
## Underlying Issue
No pre-existing public issue covers this, so per CONTRIBUTING ("Link
Issues or Describe Them In-PR") here are the `bug_report.yml` fields
inline:
- **What happened:** with `PAPERCLIP_AUTH_PUBLIC_BASE_URL` on a
tailnet-only address and `PAPERCLIP_API_URL=http://localhost:3100`
explicitly set in the server environment, every agent run still received
`PAPERCLIP_API_URL=http://100.x.y.z:3100` (the derived,
container-unreachable URL); `curl` from inside the run exits 7 and
agents can only reach the API by hand-rolling a localhost fallback
- **Expected behavior:** the run env inherits the operator-configured
`PAPERCLIP_API_URL`, as documented in
`docs/deploy/environment-variables.md` ("preserves the value", run-level
var "inherits the server-level value")
- **Steps to reproduce:** (1) set `PAPERCLIP_AUTH_PUBLIC_BASE_URL` to an
address not reachable from inside the server container, (2) set
`PAPERCLIP_API_URL=http://localhost:3100` in the server env, (3) trigger
any agent run and inspect the spawned process env: it carries the
derived URL, not the override
- **Version/commit:** reproduced on the `91e58acb` image (2026-07-19);
the precedence is unchanged on current `master` (`a3b293e`)
- **Deployment mode:** single-host Docker Compose, local adapters
(`claude_local`/`codex_local`), web UI exposed via VPN/tailnet only
## Related PRs (dedup search)
Several in-flight PRs touch the same pain point (runs receiving an
unreachable injected API URL) — linked for reviewer context; none of
them honors the documented explicit override, and the older ones appear
stale:
- #9916 — reworks `PAPERCLIP_RUNTIME_API_URL` derivation and port
preservation (server side); complementary, does not change run-env
precedence
- #8130 — honors a pre-set `PAPERCLIP_RUNTIME_API_URL` (server side); a
complementary escape hatch via the runtime var instead of the documented
`PAPERCLIP_API_URL` override
- #8025 — heuristic: prefer loopback when the runtime bind is loopback
(no activity since Jun 12)
- #5692 — heuristic loopback-safe URL inside `buildPaperclipEnv` (no
activity since May 14)
- #4877 — broader same-host injection rework across 10 files (no
activity since May 2)
- #4794 — always forces loopback for spawned agents (no activity since
Apr 30; would break split-horizon setups where a reachable non-loopback
URL is intended)
This PR intentionally takes the Path-1 route from CONTRIBUTING: the
smallest possible change (swap two lines so the documented operator
override wins) plus regression tests, rather than a new heuristic.
## What Changed
- `packages/adapter-utils/src/server-utils.ts`: `buildPaperclipEnv` now
resolves the injected URL as `PAPERCLIP_API_URL ??
PAPERCLIP_RUNTIME_API_URL ?? http://<listen-host>:<port>` (explicit
override first), with a short comment explaining why
- `packages/adapter-utils/src/server-utils.test.ts`: three new tests
covering the override precedence, the derived-URL fallback, and the
listen-host default (including the `0.0.0.0` to `localhost` mapping)
- `server/src/__tests__/paperclip-env.test.ts`: updated the expectation
that encoded the old runtime-URL-first precedence and added the
symmetric fallback case (runtime URL used when no explicit override is
set)
- No docs changes needed: `docs/deploy/environment-variables.md` already
describes the fixed behavior
## Verification
- `vitest run` on the new `buildPaperclipEnv` tests in
`packages/adapter-utils`: 3/3 pass
- `vitest run` on `server/src/__tests__/paperclip-env.test.ts` after the
expectation update: 5/5 pass (the first CI run correctly flagged the one
test that encoded the old precedence)
- Reproduced and verified on a production deployment (single-host
Docker, `PAPERCLIP_AUTH_PUBLIC_BASE_URL` on a tailnet-only address):
- Before: freshly spawned runs received
`PAPERCLIP_API_URL=http://100.x.y.z:3100` (verified in the spawned
process `/proc/<pid>/environ`); `curl` to it from inside the container
exits 7
- After (with `PAPERCLIP_API_URL=http://localhost:3100` in the compose
environment): a fresh run received `http://localhost:3100`, and `curl
$PAPERCLIP_API_URL/api/agents/me` with the run-scoped key returned HTTP
200; the run finished `succeeded` with usage telemetry recorded
## Risks
- Low. Behavior changes only for deployments that explicitly set
`PAPERCLIP_API_URL`; when unset (the default),
`PAPERCLIP_RUNTIME_API_URL` is used exactly as before
- The sandbox callback bridge (`execution-target.ts`) is intentionally
untouched: remote sandboxes genuinely need the publicly reachable URL,
and its `input.hostApiUrl || PAPERCLIP_RUNTIME_API_URL || ...` chain
still provides it
## Model Used
- Claude Fable 5 (Anthropic, model ID `claude-fable-5`), extended
thinking + agentic tool use via Claude Code, operating over SSH against
the affected deployment
## Checklist
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] If this change affects the UI, I have included before/after
screenshots
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Sergio-LPA <204395363+Sergio-LPA@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
4660562fde |
fix(opencode-local): make the model-availability probe non-fatal (#10294)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents run through adapters; the `opencode-local` adapter shells out to the OpenCode CLI and, before each run, does a pre-flight `opencode models` **availability probe** to fail fast on a misconfigured `provider/model`. > - That probe was written to **throw on any probe failure** — a timeout, a non-zero exit, or a transient `Unexpected error` from the CLI — which aborts the whole heartbeat run. > - In practice the CLI probe fails transiently (provider hiccup, cold cache, momentary CLI error). When that happens *after* the agent has already done its work, the run dies before its terminal disposition is written, so the platform reopens the issue and re-runs it — a spurious crash/re-run loop that affects every agent on the OpenCode adapter. > - This PR makes the probe **non-fatal when it cannot run**: it warns and proceeds with the configured model, letting the real invocation be authoritative. > - It deliberately **keeps** the genuine guard: when the probe *succeeds* and the configured model is absent from a non-empty list, it still throws (this is what catches misconfigured slugs). > - The benefit is that a best-effort pre-flight check can no longer take down an otherwise-healthy run, while the useful misconfiguration guard is retained. ## Linked Issues or Issue Description No public GitHub issue exists; describing inline (bug). **What happened:** an OpenCode-adapter agent run terminated at the adapter level with `` `opencode models` failed: Unexpected error ``. The failure landed after the agent had produced its work, so the terminal-status update never applied and the run was reopened and re-executed. **Expected:** a transient failure of the `opencode models` availability *probe* should not abort the run — the probe is a best-effort pre-flight guard, not a gate. **Actual:** the probe threw on timeout / non-zero exit / empty output, aborting the run and discarding the completed work + disposition. **Scope:** both the local (`models.ts`) and remote/SSH (`execute.ts`) probe paths; affects any agent on the `opencode_local` adapter. Related PRs (context / prior art): - Refs #5119 — added the remote execution-target model-probe validation this PR softens. - Refs #3291 — closed prior attempt to make the `opencode_local` model probe non-blocking (at agent-create time; different entry point). - Refs #8014 — related open work raising the probe timeout (20s → 60s); complementary, not overlapping. ## What Changed - `models.ts` (`ensureOpenCodeModelConfiguredAndAvailable`): if discovery throws (probe can't run) or returns an empty list, **warn and proceed** with the configured model instead of throwing. The "model present in a non-empty list" check is unchanged and still throws when the configured model is genuinely absent. - `execute.ts` (`ensureRemoteOpenCodeModelConfiguredAndAvailable`): remote probe **timeout / non-zero exit / empty output** now warn and return (proceed) instead of throwing. The remote model-absent guard still throws. - `models.test.ts`: the local "discovery cannot run" case now asserts the probe **proceeds** with the configured model (was: asserts it rejects). - `execute.test.ts`: added remote regression tests — non-zero exit, timeout, and empty output all proceed; a successful probe missing the configured model still rejects. ## Verification ```bash pnpm --filter @paperclipai/adapter-opencode-local typecheck # clean # opencode-local server suite (default 5s per-test timeout is too tight for the # heavy SSH tests on some machines; use a realistic timeout): node node_modules/.pnpm/vitest@*/node_modules/vitest/vitest.mjs run \ packages/adapters/opencode-local/src/server/models.test.ts \ packages/adapters/opencode-local/src/server/execute.test.ts \ packages/adapters/opencode-local/src/server/execute.remote.test.ts \ --testTimeout=45000 ``` Result: typecheck clean; all opencode-local server tests pass, including the new remote fail-open tests and the retained "model unavailable on the remote target" guard test. ## Risks - **Fail-open behavior (intentional).** When the probe can't run, a genuinely misconfigured model is no longer caught at pre-flight — it surfaces at the real invocation instead. This is the accepted tradeoff: the probe is best-effort, and the real invocation is authoritative. The high-value guard (probe succeeds + model absent from a non-empty list) is retained, so the common misconfiguration — a bad `provider/model` slug — is still caught. - No API, schema, or migration changes. Behavior change is confined to the two probe helpers. Low risk overall. ## Model Used Anthropic **Claude Opus 4.8** (`claude-opus-4-8`), used via Claude Code with agentic tool use (repo search, file editing, shell/code execution) and extended reasoning. Used to diagnose the crash, implement the fix, and write the tests; the change was reviewed before submission. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change (`fix/opencode-model-probe-non-fatal`) and contains no internal ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes — N/A (internal adapter behavior; no user-facing docs affected) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green (functional gates: tests/build/e2e/typecheck/security). Review/Greptile gate re-running after this update. - [ ] Greptile is 5/5 with no open P2s — re-triggered after addressing both P2s (remote test coverage + this template-complete description) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
2f1c0e011e |
fix(hermes): surface silent nonzero exit failures (#10107)
## Thinking Path - Followed a silent nonzero Hermes exit from child-process result parsing through heartbeat run, runtime, task-session, and agent finalization. - Found two gaps: the adapter could return `errorMessage: null` for a numeric nonzero exit, and heartbeat later reused the nullable adapter field instead of its normalized fallback. - Kept timeout, signal-cancellation, and specific parsed diagnostics authoritative. ## Linked Issue(s) / Bug Report Related to #9751 (stderr classification) and #9519 (exit-zero finalization), but this is a separate failure mode. Reproduction: run Hermes with a child result equivalent to `exitCode: 1`, `timedOut: false`, and no parsed diagnostic. The heartbeat row derives `Adapter failed`, while runtime/task-session/agent finalization can persist null diagnostics. ## What Changed - Give silent numeric nonzero Hermes exits a stable fallback such as `Hermes exited with code 1`. - Preserve specific parsed errors and timeout/signal semantics. - Reuse the normalized persisted run error for recovered runtime state, task-session `lastError`, and agent `errorReason`. - Add adapter-level and embedded-Postgres regressions. ## Verification - Hermes adapter `execute.onspawn.test.ts` — 7 passed. - Focused heartbeat normalized-error regression — 1 passed (91 skipped). - `pnpm --filter @paperclipai/hermes-paperclip-adapter typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `git diff --check origin/master...HEAD` — passed. Independent review also ran the full recovery file: the changed regression passed; one unrelated pre-existing timing-sensitive test timed out. ## Risks / Rollout Notes Low risk. Fallback text is used only when a numeric nonzero exit has no better diagnostic. Existing timeout, signal, and parsed-error precedence remains unchanged. ## Model Used OpenAI Codex `gpt-5.6-sol` with repository inspection, test execution, and independent read-only review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (not applicable: internal diagnostics only) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: cucurigoo <cucurigoo@users.noreply.github.com> |
||
|
|
8a5c0615f9 |
fix(adapter-utils): forward sandbox callback bridge traffic to the local listen origin (#10017)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents can execute in remote sandboxes, where a callback bridge relays in-sandbox Paperclip API calls back to the host server process > - The bridge worker resolves its forward target from PAPERCLIP_RUNTIME_API_URL / PAPERCLIP_API_URL, which now prefer a configured public base URL and therefore mean "the origin browsers and external agents use" > - The bridge worker runs inside the same process that serves the API, so forwarding through the public origin routes an in-process loopback hop through the network edge > - On a deployment whose public origin sits behind a session-gated edge proxy, every forwarded agent API call is rejected at the edge, so agents in sandboxes cannot read their identity, comment, or hire > - This pull request resolves the bridge forward target from the explicit hostApiUrl override or the local listen host and port only, never the public URL exports > - The benefit is that sandbox agent API calls keep working regardless of how the public base URL is configured or gated ## Linked Issues or Issue Description No existing issue. Describing in-PR following the bug report template: **What happened?** On a cloud deployment with a session-gated public edge, setting a public base URL (PAPERCLIP_PUBLIC_URL) caused every in-sandbox agent API call through the sandbox callback bridge to fail with `403 text/plain "Access denied"` from the edge proxy. With PAPERCLIP_BRIDGE_DEBUG enabled, the bridge logs show the forward target is the public origin, and every proxied request (for example `GET /api/agents/me`) returns the edge proxy's 403 instead of reaching the API. **Expected behavior** The bridge worker runs in the same server process that serves the API, so forwarded calls should target the local listen origin and succeed regardless of how the public origin is configured or gated. **Steps to reproduce** 1. Run the server with a public base URL configured, fronted by a proxy that requires a browser session on API routes. 2. Start a sandbox-executed agent run (any adapter using the sandbox callback bridge). 3. Observe every in-sandbox call to the Paperclip API fail with the proxy's 403; with PAPERCLIP_BRIDGE_DEBUG the forward URL is the public origin. **Paperclip version or commit** Current `master`. **Deployment mode** Self-hosted server behind a reverse proxy. **Agent adapter(s) involved** All sandbox-executed adapters (the bridge is adapter-agnostic). ## What Changed - `packages/adapter-utils/src/execution-target.ts`: `startAdapterExecutionTargetPaperclipBridge` now resolves its forward target as `input.hostApiUrl?.trim() || resolveDefaultPaperclipApiUrl()`. It no longer consults `PAPERCLIP_RUNTIME_API_URL` / `PAPERCLIP_API_URL`, which now describe the public origin for browsers and external agents, exactly the wrong target for an in-process loopback hop. `resolveDefaultPaperclipApiUrl()` builds `http://<PAPERCLIP_LISTEN_HOST>:<PAPERCLIP_LISTEN_PORT>` (exported by server boot before any run executes) and maps wildcard listen hosts to the loopback address of the same family (`0.0.0.0` to `127.0.0.1`, `::` to `[::1]`), so the forward target always matches the address family the server is bound to. `input.hostApiUrl` remains the explicit override seam. A comment documents the reasoning. - `packages/adapter-utils/src/execution-target-sandbox.test.ts`: two new tests. One sets both public URL env vars to an unreachable public https origin and asserts the bridge forwards to the local listen origin (fails before this fix with a 502 because the worker targets the public origin). One asserts an explicit `hostApiUrl` input still overrides everything. - The acpx-engine bridge start (`packages/adapter-utils/src/acpx-engine/execute.ts`) passes no `hostApiUrl` and goes through the same resolution site, so it is covered by the same fix. The sandbox-facing env builder in `server-utils.ts` is intentionally untouched; the bridge env overrides `PAPERCLIP_API_URL` inside the sandbox separately. ## Verification - `npx vitest run packages/adapter-utils/src/execution-target-sandbox.test.ts` (28 tests pass; the new local-origin test fails without the fix) - `pnpm --filter @paperclipai/adapter-utils typecheck` (clean) - Full adapter-utils suite run; the only failures are pre-existing environment-dependent tests (bubblewrap and shallow-clone tests on macOS) identical on a clean `master` checkout ## Risks - Low risk. Deployments where the bridge previously worked did so precisely because the forward target already resolved to the local origin (no public URL configured, so the chain fell through to the same `resolveDefaultPaperclipApiUrl()` result). The only behavioral shift is for deployments with a public URL configured, where forwarding through the edge was either wasteful (an unnecessary network round trip) or broken (session-gated edge). The explicit `hostApiUrl` override seam is preserved for callers that need a nonlocal target. ## Model Used - Claude Fable 5 (claude-fable-5), extended thinking, via Claude Code ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1f7959bc69 |
fix(codex-local): skip benign stderr warnings when deriving the fallback run error (#10003)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents execute through adapters; the codex_local adapter runs the Codex CLI and reports each run's outcome, including an error message when the CLI exits nonzero > - When no error can be parsed from the CLI's JSONL output, `toResult` in `packages/adapters/codex-local/src/server/execute.ts` falls back to the first non-empty stderr line as the run error > - The adapter itself passes the approvals-bypass flag, so the CLI's first stderr line is always the benign startup warning "YOLO mode is enabled. All tool calls will be automatically approved." > - Failed runs therefore record that warning as their error, hiding the real cause (for example an OpenAI API error further down in stderr) and making failures hard to diagnose from the run record > - This pull request derives the fallback error from the first meaningful stderr line, skipping a conservative set of known benign lines, and keeps the existing behavior when every line is benign > - The benefit is that failed Codex runs surface the actual failure reason instead of a harmless startup warning, without ever producing an emptier message than before ## Linked Issues or Issue Description No public issue exists for the codex_local case. The same bug class was fixed for gemini-local in Refs #5099 and Refs #3476; this PR applies the equivalent fix to codex_local. **What happened?** On a multi-tenant cloud deployment of Paperclip, several codex_local runs failed and their run records showed `error_code=adapter_failed` with the error text "YOLO mode is enabled. All tool calls will be automatically approved." That is a benign Codex CLI startup warning, printed on every run because the adapter passes the approvals-bypass flag itself. The real failure (an OpenAI API error printed later in stderr) was never surfaced. **Expected behavior** When the Codex CLI exits nonzero and no error was parsed from its JSONL output, the run error should be the first stderr line that actually explains the failure, not a startup warning the adapter itself provoked. **Steps to reproduce** 1. Configure a codex_local agent and make the underlying Codex CLI invocation fail after startup (for example, configure a model id the active credentials cannot use). 2. Run the agent so the CLI exits nonzero with no parsed JSONL error. 3. Inspect the run's error message: it shows the YOLO approvals warning (the first stderr line) instead of the real error printed further down in stderr. ## What Changed - Added `firstMeaningfulStderrLine` next to `firstNonEmptyLine` in `packages/adapters/codex-local/src/server/execute.ts`, with a conservative benign-line predicate covering the YOLO approvals warning and `[paperclip] ...` diagnostic lines the adapter injected (for example ACP fallback notes). - Used it only in the `toResult` fallback error derivation. If every stderr line is benign, the existing chain still applies (first non-empty line, then `Codex exited with code N`), so the message never gets emptier than today. Logging is unchanged. - Added `packages/adapters/codex-local/src/server/execute.stderr-error.test.ts`: four end-to-end cases through `execute()` with a mocked CLI process, plus unit coverage for the new helper. Tests were written first and confirmed failing before the fix. ## Verification - `pnpm exec vitest run packages/adapters/codex-local/src/server/execute.stderr-error.test.ts` (7 tests pass; 5 failed before the fix as expected) - `pnpm exec vitest run packages/adapters/codex-local` (21 files, 188 tests pass) - `pnpm run typecheck` in `packages/adapters/codex-local` (clean) ## Risks Low risk. Only the derived fallback `errorMessage` changes, and only when a benign line would otherwise have been picked; parsed JSONL errors, logging, retry/quota/auth classification inputs, and the empty-stderr exit-code fallback are untouched. The benign-line list is deliberately conservative (exact prefixes) so real errors are never skipped. ## Model Used Claude Fable 5 (claude-fable-5), extended thinking, via Claude Code ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5521d768b2 |
fix(claude-local): avoid root-only skip permissions failure (#9463)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Claude local is one of the adapter paths that lets operators run Claude Code through a local Paperclip runtime. > - Claude Code rejects `--dangerously-skip-permissions` when the process is running as root or through sudo. > - Local/self-hosted Paperclip deployments may run inside root-owned Docker/runtime processes, so the Claude local adapter can fail before it reaches the actual runtime/auth condition. > - Paperclip already uses a curated `--allowedTools` list instead of `--dangerously-skip-permissions` for remote Claude targets. > - This pull request applies the same safer permission strategy to local root processes while preserving existing local non-root and remote behavior. > - The benefit is clearer, safer Claude local diagnostics/execution in containerized setups without widening permissions beyond the existing explicit tool allowlist. ## Linked Issues or Issue Description No directly matching public issue or PR found. Bug description: - **Problem:** `claude_local` can fail its local probe/execution path when Paperclip runs from a root-owned local container/runtime because Claude Code refuses `--dangerously-skip-permissions` under root/sudo. - **Actual behavior:** The adapter may fail immediately with Claude's root/sudo guard before validating the real Claude runtime/auth state. - **Expected behavior:** Local root processes should use the same explicit allowlist strategy Paperclip already uses for remote targets, while local non-root behavior remains unchanged. - **Environment:** Local/self-hosted Docker or container-style Paperclip runtime where the app process UID is `0`. Related but different: #4926 covers MCP config propagation for the Claude local adapter, not the root/sudo permission flag behavior fixed here. ## What Changed - Added root-aware permission argument selection for the Claude local adapter. - Preserved current local non-root behavior: `--dangerously-skip-permissions` is still used when allowed. - Preserved current remote behavior: remote targets continue using explicit `--allowedTools`. - Changed local root behavior to use the explicit `--allowedTools` list instead of `--dangerously-skip-permissions`. - Threaded process UID awareness through Claude local probe and execution paths. - Added unit coverage for skip-disabled, remote, local non-root, local root, and UID-unavailable behavior. ## Verification ```sh ./node_modules/.bin/vitest run --config g15-vitest-claude-local.config.mjs \ packages/adapters/claude-local/src/server/permissions.test.ts ``` Result: ```text 1 file passed 8 tests passed ``` ```sh pnpm --filter @paperclipai/adapter-claude-local typecheck ``` Result: ```text @paperclipai/adapter-claude-local typecheck passed ``` Additional local smoke: - Ran a disposable root-container Claude adapter diagnostic against this patch. - The diagnostic no longer fails with Claude's root/sudo `--dangerously-skip-permissions` error. - It proceeds to the actual environment-specific Claude auth/runtime result. - No credentials, tokens, hostnames, private paths, or internal Paperclip issue references are included in this PR. Public duplicate checks performed: ```sh gh pr list --repo paperclipai/paperclip --state open --search 'claude local root permissions dangerously skip permissions allowedTools' gh issue list --repo paperclipai/paperclip --state open --search 'claude local root permissions dangerously skip permissions allowedTools' ``` ## Risks Low-to-medium risk adapter behavior change: - Local root Claude runs will now use explicit `--allowedTools` rather than broad skip-permissions behavior. - That is intentionally safer, but an environment depending on broader implicit tool access under root may now need the adapter allowlist to include any required tools. - Local non-root behavior is unchanged. - Remote behavior is unchanged. - No database migrations, API contract changes, or UI changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex `gpt-5.5` via Hermes Agent, with shell/file/tool use for repository inspection, patching, local verification, and GitHub CLI operations. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: LeeJ <elJayAdvisor@users.noreply.github.com> |
||
|
|
04bf7a6ab5 |
feat(observability): instrument stage.sync host steps and home the agent process span (#11301)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - It runs each agent in a remote sandbox and emits OpenTelemetry spans for the sandbox bring-up and the run. > - A real trace showed two gaps. `stage.sync` had about 3 seconds of unattributed host work before its `pack` span. The persistent agent process showed a `sandbox.exec` span that outlived its parent by about 50 seconds. > - The gaps hide real cost and make the trace read as a sequencing bug, so an operator cannot see where startup time goes. > - This pull request wraps the two pre-`pack` host steps in their own spans. It also homes the long-lived process in a run-scoped `sandbox.agentProcess` span. > - The benefit is that startup time is fully attributed and the process reads as a resource that overlaps the turn, not a child that outlives its parent. ## Linked Issues or Issue Description No public issue exists. This is an enhancement to existing telemetry. It is described inline below, following `.github/ISSUE_TEMPLATE/enhancement.yml`. Prior related work: the merged PR #10999 added the run-time wrapper spans and the telemetry data-contract section this PR extends. **What existing behavior does this improve?** The sandbox bring-up and run OpenTelemetry trace. It closes two attribution gaps in that trace. **Subsystem affected** Observability for sandbox execution. The code lives in `packages/adapter-utils`. The span contract lives in `packages/shared/src/telemetry`. **Current behavior** `stage.sync` opens a `pack` span, but the git enumeration and the baseline content-hash walk that run before `pack` have no span, so about 3 seconds read as a gap. On the streamed process-session path the agent process launches fire-and-forget inside the ~2.3 second `bridge.process-session` bring-up step, so its `sandbox.exec` span parents to that step and then runs about 50 seconds. The child dangles past its parent and overlaps `agent.turn`. **Proposed behavior** Wrap the two pre-`pack` host operations in `snapshot.git` and `snapshot.baseline` spans under `stage.sync`. Wrap the streamed launch in a run-scoped `sandbox.agentProcess` span that parents to the live run root (`task.run` at launch). **Reason and benefit** Startup time is fully attributed. The long-lived process reads as a resource that overlaps the sibling `agent.turn`, not a mis-parented child. **Breaking changes** None. The spans are opt-in and export only when an OTLP endpoint is configured. The span seam is a no-op when no runner is injected. No first-party telemetry event changes. ## What Changed - `sandbox-managed-runtime.ts`: add `snapshot.git` and `snapshot.baseline` spans around the git enumeration and the baseline content-hash walk, nested under `stage.sync`, through a shared `runStepSpan` helper that `pack` now also uses. - `execution-target.ts`: wrap the fire-and-forget streamed launch in a run-rooted `sandbox.agentProcess` span, so it parents to the live run root and holds the inner `sandbox.exec`. The `.then`/`.catch` chain became try/catch inside the span callback, with identical frame-ingestion behavior. - `packages/shared/src/telemetry/README.md`: update the span table and the parenting prose. Add `snapshot.git`, `snapshot.baseline`, `pack`, and `sandbox.agentProcess`, and document the intended `sandbox.agentProcess` / `agent.turn` overlap. - Tests: update the executor span-tree test (`childNames` and parent assertions), update the `sandbox-managed-runtime` span-set and nesting tests, and add two `execution-target-sandbox` tests (the launch opens `sandbox.agentProcess`; it parents to the run root, not the bring-up step). ## Verification - Run `npx vitest run` on the three affected test files. Result: 174 tests pass. This includes the updated executor span-tree test and the new `sandbox.agentProcess` open and parenting tests. - Run `tsc --noEmit` in `packages/adapter-utils`. Result: no errors in the changed source or test files. - The full 37-test streamed process-session suite passes unchanged. This confirms the try/catch restructure preserves frame delivery and exit/error behavior. - Pre-existing and unrelated to this PR (present on `master`): `tsc` errors in `execute.ts` / `execute.test.ts` / `remote-spawn-smoke.test.ts` (`onAgentStderr` / `spawnCwd`), and a `check:forbidden-tokens` failure from internal `PAP-###` ids in `ui/src/components/IssueRecoveryActionCard.test.tsx`. This PR does not touch those files, and its own diff is token-clean. ## Risks Low. The change adds instrumentation on the opt-in span path and does not change control flow on the default path. The one production restructure is the streamed launch, which stays fire-and-forget, so bring-up does not block on it. Only the streamed path gains `sandbox.agentProcess`; the legacy poll path launches the process detached and has no host-side long-lived span to home. ## Model Used Anthropic Claude Opus 4.8 (`claude-opus-4-8`), about 200K-token context, agentic tool use through Claude Code. The trace was reviewed through the Honeycomb MCP. The code was written and tested with the model. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e31951a17d |
feat: Claude agent setup-token login in a sandbox (#11286)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Claude agents that run in a remote sandbox need a safe in-product login path > - The existing host login route cannot open a pseudo-terminal inside that sandbox > - The login flow must protect the browser code, the login URL, and the OAuth token at every step > - This pull request adds the parser, the runner, a Daytona pseudo-terminal transport, and a guarded, owner-bound session route behind an injectable transport > - The route stays inert in the default build and fails closed until a sandbox provider binds the live transport > - The benefit is a company-scoped setup-token flow with one-time secret delivery, redaction, and fail-closed transport checks, ready for a later staged production rollout ## Linked Issues or Issue Description **Agent or provider** Claude Code setup-token login for sandbox agents. **Why this adapter is useful** Sandbox agents need a supported way to sign in without host credentials. An authorized owner completes the browser step and receives the token one time. **How the agent is invoked** When a sandbox provider binds the injectable transport, the server starts `claude setup-token` through a sandbox pseudo-terminal, sends the browser code to the matched prompt, and returns the token through the guarded session route. The default build does not bind the transport. In that state the start route fails closed with a fixed no-secret `503`. It does not start a process and it does not hold a sandbox lease. **Additional context** The transport is injectable, so each sandbox provider binds its own pseudo-terminal. This pull request adds the Daytona transport but does not bind it in the production server. A production wiring needs a lease manager, a live pseudo-terminal factory, a durable token store, and its own security review. The route keeps secrets out of logs, activity details, errors, telemetry, and non-owner responses. ## What Changed - Add strict parsers for the setup-token URL, the prompt, and the success token. - Add a login runner that drives the `claude setup-token` command through a pseudo-terminal. - Add the Daytona pseudo-terminal transport and the sandbox plugin wiring. - Add a company-scoped, owner-bound login session service with rate limits, a reaper, cleanup, and one-time token delivery. - Add the guarded session routes at `/agents/:id/setup-token-login-sessions/*` behind an injectable transport. The routes become the live login path only when a provider binds the transport. - Keep the start route fail-closed in the default build. It returns a fixed no-secret `503` and it does not bind `setupTokenLogin`. - Keep the existing host route `POST /agents/:id/claude-login` in place. This pull request does not replace it. - Keep confidential responses behind a fail-closed TLS transport guard with `Cache-Control: no-store`, and extend redaction for the new fields. - Export the parser and the runner from the Claude local server entry, and document the new session routes in the OpenAPI spec. ## Verification - `pnpm --filter @paperclipai/server exec vitest run setup-token-route setup-token-session` - `pnpm --filter @paperclipai/adapter-claude-local exec vitest run` - `pnpm --filter @paperclipai/server run typecheck` - Confirm that the pull request checks pass on GitHub. ## Risks - Low user-facing risk on merge. The default build does not bind the transport, so the production start route stays fail-closed with a `503`. The merge does not change the production login behavior. - When a provider later binds the transport, the flow starts a live sandbox process and holds a short-lived in-memory secret. Cleanup must stop the child before it releases the sandbox lease. - The transport guard fails closed when the deployment does not provide a trusted TLS path. A wrong proxy allowlist can block a valid request. - The production wiring is out of scope. It needs a lease manager, a live pseudo-terminal factory, a durable token store, and its own security review before the server binds `setupTokenLogin`. ## Model Used Anthropic Claude Opus 4.8 assisted the implementation. It used extended reasoning, code execution, repository tool use, and a 200,000-token context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (the OpenAPI spec covers the new session routes; no user-facing documentation needs changes) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c57c0f7498 |
fix(sandbox-providers): accept bsdtar listings in the syncOut tarball confinement check (#11289)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox providers (Daytona, Kubernetes) sync run results back to the host with a sandbox-authored tarball > - Before extraction, a confinement check parses the host `tar -tvf` listing and fails closed on unparseable lines > - The parser only understands the GNU tar listing dialect; macOS ships bsdtar, whose ls-style listing never matches > - Every sandbox syncOut on a macOS host therefore aborts with "refusing tarball with an unparseable entry listing", and the run fails at copy-back > - This pull request teaches the parser both dialects while keeping the fail-closed and traversal guarantees > - The benefit is that Daytona and Kubernetes sandbox runs work on macOS hosts, with no behavior change on Linux ## Linked Issues or Issue Description No existing issue. Bug description: **What happened?** On a macOS host, every Daytona sandbox run fails at syncOut. The adapter reports: `Daytona syncOut refusing tarball with an unparseable entry listing: -rw-r--r-- 0 daytona daytona 7560 Aug 11 21:43 AGENTS.md`. The Kubernetes provider has the same parser and fails the same way. **Expected behavior** The confinement check accepts a well-formed listing from the host tar, whichever dialect the host tar emits. It still rejects members that escape the extraction directory, and it still fails closed on lines it cannot parse. **Steps to reproduce** 1. Run Paperclip on macOS (system tar is bsdtar). 2. Configure an agent with the Daytona sandbox provider. 3. Trigger any run that syncs files back from the sandbox. 4. The run fails at syncOut with the unparseable-entry-listing error, because bsdtar prints `<perms> <links> <user> <group> <size> <Mon> <day> <time|year> <name>` while the parser expects the GNU `<perms> <owner>/<group> <size> <date> <time> <name>` shape. **Operating system** macOS (bsdtar 3.5.3). Linux hosts with GNU tar are unaffected. ## What Changed - Extracted the listing-line parse in both providers' `file-sync.ts` into an exported `parseTarVerboseListingLine` that accepts the GNU/busybox dialect and the bsdtar (libarchive) dialect. - The GNU shape now requires the slash-joined `<owner>/<group>` field. This keeps the two shapes mutually exclusive. Without it, a bsdtar line with numeric uid/gid satisfies the loose GNU pattern shifted by one field, which would hide a leading `../` from the traversal check. - Unparseable lines still fail closed. This includes device-node entries, whose size column is `major,minor` in both dialects. - Made the path-traversal fixture in the Daytona suite portable: GNU spells member renaming `--transform`, bsdtar spells it `-s`. - Added a Daytona test that refuses a sandbox-authored tarball carrying a symlink whose target escapes the extraction dir. - Added parser unit tests for both dialects (file, dir, symlink, hardlink, numeric owner, year-form dates, fail-closed lines) to both providers' suites. ## Verification - `pnpm test` in `packages/plugins/sandbox-providers/daytona`: 136/136 pass on a macOS host. On unpatched `master` the round-trip test fails there with the unparseable-entry-listing error. - `pnpm test` in `packages/plugins/sandbox-providers/kubernetes`: the new parser tests pass; no new failures against the `master` baseline on the same host. - `pnpm typecheck` passes in both packages. - CI runs the same suites on Linux/GNU tar and proves the GNU path is unchanged. ## Risks - Low risk. The GNU pattern is one token stricter (`<owner>/<group>` must contain `/`). GNU and busybox tar always print the slash-joined owner field, so accepted GNU listings are unchanged. - The bsdtar branch only widens acceptance on hosts that were failing 100% of syncOuts before, so no working deployment changes behavior. - The check still fails closed on anything neither pattern matches. ## Model Used Claude Fable 5 (`claude-fable-5`, Claude Code CLI, extended thinking + tool use). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
e5a7fd7038 |
Add sandbox device-login for the Codex adapter (#11237)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Agent adapters connect Paperclip to tools such as the Codex command line tool. > - A sandboxed Codex agent may start without a credential. > - The operator needs a safe sign-in flow that does not expose credentials to the shared package or the sandbox. > - This pull request adds a company-scoped device-login flow with a temporary Daytona sandbox. > - The flow promotes the credential only after readiness checks pass and removes the temporary sandbox after use. > - The result lets an operator sign in to a sandboxed Codex agent from the agent form. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** A Codex adapter that runs in a sandbox cannot authenticate when the company has no pre-provisioned Codex credential. **Proposed solution** Add a company-scoped device-login session. Start a temporary sandbox, run `codex login --device-auth`, stream the code and URL, verify readiness, promote the credential, and delete the sandbox. **Alternatives considered** Pre-provisioning a credential does not support first-time sandbox login. Keeping the credential in the login sandbox does not provide a durable company credential. **Roadmap alignment** This supports the roadmap item for cloud and sandbox agents. **Additional context** The flow uses a five-minute cleanup reaper, compare-and-set status changes, and a PostgreSQL advisory lock to protect promotion and cleanup. ## What Changed - Add the adapter login-session contract, database table, and migration. - Add company-scoped server routes and a service for sandbox device login. - Add credential promotion, readiness checks, and cleanup after login. - Add restart-safe cleanup for abandoned login sandboxes. - Add sandbox login controls to the agent creation and edit forms. - Keep device-login and vendor identifiers out of public shared and adapter UI symbols. ## Verification - `pnpm --filter @paperclipai/adapter-codex-local exec vitest run` passed with 310 tests at the submitted commit. - The server login route, service, and reaper tests passed with 45 tests at the submitted commit. - The agent form render tests passed with 26 tests at the submitted commit. - The public-symbol leak check passed at the submitted commit. - A live Daytona sign-in flow still requires confirmation by a user with a live sandbox. ## Risks The migration adds a new company-scoped table. A promotion or cleanup race could remove a credential or leave a sandbox active, so the service uses claims, compare-and-set transitions, and an advisory lock. The live Daytona flow needs operator confirmation because local tests do not provide a real browser sign-in. ## Model Used OpenAI Codex, GPT-5, tool use and code execution, extended reasoning. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
67001ec6eb |
chore(db): treat Drizzle migration snapshots as binary in diffs (#11254)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip stores its state in a Postgres database, managed in `packages/db`. > - The schema uses Drizzle. `drizzle-kit` writes a full-schema snapshot to `packages/db/src/migrations/meta/` for every migration. > - Each snapshot is a large generated JSON file. One snapshot is about 39k lines. > - PR #11240 marked these files `linguist-generated=true`. That collapses the file view and drops the files from language stats. > - But `linguist-generated` does not change the PR line-count badge. Git still counts every snapshot line, so a PR that adds one migration shows a ~39k-line badge (two migrations show ~78k). PR #11237 shows +84,665 for this reason. > - This pull request adds `-diff` to the same files, so git treats them as binary and their lines leave the +/- count. > - The benefit is a pull request badge that reflects the real code change, not generated snapshot noise. ## Linked Issues or Issue Description No existing issue. This is a small follow-up to merged PR #11240. Description follows the enhancement template: **Problem / Motivation** PR #11240 marked `packages/db/src/migrations/meta/**` as `linguist-generated=true`. That attribute collapses the diff in the Files-changed view and removes the files from language stats, but it does not remove their lines from the PR additions/deletions badge. Each Drizzle snapshot is a full copy of the schema (~39k lines), so any PR that adds a migration still shows a huge line count. PR #11237 shows +84,665, of which ~78k are two generated snapshots. **Proposed Solution** Add `-diff` to the same glob. Git then treats the snapshots as binary. `git diff --numstat` reports `-` for these files, so their lines leave the +/- badge and GitHub shows "Binary file not shown" in place of the full JSON. **Alternatives Considered** - Keep only `linguist-generated`: leaves the misleading ~39k/78k badge on every migration PR. - Use the `binary` macro (`-diff -merge -text`): also disables EOL normalization. This repo has open CRLF/LF work, so `-text` is left unset on purpose. ## What Changed - Added `-diff` to `src/migrations/meta/**` in `packages/db/.gitattributes`. - Kept `linguist-generated=true` (language stats) and `-merge` (no auto-merge of generated snapshots). - Left `-text` unset on purpose, so end-of-line normalization stays intact. - Left the `.sql` migration files untouched, so their diffs stay visible for review. ## Verification Check the attributes and confirm the snapshot is now treated as binary: ``` git check-attr linguist-generated diff merge -- \ packages/db/src/migrations/meta/0031_snapshot.json \ packages/db/src/migrations/0009_fast_jackal.sql git diff --numstat origin/master...origin/feat/adapter-sandbox-login -- \ packages/db/src/migrations/meta/0214_snapshot.json ``` Expected: ``` packages/db/src/migrations/meta/0031_snapshot.json: linguist-generated: true packages/db/src/migrations/meta/0031_snapshot.json: diff: unset packages/db/src/migrations/meta/0031_snapshot.json: merge: unset packages/db/src/migrations/0009_fast_jackal.sql: linguist-generated: unspecified packages/db/src/migrations/0009_fast_jackal.sql: diff: unspecified packages/db/src/migrations/0009_fast_jackal.sql: merge: unspecified - - packages/db/src/migrations/meta/0214_snapshot.json ``` The snapshot reports `-` in numstat (binary, not counted). The `.sql` migration keeps normal diff behavior. GitHub reads the rule from the PR tree, so the badge drops on the next PR that touches these files. ## Risks Low risk. The change only affects how git and GitHub render and count generated files. It does not touch application code, the schema, or any migration. - With `-diff`, GitHub and local `git diff` no longer show a text diff for a snapshot. This is intended; the files are generated and are not reviewed by hand. The raw file is still viewable. - `-text` is left unset, so this change does not affect the CRLF/LF handling that other PRs (for example #8922) address. ## Model Used Claude Opus 4.8 (Anthropic), model ID `claude-opus-4-8`, used through Claude Code with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — not applicable; this change touches no code path, only git/GitHub file handling - [x] I have added or updated tests where applicable — not applicable; `.gitattributes` behavior is verified with `git check-attr` and `git diff --numstat` (see Verification) - [x] I have updated relevant documentation to reflect my changes — not applicable; the `.gitattributes` file documents its own rules inline - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0aa743fc30 |
build(db): clean dist before drizzle generate (#11241)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Database migrations are generated by drizzle-kit, which reads the schema from the db package's built `dist/schema/*.js` > - `tsc` never deletes stale outputs, so a long-lived checkout keeps compiled schema files whose sources were deleted long ago > - A `generate` run in such a checkout sees those ghost tables and sweeps phantom `CREATE TABLE` statements into an unrelated migration > - This pull request makes `generate` clean `dist` before building, so drizzle always diffs against exactly the current schema sources > - The benefit is that no contributor can accidentally resurrect deleted tables inside a new migration ## Linked Issues or Issue Description **What happened?** Running `pnpm --filter @paperclipai/db generate` in a months-old checkout produced a migration re-creating `cloud_upstream_connections`, `cloud_upstream_runs`, and `company_secret_pools` — tables whose schema sources were deleted in #10507. The compiled copies were still in `dist/schema/`, and `drizzle.config.ts` reads the schema from `dist`, so drizzle treated them as new tables missing from the snapshot. **Expected behavior** `generate` diffs the current schema sources only; deleted tables can never reappear in a generated migration. **Steps to reproduce** 1. Build the db package, then delete a schema source file without cleaning `dist`. 2. Run `pnpm --filter @paperclipai/db generate`. 3. The generated migration re-creates the deleted table. ## What Changed - `packages/db/package.json`: the `generate` script runs `pnpm run clean` before `tsc`, so the drizzle-kit input is always a fresh build of the current sources. ## Verification - In a checkout carrying stale `dist/schema/cloud_upstreams.js` / `company_secret_pools.js` artifacts, `generate` produced a phantom migration before this change; after a clean build it reports "No schema changes, nothing to migrate". With this change the clean happens inside `generate` itself. ## Risks - Low risk: `generate` is a developer-only script; the change only adds the existing `clean` step ahead of the existing build, at the cost of a full rebuild per generate. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
b5ebda1dca |
fix(grok-local): report real token usage and cost instead of hardcoded zeros (#10433)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Cost/usage tracking is core to that: the dashboard shows per-agent spend so a team can see what their AI workforce is costing them > - The `grok_local` adapter (xAI's Grok Build CLI) is a newer adapter than `claude_local`/`codex_local`, and its usage/cost wiring was left incomplete > - Every `grok_local` run persists `usage.inputTokens/outputTokens/cachedInputTokens = 0` and `costUsd = null` in `heartbeat_runs`, unconditionally, even though the underlying `grok` CLI reports real, non-zero token counts and cost per turn in its own JSON stream > - This pull request wires the parser to actually read `usage`/`total_cost_usd` from the CLI's terminal `end` event, threads those values into the adapter's execution result, and marks them `usageBasis: "per_run"` so the heartbeat service doesn't incorrectly delta them against a prior run on a resumed session (matching how `claude_local`/`codex_local` already do this) > - The benefit is accurate cost/usage visibility for any self-hosted Paperclip instance running Grok Build agents, instead of a dashboard that always reads zero ## Linked Issues or Issue Description Fixes: #10432 ## What Changed - `packages/adapters/grok-local/src/server/parse.ts`: `parseGrokJsonl()` now reads `usage.input_tokens` / `usage.output_tokens` / `usage.cache_read_input_tokens` / `total_cost_usd` from the terminal `end` event and returns them on `ParsedGrokJsonl` (previously discarded entirely). - `packages/adapters/grok-local/src/server/execute.ts`: `toResult()` now populates `usage.inputTokens/outputTokens/cachedInputTokens` from the parsed values instead of hardcoded `0`, sets `usageBasis: "per_run"` (each `--single` invocation reports usage for just that process, not a running session total), and surfaces `costUsd` only when `billingType === "api"` (metered) — subscription/OAuth billing has no marginal dollar cost, so it stays `null` there, but token counts are populated for both billing types since usage visibility is useful regardless of billing model. - `packages/adapters/grok-local/src/server/parse.test.ts`: added a test asserting usage/cost extraction from a representative `end` event payload, and updated the existing exact-equality test for the new fields. - `packages/adapters/grok-local/src/server/execute.test.ts`: added a test covering both subscription billing (tokens populated, `costUsd: null`) and API-key billing (tokens populated, real `costUsd`), and asserting `usageBasis: "per_run"` in both cases. ## Verification - `pnpm vitest run packages/adapters/grok-local/src/server/parse.test.ts packages/adapters/grok-local/src/server/execute.test.ts` — 9/9 passed - `tsc --noEmit` on the `grok-local` package — clean - Verified against a real self-hosted Paperclip instance running `grok` CLI `0.2.112` with SuperGrok subscription (OAuth) auth: confirmed the raw CLI stream reports real `usage`/`total_cost_usd` (e.g. `"usage":{"input_tokens":21560,...},"total_cost_usd":0.0564448`) that was previously discarded before ever reaching `heartbeat_runs.usage_json`, which always showed all-zero tokens regardless of real usage. ## Risks - Low risk, additive change scoped entirely to the `grok_local` adapter's usage/cost reporting path — no change to control flow, session handling, or process execution. - `usageBasis: "per_run"` mirrors the existing, already-tested pattern in `claude_local`/`codex_local` execute paths, so the heartbeat service's per-run vs. session-cumulative delta logic is exercised the same way. - `costUsd` is intentionally left `null` for subscription/OAuth billing (no behavior change there beyond now-populated token counts) to avoid implying a dollar cost that doesn't exist for flat-rate billing. ## Model Used Claude Sonnet 5 (`claude-sonnet-5`), via Claude Code, no extended thinking. Root cause was found by comparing real `grok` CLI JSON stream output (captured directly from a live invocation) against the persisted `heartbeat_runs.usage_json` row for the same run on a self-hosted instance, then reading `parse.ts`/`execute.ts` source to confirm the hardcoded zero values. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (none found) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (`fix/grok-local-usage-cost-tracking`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes (no user-facing docs reference this internal usage-reporting behavior) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending at time of writing) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (addressed the one P1 raised — `usageBasis: "per_run"`) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
c0bdf26633 |
chore(db): collapse Drizzle migration snapshot diffs and block auto-merge (#11240)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip stores its state in a Postgres database, managed in `packages/db`. > - The schema uses Drizzle. `drizzle-kit` writes a full-schema snapshot to `packages/db/src/migrations/meta/` for every migration. > - Each snapshot is a large generated JSON file. One snapshot is over 200 KB. > - GitHub shows these files as 30k+ line diffs in a pull request. The diffs add no review value, because a human never edits the files. > - The files also merge badly. `drizzle-kit` computes the `id`/`prevId` chain and the full schema state, so a line-level merge of two snapshots produces a file that no real `generate` run creates. > - This pull request adds a `.gitattributes` file that marks the snapshot directory as generated and blocks its auto-merge. > - The benefit is clean pull request diffs and a loud conflict that forces the correct fix when two branches add a migration. ## Linked Issues or Issue Description No existing issue. This is a small repository-hygiene change. Description follows the enhancement template: **Problem / Motivation** Every migration adds a full-schema snapshot JSON under `packages/db/src/migrations/meta/`. These files are large and generated. GitHub renders them as 30k+ line diffs in pull requests, which buries the real change (the `.sql` migration) in noise. The files also have no meaningful line-level merge: `drizzle-kit` computes each snapshot's `id`/`prevId` chain and full schema state. **Proposed Solution** Add `packages/db/.gitattributes`: - `linguist-generated=true` on `src/migrations/meta/**` — GitHub collapses the diff and drops the files from language stats. - `-merge` on the same glob — git refuses the line-level merge and raises a conflict instead of fabricating an invalid snapshot. **Alternatives Considered** - `-diff` / `binary`: hides the diff completely and blocks text merge, but also blocks any local `git diff` and gives a worse conflict experience. `linguist-generated` keeps the file expandable and text-based, so it is the lighter option. - Do nothing: leaves the noisy diffs and the risk of a silent bad merge. ## What Changed - Added `packages/db/.gitattributes`. - Marked `src/migrations/meta/**` as `linguist-generated=true` to collapse the snapshot and journal diffs on GitHub. - Set `-merge` on the same files so git raises a conflict instead of auto-merging generated snapshots. - Left the `.sql` migration files untouched, so their diffs stay visible for review. ## Verification Run `git check-attr` against the affected files and a control `.sql` file: ``` git check-attr linguist-generated merge -- \ packages/db/src/migrations/meta/0031_snapshot.json \ packages/db/src/migrations/meta/_journal.json \ packages/db/src/migrations/0009_fast_jackal.sql ``` Expected output: ``` packages/db/src/migrations/meta/0031_snapshot.json: linguist-generated: true packages/db/src/migrations/meta/0031_snapshot.json: merge: unset packages/db/src/migrations/meta/_journal.json: linguist-generated: true packages/db/src/migrations/meta/_journal.json: merge: unset packages/db/src/migrations/0009_fast_jackal.sql: linguist-generated: unspecified packages/db/src/migrations/0009_fast_jackal.sql: merge: unspecified ``` The snapshot and journal files carry both attributes. The `.sql` migration keeps its normal diff and merge behavior. GitHub applies the rule from the pull request tree, so the collapse shows on the next pull request that touches these files. ## Risks Low risk. The change only affects git and GitHub display and merge behavior for generated files. It does not touch application code, the schema, or any migration. - `-merge` leaves the current-branch version in the working tree on conflict and marks the file conflicted. It does not insert conflict markers into the JSON. The correct resolution stays "renumber the later migration and regenerate", then commit. - Related open pull requests #879 and #8922 also add `.gitattributes` rules for migration files, but for CRLF/LF hash mismatches on Windows. If either lands, a follow-up can merge the rules into one file. There is no functional overlap with this change. ## Model Used Claude Opus 4.8 (Anthropic), model ID `claude-opus-4-8`, used through Claude Code with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — not applicable; this change touches no code path, only git/GitHub file handling - [x] I have added or updated tests where applicable — not applicable; `.gitattributes` behavior is verified with `git check-attr` (see Verification) - [x] I have updated relevant documentation to reflect my changes — not applicable; the `.gitattributes` file documents its own rules inline - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8f478242f1 |
fix(adapter-utils): close sandbox stdin file race with atomic write and fault-tolerant poller (#11235)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters use execution targets to exchange input and output with sandbox processes > - The sandbox input path can expose partial files, and the poller can delete or stop on invalid input > - These timing windows can lose agent input without a clear error > - This pull request makes host writes atomic and makes the poller retry invalid files before it drops them > - The benefit is reliable sandbox input delivery with visible failure after bounded retries ## Linked Issues or Issue Description Closes #10874 ## What Changed - Decode host input into a temporary file, then rename it onto the final JSON path. - Apply the same atomic write pattern to the filesystem client. - Parse each input file before deletion. - Retry parse failures and drop a file after the bounded retry limit with an error event. - Add regression tests for empty, partial, and permanently malformed input files. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/execution-target-stdin-race.test.ts` passes 5/5 tests. - Related sandbox callback, execution target, and sandbox execution suites pass 71/71 tests. - TypeScript checks pass for the changed files. - The regression suite fails on the old code and passes on this change. ## Risks - Low risk. The change affects sandbox input file handling and adds bounded retry behavior. - A permanently malformed file now creates an error event after the retry limit. > This bug fix does not add a core feature, so a roadmap change is not needed. ## Model Used Codex, OpenAI GPT-5, current agent runtime, large context window, tool use and code review support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |