mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 20:34:57 +02:00
656ecfa585b31938e2685ffab3db22e794474803
1400
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
656ecfa585 |
fix(server): keep Date fields intact through secret redaction; harden chat notice timestamps (#10984)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The task chat thread renders issue comments, system notices, and run
transcripts
> - The server routes comment payloads through the run-secret redaction
walker before it sends them
> - The walker rebuilds each object with `Object.entries`, and this
collapses `Date` instances to `{}`
> - The chat renderer then calls `.toISOString()` on an invalid date and
throws, and the thread falls back to the error banner
> - This pull request keeps `Date` instances intact in redacted
responses and makes the renderer safe against bad timestamps
> - The benefit is that task threads with system notices render
correctly again
## Linked Issues or Issue Description
**What happened**
Task threads that contain a system notice showed the banner "Chat
renderer hit an internal state error." in place of the conversation.
This occurred on many tasks.
**Expected behavior**
The thread renders all comments and system notices with correct
timestamps.
**Steps to reproduce**
1. Open a task that has at least one system notice comment (for example
a "Workspace ready" notice).
2. `GET /api/issues/{id}/comments` returns `createdAt: {}` for every
comment because the secret-redaction walker collapses `Date` objects.
3. The system-notice row calls `new Date({}).toISOString()`. This throws
`RangeError: Invalid time value` and trips the thread error boundary.
**Version / deployment**
Regression from #9934 (`e43f187ca`). It applies to all deployments that
include that commit.
## What Changed
- `server/src/services/run-secret-redaction.ts`:
`redactRegisteredSecretValues` now returns `Date` instances as-is. Dates
hold no redactable text, and the `Object.entries` rebuild turned them
into `{}`.
- `ui/src/components/IssueChatThread.tsx`: the system-notice row formats
its timestamp with a new `toValidIsoString` helper. A value that does
not parse as a date now degrades to "no timestamp" instead of a render
crash.
- Regression tests at three layers:
- Walker unit tests: `Date` values survive with and without registered
secret values.
- Route test: `GET /issues/:id/comments` serializes `createdAt` /
`updatedAt` as ISO strings.
- Render test: a system notice with a malformed `createdAt` renders
without the error boundary.
## Verification
- `npx vitest run --root server
src/__tests__/run-secret-redaction.test.ts` — 5 passed.
- `npx vitest run --root server
src/__tests__/issue-comment-redaction.test.ts` — 4 passed (embedded
Postgres route test).
- `cd ui && npx vitest run src/components/IssueChatThread.test.tsx
src/components/IssueChatThreadSystemNotice.test.tsx
src/lib/issue-chat-messages.test.ts` — 121 passed.
- Each new test was run against the unfixed code and failed there, which
confirms it guards the regression.
- A local sweep rendered 47 real issue threads through
`IssueChatThread`: 7 tripped the boundary before the fix, 0 after.
## Risks
- Low risk. The server change only preserves `Date` objects that the
walker destroyed before. String redaction behavior does not change, and
the registry-key stripping does not change.
- The UI change only affects the timestamp of system-notice rows and
omits it when the value is invalid.
## Model Used
- Claude Fable 5 (`claude-fable-5`), Anthropic. Agentic coding session
with tool use (file edits, shell, Vitest). No extended-context or
special reasoning mode.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
814cb33676 |
feat(server): allow agents to resolve review confirmations (#10939)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Issue reviews use thread confirmations to record explicit verdicts > - The server allowed users to resolve review confirmations but rejected all agent actors > - This one-way rule prevented an eligible agent reviewer from completing a review > - The existing review policy already defines which actor can submit a verdict > - This pull request applies that policy to agent confirmation verdicts on writable issues > - The benefit is a consistent review gate for users and agents with preserved audit attribution ## Linked Issues or Issue Description - Builds on: #10931 (merged into master before this PR) - Refs #8617 ## What Changed - Allow eligible agents to accept or reject pending review confirmations on issues they can write. - Allow a creator agent to withdraw its own pending review confirmation when the review policy permits it. - Reuse the review verdict policy check for users and agents. - Require an explicit, same-run review-confirmation binding so unrelated board-only confirmations stay protected. - Preserve board-only tool action confirmations and existing user attribution. - Add route and service tests for agent accept, reject, withdrawal, human-only denial, and user attribution. ## Verification - `pnpm exec vitest run packages/shared/src/validators/issue.test.ts server/src/__tests__/issue-execution-policy-routes.test.ts server/src/__tests__/issue-review-policy.test.ts server/src/__tests__/issue-thread-interaction-routes.test.ts server/src/__tests__/issue-thread-interactions-service.test.ts server/src/__tests__/issue-stalled-review-decision-routes.test.ts` (256 passed after rebasing onto master and the atomic binding fix) - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm -r typecheck` - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm test:run` - `pnpm build` ## Risks - The change expands who can resolve pending review confirmations. The existing issue write checks and review policy limit this access. - Tool action confirmations remain board-only. - The pull request depends on the review policy helper from #10931, which is now merged into master. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. The roadmap marks Agent Reviews and Approvals as shipped. This pull request fixes a narrow server behavior gap in that shipped capability. ## Model Used - OpenAI Codex, model `gpt-5.6-sol`, with reasoning, tool use, and code execution. The runtime does not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f554d67377 |
fix(server): add explicit review verdict policies (#10931)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Issues use `in_review` to request a final decision from an authorized writer > - The server rejected an assignee agent that tried to close its own review, even when the issue had no independent-review rule > - This rejection stopped the default agent workflow and did not represent the configured execution-stage rules > - Paperclip needs an open default and explicit issue-level constraints for teams that require an independent or human verdict > - This pull request removes the unconditional rejection and adds `anyone`, `not_creator`, and `human_only` review policies > - The benefit is a working default path with opt-in, authenticated verdict controls ## Linked Issues or Issue Description Refs #10635, #4429, and #10671. The related public work covers execution-stage independence, self-approval fallback behavior, and durable review paths. This change is distinct. It controls who can resolve an issue review verdict. It keeps configured execution stages active. ## What Changed - Added a nullable `review_policy` issue column. Null has the same meaning as `anyone`. The migration does not backfill existing issues. - Added shared create, update, response, and compact issue contracts for `anyone`, `not_creator`, and `human_only`. - Removed the unconditional agent self-approval rejection for `in_review` issues. - Added one reusable verdict-actor check for terminal status changes and pending interaction accept or reject actions. - Used the authenticated principal type for `human_only`. Agent keys and run tokens remain agent principals. - Used the latest transition into `in_review` to identify the requester for `not_creator`. - Added actionable 403 responses that name the policy, the allowed actor, and the next step. - Kept the configured execution-stage transition and signoff behavior. - Added focused contract, helper, status-route, interaction-route, and execution-stage regression tests. - Updated the implementation specification for the new issue field. ## Verification - `pnpm exec vitest run packages/shared/src/validators/issue.test.ts server/src/__tests__/issue-review-policy.test.ts server/src/__tests__/issue-stalled-review-decision-routes.test.ts --reporter=dot` passed: 42 tests. - `pnpm --filter @paperclipai/shared typecheck` passed. - `pnpm --filter @paperclipai/db typecheck` passed, including migration numbering and safety checks. - `pnpm --filter @paperclipai/server typecheck` passed. - `pnpm run typecheck:build-gaps` passed across server, CLI, plugin SDK/examples, plugin wiki, and UI. - `git diff --check origin/master...HEAD` passed. - SecurityEngineer review approved the authenticated-principal checks and accepted policy-relaxation tradeoff with no required changes. - Greptile reviewed the latest head at 5/5 with zero inline comments or follow-ups. - The latest-head GitHub rollup passed build, typecheck, server/workspace tests, serialized suites, canary, e2e, and external security checks. ## Risks - The migration adds one nullable text column. It has no default and no backfill. - `not_creator` reads the latest recorded transition into `in_review`. It denies the verdict when it cannot identify the requester. - Agents can change or relax `reviewPolicy` when they have issue write access. This is intentional for this issue-level control. - Null and `anyone` do not add a database query to the verdict path. - Configured execution-stage checks still run after the issue-level policy check. > This work aligns with the completed "Agent Reviews and Approvals" and "Enforced Outcomes" roadmap items. It does not add a new roadmap capability. ## Model Used - OpenAI Codex, GPT-5. The exact deployment ID and context-window size are not exposed to the agent. The run used reasoning, repository tools, code execution, and GitHub CLI access. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5b62a3883f |
feat(settings): add experimental Simplified English Interactions flag (#10934)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents ask humans for decisions through interaction blocks: plan confirmations, structured questions, suggested tasks, and checkbox prompts > - Each agent writes these decision prompts in its own style, so operators can get long or unclear text at the exact moment they must decide > - There is no instance-level control that makes agents use a controlled language for these decision points only > - This pull request adds an experimental setting that tells agents to write all user-interaction content in ASD-STE100 Simplified Technical English, with the context the user needs and the effect of each choice > - The benefit is faster, clearer human decisions, with agent thinking and normal responses unchanged ## Linked Issues or Issue Description Refs #10410 (optional `/simplified-english` skill in the skills catalog; this PR adds the instance-level toggle for interactions). **Problem or motivation** Agent-posted user interactions (plan confirmations, structured questions, suggested-task proposals, checkbox prompts) are written in each agent's default style. Operators who want fast, unambiguous decisions have no way to ask agents to use a controlled language for exactly those decision points. **Proposed solution** Add an experimental instance setting, `enableSimplifiedEnglishInteractions` ("Simplified English Interactions"). When it is on, the server sets `simplifiedEnglishInteractions: true` in the heartbeat wake payload. The shared wake-prompt renderer, used by every adapter, then emits a directive: write all user-interaction content in ASD-STE100 Simplified Technical English, state what information the user needs to decide, and state what happens for each choice. The directive applies to interaction content only. Thinking, comments, documents, and other responses keep their usual style. **Alternatives considered** Per-agent instructions work today, but someone must maintain them on every agent. An instance-level toggle applies uniformly and turns off in one place. Server-side rewriting of interaction payloads was rejected: post-hoc translation is lossy and cannot add the decision context that only the agent has. **Roadmap alignment** Extends the experimental settings surface with another opt-in agent-behavior refinement, consistent with existing prompt-side flags. ## What Changed - Added `enableSimplifiedEnglishInteractions` to the experimental instance-settings zod schema, mirror type, and feature catalog (default off, tier preference) in `packages/shared`. - Server `instance-settings.ts` normalizes the flag on both read branches; `heartbeat.ts` reads it once and passes `simplifiedEnglishInteractions` into `buildPaperclipWakePayload`. - Shared adapter renderer (`packages/adapter-utils/src/server-utils.ts`): added the field to `PaperclipWakePayload`, normalization, and an `- interaction language (experimental): ...` directive emitted in both fresh and resume prompt lanes, so one injection point covers all adapters. - UI: new experimental settings card "Simplified English Interactions" in `ui/src/pages/InstanceExperimentalSettings.tsx`, in alphabetical card order; fixtures updated. - Tests: renderer coverage for flag on/off in both lanes, plus schema/catalog/UI fixture updates. ## Verification - From the repo root: `node_modules/.bin/vitest run packages/adapter-utils` (88/88), `packages/shared` validators (25/25), server instance-settings + heartbeat suites (47/47 and 70/70 consumer tests), `ui` settings tests (32/32). - Typecheck is clean in all four touched packages. - Manual check: turn the flag on in Settings → Experimental, wake an agent, and confirm the wake prompt contains the interaction-language directive; turn it off and confirm the directive is absent. ## Risks - Low risk: the flag defaults to off, and the only behavior change is one extra directive line in the wake prompt when an operator turns it on. - The directive is advisory to the agent; models can still deviate from STE. No data or API shape changes; no migration. ## Model Used - Claude Fable 5 (Anthropic, model ID `claude-fable-5`), extended thinking, agentic tool use via Claude Agent SDK. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e43f187cad |
feat(secrets): add human-approved secret proposals (#9934)
## Thinking Path > - Paperclip is the control plane people use to manage AI-agent companies. > - Agents can encounter credentials during work. > - Directly creating live secrets or bindings would bypass human governance. > - Proposal records must remain inert and separate from live secret resolution until an authorized human approves them. > - Approval must reuse the existing secret-create and protected agent-config write paths. > - This pull request adds the propose, review, approve, and reject lifecycle. > - The benefit is that agents can safely hand credentials into Paperclip without exposing plaintext or gaining authority to activate them. ## Linked Issues or Issue Description Follow-on to #9921, which established run-bound agent secret access. **Problem / motivation:** Agents can receive credentials during work. There is no governed way for them to propose a credential or binding without exposing plaintext in work artifacts or immediately creating live access. **Proposed solution:** Store agent-authored proposals outside live secret tables. Encrypt each proposed value and register exact-value redaction when Paperclip receives it. Require an authorized human to approve or reject each proposal. Approval executes through the normal write paths as the human approver. Binding proposals can target only the proposer or its downward reporting chain under the restrictive V1 policy. **Alternatives considered:** We rejected live secrets with a `proposed` status. That design would put untrusted rows in resolver, list, and sync paths. It would also allow uniqueness squatting. We rejected direct agent binding writes because a binding is an agent-config write and must keep the existing human permission gate. **Roadmap alignment:** This change extends the run-bound agent secret-access foundation in #9921 with a governed proposal workflow. ## Security Verdict Q0 SecEng verdict: **PASS-with-required-changes**. The review accepted the separate proposal-table design and required the implementation to: - fail closed unless both encryption and exact-value run redaction registration succeed; - scrub ciphertext idempotently on reject, withdraw, and expiry, with audit-visible state; - treat agent justification as hostile input and foreground action, target, provenance, and approver permissions; - snapshot and re-check the target agent plus reports-to chain at approval to prevent org-chart laundering; - make cascade approval atomic and fail closed if either secret creation or binding authorization fails; - deny low-trust, `skill_test`, `task_bridge`, and non-run-bound sources consistently; and - execute approval through the normal human secret/config write paths, including protected-change gates. Those requirements are implemented and covered by focused service, route, and UI tests. Residual V1 risk remains the accepted 14-day encrypted retention window. Proposal-time redaction also cannot clean a value that leaked before the propose call. ## What Changed - Added `company_secret_proposals`, migration `0207`, shared proposal contracts, and a state-machine service for create, approve, reject, withdraw, cascade, expiry, and ciphertext scrubbing. - Added run-bound agent proposal routes and board review routes. The routes derive provenance from authentication and enforce source restrictions, company isolation, chain-of-command checks, approval-as-approver, wake-on-resolution, and dual audit trails. - Added durable per-run exact-value redaction registration so proposal values remain redacted on later read surfaces. - Added the Secrets **Proposals** tab and agent configuration **Proposed access** rows. The UI shows fingerprint and length only. It also frames agent justification as untrusted input, runs permission preflight, supports approve and reject actions, and confirms cascades. - Updated OpenAPI, agent skill guidance, API reference documentation, and focused server and UI regression coverage. - Rebased the branch onto current `master` and renumbered the proposal migration after `0206`. ## QA Acceptance Results Q5 QA verdict: **PASS — 9/9 acceptance criteria met**, with one Minor non-blocking follow-up. - **AC1:** proposed values never echo, never appear in live lists/resolvers, and expose only fingerprint + length to board reviewers. - **AC2:** restrictive `self_and_reports` matrix passes: self/downward allowed; upward/lateral denied. - **AC3:** secret approval uses the normal create path, honors rename overrides, records proposer/approver provenance, and scrubs ciphertext. - **AC4:** approved bindings materialize and resolve through the target agent's runtime list/fetch routes. - **AC5:** pending-secret bindings require cascade; cascade succeeds atomically and permission failures leave nothing applied. - **AC6:** reject, withdraw, dependent rejection, and expiry paths scrub ciphertext and preserve reasons/audit state. - **AC7:** token/source and approver denial matrix passes through live checks plus focused route tests. - **AC8:** proposal lifecycle events and reused `secret.created`/config-write events form the required dual audit trail; origin-issue notification and wake are queued. - **AC9:** both review surfaces render and execute correctly; UI approval materializes the binding. QA also confirmed zero plaintext occurrences for all exercised proposal values in server logs. The single finding is that the company-level `bindingTargetPolicy` toggle is not wired yet. V1 is hardcoded to the restrictive `self_and_reports` policy. The matrix is correct and the follow-up is tracked separately, so QA classified it as non-blocking. ## Verification - Focused server proposal and redaction suite: 83 tests pass. - Focused proposal review UI suite: 54 tests pass. - Embedded-Postgres migration reapply test: 1 test passes with the documented 30-second timeout. - `pnpm --filter @paperclipai/db typecheck` passes, including migration numbering and safety checks. - `pnpm --filter @paperclipai/shared typecheck` passes. - `pnpm --filter @paperclipai/ui typecheck` passes. - `pnpm check:token-gates` passes with all gates clean. - Q5 exercised the complete propose, review, approve, bind, and runtime-resolve flow over real HTTP, JWT, and database paths. It verified 9/9 acceptance criteria. ## Risks - Proposal ciphertext is retained encrypted for up to 14 days while pending. Terminal-state and expiry scrub paths reduce but do not remove server-compromise risk during that window. - The V1 target policy is restrictive but not yet company-configurable. A separate follow-up owns that change. - A new migration can require another renumber if another migration lands before maintainers merge this pull request. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex coding agent. The exact runtime model ID and context-window size are not exposed. The agent used reasoning, repository editing, terminal execution, Paperclip API, and GitHub CLI capabilities. Q3 UI work also records Claude Opus 4.8 assistance in its commit trailers. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f950952de7 |
fix: reliably show plans in the Plan pane and restore sticky plan confirmation CTAs (#10930)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Chat-style tasks show an agent's plan in a dedicated "Plan" pane, and a plan confirmation lets the user accept or request changes to that plan > - When an agent asked for confirmation but never actually published the plan document (it only wrote the plan in a comment or a question), the Plan pane rendered empty, and the confirmation call-to-action that used to sit pinned at the bottom of the pane had disappeared > - A user asked to confirm a plan they cannot see, with no visible CTA, is stuck — the feature silently fails > - This pull request closes the gap on both sides: it prevents plan confirmations that don't point at a real, latest plan revision, it teaches agents to publish the plan document before confirming, and it restores the sticky confirmation action bar and an explanatory empty state so the pane never goes silently blank > - The benefit is that when a plan is expected, it reliably shows up in the right pane with reachable accept/revise actions ## Linked Issues or Issue Description <!-- No public GitHub issue exists; describing in-PR per the bug template. --> **Bug report** - **What happened:** A task in planning mode could present a plan confirmation while the Plan pane stayed empty (no plan document rendered), and the plan-card confirmation CTAs that were previously pinned to the bottom of the Plan pane no longer appeared. - **Expected behavior:** When a plan is expected, the plan document appears in the Plan pane; when a plan is genuinely missing, the pane explains why rather than showing nothing; and the accept/request-changes CTAs stay visible and reachable while the plan scrolls. - **Steps to reproduce:** Put a task in planning mode with the chat-style task view enabled, have an agent create a plan confirmation without first publishing the `plan` document, and open the Plan tab — the pane is blank and the confirmation actions are missing. - **Deployment mode:** Local dev and self-hosted; UI + server. Related PR (not a duplicate): #9609 "Pin pending confirmations by composer" pins confirmations in a different surface (the composer); this PR restores the Plans-pane action bar and the server/agent guarantees behind it. ## What Changed - **Server:** Reject a `request_confirmation` whose target is a plan document unless a plan document exists and the target points at its *latest* revision, so a confirmation can never reference a plan the pane cannot render (`readPlanTarget` is now exported for reuse). - **Agent instructions:** The CEO and default agent instruction bundles now spell out a plan-publish contract — publish the `plan` document, re-`GET` it and capture `latestRevisionId`, then create the confirmation targeting that revision; never present a plan only in a thread comment or via `ask_user_questions`. - **UI — sticky CTAs:** Restore the plan confirmation action bar pinned to the bottom of the Plans tab so accept/revise stay reachable while the plan scrolls. - **UI — diagnostics:** Keep the Plan tab visible whenever an issue is in planning mode (even before a plan document exists) and show an empty state explaining why the pane is empty instead of rendering nothing. - **UI — annotations:** Add a `panelPlacement="inline"` mode so the plan-document annotation panel renders in document flow instead of as a floating side panel when hosted in the narrow task properties pane. ## Verification - `pnpm check:token-gates` → 3/3 CLEAN - `pnpm typecheck` → clean (all packages) - UI: `pnpm --filter @paperclipai/ui exec vitest run src/components/issue-properties/IssuePlanConfirmationActionBar.test.tsx src/components/IssueProperties.test.tsx src/components/IssueDocumentAnnotations.test.tsx` → 69 passed - Server: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/issue-thread-interaction-routes.test.ts src/__tests__/agent-skills-routes.test.ts` → 60 passed - Manual: with a planning-mode task, the Plan tab stays visible, shows the plan document (or a diagnostic empty state), and the confirmation CTAs stay pinned at the bottom. Visual note: snapshot baselines are intentionally not updated — per `doc/design/DECISION-SHEET.md` "Per-change snapshot verification demoted to dormant (Jul 13 2026)". The `storybook-visual` label is intentionally not added. ## Risks Low-to-moderate. The server change adds a validation gate on plan-document confirmations: an interaction that targets a stale or nonexistent plan revision is now rejected with a 422 instead of being created. This is the intended guarantee, but any caller that relied on creating such confirmations will now need to publish the plan document first (which the updated agent instructions cover). UI changes are additive to the Plans tab and gated by the existing chat-style-task experimental flag. ## Model Used Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, extended thinking enabled, with tool use (file editing, shell, test execution). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
72b509c895 |
Recognize delivered workspaces and reap terminal worktrees (#10908)
## Thinking Path > - Paperclip separates workspace provisioning lifecycle from whether the work was actually delivered. > - Git ancestry alone cannot recognize squash merges or deliveries into a branch other than the workspace base. > - A merged pull request linked from a terminal issue is stronger delivery evidence for those cases. > - The read contract should expose that evidence without changing persisted workspace schema. > - Cleanup must remain conservative: terminal descendants, delivered work, and no active run checkout are all required. > - Reusing the existing cleanup primitives keeps service shutdown, lease cleanup, activity logging, and archival behavior consistent. > - Focused regression coverage locks in both the honest read signal and the fail-closed reaper guards. ## Linked Issues or Issue Description **What existing behavior does this improve?** Execution workspace close-readiness payloads and terminal workspace cleanup. **Current behavior** Delivered squash-merged or cross-branch workspaces can remain `active` and report a permanent “not merged” warning because git ancestry does not contain their original commits. **Proposed behavior** Read payloads distinguish PR-confirmed delivery, ancestry delivery, unmerged work, and unknown state. Fully terminal delivered workspace trees are archived only when no active run holds the checkout. **Reason and benefit** Operators and automation receive an honest delivery signal, while shipped worktrees stop looking active forever and genuinely unmerged work retains its warning. **Breaking changes** The workspace payload gains a derived field. Existing fields and persistence remain unchanged; no database migration is required. **What happened?** A delivered workspace can remain `active` and warn that it is not merged forever after its issue ships through a squash or cross-branch pull request. **Expected behavior** Pull-request delivery should be represented honestly, and a fully terminal delivered workspace should become cleanup-eligible when no run holds its checkout. **Steps to reproduce** 1. Create an issue workspace with commits ahead of its configured base. 2. Deliver those commits with a squash merge or into a different target branch. 3. Mark the source issue and descendants done, then read workspace close readiness. Before this change, the workspace remains active with a “not merged” warning indefinitely. ## What Changed - Added the derived `deliveryState` workspace contract: `merged_via_pr`, `merged_by_ancestry`, `unmerged`, or `unknown`. - Extracted a shared GitHub pull-request merge classifier and reused it for merge confirmations and workspace delivery checks. - Suppressed false ancestry warnings when a terminal issue has ground-truth merged-PR evidence. - Added an idempotent terminality reaper with descendant-terminal, active-run, and delivered-work guards. - Restricted PR delivery evidence to the source issue, then required live merged state plus matching GitHub repository, head branch, and current workspace HEAD; persisted status, stale PRs, lexical mentions, inbound references, and descendant PRs cannot authorize cleanup. - Preserved workspaces with modified or untracked files even when their committed HEAD was delivered. - Bounded both long-lived pull-request state caches to 1,000 entries with oldest-entry eviction. - Routed eligible workspaces through existing runtime shutdown, lease cleanup, activity logging, and archival machinery with exclusive Git index, HEAD, and branch-ref locks plus non-forced removal. - Added regression coverage for delivery derivation, warning behavior, reaper guards, scheduler wiring, and squash/cross-branch delivery. ## Verification - `pnpm -r typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/execution-workspaces-service.test.ts src/__tests__/merged-pr-confirmation-sweep.test.ts src/__tests__/server-startup-feedback-export.test.ts --reporter=verbose` — 63 passed - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/execution-workspaces-service.test.ts --reporter=verbose` after review hardening — 43 passed - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/execution-workspaces-service.test.ts src/__tests__/merged-pr-confirmation-sweep.test.ts src/__tests__/external-objects-service.test.ts --reporter=dot` on the final local head — 73 passed - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-workspace-busy.test.ts --reporter=verbose` — 15 passed - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm test:run` — server 3,662 passed (4 skipped), UI 3,599 passed, CLI 327 passed, shared 415 passed, and skills catalog 20 passed; the aggregate DB stage ran both source and built copies of one unrelated embedded-Postgres migration test and both reached its 5-second timeout - `pnpm --filter @paperclipai/db exec vitest run src/status-card-migrations.test.ts --reporter=verbose` — isolated aggregate-timeout verification passed in 3.99 seconds - `NODE_ENV=production pnpm build` - `pnpm check:token-gates` ## Risks The reaper intentionally fails closed when issue terminality, pull-request state, git ancestry, or checkout ownership cannot be proven. GitHub lookups can delay classification and cleanup but cannot cause an unproven workspace to be archived. Automated terminal archival holds exclusive Git index, HEAD, and branch-ref locks across validation and removal, skips configured destructive hooks, and uses non-forced removal so dirty writes fail closed. Reopening a source issue does not restore an archived workspace; it emits an audit event so a human or agent can re-provision explicitly. ## Model Used OpenAI Codex, GPT-5. The runtime did not expose a more specific model ID or context-window size. Reasoning, tool use, repository editing, test execution, and GitHub CLI access were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c2b41bb7cd |
fix(issues): quiet missing-disposition warnings while a live continuation is running (#10899)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - When an agent run ends without recording a disposition, Paperclip
raises a "missing disposition" handoff so the work does not silently
stall
> - The server already tracks whether such an issue has a live
continuation (a running or queued run, or a queued wake) in
`successfulRunHandoff.hasLiveContinuation`
> - But no UI surface read that flag, so an issue that an agent was
actively working on still showed the "This task still needs a next step"
banner, a loud thread warning, and "Needs next step" badges
> - This pull request makes every missing-disposition complaint respect
liveness: warn only when no live agent is on the issue and it is really
stuck
> - The benefit is that users see the warning only when action is
needed, and the noise disappears while an agent is already handling the
issue
## Linked Issues or Issue Description
No public GitHub issue exists for this bug. Description follows the
bug-report template:
**What happened?**
An issue that a live agent run was actively working on showed the
"missing disposition" warning banner, a loud thread notice, and "Needs
next step" badges at the same time. The API payload for that issue
showed `successfulRunHandoff.required: true` together with
`hasLiveContinuation: true` and a `liveRunId`, but the UI ignored the
liveness fields.
**Expected behavior**
The missing-disposition warning appears only when the issue has no live
run or queued wake. A live agent records a disposition when its run
ends. Paperclip complains only if the run ends and no disposition
exists.
**Steps to reproduce**
1. Let a run finish on an in-progress issue without a disposition.
Paperclip raises the handoff and queues a corrective wake.
2. Open the issue page while the corrective run (or any new run) is
live.
3. See the banner, the badges, and the loud thread notice — all visible
while the agent works.
**Paperclip version or commit**
Current `master` (reproduced at commit
|
||
|
|
5888cbf72a |
fix(issues): restore checkout after accepted confirmations (#10909)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Issue thread confirmations can pause an issue until a board user makes a decision > - Atomic checkout is the only supported transition into `in_progress` > - An accepted confirmation left a creator-owned issue in `in_review` while it started a continuation worker > - The worker could run without the normal checkout state transition > - This pull request returns that narrow review state to `todo` before it queues the continuation wake > - The benefit is that the worker can check out the issue and move it to `in_progress` through the normal atomic path ## Linked Issues or Issue Description No matching public GitHub issue exists. The following pull requests are related but do not fix this case: - Refs #10376. It handles refusal paths for user-owned issues. - Refs #8516. It handles rejected confirmations for user-owned issues. - Refs #10274. It gives ownerless waking interactions an agent owner. **What happened?** An agent created a confirmation on an issue that was assigned to that same agent and had status `in_review`. A board user accepted the confirmation. Paperclip started a continuation worker, but the issue stayed `in_review`. The normal checkout fields stayed empty. **Expected behavior** Paperclip must return the issue to an actionable state before it wakes the continuation worker. The worker must then use atomic checkout to move the issue to `in_progress`. **Steps to reproduce** 1. Assign an issue to an agent and set the issue status to `in_review`. 2. Let that agent create a `request_confirmation` with `wake_assignee_on_accept`. 3. Accept the confirmation as a board user. 4. Observe that the continuation worker starts while the issue remains `in_review`. **Paperclip version or commit** The bug reproduced on master before this pull request. This branch is based on `ffd62a4cbb`. **Deployment mode** Local development. The server logic is deployment-independent. **Agent adapter(s) involved** Codex exposed the bug, but the issue-thread continuation logic is adapter-independent. **Database mode** The regression test uses embedded PostgreSQL. The logic is database-mode independent. **Access context** An agent creates the confirmation. A board user accepts it. ## What Changed - Allow an accepted agent-authored confirmation to return an agent-owned issue only when the issue is `in_review` and the owner is the creating agent. - Keep active `in_progress` work unchanged so an accepted confirmation cannot reset a running worker to `todo`. - Add embedded-PostgreSQL regression coverage for user-owned review, creator-owned review, and creator-owned active work. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/issue-thread-interactions-service.test.ts --config vitest.config.ts` — 48 passed. - `pnpm -r typecheck` — passed for all workspace projects. - `pnpm build` — passed for all workspace projects. - `pnpm test:run` — 3,411 passed. Three timing-sensitive assertions failed in the unchanged `heartbeat-workspace-busy.test.ts` suite. - Isolated rerun of `heartbeat-workspace-busy.test.ts` — 15 passed. ## Risks Low risk. The behavior change is limited to accepted confirmations on non-terminal `in_review` issues that the creating agent already owns. It does not change active work, blocked work, terminal issues, other agent owners, schemas, or public API contracts. > This is a focused bug fix. It does not add roadmap scope. ## Model Used OpenAI Codex based on GPT-5. The runtime does not expose the exact deployment ID or context-window size. The model used reasoning, repository tools, code editing, Git, and local test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6ffe9df842 |
fix(auth): clarify protected-agent assignment blocks (#10893)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task assignment policies control which agents can receive work. > - Protected-agent policy flags currently stop assignment. > - The existing error says that the assignment requires approval. > - Paperclip has no approval workflow for this policy. > - This pull request models the policy as a hard block and gives the operator an action that exists. > - The benefit is accurate API guidance without weakening the existing fail-closed behavior. ## Linked Issues or Issue Description Refs #6386 **What happened?** A protected-agent assignment denial said that approval was required. No approval record or approval action existed for this policy, so the message sent agents and operators to a dead end. **Expected behavior** The authorization result must state that protected-agent policy blocks assignment. It must tell a company administrator to remove the block before retrying. **Steps to reproduce** 1. Set `authorizationPolicy.protectedAgent.requiresApproval` to `true` on a target agent. 2. Give another agent the `tasks:assign` permission. 3. Preview or attempt assignment to the protected agent. 4. Observe that the old response promises an approval step that does not exist. **Paperclip version or commit** `c54936e2e9` on `master`. **Deployment mode** Built from source. The behavior is in the core authorization service and is not deployment-specific. **Agent adapter(s) involved** Not adapter-specific. ## What Changed - Added canonical `protectedAgent.blockAssignment` and `protectedAgent.blockReason` policy fields. - Kept the legacy approval-named flags as fail-closed compatibility aliases. - Changed denial copy to name the hard block and the administrator action. - Added authorization and plugin-host regression coverage for canonical and legacy policy data. - Updated the V1 implementation contract with the protected-assignment rule. ## Verification - `pnpm exec vitest run server/src/__tests__/authorization-service.test.ts server/src/__tests__/plugin-access-authorization-host-services.test.ts` — 2 files passed, 61 tests passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/shared build` — passed. - `pnpm --filter @paperclipai/server build` — passed. - `pnpm check:token-gates` — all gates clean. - `git diff --check public-gh/master...HEAD` — passed. The repository-wide local wrappers exceeded the execution host resource limit before they printed a final summary. The PR check loop will use GitHub CI as the complete test and build authority. ## Risks - Low: assignment remains fail-closed. The change corrects the policy name and denial guidance. - Low: legacy fields remain supported, so existing plugin-owned policy data does not change behavior. - Low: the new policy schemas allow unknown keys for forward compatibility, as the existing authorization policy schema already does. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5`, tool-enabled coding agent with reasoning, shell, Git, and GitHub CLI access. The runtime does not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ef33c1d9ed |
fix(decisions): retire completed-target decisions and link targets from the card (#10892)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The decisions desk shows pending decisions that need an operator response > - A strict decision cannot apply its effects after its target task changes > - A decision still remained pending when every target task finished after proposal > - The card also linked only the origin task, even when the decision acted on another task > - This pull request expires those moot decisions and links their target tasks > - The benefit is an accurate queue and a clear path to the work that each decision affects ## Linked Issues or Issue Description Related PR: #10801 removes the issue-page decision strip, which makes clear queue provenance more important. **What happened?** A strict decision stayed pending until its time-to-live limit after every target task reached `done`. The decision card linked only the origin task. The origin task is where the agent proposed the decision, and it can differ from the task that the decision affects. An operator could therefore open a finished task with no visible decision and no explanation of the real target. **Expected behavior** Paperclip must expire a strict decision when all of its targets finish after the decision is proposed. The card must show and link every target task that differs from the origin task. **Steps to reproduce** 1. Create a strict decision that targets an active task from a different origin task. 2. Move the target task to `done` without resolving the decision. 3. Run the decision expiry sweep. 4. Observe that the old code keeps the decision open until its time-to-live limit. 5. Observe that the old card links only the origin task. **Paperclip version or commit** The bug reproduces on upstream `master` before this pull request. **Deployment mode** Local dev and self-hosted server modes are affected because the behavior is in the shared decision service and board UI. ## What Changed - Expire an open strict decision with reason `target_completed` when every strict target reached `done` after proposal. - Keep decisions that intentionally target an already-finished task. - Keep lenient-only decisions open. - Keep continuation delivery consistent with other expiry reasons. - Add target-task links to the decision card provenance line. - Use one shared target-ID helper across signing, execution, expiry, card provenance, and resolver preloading. - Add service and UI regression tests for primary, secondary, and target-completed cases. ## Verification - `pnpm exec vitest run ui/src/components/DecisionCard.test.tsx server/src/__tests__/decisions-service.test.ts` — 51 tests passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm check:token-gates` — all gates clean. - `git diff --check origin/master...HEAD` — passed. ## Risks - Low migration risk. This change does not alter the database schema. - The expiry sweep performs the existing strict-target query and adds a snapshot comparison before expiry. - A decision remains open if any strict target is active or if a target was already `done` at proposal time. > The roadmap lists work queues as planned. This pull request fixes the existing decisions desk. It does not add a new queue subsystem. ## Model Used - Implementation: Anthropic Claude through Claude Code. The runtime did not expose the exact model snapshot or context-window size. The model used reasoning, repository tools, code execution, and test execution. - PR preparation: OpenAI Codex with GPT-5. The runtime did not expose a dated model snapshot or context-window size. The model used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
68ddd6a7a0 |
feat(activity): add two-tier all-actors audit feed (#10831)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Operators need one activity feed for human, agent, plugin, and system changes > - The existing audit endpoint returns only rows that have agent attribution > - The full audit view also requires a dedicated permission > - This pull request adds an explicit all-actors scope with basic and privileged access tiers > - The benefit is that company members can inspect the shared activity history while sensitive attribution and export controls stay protected ## Linked Issues or Issue Description **What existing behavior does this improve?** The company audit activity endpoint and the board audit route. **Subsystem affected** Server REST API and board UI routing/API contracts. **Current behavior** The agent-action audit endpoint excludes activity without an agent ID. It also rejects company members who do not have the full audit permission. **Proposed behavior** Callers can opt into `actorScope=all`. A company member receives all actor kinds with sensitive attribution fields removed. A permitted board user receives complete rows and can use attribution filters. The default scope and CSV permission remain unchanged. **Reason and benefit** The board needs one chronological activity source for user, agent, plugin, and system actions. A two-tier response keeps the feed useful without widening access to detailed attribution or export capabilities. **Breaking changes** None. The endpoint keeps the existing agent-only scope and permission behavior by default. ## What Changed - Added `actorScope=all` to the unified audit query and included activity from every actor type. - Added a company-readable basic tier that removes run, responsible-user, agent, and details attribution. - Kept attribution filters and CSV export behind `audit:view_agent_actions`. - Added route and integration coverage for basic readers, permitted readers, pagination, filter denial, and all actor kinds. - Added the missing unprefixed `/audit` redirect and company route classification. ## Verification - `pnpm exec vitest run server/src/__tests__/activity-routes.test.ts server/src/__tests__/agent-action-audit-routes.test.ts ui/src/lib/company-routes.test.ts --reporter=verbose` (35 tests passed) - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` ## Risks - The all-actors query can return more rows than the legacy agent-only query. Cursor pagination and existing limits bound each request. - The basic tier intentionally exposes action and actor-kind context. It removes detailed run, agent, responsible-user, and details attribution. - The legacy endpoint behavior remains the default, which reduces compatibility risk. > The roadmap marks activity log and action attribution as shipped. This change improves that existing capability and does not introduce a separate workflow system. ## Model Used - OpenAI Codex, `gpt-5.6-sol`, 114K context, agentic reasoning with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e8ae5286eb |
feat(issues): explain cross-task agent writes with attribution, audit receipts, and actionable denials (#10843)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents write to tasks they do not own. They comment, they change
fields, and the control plane now permits this by default for
standard-trust agents on any task they can read
> - This makes a task thread ambiguous. A reader sees a comment from an
agent that is not the assignee, but no surface says whose authority that
write rode
> - The same gap applies to field edits. The activity stream named the
verb, but it did not show the before value, the after value, or the
reason the write was permitted
> - The remaining refusals are also opaque. An agent that hits a wall
receives a 403 with no boundary name, no actor who can act, and no
sanctioned path. One real incident spent a full detour to find the
workaround
> - This pull request adds the three surfaces that make open cross-task
writes legible: an attribution chip, a field-level audit receipt, and an
actionable denial contract shared by the API and the UI
> - The benefit is that a reader can answer "who did this, on whose
authority, and was it allowed?" on the task itself, and a blocked writer
is told what to do next
## Linked Issues or Issue Description
No public issue exists for this work, so the enhancement is described
here.
**What existing behavior does this improve?**
Cross-task agent writes are permitted, but they are not explained. A
task thread can hold comments from agents that are not the assignee, and
the activity stream can hold field changes made by those agents. Neither
surface names the responsible user behind the write. When a write is
refused, the error text does not name the boundary or the way forward.
**Subsystem affected**
Issue detail UI (comment thread and activity stream), the issue write
authorization responses in the server, and the shared copy contract that
both consume.
**Current behavior**
- An agent comment on a task the agent does not own looks the same as an
assignee comment.
- An `issue.updated` activity row states the verb only. It does not show
the field-level before and after values, the responsible user, or the
authorization reason.
- A refused write returns a short message such as an ownership error.
The message does not state which rule fired, who is able to perform the
action, or which alternative path is sanctioned.
**Proposed behavior**
- An agent comment on a task the agent does not own carries a chip that
reads "for {user}". The chip names the responsible user. Its tooltip
states that the author is not the assignee and cannot exceed that user's
permissions.
- Each `issue.updated` row shows a receipt: the changed fields with
before and after values, the responsible user, and the authorization
reason. This applies to board edits as well as agent edits.
- Each refusal states three things: the boundary that fired, who is able
to act, and the sanctioned path. The API error body and the in-app
notice use the same words, because both read one shared contract.
Related pull requests, found by searching this repository:
- Refs #10837 — merged. It added the default-open cross-task write rule,
the comment attribution data, and the per-run containment cap that this
pull request makes visible.
- Refs #10114 — open. It proposes a narrower authorization change in the
same area.
- Refs #7998 — open. It proposes append-only cross-assignee comments as
an alternative to opening writes.
## What Changed
- Adds `packages/shared/src/issue-write-denial.ts`. This is one copy
contract for eight ways an issue write can be refused: not visible,
responsible-user ceiling, responsible user unavailable, excluded actor
class, assignee run lock, per-run cross-task cap, missing run context,
and rejected attribution. Each entry names the boundary, who can act,
and the sanctioned path.
- Maps server authorization decisions onto that contract in
`server/src/routes/issues.ts` and
`server/src/services/cross-issue-influence-limit.ts`. The flattened
`error` string carries all three obligations, and `details.code` lets
the UI render the same words. The two cap codes keep the names they
already ship under.
- Adds `CommentAttributionChip`. It renders "for {user}" beside the
author name on agent comments where the author is not the assignee. It
renders nothing when no responsible user is recorded, so older rows stay
clean. It is wired into both `IssueChatThread` and the flagged
`TaskChatThread` redesign.
- Adds `IssueFieldChangeReceipt`. It renders the change receipt under
`issue.updated` rows in the activity stream. Ids resolve to agent and
user names where the directory is loaded. Server-truncated text is
labelled as a preview, so the receipt never implies that it shows a
whole value.
- Adds `IssueWriteDenialNotice`. It renders the shared copy in the app,
keyed off the denial events the server logs on a task.
- Adds a public `/ux-lab/cross-issue-collaboration` page. It renders all
three surfaces and their edge cases for review without a seeded thread.
This follows the existing `ux-lab` pages.
## Verification
Automated, all green:
```
pnpm --filter @paperclipai/shared exec vitest run src/issue-write-denial.test.ts # 17 tests
pnpm --filter @paperclipai/ui exec vitest run src/components/IssueWriteDenialNotice.test.tsx \
src/components/IssueFieldChangeReceipt.test.tsx src/components/CommentAttributionChip.test.tsx \
src/lib/issue-change-receipt.test.ts src/lib/comment-attribution.test.ts # 46 tests
pnpm --filter @paperclipai/server exec vitest run src/__tests__/cross-issue-influence-limit.test.ts \
src/__tests__/issue-comment-attribution-audit-routes.test.ts \
src/__tests__/issue-agent-mutation-ownership-routes.test.ts \
src/__tests__/low-trust-red-team-routes.test.ts # 98 tests
```
`tsc --noEmit` passes for the shared, ui, and server packages.
Manual, in a browser:
1. Start the UI only: `pnpm --filter @paperclipai/ui exec vite`.
2. Open `/ux-lab/cross-issue-collaboration`. No session is needed,
because `ux-lab` routes are public.
3. All three surfaces were captured at 1440x900 in light mode and dark
mode, and at 390x844. The page reported no errors.
4. The chip tooltip was opened by a hover and by a keyboard focus.
Rendering the page found defects that the tests had missed. Three copy
and contrast defects were fixed, and two of them are now pinned by a
test. A design review then found three layout defects, which are also
fixed: the denial notice orphaned its label when a value wrapped, the
receipt icon wrapped onto its own line at narrow widths, and the chip
tooltip was reachable by hover only.
## Risks
Low risk, and additive.
- Every new surface renders nothing when its data is absent. Comments
without a recorded responsible user show no chip, and activity events
without a receipt show no receipt, so existing rows do not change.
- No migration is included. The data these surfaces read already ships.
- The wire values of the two per-run cap denial codes are unchanged.
Only the human-readable text changes, plus six codes that had no
`details.code` before.
- The denial copy is read by agents as well as people. If wording must
change later, one shared module is the only place to change it.
- Roadmap check: this extends the completed "Activity log & action
attribution" area rather than duplicating planned core work.
## Model Used
Claude Opus 5 (Anthropic), model id `claude-opus-5[1m]`, 1M context
window, extended thinking, with tool use and code execution. It ran as
an agent in Claude Code and drove a real browser to capture the review
screenshots.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
5858ccb981 |
feat: make in-app features cloud-aware (#10850)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators use the same board application in self-hosted and Paperclip Cloud deployments. > - A Cloud tenant contains one company, so an in-app company switch does not change the active Cloud stack. > - Cloud operators need the sidebar and company surfaces to use the signed-in user's stack portfolio. > - The server must derive Cloud identity and links from trusted instance context instead of client input. > - This pull request adds canonical Cloud context, a trusted stack portfolio proxy, and Cloud-aware navigation. > - The benefit is consistent stack switching on Cloud while self-hosted company behavior stays unchanged. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: server REST routes and the React board UI. **Problem or motivation** A Cloud-managed instance contains one company. The existing company switcher could only switch records inside that tenant. It could not move the operator to another Cloud stack. The existing header also gave long organization names too little width. **Proposed solution** Expose a canonical public Cloud context in health data. Add a trusted server proxy for the current user's stack portfolio. Use that data in the board UI to switch stacks with top-level navigation. Keep the existing company behavior on self-hosted instances. Move search into the navigation and keep long organization names inside the sidebar panel. **Alternatives considered** An in-app `/stacks` route was rejected because Cloud tenant hosts reserve that path and stack selection must wake or authenticate another tenant. Client-supplied user identity was rejected because the server can derive the trusted Cloud actor. **Roadmap alignment** This change advances the Cloud deployments milestone. It keeps the product local-first and Cloud-ready without changing the self-hosted mental model. ## What Changed - Added canonical Cloud instance context and public health metadata. - Added a Cloud-only stack portfolio proxy with trusted actor forwarding and per-user caching. - Prevented normal company creation on Cloud-managed instances. - Switched the sidebar and Companies page from company actions to stack actions on Cloud. - Added full-page stack navigation and Cloud create-stack links. - Moved search into the sidebar navigation so the organization name keeps more width. - Added truncation and hover recovery for long organization and stack names. - Added server and UI regression coverage for Cloud and self-hosted behavior. - Updated the implementation specification for the Cloud contracts. ## Verification - `node scripts/check-token-gates.mjs` passed. All three token gates are clean. - `pnpm --dir server exec vitest run src/__tests__/health.test.ts src/__tests__/cloud-instance.test.ts src/__tests__/cloud-routes.test.ts src/__tests__/company-cloud-floor.test.ts src/__tests__/company-portability-routes.test.ts` passed: 5 files and 66 tests. - `pnpm --dir ui exec vitest run src/components/SidebarCompanyMenu.test.tsx` passed: 1 file and 11 tests. - Pre-PR QA report `7da87ca7` passed all 8 acceptance criteria with real HTTP route factories and real Chromium screenshots in Cloud and self-hosted modes. - Security reviews passed for the canonical Cloud context and stack portfolio proxy. ## Risks - Cloud stack switching depends on the configured Cloud application and tenant portfolio URLs. - The new health `cloud` block is public by design, but it contains only canonical public instance metadata. - The stack proxy fails closed on self-hosted instances and derives the user identity from the trusted actor. - Self-hosted navigation and company creation retain their existing paths and behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, model `gpt-5`. The run used reasoning, repository tools, shell execution, and GitHub integration. The deployment did not expose its context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d114c4925e |
feat(observability): give sandbox sync spans true wall-clock width (#10864)
## Thinking Path > - Paperclip keeps company work visible and governed. > - Sandbox agents run serial sync work across worker and host boundaries. > - The current span path hid real wall-clock time for that sync work. > - The host needs safe timestamps if it wants true span width. > - This pull request carries worker timestamps, validates them, and records the real duration. > - The benefit is clearer operator visibility for sandbox sync work. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This touches `packages/plugins`, `server`, and the Daytona plugin test surface. **Problem or motivation** Sandbox sync spans opened and closed in one host call. The native width stayed near zero, so the real time spent in serial round trips was hard to see. **Proposed solution** Carry worker start and end times across the span record protocol. Validate the pair at the host boundary. Record the host span with the true duration when the pair is safe. **Alternatives considered** Keep the numeric duration only. That keeps the data, but it does not widen the span and it does not show the real wall-clock time. **Roadmap alignment** This fits the `Cloud / Sandbox agents` and `Artifacts & Work Products` areas in `ROADMAP.md`. I found no other roadmap item that covers this span-width gap. **Additional context** The host allowlist stays narrow. Unknown names still map to `sandbox.provider.other`. Invalid timestamp pairs still fall back to the synchronous path. Related public PRs: none found. ## What Changed - Added optional `startTimeMs` and `endTimeMs` fields to the `span.record` protocol. - Captured start and end times in the worker tracer and sent them to the host. - Validated host timestamps with finite, ordered, bounded checks before span reconstruction. - Extended the host allowlist to the sandbox sync command names. - Wrapped each inbound sync round trip in its own named span. - Added tests for the worker path, host boundary, host recorder, and Daytona sync flow. ## Verification - `pnpm --filter @paperclipai/plugins-sdk test` - `pnpm --filter @paperclipai/server test` - `pnpm --filter @paperclipai/daytona-plugin test` - `pnpm --filter @paperclipai/server tsc --noEmit` still shows pre-existing `drizzle-orm` duplicate-declaration errors in this sandbox. The changed files do not touch those lines. - GitHub checks are green. - Greptile review is 5/5. - No open review threads remain. ## Risks - A bad timestamp pair can fall back to the synchronous path. - The host clock gate can reject spans if the pair is stale, reversed, or too large. - The new worker fields change the wire protocol, but the public plugin tracer contract stays the same. ## Model Used OpenAI GPT-5, tool-enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues with `Fixes: #` / `Closes #` / `Refs #` or described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2ebc236b08 |
fix(server): accept parentIssueId alias in GET /issues (#4032)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents coordinate through the server API. They find sub-tasks by filtering the company issues list by parent. > - `GET /api/companies/:companyId/issues` accepts `?parentId=`. Many callers send `?parentIssueId=` instead, which the handler never read. > - The mismatch is silent. The filter is dropped and the full company list comes back, so agents fetch everything and filter client-side. Issue #3846 reports this. > - `parentIssueId` is not an arbitrary spelling. It is the field name the wakeup payloads in this same route file already use, so callers expect it. > - This pull request accepts `parentIssueId` as an alias for `parentId` at the route boundary, on both the issues list and `issues/count`. > - The benefit is that parent filtering works for both spellings, and the list and its count cannot disagree. ## Linked Issues or Issue Description Fixes #3846 Related: #3870 proposes the same alias for the list route. ## What Changed - `server/src/routes/issues.ts`: `listFilters.parentId` in `GET /companies/:companyId/issues` now reads `req.query.parentId ?? req.query.parentIssueId`. - `server/src/routes/issues.ts`: `blockedCountFilters.parentId` in `GET /companies/:companyId/issues/count` reads the same alias, so the list and its count agree. - `server/src/__tests__/issues-parent-id-alias.test.ts`: new regression test for alias resolution, precedence, and absence. ## Verification - Run `pnpm run test:run -- server/src/__tests__/issues-parent-id-alias.test.ts`. - The test covers four query shapes: `?parentId=`, `?parentIssueId=`, both present (short form wins), and neither present (filter unset). - Existing callers are unaffected. The UI client `ui/src/api/issues.ts` only sets `parentId`. Nullish coalescing falls back only when the primary key is absent. - The service layer applies the filter with `if (filters?.parentId)` in `server/src/services/issues.ts`. This pull request does not change it. ## Risks - Low risk. The change only widens accepted query input. Both spellings resolve, and the short form still wins. - `?parentId=` with an empty value stays falsy and unfiltered, exactly as before. - This route has no validation middleware, and these list filters are not in the published OpenAPI surface. No contract needs an update. ## Model Used - Claude Opus 5 (`claude-opus-5`), extended thinking with tool use, run by the maintainer's triage agent. It rebased the original commit onto current `master`, extended the alias to `issues/count`, and wrote the regression test. @scokeepa authored the original one-line route change. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: josangmun <cmeia.ai02@cmeia.co.kr> Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com> |
||
|
|
2495e29f7f |
Preserve company display names in cloud tenants (#10845)
Prefer the trusted organization name, repair known machine-generated legacy names with compare-and-set safety, and preserve the audited fallback behavior required by PAP-16331. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
c647b8cc2e |
feat(acpx-engine): give sandbox.exec spans real parents (#10852)
## Thinking Path > - Paperclip uses spans and traces to show how work moves through agents and tools > - sandbox.exec spans need a real parent so the trace tree matches the work tree > - Wrong parent links make execution history hard to read and hard to debug > - This pull request adds a single task.run root span and re-parents live work to the nearest active span > - The change keeps detached work under the closest live span instead of the HTTP root > - The benefit is a clear trace tree for sandbox.exec work and better execution diagnosis ## Linked Issues or Issue Description **What happened?** sandbox.exec spans attached to the wrong parent or to no live parent in some paths. **Expected behavior** Each sandbox.exec span should attach to the nearest live span. **Steps to reproduce** 1. Run work that creates sandbox.exec spans during startup and callback bridge paths. 2. Inspect the trace tree. 3. Observe an orphaned span or a span with the wrong parent. **Paperclip version or commit** `672e9de9c8b004aebc1f08e24b612ab067735ad1` **Deployment mode** Local dev. **Additional context** The branch adds the task.run root span, parents sandbox.startup to it, and re-parents detached bridge work to the nearest live span. ## What Changed - Added a task.run root span for the run tree. - Re-parented sandbox.startup, agent.turn, and detached bridge work to the nearest live span. - Added end-to-end trace-tree assertions for the full parent chain. - Added negative coverage so sandbox.exec does not parent to the HTTP root. ## Verification - Focused Vitest suite passed: `packages/adapter-utils/src/acpx-engine/execute.test.ts`, `packages/adapter-utils/src/acpx-engine/startup-timing.test.ts`, `packages/adapter-utils/src/execution-target-sandbox.test.ts`, `packages/adapter-utils/src/sandbox-callback-bridge.test.ts`, and `server/src/__tests__/environment-execution-target.test.ts`. - Result: 5 files passed, 204 tests passed. - The submitted branch also reported `adapter-utils` checks, `server` seam checks, and `tsc` exit 0 in the handoff state. ## Risks - This change can alter trace tree shape in tools that read parent spans. - A missed bridge path could still point to the wrong live span. - Low risk for runtime behavior, because the change only changes span parent attribution. ## Model Used OpenAI GPT-5, tool-use capable. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
74416e9cd1 |
fix(server): render company-export YAML iteratively to stop stack overflow (#10854)
## Thinking Path > - Paperclip is the control plane for autonomous AI companies. > - The company export path writes `.paperclip.yaml` data for large companies. > - The YAML renderer used a spread append that can overflow the call stack on large arrays. > - That failure turns a normal export into a 500 for large companies. > - This pull request rewrites the renderer to use an iterative stack and removes the last spread append. > - The benefit is that large exports finish without a RangeError and keep the same output. ## Linked Issues or Issue Description I searched GitHub for related work. I found PR #7506. This pull request closes the last spread site that PR left open. **What happened?** The company export failed with `RangeError: Maximum call stack size exceeded` on large YAML output. **Expected behavior** The export should finish without a stack overflow. **Steps to reproduce** 1. Export a company with a very large YAML payload. 2. Render the export through `renderYamlBlock` or `renderFrontmatter`. 3. Observe that the old spread append can overflow the call stack. **Paperclip version or commit** `79f3a216215500e2ec1a928d5eb5c09364c2abf5` **Deployment mode** Local dev (`pnpm dev`) or built from source. **Additional context** Related public PR: #7506. This change keeps the YAML shape, scalar format, and key order the same. ## What Changed - Reworked `renderYamlBlock` to render iteratively. - Replaced the last spread append in `renderFrontmatter` with a loop. - Added regression tests for high-volume block and frontmatter arrays. ## Verification - `node_modules/.bin/vitest run server/src/__tests__/company-portability.test.ts` - The two new overflow tests pass. - The existing round-trip tests still pass. - `tsc --noEmit` is clean for `server/src/services/company-portability.ts`. ## Risks Low risk. The change keeps exported YAML content and ordering the same. ## Model Used OpenAI GPT-5, tool-using coding agent. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
678728f650 |
feat: maintained in_review review-path contract + stalled-review actions (#10675)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents move issues to `in_review` and rely on a "review path" (an interaction, an approval, a monitor, or a named reviewer) to tell them who decides next. > - That review path can silently disappear. A user comment supersedes the pending interaction, a monitor is exhausted, or a run ends without restoring a path. The issue then sits in `in_review` with nobody reviewing it and no visible action. > - Such issues become invisible zombies. Nobody knows a decision is owed, so the work stalls forever. > - This pull request makes the review path a maintained invariant, exposes a `reviewAttention` surface, and gives every stalled review three inline actions in the UI. > - The benefit is that an `in_review` issue always shows who reviews it, or shows an amber "nobody is reviewing this" notice with one-click Approve, Request changes, and Send back to work. ## Linked Issues or Issue Description This pull request describes the problem inline. The tracking issue is internal. **Subsystem affected** The review and attention loop that agents and humans share: the `in_review` status, the `reviewAttention` surface, the /decisions attention feed, and the issue-page review panel. **Problem or motivation** Agent-owned issues in `in_review` can lose their last review path. A user comment supersedes the pending interaction. A monitor is exhausted. A run ends without restoring a path. The issue then sits in `in_review` with no reviewer and no visible action. It becomes an invisible zombie and the work never progresses. **Proposed solution** Maintain the review path as a server invariant. Expose a `reviewAttention` field that says what is under review, who decides, and since when. Render a persistent review panel on the issue page and inline actions on the /decisions feed. Keep human PATCHes into `in_review` ungated, but record the requesting user so the panel never renders empty. **Alternatives considered** A pure background auto-recovery sweep. This stays opt-in and is not enough on its own, because it is invisible to the human. A bare status banner. This is rejected, because it gives no action to resolve the stall. **Roadmap alignment** This improves the core review and attention loop that both agents and humans use every day. ## What Changed - **Server — maintained review-path invariant:** when an issue enters or sits in `in_review`, the server derives and persists a review path (interaction, approval, monitor, or the requesting user) and recovers a stale path with one bounded wake instead of leaving the issue pathless. - **Server — `reviewAttention` surface:** a new field describes what is under review (bound target with links), who decides, since when, and whether the review is stalled. Stalled agent-assigned reviews are now included in the attention feed. - **Server — inline stalled-review decisions:** secured routes let a permitted responder Approve (→ `done`), Request changes (→ `todo` + wake carrying the note), or Send back to work (→ `todo` + wake) directly from the attention feed. - **Server — resume-intent wake:** an `in_review -> todo` transition now wakes the assigned agent so a resumed review is not dropped. - **Server — user-entry symmetry:** user PATCHes into `in_review` stay ungated (no 422 for humans) and record the requesting user, who becomes the named responder when no other path exists. - **UI — review panel:** a persistent `IssueReviewPanel` renders above the thread whenever status is `in_review`. The covered state shows the bound target, responder, and outcomes and hoists the pending interaction/approval card. The stalled state shows the amber notice plus the three actions. - **UI — decisions card actions:** the same three actions render inline on the /decisions `AttentionQueueRow`. - **UI — responsive fix:** the stalled action row stacks to full-width buttons at phone width and returns to a horizontal row at `sm` and up. New 390px stories capture the phone layout. ## Verification - `cd ui && npx vitest run src/components/IssueReviewPanel.test.tsx src/components/AttentionQueueRow.test.tsx src/lib/attention.test.ts src/api/issues.test.ts` — 91 tests pass. - Server suites added and updated: `issue-review-attention`, `issue-stalled-review-decision-routes`, `review-path-recovery`, `recovery-observability`, and related route/liveness tests (run by CI). - A designer reviewed the UI at 390px and desktop in light and dark themes on both the issue-page panel and the /decisions card. The stalled action row stacks cleanly at phone width with no overlap and keeps the horizontal row on desktop. ## Risks - **Migration:** adds migration `0200` (next after master `0199`, no renumber). It extends the agent-wakeup-requests schema and is additive. - **Behavioral shift:** `in_review -> todo` now dispatches a wake. This is intended (resume intent) and covered by tests. - **Authz:** the inline decision routes are permission-gated. Only a permitted responder sees and can trigger the actions. - Overall risk is moderate and contained to the review and attention loop. ## Model Used - Claude, Opus 4.8 (`claude-opus-4-8`), extended thinking, tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f91a6e27c0 |
feat(issues): contain cross-issue agent side effects (#10837)
## Thinking Path > - Paperclip is the control plane that coordinates autonomous agent work. > - Agents need to collaborate on issues beyond their current assignment. > - Cross-issue comments and updates are useful, but an unbounded run can create cascading side effects. > - The control plane must preserve company-wide collaboration while containing each run's influence. > - Comment attribution must also show the responsible user and the acting agent in audits. > - This pull request adds run-bound cross-issue containment, attribution, and agent-class wake rules. > - The benefit is safer collaboration without restoring issue-assignee ownership restrictions. ## Linked Issues or Issue Description **What existing behavior does this improve?** Agent-authenticated issue comments, updates, reopen behavior, and assignee wake routing. **Subsystem affected** Cross-cutting: server routes and services, shared contracts, database schema and migration, and implementation documentation. **Current behavior** An authenticated agent can collaborate across company issues, but one heartbeat run has no per-run side-effect boundary. Comment records also do not persist the responsible user separately from the acting agent. **Proposed behavior** Require a valid heartbeat run for agent cross-issue comments and updates. Audit each attempt and cap a run at 20 cross-issue effects. Keep the cap in log-only mode until it automatically changes to enforcement at 2026-08-11 00:00 UTC. Preserve same-issue writes. Use agent-class wakes for agent comments. Keep same-run completion comments from reopening completed work. Record the responsible user on agent-authored comments and activity. **Reason and benefit** Agents can collaborate on other issues without an assignment gate, while each run has an atomic and inspectable side-effect limit. Operators can identify both the acting agent and the responsible user. **Breaking changes** After 2026-08-11 00:00 UTC, the twenty-first cross-issue comment or update from one heartbeat run returns a containment error. Agent cross-issue writes without valid run context are rejected. The migration is additive and backfills existing agent-authored comment attribution where the source data is available. ## What Changed - Added an atomic per-run counter for cross-issue agent comments and updates. - Added audit events for allowed and rejected cross-issue effects. - Added the automatic log-only to enforcement flip at 2026-08-11 00:00 UTC. - Added responsible-user attribution to agent-authored comments, activity records, shared types, and validators. - Added an additive migration and migration coverage for existing comments. - Updated reopen, resume, and wake behavior so agent comments create agent-class wakes and same-run completion comments remain inert. - Updated the implementation specification and regression coverage. ## Verification - `pnpm exec vitest run server/src/__tests__/cross-issue-influence-limit.test.ts server/src/__tests__/issue-comment-attribution-audit-routes.test.ts server/src/__tests__/issue-comment-reopen-routes.test.ts packages/db/src/issue-comment-on-behalf-migration.test.ts` — 97 tests passed. - `pnpm -r typecheck` — passed, including migration safety checks. - `pnpm test:run` — server batch: 3,364 passed and 2 skipped; UI batch: 3,504 passed. One unrelated CLI doctor test warned because this agent runtime injects static AWS credentials. - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts` — 8 tests passed and confirmed the CLI failure was ambient-environment sensitive. - `pnpm build` — passed. ## Risks - The fixed enforcement timestamp changes production behavior automatically on 2026-08-11 00:00 UTC. Audit logs before that time provide rollout visibility. - The per-run counter serializes on the heartbeat-run row. This prevents concurrent attempts from racing past the cap but adds a small lock scope for cross-issue writes. - Existing comments can only be backfilled when their acting run or agent attribution is recoverable. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 in the Codex agent runtime. The runtime did not expose a context-window size. Reasoning, shell tools, code editing, and test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ded813ad6f |
feat(interactions): add governed agent addressees (#10252)
## Thinking Path > - Paperclip is the control plane that lets humans govern companies of AI agents. > - Issue-thread interactions are the structured handoff point for confirmations, questions, suggested tasks, and other governed decisions. > - Those interactions previously assumed that only board users could resolve them, preventing one agent from explicitly addressing another agent for a response. > - Agent resolution needs company-level governance, auditable resolver identity, safe terminal-state handling, and attention routing so authorization is enforced server-side rather than inferred from UI behavior. > - This pull request adds governed agent resolution, withdrawal and terminal expiry semantics, explicit agent addressees, lifecycle reconciliation, and attention-feed filtering. > - The benefit is that agents can participate in structured decisions without weakening board control, company isolation, wake behavior, or audit invariants. ## Linked Issues or Issue Description ### Subsystem affected Issue-thread interactions across database, shared contracts, server authorization/services, adapter callbacks, agent skill guidance, API docs, and UI governance surfaces. ### Problem or motivation Structured interactions were board-only, had no explicit agent addressee, and lacked durable withdrawal/terminal-expiry semantics. That made peer-agent decisions impossible to authorize and audit safely. ### Proposed solution Persist requested/effective resolver policy and addressee identity, enforce company governance and eligible agent resolution, reconcile addressee lifecycle changes, expose withdrawal and terminal expiry, and route attention to the intended active agent with board fallback. ### Alternatives considered Implicitly authorizing the issue assignee or mentioned agents was rejected as ambiguous and difficult to audit. Using comments alone was rejected because it loses structured outcomes and continuation behavior. ### Roadmap alignment Supports the ROADMAP direction for lightweight leadership-agent communication that still resolves into governed decisions and work objects. ### Additional context Public GitHub issue/PR search found no duplicate implementation; open PR search for interaction resolver governance and agent addressees only returned this PR. ## What Changed - Add company-scoped interaction resolver governance contracts and persistence. - Add requested/effective resolver policy, resolver identity, withdrawal, and terminal-expiry behavior. - Add explicit `addresseeAgentId` validation, authorization, persistence, lifecycle reconciliation, API documentation, and skill guidance. - Route pending addressed interactions to the intended invokable agent and fall back to board attention when that agent becomes ineligible or is deleted. - Preserve sandbox callback identity fields required by governed resolution paths. - Add migrations `0193` and `0194` plus route, service, attention, adapter, CLI, and UI coverage. - Add governance state and company settings UI, including responsive mobile behavior and distinct withdrawn/expired audit presentation. ## Verification - `pnpm check:token-gates` — passed. - `pnpm -r typecheck` — passed, including migration numbering and safety checks. - `pnpm test:run` — feature/server and UI workspace suites passed; one unrelated CLI AWS doctor test observed injected static AWS credentials and warned instead of passing. - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts --project paperclipai` — 8 tests passed, confirming the failure was environment-sensitive. - `pnpm build` — passed. - Latest rebased head `e24cece6be9f1877bdbac7691bcb44fd583c0161` completed all GitHub CI jobs successfully. ## Risks - Migrations add interaction and company-governance fields; numbering is conflict-free on current `master`, additive statements are idempotent, and migration safety checks pass. - Agent authorization behavior expands beyond board-only resolution, but defaults remain board-only and coverage exercises company boundaries, resolver eligibility, lifecycle invalidation, wake behavior, withdrawal, expiry, and attention fallback. - Attention routing depends on current agent invokability; reconciliation and read-time filtering prevent stale addressees from retaining visibility or resolution authority. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using `gpt-5.6-sol` with reasoning, terminal tool use, code execution, Git/GitHub integration, and Paperclip control-plane tools. Context-window metadata was not reported by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f0b06d2de9 |
feat(claude): environment-aware test-environment probe and claude-local CI coverage (#10833)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The claude_local adapter runs Claude Code on sandbox execution targets, and operators verify an agent's configuration with the test-environment probe before running it > - Real runs merge the selected environment's env vars (secret refs included) under the agent's adapter config env, but the probe built its config from the adapter config alone — so environment-level auth worked in runs while the Test button reported missing auth, and a dropped secret binding passed silently > - The claude env-test hints also did not recognize `CLAUDE_CODE_OAUTH_TOKEN` even though the CLI accepts it, and a hello probe that hit the subscription usage limit reported a hard failure although authentication worked > - Separately, the claude-local package test suites were absent from the CI project list, so two suites drifted broken without notice > - This pull request makes the probe resolve the same layered env as a real run, adds the missing auth hint, classifies usage-limit probe results as a warning, repairs the drifted suites, and turns the claude-local project on in CI > - The benefit is a Test button that tells the truth about environment-level configuration, and a test suite that actually gates the claude-local adapter ## Linked Issues or Issue Description No public issue exists; related open PRs: Refs #9488 (recognizes CLAUDE_CODE_OAUTH_TOKEN in environment checks — overlaps with the auth-hint portion of this PR via a differently named check; it does not cover the environment-envVars probe merge, the usage-limit classification, or the CI coverage), Refs #9933 (live credential validation in environment checks — complementary, no file-level conflict with the route change). The underlying problem, following the enhancement template: **Current behavior** The test-environment route builds the probe config from the agent's adapterConfig only. Real runs merge the selected environment's envVars under the agent env, so environment-level env vars (including auth such as `ANTHROPIC_API_KEY` or `CLAUDE_CODE_OAUTH_TOKEN` bound as environment secrets) work in runs while "Test environment" cannot see them, and a missing secret binding passes silently. The claude env-test hints do not recognize `CLAUDE_CODE_OAUTH_TOKEN`. A hello probe that hits the subscription usage limit reports a hard `claude_hello_probe_failed`. The claude-local package test suites do not run in CI, and two of them are stale. **Proposed behavior** The probe resolves the selected environment's envVars (environment-consumer secret bindings included) and merges them under the agent config env with the run-path precedence; missing bindings surface as an explicit error check that fails the test. The env-test emits a `claude_oauth_token_configured` info check when that variable is set. Usage-limit probe results classify as a `claude_hello_probe_usage_limited` warning because auth works and only the usage window is spent. The claude-local suites run in CI. Docs state the resulting facts. **Reason and benefit** The Test button should tell the truth: it previously contradicted run behavior for environment-level configuration and hid broken secret bindings. Enabling the package suites in CI prevents further silent drift — two suites were already broken on master without anyone noticing. **Breaking changes** None. Runs are unchanged. The probe route only adds env layers and checks; setups without environment envVars behave exactly as before. ## What Changed - `server/src/routes/agents.ts`: the test-environment route resolves the selected environment's envVars (forbidden keys stripped, environment-consumer secret context) and merges them under the agent adapterConfig env, mirroring `resolveExecutionRunAdapterConfig` precedence. Missing secret bindings are skipped, reported as an `environment_env_binding_missing` error check, and fail the test — matching the `ConfigurationIncompleteFailure` a real dispatch would raise. - `packages/adapters/claude-local/src/server/test.ts`: new `claude_oauth_token_configured` info hint between the API-key warning and the subscription fallback; hello-probe classification gains a `claude_hello_probe_usage_limited` warning for provider-quota results (previously a hard `claude_hello_probe_failed`). - `scripts/run-vitest-stable.mjs`: add `@paperclipai/adapter-claude-local` to `nonServerProjects` so CI runs the package suites. - `packages/adapters/claude-local/src/server/execute.remote.test.ts`: assert both runtime asset syncs (skills and mcp-config); the suite predated the mcp-config asset. - `packages/adapters/claude-local/src/server/test.probe.test.ts`: usage-limit fixture now expects the usage-limited warning; new fixture covers the genuine transient path (529 overloaded); new tests cover the token hint and API-key precedence. - `server/src/__tests__/agent-test-environment-routes.test.ts`: new tests for the env merge (agent wins on conflict, forbidden key filtered), missing-binding reporting, and the no-execution-target fallback path. - `docs/adapters/claude-local.md`, `docs/adapters/overview.md`: state the auth-input facts (API key or oauth token wins over stored logins; snapshot-owns-auth applies when neither is configured) and describe the environment-aware Test behavior. ## Verification - `npx vitest run --project @paperclipai/adapter-claude-local` — 131 tests pass (both drifted suites repaired; they fail on master today). - `npx vitest run server/src/__tests__/agent-test-environment-routes.test.ts` — 7 tests pass. - `node --test ./scripts/__tests__/run-vitest-stable-shard.test.mjs` — passes with the added project. - `pnpm typecheck` in `server/` and `packages/adapters/claude-local/` — clean. ## Risks - Low risk. The run path is untouched; the probe route change is additive and inert when the environment has no envVars. - The probe now performs environment-consumer secret resolution at test time; access is authorized per binding exactly as at run time, and the audit consumer is the environment (as before for adapter-config resolution). - Enabling the claude-local project in CI adds about 2 seconds of vitest wall time to the general workspaces group and could surface future regressions in that package — which is the point. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking enabled, agentic tool use via Claude Code (CLI). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ca6416da81 |
fix(server): preserve hot-restart shutdown snapshots (#10815)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A guarded hot restart must preserve or finalize every active agent run. > - The server uses embedded PostgreSQL when `DATABASE_URL` is not set. > - The database dependency installs signal handlers before Paperclip installs its coordinated shutdown handler. > - Those handlers can stop PostgreSQL before Paperclip writes the shutdown snapshot. > - ACP runs also use server-owned stdio and cannot be adopted after that server exits. > - This pull request keeps PostgreSQL available through snapshot and drain, then uses the existing ordered stop. > - The benefit is a complete restart report with no false adoption and no missing snapshot loss. ## Linked Issues or Issue Description **What happened?** A guarded hot restart with a valid marker can report a live preflight run as lost with reason `missing_shutdown_snapshot`. The `embedded-postgres` package imports `async-exit-hook`. That package registers `SIGINT` and `SIGTERM` listeners before Paperclip registers its own shutdown listener. The dependency can close PostgreSQL while Paperclip queries active heartbeat runs and writes the snapshot. **Expected behavior** Paperclip must keep its database available until it persists the shutdown snapshot and completes any required run drain. A detached CLI run must remain eligible for adoption. An ACP run must finish as interrupted and queue a retry because its server-owned stdio cannot survive the server. **Steps to reproduce** 1. Run Paperclip from source with embedded PostgreSQL. 2. Start a local ACP-backed agent run. 3. Write a valid hot-restart marker for the current server process. 4. send `SIGTERM` through the service manager. 5. Inspect the restart report and server log. 6. Observe that PostgreSQL can close before the shutdown snapshot query completes. **Paperclip version or commit** The defect reproduces on `2ab797dcbed0031c45c7335a0f497fea2a20bd9a`. **Deployment mode** Self-hosted server built from source, with embedded PostgreSQL and a systemd service. Related work: #9628 introduced hot-restart continuity. #10556 explores a broader database ownership transfer. #10775 addresses ACP continuity after replacement startup. This pull request uses a smaller path: it keeps the current database owner alive through snapshot and drain, then performs the existing explicit database stop. ## What Changed - Remove only the `SIGINT` and `SIGTERM` listeners added by the embedded PostgreSQL import. - Preserve Paperclip's existing ordered database stop after heartbeat snapshot and drain. - Detect active ACP and server-stdio local runs before shutdown. - Persist their complete snapshot before changing the marker to an ACP drain request. - Drain only ACP runs to an interrupted terminal state and queue their retry. - Keep detached CLI runs eligible for adoption in the same mixed restart. - Quiesce already-running scheduler queue claims before capturing the snapshot and selective drain set. - Report a selected ACP run as lost if process termination succeeds but its terminal database write does not persist. - Add the drain reason to the restart report. - Document the normal path and the one-time recovery path across an older affected build. ## Verification - `PAPERCLIP_TEST_DATABASE_MODE=native pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-process-recovery.test.ts` — 100 passed. - `pnpm exec vitest run server/src/shutdown.test.ts server/src/services/hot-restart.test.ts` — 24 passed. - The shutdown suite imports the real `embedded-postgres` package and verifies that its eager signal listeners are absent after the guarded import. - The embedded PostgreSQL recovery suite verifies snapshot, pre-snapshot scheduler quiescence, selective ACP drain, detached CLI adoption, queued retry, original-run finalization, `lostRunIds=[]` in a mixed restart, and fail-closed reporting when terminal persistence fails. - `pnpm --filter @paperclipai/server typecheck` — passed. - `git diff --check` — passed. - A full workspace typecheck reached the UI and stopped because the shared local install does not contain its declared `@base-ui/react` dependency. All server and preceding package checks passed. CI uses a clean install and remains the authoritative full gate. ## Risks - Low to moderate risk. This changes shutdown signal ownership and local run behavior during guarded restarts. - Paperclip already stops its managed embedded database explicitly. The change removes only the dependency listeners that race the coordinated path. - ACP runs now retry instead of receiving an unsafe bare-process adoption. Detached CLI runs keep their existing adoption behavior. - The report adds one field. There is no schema migration or breaking API change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The runtime did not expose a more specific model revision or context-window size. Reasoning, repository editing, shell execution, and test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dfcda67650 |
feat(auth): default-open visible issue writes (#10804)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents coordinate through company-visible issues, comments, child tasks, and assignments. > - The current authorization rules give these write channels different and narrow ownership grants. > - Those differences prevent standard-trust agents from coordinating on work that they can already read. > - A responsible human user must still bound every agent action. > - This pull request gives the four issue-write channels one default-open rule based on issue visibility. > - The benefit is consistent multi-agent coordination without weakening company, user, trust-scope, or run-lifecycle controls. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves authorization for comments, issue updates, child creation, and assignment on company-visible issues. **Subsystem affected** `server/` REST API authorization and issue routes. **Current behavior** Standard-trust agents can read company-visible issues, but narrow ownership, parent, or mention grants can still deny related writes. Each write channel also applies a different rule. **Proposed behavior** Allow standard-trust agents to comment, update fields, create child issues, and assign work when they can read the target issue and the responsible user is also authorized. Keep company boundaries, low-trust scopes, checkout conflicts, status rules, pause gates, budget gates, and explicit reopen rules unchanged. **Reason and benefit** Agents can coordinate on visible company work without relay issues or unnecessary manager runs. One shared rule also makes the authorization model easier to test and maintain. **Breaking changes** This intentionally broadens write access for standard-trust agents on visible issues. Existing company boundaries and governance controls remain in force. Related prior approaches: Refs #10233 and Refs #9768. This change unifies the comment case with visible issue updates, child creation, and assignment while preserving the responsible-user ceiling and excluded trust scopes. ## What Changed - Added a shared default-open authorization decision for visible issue writes. - Applied the shared rule to comments, issue updates, child creation, and assignment. - Preserved low-trust, `skill_test`, `task_bridge`, responsible-user, checkout, lifecycle, pause, and budget controls. - Preserved explicit resume/restore authority for direct peer lifecycle transitions on blocked, completed, and cancelled issues. - Added regression coverage for cross-company denial, user intersection, excluded scopes, comment-read structure, closed issues, child creation, assignment, and peer updates. - Updated the V1 implementation contract for the shared rule. ## Verification - `pnpm exec vitest run server/src/__tests__/authorization-service.test.ts server/src/__tests__/issue-agent-mutation-ownership-routes.test.ts server/src/__tests__/issue-comment-reopen-routes.test.ts server/src/__tests__/low-trust-red-team-routes.test.ts --reporter=dot` — 217 tests passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `git diff --check public-gh/master...HEAD` — passed. - Independent security review covered broken access control, object-level authorization, excessive agency, cross-company access, responsible-user intersection, excluded scopes, and lifecycle controls; its peer lifecycle-transition finding is fixed with regression coverage. ## Risks - Standard-trust agents gain broader write influence on issues that they can already read. - Future issue-visibility controls must keep `issue:read` as the canonical authorization hook. - Regression tests cover company boundaries, responsible-user intersection, excluded scopes, active checkout conflicts, closed-issue behavior, and non-transitive mention authority. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5 through Codex. The runtime does not expose the exact deployment revision or context-window size. Agentic reasoning, repository tools, shell execution, and test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cb52f0b750 |
db: env-configurable client options; parallelize attention feed queries (#10795)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The server stores all state in PostgreSQL through Drizzle and the
postgres.js driver
> - Self-hosted installs run Postgres on localhost, so per-query latency
is near zero; hosted installs often attach Postgres over a network,
sometimes through a transaction-mode pooler
> - The DB client passes no options to the driver, so operators cannot
disable prepared statements or tune the pool without a source edit, and
the deploy docs told them to edit `client.ts`
> - The attention feed also runs its related-data lookups one after
another, so its latency grows as queries × network round trip
> - This pull request adds optional environment configuration for the DB
client and batches the independent attention-feed lookups with
`Promise.all`
> - The benefit is that network-attached deployments get correct pooler
support and a much faster attention feed, while self-hosted behavior
does not change
## Linked Issues or Issue Description
No public issue exists for this; description follows the bug report
template:
**What happened?**
On deployments where PostgreSQL is network-attached (managed providers,
pooled endpoints), the attention feed endpoint is slow:
`attentionService.list()` awaits ~15–20 queries strictly in sequence, so
a 70ms round trip turns into more than one second of pure network wait
per call. Separately, connecting through a transaction-mode pooler
(pgbouncer, Supavisor port 6543, Neon `-pooler` hosts) requires
disabling prepared statements, and the only documented way was to
hand-edit `packages/db/src/client.ts` — which `doc/DATABASE.md` itself
tells operators not to do.
**Expected behavior**
The DB client is configurable from the environment (prepared statements,
pool size, timeouts) with driver defaults when unset, and hot read paths
do not multiply network latency by issuing independent queries
sequentially.
**Steps to reproduce**
1. Run the server with `DATABASE_URL` pointing at a Postgres instance
with ~70ms round-trip latency.
2. Open the attention feed (`GET /companies/:companyId/attention`) and
measure response time — it exceeds one second even with little data.
3. Try to connect through a transaction-mode pooler: there is no
supported configuration to disable prepared statements.
## What Changed
- `packages/db/src/client.ts`: `createDb` accepts a
`DatabaseClientOptions` argument and reads optional env config —
`DATABASE_PREPARED_STATEMENTS`, `DATABASE_POOL_MAX`,
`DATABASE_IDLE_TIMEOUT_SECONDS`, `DATABASE_CONNECT_TIMEOUT_SECONDS`.
When nothing is set, no option is passed to the driver and behavior is
identical to the previous bare `postgres(url)`.
- `packages/db/src/client-options.test.ts` (new): env parsing and
driver-option mapping tests, including malformed-value rejection.
- `server/src/services/attention.ts`: the independent related-data
lookups in each feed section now run under `Promise.all` (issue
summary/image/plan-document maps, decision bundle titles, blocked-issue
maps, the newer-runs scan). Section order, item assembly, and query
shapes are unchanged.
- `doc/DATABASE.md` and `docs/deploy/database.md`: the edit-source
pooling instruction is replaced with the env toggle, plus a short
client-tuning reference.
## Verification
- `pnpm --filter @paperclipai/db exec vitest run
src/client-options.test.ts` — 6 tests pass.
- `pnpm --filter server exec vitest run
src/__tests__/attention-service.test.ts` — 22 tests pass.
- `pnpm --filter server exec vitest run
src/__tests__/decisions-service.test.ts
src/__tests__/decision-training.test.ts` — 45 tests pass; this covers
the call path that runs `attentionService.list()` inside
`db.transaction`, where postgres.js serializes queries on the reserved
connection.
- `tsc` reports no errors in the changed files.
## Risks
- Low risk for self-hosted installs: with no env vars set,
`postgres(url, {})` receives an empty options object, which postgres.js
treats the same as no options — driver defaults throughout.
- The `Promise.all` batches only group queries that had no data
dependency on each other; on the transaction call path the driver still
executes them one at a time on the reserved connection, so transactional
semantics are unchanged.
- Malformed env values now fail fast at startup with a clear message
instead of being silently ignored; this is intentional and only affects
operators who set the new variables.
## Model Used
Claude Fable 5 (`claude-fable-5`), Anthropic — via Claude Code CLI,
extended thinking enabled, tool use (test execution, live latency
measurement against a network-attached Postgres to size the problem).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (searched "prepared statements", "pgbouncer", "pool",
"attention feed", "lockfile" — closest matches are #10573/#10787
lockfile chores, unrelated to this change)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
2ab797dcbe |
fix(sandbox): reset step store for long-lived bridge work (#10813)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Bridge workers keep startup setup separate from long-lived runtime
work
> - The callback bridge must keep queue-directory setup inside the
startup step
> - The long-lived poll loop must run with no active startup step
> - This pull request keeps that boundary in the right place
> - The benefit is correct span parents and correct runtime exec
metadata
## Linked Issues or Issue Description
**Bug**
**What happened?**
Long-lived bridge continuations kept a stale startup step store during
the queue-directory setup path.
**Expected behavior**
Runtime exec spans should start with no active startup step.
**Steps to reproduce**
1. Start a bridge lane.
2. Let the startup step end.
3. Run later runtime exec work on the same lane.
**Paperclip version or commit**
|
||
|
|
8e7f1c03eb |
feat(decisions): improve desk triage and queue parity (#10785)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The decisions desk and queue help operators find work that needs a human decision. > - The current views use different grouping, sorting, and labels. > - Repeated confirmation requests can also leave stale pending actions in the queue. > - Blocked-work attention can point at an intermediate issue instead of the terminal blocker. > - This pull request aligns the server contract and both user interfaces. > - The benefit is a smaller, clearer queue that ranks the decisions with the largest impact. ## Linked Issues or Issue Description Related PR: #10774 **What existing behavior does this improve?** The decisions desk and queue currently use different triage rules. They can show stale repeated confirmations and can rank blocked work by an intermediate issue. **Subsystem affected** This change affects attention aggregation, issue thread interactions, shared attention contracts, and the decisions user interface. **Current behavior** The desk uses a can-wait group that has no clear arrival meaning. The queue has fewer controls than the desk. Repeated pending confirmations remain actionable. Blocked-work rows do not always identify the terminal actionable blocker. **Proposed behavior** Group desk items by arrival date, and reserve Decide now for explicit due dates. Use one toolbar and shelf model on both pages. Supersede older repeated pending confirmations. Aggregate blocked work under the terminal actionable blocker and rank it by impact. **Reason and benefit** Operators get one consistent triage model. The badge reflects new and overdue work. High-impact blockers move to the top. Duplicate confirmation work no longer consumes attention. **Breaking changes** The attention summary field `decideNowCount` changes to `deskBadgeCount`. Consumers must use the new field. Older repeated confirmation interactions can now finish with the `superseded_by_newer_request` outcome. ## What Changed - Supersede older pending confirmation requests for the same issue and record the mutation in activity history. - Resolve blocked-work attention to actionable terminal blockers, suppress live blocker trees, and rank rows by blocked-work impact. - Group the decisions desk into New today and Earlier, and count new plus overdue work in the desk badge. - Share the decision toolbar and shelf components across the desk and queue. - Add queue grouping, sorting, filtering, aging, visible training controls, and clearer recommendation copy. - Add server, shared-contract, and user-interface tests for the new behavior. ## Verification - `pnpm check:token-gates` - `pnpm -r typecheck` - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm test:run` (all server, UI, CLI, shared, and catalog tests passed; one fixed five-second DB timeout flaked under full-suite load) - `NODE_ENV=test pnpm --filter @paperclipai/db exec vitest run src/status-card-migrations.test.ts` (passed in isolation) - `pnpm build` ## Risks - The attention summary field rename requires synchronized consumers. - Terminal-blocker traversal uses cycle and depth guards. A malformed dependency graph can stop at the last safe node. - The new arrival grouping changes which items contribute to the decisions badge. - Superseding repeated confirmations changes the terminal state of older pending interactions. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5. The runtime did not expose a dated model snapshot or context-window size. The model used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6a69a3d6b0 |
refactor(observability): stop writing detailed per-step timing to the run log (#10776)
## Thinking Path > - Paperclip tracks agent work for the team. > - The startup timing path already emits detailed spans. > - The run log repeats the same detailed timing. > - That makes the event larger than it needs to be. > - This pull request removes the redundant per-step timing fields from the run-log payload. > - The benefit is a smaller log with the same detail still available in traces. ## Linked Issues or Issue Description No public GitHub issue or open pull request matches this change. I described the enhancement below. **What existing behavior does this improve?** `run.startup.step` event output. **Subsystem affected** Cross-cutting (multiple of the above) **Current behavior** `run.startup.step` writes `roundTrips`, `providerExecMs`, `providerGetMs`, `createRuntimeMs`, and `ensureSessionMs`. The same detail already exists in the spans. **Proposed behavior** `run.startup.step` keeps only `step`, `durationMs`, and `outcome`. The heartbeat lifecycle timestamps stay unchanged. **Reason and benefit** This change removes redundant data from the run log. It keeps the useful detail in trace spans. It also makes the payload smaller and easier to read. The revert path stays clear because the removed data has one producer chain. **Breaking changes** Yes. Consumers that read the removed fields must switch to span data or the remaining event fields. The heartbeat lifecycle timestamps do not change. **Additional context** I searched GitHub for related open PRs and issues. I found no match. ## What Changed - Removed the redundant per-step timing fields from the run-log payload. - Removed the now-dead producer chain that fed those fields. - Kept the span-level timing data and the heartbeat lifecycle timestamps. ## Verification - `adapter-utils` typecheck clean - `server` typecheck clean - `startup-timing` suite: 30 passed - `adapter-utils` `execute` suite: 99 passed - `server` `environment-execution-target` suite: 21 passed ## Risks - Consumers that still read the removed fields will need a code change. - This is a one-way-door data removal for the run log. ## Model Used OpenAI Codex, GPT-5, tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
af6b32d82a |
feat(observability): add provider OTel spans - cache-hit flag, plugin tracer, Daytona pack/transfer spans (#10764)
## Thinking Path > - Paperclip uses pull requests to ship control-plane changes with review and audit. > - This change extends sandbox provider telemetry. > - The first PR added startup and exec spans. > - This PR adds provider spans for cache hits, plugin tracer flow, and Daytona file sync steps. > - The host must keep the trust boundary and reject unsafe span data. > - The result is fuller OTel coverage with bounded labels and safe parentage. ## Linked Issues or Issue Description **Problem or motivation** Sandbox provider telemetry lacks an explicit cache-hit signal, plugin span context, and file-sync pack and transfer spans. The current proxy for a cache hit is fragile. The plugin SDK also needs a safe tracer surface and a per-call parent context. **Proposed solution** Add an explicit `cacheHit` flag from the Daytona sandbox handle lookup. Expose `ctx.tracer` on the plugin context and pass a `traceparent` string with each call. Record provider spans on the host with allowlist clamping and capability checks. Wrap Daytona file sync pack and transfer steps in spans. **Alternatives considered** Keep the old `providerGetMs == 0` proxy. Reject that path because the cache-hit decision belongs at the handle lookup, not in a timing proxy. The host stays the trust boundary for worker span data. It clamps labels, rejects bad parent data, and gates the worker-to-host span RPC by capability. A security review completed before this PR opened. ## What Changed - Added an explicit `cacheHit` flag from the Daytona sandbox handle lookup. - Added `ctx.tracer` on the plugin context and a `traceparent` field in the per-call context. - Added host-side span recording with attribute clamping and capability checks. - Wrapped Daytona file sync pack and transfer steps in spans. - Added tests for the host trust boundary, the plugin tracer no-op path, and the Daytona span paths. ## Verification - `pnpm --filter @paperclipai/plugin-daytona exec vitest run` - `pnpm --filter @paperclipai/plugin-sdk exec vitest run` - `pnpm --filter @paperclipai/server exec vitest run plugin-host-services environment-execution-target instrumentation plugin-worker-manager` - `tsc --noEmit` for server, plugin-sdk, and adapter-utils ## Risks - Span data now crosses the worker boundary, so the host allowlist and capability gate must stay strict. - The `traceparent` path must stay valid and host-owned. - Daytona file sync spans must not add new round trips. ## Model Used OpenAI Codex, GPT-5, tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues with `Fixes:` / `Closes:` / `Refs:` or described the issue in the PR body - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ba396c608c |
feat(server): configure shared workspace concurrency (#10759)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat service starts agent runs in local, SSH, sandbox, plugin, or Kubernetes environments. > - Runs can share one project working tree. > - PR #10699 made every shared working tree single-file, including trusted local and SSH hosts that support coordinated concurrent work. > - Operators need a policy that keeps remote environments safe and restores local multi-agent work. > - This pull request adds an `auto`, `serialize`, or `allow` concurrency policy and applies it to the final execution environment. > - The benefit is safe serialization by default for sandboxed targets and useful concurrency by default for persistent local targets. ## Linked Issues or Issue Description Related work: #10699 introduced the busy gate that this change makes configurable. #7852 covers a separate environment-lease race and does not provide this dispatch policy. **Subsystem affected** `server/` heartbeat orchestration and `packages/shared/` workspace policy contracts. **Problem or motivation** The shared-workspace busy gate always defers a second run. This behavior prevents local multi-agent projects from running concurrently even when operators expect agents to coordinate through commits. **Proposed solution** Add `sharedWorkspaceConcurrency` with `auto`, `serialize`, and `allow` values. Default `auto` permits local and SSH concurrency. It serializes sandbox, plugin, and forced Kubernetes execution. **Alternatives considered** Keeping unconditional serialization is too restrictive for persistent host working trees. Always allowing overlap removes the protection that sandboxed and remote targets need. **Roadmap alignment** This is a focused correction to the shipped cloud and sandbox execution milestone. It does not add a new roadmap feature. ## What Changed - Added the optional tri-state field to project policy and issue override contracts and validators. - Added a pure resolver with issue override, project policy, and `auto` default precedence. - Moved final environment and Kubernetes resolution before the shared-workspace busy gate. - Kept the existing deferral and retry behavior for every path that resolves to serialization. - Added a task-context warning and a structured log when a run dispatches beside a live holder. - Added policy and heartbeat coverage for all requested policy and environment combinations. - Documented the new policy and its default behavior. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/execution-workspace-policy.test.ts src/__tests__/heartbeat-workspace-busy.test.ts` — 32 tests passed. - `pnpm -r typecheck` — passed. - `PAPERCLIP_IN_WORKTREE=false PAPERCLIP_DATABASE_RESTORE_IN_PROGRESS=false PAPERCLIP_RESTORE_IN_PROGRESS=false pnpm test:run` — all 3,528 server tests passed. The later UI/workspace phase passed 3,384 tests and hit one unrelated `CompanyEnvironments` navigation timing failure. - `pnpm --filter @paperclipai/ui exec vitest run src/pages/CompanyEnvironments.test.tsx -t "opens the edit form on a standalone page with existing values and closes after save"` — the unrelated UI test passed in isolation. - `pnpm build` — passed. Policy matrix: | Policy | Final target | Expected result | Result | | --- | --- | --- | --- | | `auto` | local | Dispatch with holder note | Passed | | `auto` | sandbox | Defer with `workspace_busy` | Passed | | `auto` | instance-forced Kubernetes | Defer with `workspace_busy` | Passed | | `serialize` | local | Defer with `workspace_busy` | Passed | | `allow` | sandbox | Dispatch with holder note | Passed | Existing serialization, retry, and stale-holder tests also pass. ## Risks - `auto` changes the post-#10699 local and SSH behavior back to concurrent dispatch. Concurrent agents can mutate the same working tree, so each dispatched run receives an explicit coordination warning. - `allow` is an operator override and can permit overlap in sandbox or plugin environments. - Unknown environment drivers serialize in `auto` mode. This keeps the fallback conservative. - There is no database migration. An absent field resolves to `auto`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 (Codex). The exact API snapshot and context-window size are not exposed to the agent. Reasoning, repository editing, terminal tool use, and local code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c09ea7112b |
feat(observability): add granular OpenTelemetry spans for sandbox startup and execution (#10758)
## Thinking Path > - Paperclip coordinates AI agent work and records execution data. > - The sandbox startup path and the host-to-sandbox execution path need clearer OpenTelemetry spans. > - The trace data now shows wait time, critical path data, and bounded labels. > - This pull request adds bounded spans and closed attribute helpers for sandbox startup and execution. > - The result is better trace data with low cardinality and no secret leakage. ## Linked Issues or Issue Description - No public GitHub issue exists for this branch. - Related PR: #10536 - This PR extends the sandbox OpenTelemetry work with exec spans, root timing, and bounded labels. ## What Changed - Added a closed span-attribute contract for sandbox startup data. - Added a bounded label helper for command and region values. - Threaded the active step context into the execution path. - Added a `sandbox.exec` span with true timestamps and bounded attributes. - Added skipped-step and root-span timing data with low-cardinality context. - Moved handshake sub-times and bridge batch data onto spans. ## Verification - `pnpm --filter @paperclipai/adapter-utils run typecheck` - `pnpm --filter @paperclipai/server run typecheck` - `pnpm --filter @paperclipai/adapter-utils exec vitest run` - `pnpm --filter @paperclipai/server exec vitest run environment-execution-target` - `git log --oneline origin/master..origin/feat/sandbox-startup-otel-spans` - `git diff origin/master...origin/feat/sandbox-startup-otel-spans --stat` ## Risks - Span attribute rules could still miss a future field if new code skips the shared helper. - Low-cardinality labels may hide some detail, but that is the intended tradeoff. - Tracing stays fail-open, so a missing tracer still hides data rather than stopping work. ## Model Used - OpenAI Codex, GPT-5, tool use enabled. The runtime did not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge The current head still has one failing required check: `e2e`. The PR stays unready for board handoff until that check passes. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bc5c392331 |
fix(server): stop inferring PR credential preflight from issue text (#10755)
## Thinking Path > - Paperclip coordinates agent runs and checks required runtime credentials before dispatch. > - The push-capability preflight protects runs that use the GitHub PR workflow skill. > - PR #10659 also made issue title and description text trigger this preflight. > - That text heuristic blocks tasks that mention a pull request but do not need a bound GitHub token at dispatch time. > - This pull request removes the text heuristic and keeps the explicit skill trigger. > - The benefit is that ordinary task wording no longer causes a false configuration failure. ## Linked Issues or Issue Description Refs: #10659 **What happened?** An issue title or description that said to open a pull request or push a branch could trigger the push-capability credential preflight. The run then failed before dispatch when no project-level or agent-level GitHub token was bound, even when the task could proceed without that preflight. **Expected behavior** The preflight must run only when the issue explicitly selects the GitHub PR workflow skill. Issue prose alone must not enable the guard. **Steps to reproduce** 1. Assign a local Codex or Claude agent an issue that says to open a pull request. 2. Do not attach the GitHub PR workflow skill to the issue. 3. Start the run without a project-level or agent-level GitHub token binding. 4. Observe the false `push_write_credential_missing` failure before this fix. **Paperclip version or commit** `master` after PR #10659. **Deployment mode** Local dev with a git-sensitive local adapter. ## What Changed - Removed `issueTextImpliesPrDeliverable` and its text-pattern matcher. - Restored `requiresPushCapabilityPreflight` to use only explicit GitHub PR workflow skill keys. - Removed the obsolete issue-text tests while keeping coverage for the explicit skill, adapter, and issue gates. ## Verification - `PAPERCLIP_LOG_DIR=<run-owned-dir> pnpm vitest run server/src/__tests__/heartbeat-workspace-session.test.ts` — 134 tests passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `git diff --check origin/master...HEAD` — passed. ## Risks - Low risk. A task that states a PR deliverable but does not select the GitHub PR workflow skill will no longer receive the early credential preflight. This is the intended temporary behavior. - Tasks that explicitly select the skill still receive the existing credential and checkout checks. > This change does not add a core feature and does not overlap with ROADMAP.md work. ## Model Used OpenAI Codex with model ID `gpt-5`. The runtime did not expose the context-window size. The model used agentic reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2c90cf0f2c |
fix(server): serialize managed-checkout materialization and stop misattributing clone failures (#10723)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Repo-only project workspaces are materialized by a server-side
managed `git clone` into a per-project directory (#10720 added
credentials for private repos)
> - Two issues on the same project routinely wake seconds apart, and
both runs race the same clone target
> - The loser fails with "destination path already exists", and its
failure cleanup removes the directory out from under the winner's
in-progress clone — both runs then fail every round
> - This pull request serializes materialization per target directory
and makes the clone land atomically via a temp sibling + rename, so the
shared target is never partial and never removed
> - The benefit is that concurrent runs on the same project converge:
one clone happens, everyone adopts it, and unrelated failures no longer
blame the GitHub credential
## Linked Issues or Issue Description
**What happened?**
With isolated workspaces enabled on a project whose only workspace is
repo-only, unblocking two issues at once produced lockstep mutual
destruction (observed live, two consecutive rounds): both runs called
the managed-checkout materialization concurrently; one clone created the
target directory, the other's `git clone` failed with `fatal:
destination path '…' already exists and is not an empty directory`, and
that run's failure cleanup deleted the directory while the first clone
was still writing into it (`fatal: could not set 'core…'`). Both runs
failed `workspace_validation_failed`; their staggered retries could race
again. The failure message also wrongly claimed the GitHub credential
"was rejected or lacks access" — the collision had nothing to do with
auth.
**Expected behavior**
Concurrent materializations of the same project checkout share one
clone; a completed checkout is never removed by a failing sibling; the
credential is only blamed for auth-shaped failures.
**Steps to reproduce**
1. Project with a repo-only workspace (private repo, isolated workspaces
on).
2. Move two issues on that project to `todo` at the same time so both
runs start within seconds.
3. Both runs fail workspace validation with "destination path already
exists" / "could not set 'core…'" instead of one clone succeeding.
**Paperclip version or commit**
`master` (
|
||
|
|
75f6256b76 |
fix(server): resolve duplicate scrubGitCredentialText declaration on master (#10722)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - #10719 and #10720 both hardened credential scrubbing for server-side git operations > - #10719 defined `scrubGitCredentialText` locally in `heartbeat.ts`; #10720 imported the identical function from the new `git-credentials.ts` > - The two merged cleanly textually, but together they leave `heartbeat.ts` with an import that conflicts with a local declaration (TS2440) > - This pull request keeps the `git-credentials.ts` copy as canonical, drops the heartbeat-local duplicate, and re-exports the import so existing importers are unchanged > - The benefit is that master typechecks, builds, and produces Docker images again ## Linked Issues or Issue Description **What happened?** After #10719 and #10720 merged, master fails typecheck, the server build, and both Docker image builds with `src/services/heartbeat.ts(76,3): error TS2440: Import declaration conflicts with local declaration of 'scrubGitCredentialText'`. The canary release run for |
||
|
|
e0c2448267 |
feat(server): authenticate server-side git clone and fetch with a company-secret GitHub token (#10720)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Repo-only project workspaces are materialized by a server-side `git clone`, and isolated `git_worktree` runs refresh their base ref with server-side `git fetch` > - Both operations run outside the agent process with no credentials, so private GitHub repositories can never be cloned or refreshed — agent-scoped credential env bindings do not reach them > - The company secret store already has a well-known GitHub token convention (`GITHUB_TOKEN` / `GH_TOKEN` / `PAPERCLIP_GITHUB_TOKEN`, consumed by the external-object provider for API reads), but nothing server-side consults it for git > - This pull request resolves that token per run and authenticates the managed clone and every base-ref refresh with it through an ephemeral credential helper > - The benefit is that isolated workspaces work on private repositories with one company secret, while public repositories and self-hosted ambient git configuration keep working unchanged ## Linked Issues or Issue Description **Subsystem affected** Server workspace materialization (`server/src/services/heartbeat.ts`) and execution-workspace realization (`server/src/services/workspace-runtime.ts`). **Problem or motivation** A project workspace configured with only a private GitHub `repoUrl` cannot be used for isolated `git_worktree` runs: the managed `git clone` runs with a sanitized, credential-less environment, and plain git cannot consume a bare token env variable without a credential helper. There is no way to give the server a git credential — storing a `GH_TOKEN` company secret has no effect on server-side git, and a credential-less private clone hangs on a terminal prompt until the ten-minute clone timeout. Base-ref refreshes (`git fetch`) during worktree realization have the same gap. **Proposed solution** A `git-credentials` module resolves a token per run — company secret by well-known name (`GITHUB_TOKEN`, `GH_TOKEN`, `PAPERCLIP_GITHUB_TOKEN`), then `GITHUB_TOKEN`/`GH_TOKEN` in the server process environment for self-hosted deployments, then none — and builds a git invocation that authenticates via an inline credential helper. The token travels in an env variable; it never appears in argv, URLs, or on disk. Only `https://github.com` remotes are authenticated; everything else keeps ambient behavior. The provider is a single factory seam so a future brokered credential source can replace it without touching call sites. **Alternatives considered** - A GitHub OAuth "connect your account" flow: heavier product surface, needs app registration and callback custody; out of scope for a server credential and better served by a dedicated connector later. The provider seam keeps that path open. - `gh auth setup-git`: writes helper configuration to disk and requires a global token env; rejected in favor of per-invocation config with no persistent state. - Embedding the token in the clone URL: leaks into argv, error messages, and `.git/config`; rejected. ## What Changed - New `server/src/services/git-credentials.ts`: `createGitRemoteAuthProvider` (memoized per run, one secret resolution and one audit event), `buildGitAuthInvocation` (helper-reset + inline helper, `x-access-token` username, `GIT_TERMINAL_PROMPT=0`), `isGitHubHttpsRemoteUrl` host gating (rejects ssh/GHES/http/other hosts/userinfo URLs), `describeGitAuthFailure`, and the canonical `scrubGitCredentialText`. Secret resolutions pass a `system` consumer access context so they are recorded as secret access events. - `ensureManagedProjectWorkspace` (now exported) accepts an optional auth provider; the clone env spreads the token after `sanitizeRuntimeServiceBaseEnv` (which strips `PAPERCLIP_*`), always sets `GIT_TERMINAL_PROMPT=0`, distinguishes "credential rejected" from "no credential configured — add a GITHUB_TOKEN or GH_TOKEN company secret" in the error, and removes the partially created directory on clone failure so a timeout-killed clone cannot be adopted as a broken checkout by the next run. - `refreshRemoteTrackingBaseRef` (now exported) captures the remote URL it already looked up, asks the provider for an invocation, and attributes failed authenticated fetches to the credential in a scrubbed warning. The optional provider threads through `detectDefaultBranch`, `resolveAuthoritativeBaseRef`, `inspectExecutionWorkspaceBaseDrift`, `realizeExecutionWorkspace`, and `ensurePersistedExecutionWorkspaceAvailable`; heartbeat builds one provider per run for both the anchor-resolution clone path and workspace realization/restore. - `github-external-object-provider.ts` imports the shared secret-name list; `isGitHubDotCom` is exported from `github-fetch.ts`. - Docs: "Private repositories and repo-only project workspaces" section in the execution-workspaces guide, cross-linked from the secrets deploy doc. ## Verification - `cd server && npx vitest run src/__tests__/git-credentials.test.ts` — resolution chain order and precedence, env fallback, memoization, audited access context, host-gating matrix, invocation shape (token absent from argv), scrubber, failure descriptions, and a real-git `git credential fill` round trip that proves the helper executes and answers with the env-carried token (no network). - `cd server && npx vitest run src/__tests__/heartbeat-managed-clone-credentials.test.ts` — clones behave byte-identically with no provider or a null-returning provider (local repos, no network), authenticated-failure errors name the credential, non-auth failures do not mention credentials, partial clone directories are removed, pre-existing non-git directories keep the "Using it as-is" path, and the sanitizer spread order keeps the token env alive. - `cd server && npx vitest run src/__tests__/workspace-runtime.test.ts` — new `refreshRemoteTrackingBaseRef` cases: provider offered the remote URL and null keeps behavior identical; failed authenticated fetch warning names the credential; unauthenticated failure warning stays credential-free. - `pnpm --filter @paperclipai/server typecheck` is clean. - Manual (optional, networked): store a `GH_TOKEN` company secret, configure a repo-only project workspace pointing at a private GitHub repository, run an isolated-workspace issue — the managed clone succeeds and the worktree run proceeds. ## Risks - Every new parameter is optional; with no provider the git invocations are byte-identical to before. Public repos and ambient credential helpers keep working whenever no token resolves. - Precedence change when a token exists: a stored company secret now wins over ambient helpers for `https://github.com` remotes (the helper list is reset for that invocation). The rejected-credential error names the secret so an operator can fix or remove it. - `GIT_TERMINAL_PROMPT=0` on the managed clone is the one always-on change: a credential-less private clone now fails fast with a clear message instead of hanging until the ten-minute timeout (it could only ever "succeed" interactively on a TTY dev server). - The token is scoped to the git process env for one invocation; it is never written to agent env, run context, disk, or logs, and error text is scrubbed of URL userinfo. - No migrations, no image changes (git ships in the image). ## Model Used Claude Fable 5 (`claude-fable-5`, extended thinking, agentic tool use via Claude Code CLI). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
185515c97b |
fix(external-objects): refresh PR status labels (#10704)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The issue properties panel can show external objects such as GitHub pull requests. > - Those objects are resolved by external-object providers and then displayed as compact status labels. > - A GitHub pull request could remain in the fallback `unknown` state and appear as `Not yet resolved`. > - That label is confusing when the object is known but has not been refreshed yet. > - This pull request refreshes due external objects from the heartbeat scheduler and improves the unknown-status copy. > - The benefit is a properties panel that moves from pending refresh to the real pull request state without a manual refresh. ## Linked Issues or Issue Description No public GitHub issue exists for this bug. I searched for related public issues and pull requests using the terms `Not yet refreshed`, `external objects refresh`, and `external PR status`, and did not find a duplicate implementation. **What happened?** The issue properties panel could show a GitHub pull request as `Not yet resolved` even when the referenced pull request was valid. The object stayed stale unless a manual refresh path ran. **Expected behavior** A known external object should show pending-refresh copy while it waits for provider data. When the scheduler refreshes it, the properties panel should show the provider status such as open, merged, or closed. **Steps to reproduce** 1. Create or view an issue that references a GitHub pull request. 2. Open the issue properties panel. 3. Observe the external object row before a manual refresh has run. **Paperclip version or commit** Current `master` before this pull request. **Deployment mode** Local dev and self-hosted server. **Installation method** Built from source. **Agent adapter(s) involved** Not adapter-specific. **Database mode** Not database-related. **Access context** Board view. **Privacy checklist** I reviewed this description and did not include logs, credentials, private URLs, internal issue IDs, or PII. ## What Changed - Added a heartbeat scheduler tick that refreshes due external objects for active companies. - Kept manual external-object refresh behavior on the same service path. - Changed display copy so known provider objects use liveness labels such as `Not yet refreshed`, while fresh unknown provider statuses show `Status unavailable`. - Added server and UI tests for scheduled refresh and label behavior. ## Verification - `corepack pnpm install --frozen-lockfile` - `pnpm check:token-gates` - `pnpm exec vitest run server/src/__tests__/external-objects-service.test.ts server/src/__tests__/server-startup-feedback-export.test.ts ui/src/components/ExternalObjectPill.test.tsx ui/src/components/IssueProperties.test.tsx ui/src/lib/external-objects.test.ts` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/server build` - `pnpm --filter @paperclipai/ui build` - `pnpm run typecheck:build-gaps` - GitHub PR checks passed on head `0e7fcd30` - Greptile reported 5/5 on head `0e7fcd30` with no unresolved review threads Notes: - I ran recursive typecheck and build first. Both hit container resource limits with exit 137 during concurrent package work, so I reran the affected server and UI targets separately. - An unrelated workspace-runtime auto-port test fails in this container with a PID ownership mismatch. It is outside the files changed here. ## Risks Low to medium risk. The scheduler does more periodic external-object work, so the main risk is extra provider refresh load. The implementation bounds the work to active companies, due non-terminal objects, and 50 objects per company per tick. The path also stays behind the external-objects experimental setting. ## Model Used OpenAI GPT-5 Codex in the Codex execution environment, with shell and GitHub CLI tool use. The runtime did not expose a more specific internal model ID or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7a3815eb9a |
fix(heartbeat): surface the real cause when a git_worktree base cannot be materialized (#10719)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Heartbeat runs prepare an execution workspace for each issue; with
isolated workspaces (or low-trust runs), the `git_worktree` strategy
needs a real git checkout as its base
> - For repo-only project workspaces, the server materializes the
checkout with a managed `git clone`; when that clone fails, the resolver
drops the error and silently falls back to the agent home directory with
`source: "project_primary"`
> - The pre-dispatch guard then reports
`git_worktree_base_not_git_checkout`, which hides the real cause; if the
fallback directory happens to be a git checkout, the run silently builds
worktrees off the wrong repository
> - This pull request records materialization failures on the resolved
workspace, marks the fallback explicitly, and fails the guard with a
truthful `git_worktree_base_materialization_failed` reason that carries
the clone error
> - The benefit is that operators see the real cause (a failed clone) in
the run error, the blocked-issue comment, and the recovery next action,
instead of a misleading symptom
## Linked Issues or Issue Description
**What happened?**
An issue configured for isolated `git_worktree` execution on a project
whose only workspace is repo-only (a `repoUrl` with no local path) fails
every run with `workspace_validation_failed` and reason
`git_worktree_base_not_git_checkout`, pointing at the agent home
directory. The message does not mention that the managed `git clone` of
the project repository failed (for a private repository the clone can
never succeed without credentials). The recovery flow then blocks the
issue with the same misleading explanation. Run warnings claim "Project
workspace has no local cwd configured" even though the workspace is
configured and the clone failed.
**Expected behavior**
The run failure, the blocked-issue comment, and the recovery next action
should state the real cause: the project workspace checkout could not be
prepared, including the clone error, so the operator can repair the
repository URL, clone access, or configured local cwd. A fallback
directory that happens to be a git checkout must not let the run proceed
against the wrong repository.
**Steps to reproduce**
1. Create a project whose primary workspace has a `repoUrl` pointing at
a private GitHub repository and no local path.
2. Enable the Isolated Workspaces experimental setting (or use a
low-trust run, which forces isolation).
3. Run any issue in that project.
4. The run fails with `git_worktree_base_not_git_checkout` on the agent
home directory; the clone failure appears nowhere.
**Paperclip version or commit**
`master` (
|
||
|
|
8b83d69e3c |
feat(heartbeat): serialize shared-workspace issue runs with bounded busy deferrals (#10699)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat service dispatches agent runs, and issues in a project can share one project workspace (one working tree on disk). > - Two runs can execute in the same shared workspace at the same time. Each run mutates the same uncommitted files and branches, and the runs corrupt each other's state. > - Multi-agent projects hit this as soon as two issues in one project become active together, so the platform needs to serialize shared-workspace execution instead of relying on luck. > - This pull request adds a pre-dispatch gate: a run whose issue targets a busy shared workspace is deferred with a bounded scheduled retry instead of dispatched. > - The benefit is that concurrent issue runs in one project take turns in the shared working tree, while isolated-workspace runs and unrelated workspaces stay fully parallel. ## Linked Issues or Issue Description Fixes #10645 ## What Changed - `server/src/services/heartbeat.ts`: - New pre-dispatch gate in the run executor. Before adapter dispatch, when the run's issue has a `projectWorkspaceId` and the effective execution workspace mode is `shared_workspace`, the executor looks for a holder: another `running` run whose context issue shares the same project workspace. The gate covers every run shape that reaches adapter dispatch with issue context — assignee execution runs, comment/mention interaction wakes, and review-participant runs. - When a holder exists, the run throws `WorkspaceBusyDeferral` instead of dispatching. The outer catch recognizes the deferral: it cancels the run with `errorCode: "workspace_busy"` (contention is not a failure), cancels its wakeup, schedules a retry through the existing `scheduleBoundedRetryForRun` primitive (`workspace_busy` reason, 60–120 s jittered delay), and returns the agent to idle. The issue execution lock transfers to the scheduled retry run, so the issue keeps an active execution path and stranded-issue recovery does not fire. - An adapter never dispatches alongside a live holder: deferral has no attempt ceiling, so a deferred run keeps rescheduling until the workspace frees. Deadlock safety comes from holder liveness, not a counter — a holder silent past `ACTIVE_RUN_OUTPUT_SUSPICION_THRESHOLD_MS` (recovery's own "suspicious silence" bar, 1 h) stops counting as a holder, so a zombie run can only delay work, never park it forever, and recovery's silent-run escalation is already reaping it in parallel. If no retry can be scheduled (agent paused, issue reassigned), the deferral releases the issue execution lock so the issue does not strand. - Holder detection honors isolation: when the isolated-workspaces experiment is enabled, holders whose issue settings select `isolated_workspace` / `operator_branch` (or the legacy `isolated` alias) are not counted, because they never touch the shared tree. A NULL or `agent_default` mode counts as a holder — over-serializing is the safe direction. - Non-assignee deferrals survive replay: the deferral stamps `workspaceBusyDeferredWhileAssignee` into the run context (inherited by the scheduled retry), and both the retry promotion gate and the claim-time staleness check exempt a non-assignee `workspace_busy` retry from the reassignment cancellation — for such a retry an assignee mismatch is the expected state, not a reassignment race. An assignee run's retry keeps the full protection: if the issue is reassigned while the retry pends, it still cancels with `issue_reassigned`. - `server/src/__tests__/heartbeat-workspace-busy.test.ts` (new): embedded-Postgres coverage of the full lifecycle plus unit coverage of the delay window. ## Verification - `cd server && pnpm vitest run src/__tests__/heartbeat-workspace-busy.test.ts` — 10 tests: - a run whose issue targets a busy shared workspace is cancelled with `workspace_busy`, its adapter never executes, a `scheduled_retry` run exists with the 60–120 s window, the issue execution lock points at the retry run, the holder run is untouched, and the agent returns to idle; - after the holder finishes, `promoteDueScheduledRetries` + `resumeQueuedRuns` execute the retry run to success; - a non-assignee comment-mention wake defers, does not touch the issue execution lock, and its retry promotes, survives the claim-time staleness check, and executes despite the assignee mismatch; - an assignee retry is still cancelled with `issue_reassigned` when the issue is reassigned while the retry pends; - a holder issue with `executionWorkspaceSettings.mode = "isolated_workspace"` does not cause deferral; - a running run in a different project workspace does not cause deferral; - a holder silent past the staleness threshold does not cause deferral (the run executes); - a retry with ten prior deferrals still defers again — never dispatches — while the holder is live; - delay jitter stays inside the base-to-base-plus-jitter window and clamps out-of-range random sources. - `cd server && pnpm vitest run src/__tests__/heartbeat-` — full heartbeat suite sweep. - `cd server && pnpm run typecheck`. ## Risks - Behavioral shift: shared-workspace runs that used to start immediately now wait for the workspace to free. Against a long-running live holder the wait is unbounded by design — the alternative is dispatching into a held working tree, which is the corruption this PR removes. Every deferral is visible in the run timeline (lifecycle event with the holder run, issue, and attempt number), and the wake is parked, never dropped. - A zombie holder (a `running` row whose process died) delays contending runs by up to the 1 h staleness threshold before it stops counting. Recovery's silent-run escalation targets the same run on the same clock, so this window matches what the system already tolerates for silent active runs. - The holder check and the dispatch are not atomic; two runs that pass the gate in the same instant can still race. The gate closes the common window (a second run waking while the first is mid-execution); the pre-existing sync-conflict handling remains the backstop for the rare simultaneous start. - No schema change, no API change, no new configuration. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, tool use, full repository access; implementation, tests, and verification runs. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
6008482aab |
fix(decisions): remove clock race from sweep-expiry tests (#10701)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The decisions service lets agents propose decisions with a TTL, and a sweep expires them. > - Four sweep tests create decisions that expire 5 ms in the future, then race the service's own clock read. > - On a loaded CI runner more than 5 ms routinely elapse before validation, so `create` itself rejects the decision and the test fails on unrelated PRs. > - This pull request makes the expiry deterministic: create with a comfortable future TTL, then move `expiresAt` into the past directly in the store. > - The benefit is that the decisions suite stops failing intermittently and stops blocking unrelated PRs. ## Linked Issues or Issue Description No open issue exists; the defect is described here following the bug template. **What happened?** `decisions-service.test.ts` fails intermittently in CI with `expiresAt must be within 30 days` in `bounds expiration work to the configured batch size` and `falls back to the default sweep batch size for invalid configuration`. The failure hits unrelated PRs — for example the `PR` workflow runs for #10699 failed three times on this suite while the same suite passes locally. **Steps to reproduce** 1. Run `pnpm vitest run src/__tests__/decisions-service.test.ts` on a machine under load (or add a ~10 ms delay inside `decisionService.create` before the expiry validation). 2. The test builds `expiresAt: new Date(Date.now() + 5)`; by the time `create` validates, `expiresAt.getTime() <= Date.now()` is true. 3. `create` throws `expiresAt must be within 30 days` (the past-expiry branch of the validator) and the test fails before the sweep runs. **Expected behavior** The sweep tests exercise expiry deterministically and never depend on fewer than 5 ms elapsing between two clock reads in different modules. **Paperclip version** master (`717684ad8f`); the tests landed with the decisions desk workflow in #10672. **Deployment mode** Not deployment-specific — CI and local test runs. ## What Changed - `server/src/__tests__/decisions-service.test.ts`: added two helpers — `nearFutureExpiry()` (a 60 s TTL that passes validation with a wide margin) and `expireDecisionNow(id)` (moves the stored `expiresAt` into the past). The four affected tests create decisions with the future TTL, force-expire them through the store, and drop the 10 ms sleeps. The sweep observes the same expired state as before with no scheduler-timing dependence. ## Verification - `cd server && pnpm vitest run src/__tests__/decisions-service.test.ts` — five consecutive local runs, 31/31 passing each. - No production code changed; the diff is test-only. ## Risks - Low risk: test-only change. The force-expire helper writes the store directly, which is the same technique other TTL suites use to avoid sleeping through real time. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, tool use; diagnosis, fix, and verification runs. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
717684ad8f |
Add project folder browsing to skill imports (#9930)
## Thinking Path > - Paperclip is the control plane teams use to manage AI agents and their reusable capabilities. > - Skills Manager lets operators discover and import skills from project workspaces. > - Automatic discovery only surfaces skills in conventional locations, so valid skills stored elsewhere in a project are invisible. > - Operators need a safe way to navigate project folders without exposing paths outside the selected workspace. > - This pull request adds company-scoped workspace folder browsing and selection to the project skill import flow. > - The benefit is that operators can find and import valid skill folders regardless of repository layout while preserving workspace boundaries. ## Linked Issues or Issue Description - **Subsystem affected:** Cross-cutting (`server/`, `ui/`, and `packages/shared`). - **Problem or motivation:** Project skill imports rely on conventional directory discovery, which prevents operators from selecting valid `SKILL.md` folders stored in atypical locations. - **Proposed solution:** Add a company-scoped browse endpoint and a folder browser in the import dialog. The server resolves real paths, rejects traversal outside the workspace, skips symlinks and high-noise directories, identifies skill directories/files, and caps listings at 250 entries. - **Alternatives considered:** Expanding the automatic scan to every directory would be slower and noisier, while accepting arbitrary filesystem paths would weaken project/workspace scoping. - **Roadmap alignment:** This extends the completed “Skills Manager, Skill Studio & Skills Store” capability in `ROADMAP.md` without duplicating planned core work. - **Additional context:** GitHub search found no duplicate or closely related public issues or pull requests. ## What Changed - Added shared browse request/result contracts and validation for project workspace navigation. - Added a company-scoped API route and service that safely lists local workspace folders and detects `SKILL.md` entries. - Added project workspace/folder navigation to the import dialog, including parent navigation, workspace switching, truncation feedback, and direct skill selection. - Added service and route regression tests for browsing, skill detection, company isolation, and traversal rejection. - Added shared response schemas and OpenAPI documentation for the browse endpoint. - Hardened explicit skill selections with realpath containment so symlinked directories cannot escape the project workspace. ## Verification - `pnpm exec vitest run server/src/__tests__/company-skills-service.test.ts server/src/__tests__/company-skills.test.ts` — 64 tests passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm exec vitest run server/src/__tests__/openapi-routes.test.ts` — 3 tests passed. - Focused post-review reruns: `company-skills-service.test.ts` — 45 tests passed; shared/server typechecks passed. - GitHub latest-head checks — all green after one transient e2e rerun; no pending or failing checks. - `pnpm check:token-gates` — all gates clean on the rebased head. ## Risks - Low-to-moderate risk: this adds a filesystem browsing surface. Realpath containment checks prevent workspace escape, symlinks are excluded, remote-managed workspaces are rejected, and directory listings are capped. - The browser intentionally hides `.git` and `node_modules`; skills inside those directories cannot be selected through this flow. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.5`, reasoning-enabled with terminal/tool use and code execution; runtime context-window size is not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details — the execution harness requires preserving the assigned branch name. - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes — no user-facing docs changes are needed beyond this PR description because the flow is self-explanatory UI behavior. - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0a09e4d975 |
feat(decisions): add desk workflow and retention (#10672)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Decisions desk shows work that needs a human decision > - The queue foundation can group and rank decision work > - Operators also need a focused daily view and a safe way to handle old work > - This pull request adds the desk controls, the aging shelf, and reversible retention > - It also binds bulk archive decisions to the exact reviewed item set > - The benefit is a smaller daily queue without lost or orphaned work ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: server, UI, database, and shared contracts. **Problem or motivation** Decision work can grow into one large company list. Operators need quick queue and date controls. Old items also need a safe retention path that does not delete work. **Proposed solution** Add queue and date controls to the Decisions desk. Compute the aging shelf on the server. Archive idle items after 90 days unless an operator keeps them. Keep archived items searchable and revivable. Notify origin agents in one batch per sweep. Bind bulk archive proposals to a signed, exact item manifest. **Alternatives considered** Client-only aging can drift across browsers and source kinds. Deleting old rows removes audit and recovery paths. An unsigned dynamic bulk query can archive items that the reviewer did not inspect. **Roadmap alignment** This work supports the Work Queues and decision-memory directions in `ROADMAP.md`. It extends the Decisions and attention-feed foundation from #10651. Related earlier work includes #9380, #10010, and #10474. ## What Changed - Added the queue rail, date chips, decide split, triage strip, queue page, and aging shelf UI. - Added server-owned shelf state with per-queue retention overrides. - Added reversible retention state, archive history, and an idempotent notification outbox. - Added the 90-day archive sweeper, Keep exemption, archived feed query, and revive actions. - Added one origin-agent notification per agent and sweep. - Added signed bulk archive proposals with exact-set and version checks. - Persisted queue-exclusion reasons atomically and kept cross-domain source resolution per-item until it has an exact-set transaction contract. - Added API contracts, OpenAPI entries, migration coverage, focused tests, and Storybook screens. ## Verification - `pnpm -r typecheck` - `pnpm test:run` (server: 329 files and 3,453 tests passed; UI: 408 files and 3,362 tests passed) - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts` - `pnpm build` - `pnpm check:token-gates` - Focused retention, attention, decisions, migration replay, startup, and UI API tests. - Complete queue snapshot regression with 51 items across the normal 50-item page boundary. The unmodified CLI test reports one warning assertion in this runtime because the harness injects static AWS credential variables. The isolated test passes when those two variables are removed. ## Risks - The migration adds retention and notification outbox tables. It uses idempotent table, index, and foreign-key creation. - Retention runs on the heartbeat scheduler interval. Compare-and-set version checks prevent stale archive writes. - Bulk archive acceptance fails closed when authority, activity, version, or the reviewed set changes. - Archive is reversible and does not delete source records. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with `gpt-5.6-sol`. The model used tool calls, code execution, database migration generation, and test execution. The context-window size is not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
95d33e1788 |
fix(inbox): archive tasks completed by human users (#10668)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Mine inbox gives each human user a personal work queue. > - A human user can complete a task from the board. > - The status update did not update the user's inbox archive state. > - The completed task therefore stayed in the user's Mine inbox. > - This pull request archives the task for the human user who completes it. > - Agent completion does not change another user's inbox archive state. > - The benefit is that manual completion removes the task from Mine without an extra action or comment. ## Linked Issues or Issue Description **What happened?** When a human user changed a task status to `done`, the task stayed in that user's Mine inbox. **Expected behavior** The status change should archive the task from that user's Mine inbox. The operation should not add a task comment. **Steps to reproduce** 1. Open a task that appears in Mine. 2. Change the task status to `done`. 3. Return to Mine. 4. Observe that the completed task is still present. **Paperclip version or commit** `master` before this change. **Deployment mode** All deployment modes with a board user and the Mine inbox. **Access context** Board (human operator). ## What Changed - Archive the completed task for the board user who changes its status to `done`. - Persist the status update, inbox archive, and archive audit atomically. - Publish live and plugin activity only after the owning transaction commits, including recovery, decision, and approval completion paths. - Keep agent-driven completion from changing a human user's inbox archive state. - Add database-backed regression tests for the archive, audit, Mine filter, rollback, event publication, recovery, and agent paths. ## Verification - `pnpm exec vitest run server/src/__tests__/inbox-archive-routes.test.ts` passed all 7 tests. - `pnpm exec vitest run server/src/__tests__/issue-recovery-actions.test.ts` passed all 44 tests. - Five directly affected server test files passed all 74 tests; decision and comment-route suites passed all 105 tests. - `pnpm --filter @paperclipai/server typecheck` passed after the final fix. - `pnpm -r typecheck` and `pnpm build` passed. - The initial repo-wide `pnpm test:run` passed 3,223 tests; one unrelated plugin orchestration wake-reason test failed and reproduced in isolation. - All latest-head GitHub Actions gates are green, including all server and E2E shards. - Greptile reviewed the latest commit at 5/5 with no unresolved review threads. - Public GitHub search found no duplicate issue or pull request. ## Risks - Low risk. The behavior only runs on a board user's transition into `done`. - Reopening and completing the task again updates the existing per-user archive row. - Agent status updates do not archive a human user's inbox. - Transactional callers must supply a post-commit activity queue; covered completion paths do so and tests exercise rollback and publication order. - This bug fix does not duplicate planned core work in `ROADMAP.md`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex from the GPT-5 model family assisted with this change. The runtime does not expose the exact serving model ID or context window. Reasoning, repository tools, code execution, GitHub CLI, and Paperclip API access were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dcac49a4fd |
feat(workspaces): defer isolated setup until runtime start (#10653)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Isolated workspaces give each task a safe and reproducible checkout. > - The existing setup cloned the development database before an agent needed to run the app. > - This made worktree creation slower and heavier for tasks that never start a service. > - Runtime services already use one server start path for heartbeat, operator, and startup recovery flows. > - This pull request moves heavy setup to that start path and keeps worktree creation lean. > - The benefit is faster isolated workspace creation with the same reliable runtime setup when a service starts. ## Linked Issues or Issue Description Related pull request: #10652 covers the initial deferred database-seeding slice. This pull request supersedes it with end-to-end runtime provisioning and safe cleanup. **What existing behavior does this improve?** This improves isolated worktree creation, runtime service startup, and isolated instance cleanup. **Subsystem affected** Cross-cutting: CLI worktree setup, server runtime orchestration, shared workspace contracts, and development scripts. **Current behavior** Paperclip seeds an isolated development database during worktree creation. It can also leave an isolated instance directory after workspace teardown. This work happens even when no runtime service starts. **Proposed behavior** Paperclip creates the worktree with a lean eager setup. It runs an idempotent runtime provision command before the first managed service spawn. Concurrent starts share one provision attempt. Teardown removes the isolated instance safely. **Reason and benefit** Many agent tasks only edit and test code. They do not need a running Paperclip instance. Deferring the database seed reduces workspace startup cost while preserving automatic setup for tasks that start the app. **Breaking changes** None. The new runtime provision command is optional. Existing workspace behavior is unchanged when it is absent. ## What Changed - Split Paperclip worktree setup into a lean eager script and an idempotent runtime provision script. - Added `runtimeProvisionCommand` to project, issue, realized workspace, and persisted workspace contracts. - Added a per-workspace provision mutex before local service spawn for heartbeat, operator, and startup recovery flows. - Added a persisted `provisioning` service state and the `workspace_runtime_provision` operation phase. - Kept provision time outside the service readiness timeout and made failed attempts visible and retryable. - Reclaimed isolated instance data during safe workspace teardown. - Serialized deferred database seeding across processes and bound teardown to the instance root captured in persisted workspace metadata. - Added tests for config flow, concurrency, retry, no-op behavior, readiness timing, scripts, CLI commands, and cleanup. - Documented the eager and runtime provisioning contracts. ## Verification - `pnpm -r typecheck` - `pnpm build` - `pnpm test:run` (server: 3,201 passed; UI: 3,345 passed; the CLI phase exposed one environment-sensitive AWS doctor assertion because the agent runtime injects static AWS credentials) - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts -t 'passes AWS doctor checks when non-secret provider config is present'` - Focused runtime tests cover serialized provisioning, retry after stderr failure, absent-command no-op behavior, operation logging, persisted state order, and readiness timeout exclusion. - Focused CLI and cleanup tests cover concurrent seed serialization, stale-lock fail-closed behavior, persisted instance ownership, and rewritten sibling pointers. ## Risks - A faulty runtime provision script blocks service startup. Paperclip records stderr, marks the service failed, and retries on the next start. - Concurrent service requests share an in-process provision attempt, while the seed command uses an atomic filesystem lock across processes. A stale lock fails closed and requires an operator to verify no seed is running before removing it. - Isolated instance cleanup is destructive. The cleanup service validates ownership and path containment before removal. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, `gpt-5.6-sol`, with agentic reasoning, tool use, and code execution. The service does not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
173d6d2a71 |
Reclaim isolated worktree instances during teardown (#10649)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip creates isolated instances for server-managed git worktrees. > - The worktree teardown path removes the git worktree but leaves its isolated instance directory behind. > - The leaked directory can retain an embedded PostgreSQL process and database files. > - Teardown must remove only the collision-resistant instance assigned to that exact worktree path. > - This pull request stops the verified embedded PostgreSQL process and removes the guarded instance directory. > - The benefit is complete worktree cleanup without risk to another, default, or live Paperclip instance. ## Linked Issues or Issue Description **What happened?** Closing a server-managed git worktree removed the git worktree and branch, but it left the isolated Paperclip instance directory behind. A live embedded PostgreSQL process could also keep running against that directory. **Expected behavior** Worktree teardown must stop the isolated embedded PostgreSQL process and remove only the instance assigned to that exact worktree. It must refuse mismatched instance IDs and all paths outside `PAPERCLIP_WORKTREES_DIR/instances/`. **Steps to reproduce** 1. Create a server-managed git worktree with a repo-local `.paperclip/.env` file. 2. Start its isolated embedded PostgreSQL instance. 3. Close the execution workspace. 4. Observe that the git worktree is removed but the isolated instance directory remains. **Paperclip version or commit** The bug reproduces on `master` before this change. **Deployment mode** Local development with a server-managed git worktree and embedded PostgreSQL. ## What Changed - Give server-managed worktrees collision-resistant instance IDs derived from their resolved absolute paths. - Capture the repo-local instance pointer before custom teardown commands can remove it. - Require the pointer's instance ID to match the exact worktree-derived ID. - Resolve and validate the instance path against the canonical managed worktree instance root. - Verify and stop the matching embedded PostgreSQL process before directory removal, including process-exit races. - Record successful and refused cleanup operations in the workspace operation log. - Add focused ownership, process-race, path-safety, and runtime integration tests. - Document automatic isolated-instance cleanup for server-managed worktrees. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/workspace-instance-cleanup.test.ts` — 9 tests passed. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/workspace-runtime.test.ts -t "records teardown and cleanup operations when a recorder is provided"` — 1 test passed and 99 tests skipped. - `node scripts/__tests__/provision-worktree-self-heal.test.mjs` — 4 tests passed. - `bash -n scripts/provision-worktree.sh` — passed. - `pnpm --filter @paperclipai/server build` — passed. - `git diff --check` — passed. ## Risks The main risk is removal of the wrong instance directory. Provisioning assigns a path-derived ID with a SHA-256 suffix, and cleanup requires that exact ID in addition to a safe instance identifier, an absolute configured home, a strict child path, canonical path checks, and a second canonical path check immediately before removal. It refuses legacy or mismatched IDs, symlink escapes, and all paths outside the managed worktree instance root. Cleanup failures become visible warnings and do not delete an unverified path. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5. The runtime does not expose the exact model snapshot or context-window size. The agent used reasoning, repository tools, GitHub tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Paperclip <paperclip@paperclip.ing> |
||
|
|
30c49c8327 |
feat(decisions): add queues and prioritized attention feed (#10651)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source control plane for companies of AI agents. > - Operators use the attention feed to find decisions that need action. > - The feed has eleven source kinds, but it has no durable queue or triage state. > - The feed also returns every item and lacks decision deadlines, snooze state, and decision-focused ordering. > - This pull request adds secure queue sidecars and enriches the attention feed with triage data, filters, cursor pagination, and decide-now ranking. > - The benefit is a bounded feed that can show the most urgent decisions first without weakening source visibility rules. ## Linked Issues or Issue Description This pull request replaces the closed [#10634](https://github.com/paperclipai/paperclip/pull/10634). It combines that queue foundation with the dependent attention-feed change as one review unit. **Subsystem affected** Database schema, shared contracts, server authorization and REST APIs, and the UI attention client library. **Problem or motivation** The attention feed can contain hundreds of mixed decision items. Operators cannot group them into durable queues, set a decision deadline, snooze an item, or request a bounded page ordered by urgency. The current client must download the full feed on each refresh. **Proposed solution** Store queue membership and triage state by stable attention identity. Re-authorize each source during queue reads and writes. Enrich attention items with queue, deadline, snooze, expiry, rule, and origin data. Add activity and queue filters, opaque cursor pagination, decide-focused ordering, and a decide-now count. **Alternatives considered** Adding queue fields to every source would duplicate schema and authorization logic across eleven source kinds. Client-only filtering and sorting would still transfer the full feed and would make pagination unstable. **Roadmap alignment** This change improves the core decision-attention surface and operator oversight. It does not implement the separate general-purpose work queue milestone in `ROADMAP.md`. ## What Changed - Added company-scoped queue, membership, triage, and append-only event tables with actor and run provenance. - Added queue CRUD, item membership, starter-rule discovery, and decide-by and snooze endpoints. - Kept source authorization on each queue mutation, read, and count. - Added attention fields for expiry, rule, origin agent, queues, decide-by attribution, and snooze state. - Added activity date filters, queue filters, opaque cursor pagination, and configurable page limits. - Added decide-now ordering by deadline, expiry, severity, and activity. - Added `decideNowCount` and excluded actively snoozed items from the default feed. - Updated the shared and UI client contracts. - Added focused server, route, OpenAPI, and UI client tests. ## Verification - `pnpm exec vitest run server/src/__tests__/attention-service.test.ts server/src/__tests__/decision-queues-routes.test.ts server/src/__tests__/openapi-routes.test.ts ui/src/api/attention.test.ts ui/src/lib/attention.test.ts` (72 tests passed) - `pnpm --filter @paperclipai/db check:migrations` - `pnpm -r --filter @paperclipai/db --filter @paperclipai/shared --filter @paperclipai/server --filter @paperclipai/ui typecheck` - `git diff --check origin/master...HEAD` ## Risks - The migration adds four company-scoped tables and provenance foreign keys. Migration numbering and safety checks pass. - Attention reads can lazily create starter queues and memberships. Inserts are idempotent, audited, and transactional. - Cursor validity depends on the filtered feed. The API returns a clear validation error when the cursor item no longer exists in that feed. - Queue reads re-check source visibility. This favors correct authorization over fewer queries. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, model `gpt-5`. The runtime used agentic reasoning, repository tools, code execution, and test execution. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
14d4db6330 |
fix(inbox): keep passive issue views out of Mine (#10581)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The inbox shows work that a board user owns or has joined > - Issue detail views update a per-user read receipt > - The Mine query treated that read receipt as user participation > - A passive view therefore added the issue to Mine > - This pull request separates passive reads from audited user mutations > - The benefit is a Mine inbox that reflects ownership and real participation ## Linked Issues or Issue Description **What happened?** Opening an issue detail marks the issue as read. The Mine query treated the read receipt as participation. The viewed issue then appeared in Mine even when the user did not change it. **Expected behavior** Viewing an issue can update unread state. A view alone must not add the issue to Mine. Issue creation, assignment, comments, and audited user mutations must add it. **Steps to reproduce** 1. Open an issue that you did not create and that is not assigned to you. 2. Do not comment or change the issue. 3. Open the Mine inbox. 4. Observe that the issue appears in Mine on the previous implementation. **Paperclip version or commit** `90ead239a8` **Deployment mode** Local dev from source. **Additional context** Related approach: #3421 changes Mine to an assignee-only filter. This change keeps participation-based Mine behavior and corrects the participation signal. ## What Changed - Use an explicit audited user-mutation allowlist for Mine participation, including comment cancellation. - Keep passive reads, previews, denied resource requests, and archive bookkeeping out of Mine participation. - Record manual routine reuse as an explicit audited inbox touch instead of a read receipt. - Keep that inbox bookkeeping from satisfying routine activity gates. - Add focused regression coverage for passive views, real mutations, comment cancellation, and manual routine runs. - Document the Mine participation contract. ## Verification - `pnpm exec vitest run server/src/__tests__/issues-service.test.ts -t "does not treat passive issue activity"` - `pnpm exec vitest run server/src/__tests__/routines-service.test.ts -t "touches a (coalesced|skipped active) routine issue|ignores inbox bookkeeping activity"` - `pnpm --filter @paperclipai/server typecheck` - GitHub latest-head CI: all required checks passed, including build, typecheck, server suites, and e2e shards. ## Risks - Low risk. The change affects only the server query that defines user participation in Mine and the manual routine touch signal. - Historical audited user mutations can now qualify an issue for Mine. Passive read and archive actions remain excluded. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5 family. The runtime did not expose the exact deployment ID or context window. Agentic reasoning, tool use, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e4b0152ca3 |
fix(server): keep the agent invokable when a run fails on a workspace sync conflict (#10660)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents on the same project can share one `shared_workspace` clone, and each sandbox run reconciles its git history back into it > - When histories written by different runs genuinely diverge, reconciliation can hit a real merge conflict — a property of the *workspace*, not of whichever agent happened to run last > - That agent nonetheless finalized into a sticky `error` state, removing a healthy agent from rotation while the workspace stayed broken — and on a shared workspace this serially knocks out every agent that touches it > - This pull request classifies workspace-reconciliation failure signatures as workspace-scoped, so the run still fails with the full message but the agent stays invokable > - The benefit is that one bad workspace state no longer disables agents one by one ## Linked Issues or Issue Description Refs #10645 — this addresses the sticky-agent-error clause of that issue. Workspace run serialization / per-agent worktrees remain tracked there (design sketch on the issue). ## What Changed - New exported `isWorkspaceSyncConflictFailure(message)` matching the reconciliation failure signatures: `merge-tree` conflict ("Failed to merge concurrent remote git histories"), integrate-retry exhaustion ("Failed to integrate concurrent remote git history"), and bundle prerequisite failures ("did not send all necessary objects", "lacks these prerequisite commits"). - Both run-failure finalization paths (adapter returned a failed result; adapter threw) pass `keepIdleOnFailure` for these signatures — the same mechanism already used for provider-quota failures — so the agent finalizes to `idle` instead of `error`. The run itself still fails and carries the full message; nothing about run reporting changes. ## Verification - `pnpm vitest run server/src/__tests__/heartbeat-workspace-session.test.ts` — new signature matrix (4 positive signatures, negatives for unrelated adapter failures and null/empty); 121 tests total. - `cd server && pnpm run typecheck`. ## Risks - Low. The only change is which failure families put the agent into `error`; behavior for every other failure is untouched. A workspace stuck in conflict still fails every run against it (visible on the runs surface) — it just no longer takes agents down with it. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use. No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
592cade5a6 |
feat(server): trigger the push-capability preflight from the issue's stated PR deliverable (#10659)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Runs that must push to GitHub have a pre-dispatch credential preflight (`push_write_credential_missing`) so the missing token surfaces as a configuration-incomplete blocker instead of a late runtime failure > - The preflight only triggers when the issue mentions the GitHub PR workflow *skill* — routine-created issues and agent-to-agent handoffs rarely do, even when their text literally says "push the branch and open a PR" > - In practice the credential gap then surfaced only after implementation and review were complete, stranding finished work > - This pull request adds a conservative, verb-anchored text heuristic over the issue title and description as a second preflight trigger > - The benefit is that the credential ask reaches the human before any work is burned ## Linked Issues or Issue Description Fixes #10644 (completes the prevention set with #10648, #10650, #10658) ## What Changed - `issueTextImpliesPrDeliverable(text)`: matches verb-anchored deliverable statements — "open/create/raise/submit a (draft) pull request/PR", "push … branch/remote/origin/upstream". Verb anchoring deliberately ignores passing mentions ("the PR merged yesterday", "PR feedback addressed"). - `requiresPushCapabilityPreflight` takes the issue's title+description and ORs the text heuristic with the existing skill-mention trigger; adapter-type and issue gating are unchanged. The run-dispatch call site threads the already-loaded issue text — no extra query. ## Verification - `pnpm vitest run server/src/__tests__/heartbeat-workspace-session.test.ts` — new heuristic matrix (4 positive, 6 negative including null/empty) and preflight-by-text cases (text triggers, passing mention does not, no issue → no preflight); 122 tests total. - `cd server && pnpm run typecheck`. ## Risks - A false positive turns into a configuration-incomplete blocker asking for a GitHub token on an issue that didn't need one — the heuristic is intentionally conservative (verb-anchored) to keep that rare, and the blocker names the exact remediation. - No behavior change for issues that neither mention the skill nor state a PR deliverable. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use. No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ddbcf53e31 |
fix(server): refuse agent delegation cycles back to an open ancestor's creator (#10658)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents decompose work by creating child issues assigned to other agents > - When two agents each lack a capability the other assumed (e.g. neither can push to GitHub), each can "resolve" its blocker by delegating the same step to the other: A creates a child for B, B creates a grandchild back for A > - Nothing detects the cycle; the chain of blocked issues grows and no signal reaches the human who could actually fix the capability gap > - This pull request refuses agent-initiated child creation when the child's assignee is the creator of a still-open ancestor in the same chain — a mechanical, semantics-free cycle signal > - The benefit is that the hot-potato dies at creation time with an actionable error instead of growing a dead chain ## Linked Issues or Issue Description Fixes #10642 (write-time counterpart: #10648 refuses assignment to paused agents; the credential-gap *preflight* side is tracked separately in #10644) ## What Changed - `issueService.findOpenAncestorCreatedByAgent(parentIssueId, agentId, {maxDepth})`: bounded walk up the parent chain looking for a still-open (not done/cancelled) ancestor created by the given agent. - Agent-initiated issue creation with a parent (both the create-with-`parentId` route and `POST /issues/:id/children`) now refuses with a structured 409 (`code: delegation_cycle`, naming the ancestor) when the new child would be assigned to the agent that created a still-open ancestor: that agent delegated the work into this chain, so assigning it back is a cycle. The message states the alternatives — complete the work, leave the child unassigned, or escalate to a board operator. - Deliberately unaffected: human actors (deliberate re-routing is their call), closed ancestors (re-engaging the creator of finished work is normal), and accepted-plan decomposition (its children come from a human-approved plan). ## Verification - `pnpm vitest run server/src/__tests__/issue-assignee-invokability-routes.test.ts` — cycle refused with 409 and no create call; the same child allowed when no open ancestor matches; board actors never consult the guard. - `pnpm vitest run server/src/__tests__/issues-service.test.ts` — new embedded-Postgres coverage: ancestor found through the chain, closed ancestors ignored, depth bound honored (114 total). - `pnpm vitest run server/src/__tests__/issue-create-deduplication-routes.test.ts server/src/__tests__/issue-agent-mutation-ownership-routes.test.ts` — unchanged (79). - `cd server && pnpm run typecheck`. ## Risks - Low-to-moderate: a new 409 for a creation shape that previously succeeded. The blocked shape (agent assigns new work to the creator of an open ancestor) is the cycle signature; the legitimate "hand a subtask to the parent's assignee" pattern is unaffected because it keys on assignee, not creator. Watchdog and plan-decomposition flows are exempt or unaffected as described. - The walk adds at most `maxDepth` (10) single-row lookups per agent child creation with an assignee. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use. No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |