mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 16:11:46 +02:00
00a24d7e8f50d5eafd4ecb6b7accb46f1cf55d1d
438
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
72b509c895 |
Recognize delivered workspaces and reap terminal worktrees (#10908)
## Thinking Path > - Paperclip separates workspace provisioning lifecycle from whether the work was actually delivered. > - Git ancestry alone cannot recognize squash merges or deliveries into a branch other than the workspace base. > - A merged pull request linked from a terminal issue is stronger delivery evidence for those cases. > - The read contract should expose that evidence without changing persisted workspace schema. > - Cleanup must remain conservative: terminal descendants, delivered work, and no active run checkout are all required. > - Reusing the existing cleanup primitives keeps service shutdown, lease cleanup, activity logging, and archival behavior consistent. > - Focused regression coverage locks in both the honest read signal and the fail-closed reaper guards. ## Linked Issues or Issue Description **What existing behavior does this improve?** Execution workspace close-readiness payloads and terminal workspace cleanup. **Current behavior** Delivered squash-merged or cross-branch workspaces can remain `active` and report a permanent “not merged” warning because git ancestry does not contain their original commits. **Proposed behavior** Read payloads distinguish PR-confirmed delivery, ancestry delivery, unmerged work, and unknown state. Fully terminal delivered workspace trees are archived only when no active run holds the checkout. **Reason and benefit** Operators and automation receive an honest delivery signal, while shipped worktrees stop looking active forever and genuinely unmerged work retains its warning. **Breaking changes** The workspace payload gains a derived field. Existing fields and persistence remain unchanged; no database migration is required. **What happened?** A delivered workspace can remain `active` and warn that it is not merged forever after its issue ships through a squash or cross-branch pull request. **Expected behavior** Pull-request delivery should be represented honestly, and a fully terminal delivered workspace should become cleanup-eligible when no run holds its checkout. **Steps to reproduce** 1. Create an issue workspace with commits ahead of its configured base. 2. Deliver those commits with a squash merge or into a different target branch. 3. Mark the source issue and descendants done, then read workspace close readiness. Before this change, the workspace remains active with a “not merged” warning indefinitely. ## What Changed - Added the derived `deliveryState` workspace contract: `merged_via_pr`, `merged_by_ancestry`, `unmerged`, or `unknown`. - Extracted a shared GitHub pull-request merge classifier and reused it for merge confirmations and workspace delivery checks. - Suppressed false ancestry warnings when a terminal issue has ground-truth merged-PR evidence. - Added an idempotent terminality reaper with descendant-terminal, active-run, and delivered-work guards. - Restricted PR delivery evidence to the source issue, then required live merged state plus matching GitHub repository, head branch, and current workspace HEAD; persisted status, stale PRs, lexical mentions, inbound references, and descendant PRs cannot authorize cleanup. - Preserved workspaces with modified or untracked files even when their committed HEAD was delivered. - Bounded both long-lived pull-request state caches to 1,000 entries with oldest-entry eviction. - Routed eligible workspaces through existing runtime shutdown, lease cleanup, activity logging, and archival machinery with exclusive Git index, HEAD, and branch-ref locks plus non-forced removal. - Added regression coverage for delivery derivation, warning behavior, reaper guards, scheduler wiring, and squash/cross-branch delivery. ## Verification - `pnpm -r typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/execution-workspaces-service.test.ts src/__tests__/merged-pr-confirmation-sweep.test.ts src/__tests__/server-startup-feedback-export.test.ts --reporter=verbose` — 63 passed - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/execution-workspaces-service.test.ts --reporter=verbose` after review hardening — 43 passed - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/execution-workspaces-service.test.ts src/__tests__/merged-pr-confirmation-sweep.test.ts src/__tests__/external-objects-service.test.ts --reporter=dot` on the final local head — 73 passed - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-workspace-busy.test.ts --reporter=verbose` — 15 passed - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm test:run` — server 3,662 passed (4 skipped), UI 3,599 passed, CLI 327 passed, shared 415 passed, and skills catalog 20 passed; the aggregate DB stage ran both source and built copies of one unrelated embedded-Postgres migration test and both reached its 5-second timeout - `pnpm --filter @paperclipai/db exec vitest run src/status-card-migrations.test.ts --reporter=verbose` — isolated aggregate-timeout verification passed in 3.99 seconds - `NODE_ENV=production pnpm build` - `pnpm check:token-gates` ## Risks The reaper intentionally fails closed when issue terminality, pull-request state, git ancestry, or checkout ownership cannot be proven. GitHub lookups can delay classification and cleanup but cannot cause an unproven workspace to be archived. Automated terminal archival holds exclusive Git index, HEAD, and branch-ref locks across validation and removal, skips configured destructive hooks, and uses non-forced removal so dirty writes fail closed. Reopening a source issue does not restore an archived workspace; it emits an audit event so a human or agent can re-provision explicitly. ## Model Used OpenAI Codex, GPT-5. The runtime did not expose a more specific model ID or context-window size. Reasoning, tool use, repository editing, test execution, and GitHub CLI access were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6ffe9df842 |
fix(auth): clarify protected-agent assignment blocks (#10893)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task assignment policies control which agents can receive work. > - Protected-agent policy flags currently stop assignment. > - The existing error says that the assignment requires approval. > - Paperclip has no approval workflow for this policy. > - This pull request models the policy as a hard block and gives the operator an action that exists. > - The benefit is accurate API guidance without weakening the existing fail-closed behavior. ## Linked Issues or Issue Description Refs #6386 **What happened?** A protected-agent assignment denial said that approval was required. No approval record or approval action existed for this policy, so the message sent agents and operators to a dead end. **Expected behavior** The authorization result must state that protected-agent policy blocks assignment. It must tell a company administrator to remove the block before retrying. **Steps to reproduce** 1. Set `authorizationPolicy.protectedAgent.requiresApproval` to `true` on a target agent. 2. Give another agent the `tasks:assign` permission. 3. Preview or attempt assignment to the protected agent. 4. Observe that the old response promises an approval step that does not exist. **Paperclip version or commit** `c54936e2e9` on `master`. **Deployment mode** Built from source. The behavior is in the core authorization service and is not deployment-specific. **Agent adapter(s) involved** Not adapter-specific. ## What Changed - Added canonical `protectedAgent.blockAssignment` and `protectedAgent.blockReason` policy fields. - Kept the legacy approval-named flags as fail-closed compatibility aliases. - Changed denial copy to name the hard block and the administrator action. - Added authorization and plugin-host regression coverage for canonical and legacy policy data. - Updated the V1 implementation contract with the protected-assignment rule. ## Verification - `pnpm exec vitest run server/src/__tests__/authorization-service.test.ts server/src/__tests__/plugin-access-authorization-host-services.test.ts` — 2 files passed, 61 tests passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/shared build` — passed. - `pnpm --filter @paperclipai/server build` — passed. - `pnpm check:token-gates` — all gates clean. - `git diff --check public-gh/master...HEAD` — passed. The repository-wide local wrappers exceeded the execution host resource limit before they printed a final summary. The PR check loop will use GitHub CI as the complete test and build authority. ## Risks - Low: assignment remains fail-closed. The change corrects the policy name and denial guidance. - Low: legacy fields remain supported, so existing plugin-owned policy data does not change behavior. - Low: the new policy schemas allow unknown keys for forward compatibility, as the existing authorization policy schema already does. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5`, tool-enabled coding agent with reasoning, shell, Git, and GitHub CLI access. The runtime does not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ef33c1d9ed |
fix(decisions): retire completed-target decisions and link targets from the card (#10892)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The decisions desk shows pending decisions that need an operator response > - A strict decision cannot apply its effects after its target task changes > - A decision still remained pending when every target task finished after proposal > - The card also linked only the origin task, even when the decision acted on another task > - This pull request expires those moot decisions and links their target tasks > - The benefit is an accurate queue and a clear path to the work that each decision affects ## Linked Issues or Issue Description Related PR: #10801 removes the issue-page decision strip, which makes clear queue provenance more important. **What happened?** A strict decision stayed pending until its time-to-live limit after every target task reached `done`. The decision card linked only the origin task. The origin task is where the agent proposed the decision, and it can differ from the task that the decision affects. An operator could therefore open a finished task with no visible decision and no explanation of the real target. **Expected behavior** Paperclip must expire a strict decision when all of its targets finish after the decision is proposed. The card must show and link every target task that differs from the origin task. **Steps to reproduce** 1. Create a strict decision that targets an active task from a different origin task. 2. Move the target task to `done` without resolving the decision. 3. Run the decision expiry sweep. 4. Observe that the old code keeps the decision open until its time-to-live limit. 5. Observe that the old card links only the origin task. **Paperclip version or commit** The bug reproduces on upstream `master` before this pull request. **Deployment mode** Local dev and self-hosted server modes are affected because the behavior is in the shared decision service and board UI. ## What Changed - Expire an open strict decision with reason `target_completed` when every strict target reached `done` after proposal. - Keep decisions that intentionally target an already-finished task. - Keep lenient-only decisions open. - Keep continuation delivery consistent with other expiry reasons. - Add target-task links to the decision card provenance line. - Use one shared target-ID helper across signing, execution, expiry, card provenance, and resolver preloading. - Add service and UI regression tests for primary, secondary, and target-completed cases. ## Verification - `pnpm exec vitest run ui/src/components/DecisionCard.test.tsx server/src/__tests__/decisions-service.test.ts` — 51 tests passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm check:token-gates` — all gates clean. - `git diff --check origin/master...HEAD` — passed. ## Risks - Low migration risk. This change does not alter the database schema. - The expiry sweep performs the existing strict-target query and adds a snapshot comparison before expiry. - A decision remains open if any strict target is active or if a target was already `done` at proposal time. > The roadmap lists work queues as planned. This pull request fixes the existing decisions desk. It does not add a new queue subsystem. ## Model Used - Implementation: Anthropic Claude through Claude Code. The runtime did not expose the exact model snapshot or context-window size. The model used reasoning, repository tools, code execution, and test execution. - PR preparation: OpenAI Codex with GPT-5. The runtime did not expose a dated model snapshot or context-window size. The model used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e8ae5286eb |
feat(issues): explain cross-task agent writes with attribution, audit receipts, and actionable denials (#10843)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents write to tasks they do not own. They comment, they change
fields, and the control plane now permits this by default for
standard-trust agents on any task they can read
> - This makes a task thread ambiguous. A reader sees a comment from an
agent that is not the assignee, but no surface says whose authority that
write rode
> - The same gap applies to field edits. The activity stream named the
verb, but it did not show the before value, the after value, or the
reason the write was permitted
> - The remaining refusals are also opaque. An agent that hits a wall
receives a 403 with no boundary name, no actor who can act, and no
sanctioned path. One real incident spent a full detour to find the
workaround
> - This pull request adds the three surfaces that make open cross-task
writes legible: an attribution chip, a field-level audit receipt, and an
actionable denial contract shared by the API and the UI
> - The benefit is that a reader can answer "who did this, on whose
authority, and was it allowed?" on the task itself, and a blocked writer
is told what to do next
## Linked Issues or Issue Description
No public issue exists for this work, so the enhancement is described
here.
**What existing behavior does this improve?**
Cross-task agent writes are permitted, but they are not explained. A
task thread can hold comments from agents that are not the assignee, and
the activity stream can hold field changes made by those agents. Neither
surface names the responsible user behind the write. When a write is
refused, the error text does not name the boundary or the way forward.
**Subsystem affected**
Issue detail UI (comment thread and activity stream), the issue write
authorization responses in the server, and the shared copy contract that
both consume.
**Current behavior**
- An agent comment on a task the agent does not own looks the same as an
assignee comment.
- An `issue.updated` activity row states the verb only. It does not show
the field-level before and after values, the responsible user, or the
authorization reason.
- A refused write returns a short message such as an ownership error.
The message does not state which rule fired, who is able to perform the
action, or which alternative path is sanctioned.
**Proposed behavior**
- An agent comment on a task the agent does not own carries a chip that
reads "for {user}". The chip names the responsible user. Its tooltip
states that the author is not the assignee and cannot exceed that user's
permissions.
- Each `issue.updated` row shows a receipt: the changed fields with
before and after values, the responsible user, and the authorization
reason. This applies to board edits as well as agent edits.
- Each refusal states three things: the boundary that fired, who is able
to act, and the sanctioned path. The API error body and the in-app
notice use the same words, because both read one shared contract.
Related pull requests, found by searching this repository:
- Refs #10837 — merged. It added the default-open cross-task write rule,
the comment attribution data, and the per-run containment cap that this
pull request makes visible.
- Refs #10114 — open. It proposes a narrower authorization change in the
same area.
- Refs #7998 — open. It proposes append-only cross-assignee comments as
an alternative to opening writes.
## What Changed
- Adds `packages/shared/src/issue-write-denial.ts`. This is one copy
contract for eight ways an issue write can be refused: not visible,
responsible-user ceiling, responsible user unavailable, excluded actor
class, assignee run lock, per-run cross-task cap, missing run context,
and rejected attribution. Each entry names the boundary, who can act,
and the sanctioned path.
- Maps server authorization decisions onto that contract in
`server/src/routes/issues.ts` and
`server/src/services/cross-issue-influence-limit.ts`. The flattened
`error` string carries all three obligations, and `details.code` lets
the UI render the same words. The two cap codes keep the names they
already ship under.
- Adds `CommentAttributionChip`. It renders "for {user}" beside the
author name on agent comments where the author is not the assignee. It
renders nothing when no responsible user is recorded, so older rows stay
clean. It is wired into both `IssueChatThread` and the flagged
`TaskChatThread` redesign.
- Adds `IssueFieldChangeReceipt`. It renders the change receipt under
`issue.updated` rows in the activity stream. Ids resolve to agent and
user names where the directory is loaded. Server-truncated text is
labelled as a preview, so the receipt never implies that it shows a
whole value.
- Adds `IssueWriteDenialNotice`. It renders the shared copy in the app,
keyed off the denial events the server logs on a task.
- Adds a public `/ux-lab/cross-issue-collaboration` page. It renders all
three surfaces and their edge cases for review without a seeded thread.
This follows the existing `ux-lab` pages.
## Verification
Automated, all green:
```
pnpm --filter @paperclipai/shared exec vitest run src/issue-write-denial.test.ts # 17 tests
pnpm --filter @paperclipai/ui exec vitest run src/components/IssueWriteDenialNotice.test.tsx \
src/components/IssueFieldChangeReceipt.test.tsx src/components/CommentAttributionChip.test.tsx \
src/lib/issue-change-receipt.test.ts src/lib/comment-attribution.test.ts # 46 tests
pnpm --filter @paperclipai/server exec vitest run src/__tests__/cross-issue-influence-limit.test.ts \
src/__tests__/issue-comment-attribution-audit-routes.test.ts \
src/__tests__/issue-agent-mutation-ownership-routes.test.ts \
src/__tests__/low-trust-red-team-routes.test.ts # 98 tests
```
`tsc --noEmit` passes for the shared, ui, and server packages.
Manual, in a browser:
1. Start the UI only: `pnpm --filter @paperclipai/ui exec vite`.
2. Open `/ux-lab/cross-issue-collaboration`. No session is needed,
because `ux-lab` routes are public.
3. All three surfaces were captured at 1440x900 in light mode and dark
mode, and at 390x844. The page reported no errors.
4. The chip tooltip was opened by a hover and by a keyboard focus.
Rendering the page found defects that the tests had missed. Three copy
and contrast defects were fixed, and two of them are now pinned by a
test. A design review then found three layout defects, which are also
fixed: the denial notice orphaned its label when a value wrapped, the
receipt icon wrapped onto its own line at narrow widths, and the chip
tooltip was reachable by hover only.
## Risks
Low risk, and additive.
- Every new surface renders nothing when its data is absent. Comments
without a recorded responsible user show no chip, and activity events
without a receipt show no receipt, so existing rows do not change.
- No migration is included. The data these surfaces read already ships.
- The wire values of the two per-run cap denial codes are unchanged.
Only the human-readable text changes, plus six codes that had no
`details.code` before.
- The denial copy is read by agents as well as people. If wording must
change later, one shared module is the only place to change it.
- Roadmap check: this extends the completed "Activity log & action
attribution" area rather than duplicating planned core work.
## Model Used
Claude Opus 5 (Anthropic), model id `claude-opus-5[1m]`, 1M context
window, extended thinking, with tool use and code execution. It ran as
an agent in Claude Code and drove a real browser to capture the review
screenshots.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
678728f650 |
feat: maintained in_review review-path contract + stalled-review actions (#10675)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents move issues to `in_review` and rely on a "review path" (an interaction, an approval, a monitor, or a named reviewer) to tell them who decides next. > - That review path can silently disappear. A user comment supersedes the pending interaction, a monitor is exhausted, or a run ends without restoring a path. The issue then sits in `in_review` with nobody reviewing it and no visible action. > - Such issues become invisible zombies. Nobody knows a decision is owed, so the work stalls forever. > - This pull request makes the review path a maintained invariant, exposes a `reviewAttention` surface, and gives every stalled review three inline actions in the UI. > - The benefit is that an `in_review` issue always shows who reviews it, or shows an amber "nobody is reviewing this" notice with one-click Approve, Request changes, and Send back to work. ## Linked Issues or Issue Description This pull request describes the problem inline. The tracking issue is internal. **Subsystem affected** The review and attention loop that agents and humans share: the `in_review` status, the `reviewAttention` surface, the /decisions attention feed, and the issue-page review panel. **Problem or motivation** Agent-owned issues in `in_review` can lose their last review path. A user comment supersedes the pending interaction. A monitor is exhausted. A run ends without restoring a path. The issue then sits in `in_review` with no reviewer and no visible action. It becomes an invisible zombie and the work never progresses. **Proposed solution** Maintain the review path as a server invariant. Expose a `reviewAttention` field that says what is under review, who decides, and since when. Render a persistent review panel on the issue page and inline actions on the /decisions feed. Keep human PATCHes into `in_review` ungated, but record the requesting user so the panel never renders empty. **Alternatives considered** A pure background auto-recovery sweep. This stays opt-in and is not enough on its own, because it is invisible to the human. A bare status banner. This is rejected, because it gives no action to resolve the stall. **Roadmap alignment** This improves the core review and attention loop that both agents and humans use every day. ## What Changed - **Server — maintained review-path invariant:** when an issue enters or sits in `in_review`, the server derives and persists a review path (interaction, approval, monitor, or the requesting user) and recovers a stale path with one bounded wake instead of leaving the issue pathless. - **Server — `reviewAttention` surface:** a new field describes what is under review (bound target with links), who decides, since when, and whether the review is stalled. Stalled agent-assigned reviews are now included in the attention feed. - **Server — inline stalled-review decisions:** secured routes let a permitted responder Approve (→ `done`), Request changes (→ `todo` + wake carrying the note), or Send back to work (→ `todo` + wake) directly from the attention feed. - **Server — resume-intent wake:** an `in_review -> todo` transition now wakes the assigned agent so a resumed review is not dropped. - **Server — user-entry symmetry:** user PATCHes into `in_review` stay ungated (no 422 for humans) and record the requesting user, who becomes the named responder when no other path exists. - **UI — review panel:** a persistent `IssueReviewPanel` renders above the thread whenever status is `in_review`. The covered state shows the bound target, responder, and outcomes and hoists the pending interaction/approval card. The stalled state shows the amber notice plus the three actions. - **UI — decisions card actions:** the same three actions render inline on the /decisions `AttentionQueueRow`. - **UI — responsive fix:** the stalled action row stacks to full-width buttons at phone width and returns to a horizontal row at `sm` and up. New 390px stories capture the phone layout. ## Verification - `cd ui && npx vitest run src/components/IssueReviewPanel.test.tsx src/components/AttentionQueueRow.test.tsx src/lib/attention.test.ts src/api/issues.test.ts` — 91 tests pass. - Server suites added and updated: `issue-review-attention`, `issue-stalled-review-decision-routes`, `review-path-recovery`, `recovery-observability`, and related route/liveness tests (run by CI). - A designer reviewed the UI at 390px and desktop in light and dark themes on both the issue-page panel and the /decisions card. The stalled action row stacks cleanly at phone width with no overlap and keeps the horizontal row on desktop. ## Risks - **Migration:** adds migration `0200` (next after master `0199`, no renumber). It extends the agent-wakeup-requests schema and is additive. - **Behavioral shift:** `in_review -> todo` now dispatches a wake. This is intended (resume intent) and covered by tests. - **Authz:** the inline decision routes are permission-gated. Only a permitted responder sees and can trigger the actions. - Overall risk is moderate and contained to the review and attention loop. ## Model Used - Claude, Opus 4.8 (`claude-opus-4-8`), extended thinking, tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f91a6e27c0 |
feat(issues): contain cross-issue agent side effects (#10837)
## Thinking Path > - Paperclip is the control plane that coordinates autonomous agent work. > - Agents need to collaborate on issues beyond their current assignment. > - Cross-issue comments and updates are useful, but an unbounded run can create cascading side effects. > - The control plane must preserve company-wide collaboration while containing each run's influence. > - Comment attribution must also show the responsible user and the acting agent in audits. > - This pull request adds run-bound cross-issue containment, attribution, and agent-class wake rules. > - The benefit is safer collaboration without restoring issue-assignee ownership restrictions. ## Linked Issues or Issue Description **What existing behavior does this improve?** Agent-authenticated issue comments, updates, reopen behavior, and assignee wake routing. **Subsystem affected** Cross-cutting: server routes and services, shared contracts, database schema and migration, and implementation documentation. **Current behavior** An authenticated agent can collaborate across company issues, but one heartbeat run has no per-run side-effect boundary. Comment records also do not persist the responsible user separately from the acting agent. **Proposed behavior** Require a valid heartbeat run for agent cross-issue comments and updates. Audit each attempt and cap a run at 20 cross-issue effects. Keep the cap in log-only mode until it automatically changes to enforcement at 2026-08-11 00:00 UTC. Preserve same-issue writes. Use agent-class wakes for agent comments. Keep same-run completion comments from reopening completed work. Record the responsible user on agent-authored comments and activity. **Reason and benefit** Agents can collaborate on other issues without an assignment gate, while each run has an atomic and inspectable side-effect limit. Operators can identify both the acting agent and the responsible user. **Breaking changes** After 2026-08-11 00:00 UTC, the twenty-first cross-issue comment or update from one heartbeat run returns a containment error. Agent cross-issue writes without valid run context are rejected. The migration is additive and backfills existing agent-authored comment attribution where the source data is available. ## What Changed - Added an atomic per-run counter for cross-issue agent comments and updates. - Added audit events for allowed and rejected cross-issue effects. - Added the automatic log-only to enforcement flip at 2026-08-11 00:00 UTC. - Added responsible-user attribution to agent-authored comments, activity records, shared types, and validators. - Added an additive migration and migration coverage for existing comments. - Updated reopen, resume, and wake behavior so agent comments create agent-class wakes and same-run completion comments remain inert. - Updated the implementation specification and regression coverage. ## Verification - `pnpm exec vitest run server/src/__tests__/cross-issue-influence-limit.test.ts server/src/__tests__/issue-comment-attribution-audit-routes.test.ts server/src/__tests__/issue-comment-reopen-routes.test.ts packages/db/src/issue-comment-on-behalf-migration.test.ts` — 97 tests passed. - `pnpm -r typecheck` — passed, including migration safety checks. - `pnpm test:run` — server batch: 3,364 passed and 2 skipped; UI batch: 3,504 passed. One unrelated CLI doctor test warned because this agent runtime injects static AWS credentials. - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts` — 8 tests passed and confirmed the CLI failure was ambient-environment sensitive. - `pnpm build` — passed. ## Risks - The fixed enforcement timestamp changes production behavior automatically on 2026-08-11 00:00 UTC. Audit logs before that time provide rollout visibility. - The per-run counter serializes on the heartbeat-run row. This prevents concurrent attempts from racing past the cap but adds a small lock scope for cross-issue writes. - Existing comments can only be backfilled when their acting run or agent attribution is recoverable. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 in the Codex agent runtime. The runtime did not expose a context-window size. Reasoning, shell tools, code editing, and test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ded813ad6f |
feat(interactions): add governed agent addressees (#10252)
## Thinking Path > - Paperclip is the control plane that lets humans govern companies of AI agents. > - Issue-thread interactions are the structured handoff point for confirmations, questions, suggested tasks, and other governed decisions. > - Those interactions previously assumed that only board users could resolve them, preventing one agent from explicitly addressing another agent for a response. > - Agent resolution needs company-level governance, auditable resolver identity, safe terminal-state handling, and attention routing so authorization is enforced server-side rather than inferred from UI behavior. > - This pull request adds governed agent resolution, withdrawal and terminal expiry semantics, explicit agent addressees, lifecycle reconciliation, and attention-feed filtering. > - The benefit is that agents can participate in structured decisions without weakening board control, company isolation, wake behavior, or audit invariants. ## Linked Issues or Issue Description ### Subsystem affected Issue-thread interactions across database, shared contracts, server authorization/services, adapter callbacks, agent skill guidance, API docs, and UI governance surfaces. ### Problem or motivation Structured interactions were board-only, had no explicit agent addressee, and lacked durable withdrawal/terminal-expiry semantics. That made peer-agent decisions impossible to authorize and audit safely. ### Proposed solution Persist requested/effective resolver policy and addressee identity, enforce company governance and eligible agent resolution, reconcile addressee lifecycle changes, expose withdrawal and terminal expiry, and route attention to the intended active agent with board fallback. ### Alternatives considered Implicitly authorizing the issue assignee or mentioned agents was rejected as ambiguous and difficult to audit. Using comments alone was rejected because it loses structured outcomes and continuation behavior. ### Roadmap alignment Supports the ROADMAP direction for lightweight leadership-agent communication that still resolves into governed decisions and work objects. ### Additional context Public GitHub issue/PR search found no duplicate implementation; open PR search for interaction resolver governance and agent addressees only returned this PR. ## What Changed - Add company-scoped interaction resolver governance contracts and persistence. - Add requested/effective resolver policy, resolver identity, withdrawal, and terminal-expiry behavior. - Add explicit `addresseeAgentId` validation, authorization, persistence, lifecycle reconciliation, API documentation, and skill guidance. - Route pending addressed interactions to the intended invokable agent and fall back to board attention when that agent becomes ineligible or is deleted. - Preserve sandbox callback identity fields required by governed resolution paths. - Add migrations `0193` and `0194` plus route, service, attention, adapter, CLI, and UI coverage. - Add governance state and company settings UI, including responsive mobile behavior and distinct withdrawn/expired audit presentation. ## Verification - `pnpm check:token-gates` — passed. - `pnpm -r typecheck` — passed, including migration numbering and safety checks. - `pnpm test:run` — feature/server and UI workspace suites passed; one unrelated CLI AWS doctor test observed injected static AWS credentials and warned instead of passing. - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts --project paperclipai` — 8 tests passed, confirming the failure was environment-sensitive. - `pnpm build` — passed. - Latest rebased head `e24cece6be9f1877bdbac7691bcb44fd583c0161` completed all GitHub CI jobs successfully. ## Risks - Migrations add interaction and company-governance fields; numbering is conflict-free on current `master`, additive statements are idempotent, and migration safety checks pass. - Agent authorization behavior expands beyond board-only resolution, but defaults remain board-only and coverage exercises company boundaries, resolver eligibility, lifecycle invalidation, wake behavior, withdrawal, expiry, and attention fallback. - Attention routing depends on current agent invokability; reconciliation and read-time filtering prevent stale addressees from retaining visibility or resolution authority. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using `gpt-5.6-sol` with reasoning, terminal tool use, code execution, Git/GitHub integration, and Paperclip control-plane tools. Context-window metadata was not reported by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8e7f1c03eb |
feat(decisions): improve desk triage and queue parity (#10785)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The decisions desk and queue help operators find work that needs a human decision. > - The current views use different grouping, sorting, and labels. > - Repeated confirmation requests can also leave stale pending actions in the queue. > - Blocked-work attention can point at an intermediate issue instead of the terminal blocker. > - This pull request aligns the server contract and both user interfaces. > - The benefit is a smaller, clearer queue that ranks the decisions with the largest impact. ## Linked Issues or Issue Description Related PR: #10774 **What existing behavior does this improve?** The decisions desk and queue currently use different triage rules. They can show stale repeated confirmations and can rank blocked work by an intermediate issue. **Subsystem affected** This change affects attention aggregation, issue thread interactions, shared attention contracts, and the decisions user interface. **Current behavior** The desk uses a can-wait group that has no clear arrival meaning. The queue has fewer controls than the desk. Repeated pending confirmations remain actionable. Blocked-work rows do not always identify the terminal actionable blocker. **Proposed behavior** Group desk items by arrival date, and reserve Decide now for explicit due dates. Use one toolbar and shelf model on both pages. Supersede older repeated pending confirmations. Aggregate blocked work under the terminal actionable blocker and rank it by impact. **Reason and benefit** Operators get one consistent triage model. The badge reflects new and overdue work. High-impact blockers move to the top. Duplicate confirmation work no longer consumes attention. **Breaking changes** The attention summary field `decideNowCount` changes to `deskBadgeCount`. Consumers must use the new field. Older repeated confirmation interactions can now finish with the `superseded_by_newer_request` outcome. ## What Changed - Supersede older pending confirmation requests for the same issue and record the mutation in activity history. - Resolve blocked-work attention to actionable terminal blockers, suppress live blocker trees, and rank rows by blocked-work impact. - Group the decisions desk into New today and Earlier, and count new plus overdue work in the desk badge. - Share the decision toolbar and shelf components across the desk and queue. - Add queue grouping, sorting, filtering, aging, visible training controls, and clearer recommendation copy. - Add server, shared-contract, and user-interface tests for the new behavior. ## Verification - `pnpm check:token-gates` - `pnpm -r typecheck` - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm test:run` (all server, UI, CLI, shared, and catalog tests passed; one fixed five-second DB timeout flaked under full-suite load) - `NODE_ENV=test pnpm --filter @paperclipai/db exec vitest run src/status-card-migrations.test.ts` (passed in isolation) - `pnpm build` ## Risks - The attention summary field rename requires synchronized consumers. - Terminal-blocker traversal uses cycle and depth guards. A malformed dependency graph can stop at the last safe node. - The new arrival grouping changes which items contribute to the decisions badge. - Superseding repeated confirmations changes the terminal state of older pending interactions. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5. The runtime did not expose a dated model snapshot or context-window size. The model used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6a69a3d6b0 |
refactor(observability): stop writing detailed per-step timing to the run log (#10776)
## Thinking Path > - Paperclip tracks agent work for the team. > - The startup timing path already emits detailed spans. > - The run log repeats the same detailed timing. > - That makes the event larger than it needs to be. > - This pull request removes the redundant per-step timing fields from the run-log payload. > - The benefit is a smaller log with the same detail still available in traces. ## Linked Issues or Issue Description No public GitHub issue or open pull request matches this change. I described the enhancement below. **What existing behavior does this improve?** `run.startup.step` event output. **Subsystem affected** Cross-cutting (multiple of the above) **Current behavior** `run.startup.step` writes `roundTrips`, `providerExecMs`, `providerGetMs`, `createRuntimeMs`, and `ensureSessionMs`. The same detail already exists in the spans. **Proposed behavior** `run.startup.step` keeps only `step`, `durationMs`, and `outcome`. The heartbeat lifecycle timestamps stay unchanged. **Reason and benefit** This change removes redundant data from the run log. It keeps the useful detail in trace spans. It also makes the payload smaller and easier to read. The revert path stays clear because the removed data has one producer chain. **Breaking changes** Yes. Consumers that read the removed fields must switch to span data or the remaining event fields. The heartbeat lifecycle timestamps do not change. **Additional context** I searched GitHub for related open PRs and issues. I found no match. ## What Changed - Removed the redundant per-step timing fields from the run-log payload. - Removed the now-dead producer chain that fed those fields. - Kept the span-level timing data and the heartbeat lifecycle timestamps. ## Verification - `adapter-utils` typecheck clean - `server` typecheck clean - `startup-timing` suite: 30 passed - `adapter-utils` `execute` suite: 99 passed - `server` `environment-execution-target` suite: 21 passed ## Risks - Consumers that still read the removed fields will need a code change. - This is a one-way-door data removal for the run log. ## Model Used OpenAI Codex, GPT-5, tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
af6b32d82a |
feat(observability): add provider OTel spans - cache-hit flag, plugin tracer, Daytona pack/transfer spans (#10764)
## Thinking Path > - Paperclip uses pull requests to ship control-plane changes with review and audit. > - This change extends sandbox provider telemetry. > - The first PR added startup and exec spans. > - This PR adds provider spans for cache hits, plugin tracer flow, and Daytona file sync steps. > - The host must keep the trust boundary and reject unsafe span data. > - The result is fuller OTel coverage with bounded labels and safe parentage. ## Linked Issues or Issue Description **Problem or motivation** Sandbox provider telemetry lacks an explicit cache-hit signal, plugin span context, and file-sync pack and transfer spans. The current proxy for a cache hit is fragile. The plugin SDK also needs a safe tracer surface and a per-call parent context. **Proposed solution** Add an explicit `cacheHit` flag from the Daytona sandbox handle lookup. Expose `ctx.tracer` on the plugin context and pass a `traceparent` string with each call. Record provider spans on the host with allowlist clamping and capability checks. Wrap Daytona file sync pack and transfer steps in spans. **Alternatives considered** Keep the old `providerGetMs == 0` proxy. Reject that path because the cache-hit decision belongs at the handle lookup, not in a timing proxy. The host stays the trust boundary for worker span data. It clamps labels, rejects bad parent data, and gates the worker-to-host span RPC by capability. A security review completed before this PR opened. ## What Changed - Added an explicit `cacheHit` flag from the Daytona sandbox handle lookup. - Added `ctx.tracer` on the plugin context and a `traceparent` field in the per-call context. - Added host-side span recording with attribute clamping and capability checks. - Wrapped Daytona file sync pack and transfer steps in spans. - Added tests for the host trust boundary, the plugin tracer no-op path, and the Daytona span paths. ## Verification - `pnpm --filter @paperclipai/plugin-daytona exec vitest run` - `pnpm --filter @paperclipai/plugin-sdk exec vitest run` - `pnpm --filter @paperclipai/server exec vitest run plugin-host-services environment-execution-target instrumentation plugin-worker-manager` - `tsc --noEmit` for server, plugin-sdk, and adapter-utils ## Risks - Span data now crosses the worker boundary, so the host allowlist and capability gate must stay strict. - The `traceparent` path must stay valid and host-owned. - Daytona file sync spans must not add new round trips. ## Model Used OpenAI Codex, GPT-5, tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues with `Fixes:` / `Closes:` / `Refs:` or described the issue in the PR body - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ba396c608c |
feat(server): configure shared workspace concurrency (#10759)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat service starts agent runs in local, SSH, sandbox, plugin, or Kubernetes environments. > - Runs can share one project working tree. > - PR #10699 made every shared working tree single-file, including trusted local and SSH hosts that support coordinated concurrent work. > - Operators need a policy that keeps remote environments safe and restores local multi-agent work. > - This pull request adds an `auto`, `serialize`, or `allow` concurrency policy and applies it to the final execution environment. > - The benefit is safe serialization by default for sandboxed targets and useful concurrency by default for persistent local targets. ## Linked Issues or Issue Description Related work: #10699 introduced the busy gate that this change makes configurable. #7852 covers a separate environment-lease race and does not provide this dispatch policy. **Subsystem affected** `server/` heartbeat orchestration and `packages/shared/` workspace policy contracts. **Problem or motivation** The shared-workspace busy gate always defers a second run. This behavior prevents local multi-agent projects from running concurrently even when operators expect agents to coordinate through commits. **Proposed solution** Add `sharedWorkspaceConcurrency` with `auto`, `serialize`, and `allow` values. Default `auto` permits local and SSH concurrency. It serializes sandbox, plugin, and forced Kubernetes execution. **Alternatives considered** Keeping unconditional serialization is too restrictive for persistent host working trees. Always allowing overlap removes the protection that sandboxed and remote targets need. **Roadmap alignment** This is a focused correction to the shipped cloud and sandbox execution milestone. It does not add a new roadmap feature. ## What Changed - Added the optional tri-state field to project policy and issue override contracts and validators. - Added a pure resolver with issue override, project policy, and `auto` default precedence. - Moved final environment and Kubernetes resolution before the shared-workspace busy gate. - Kept the existing deferral and retry behavior for every path that resolves to serialization. - Added a task-context warning and a structured log when a run dispatches beside a live holder. - Added policy and heartbeat coverage for all requested policy and environment combinations. - Documented the new policy and its default behavior. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/execution-workspace-policy.test.ts src/__tests__/heartbeat-workspace-busy.test.ts` — 32 tests passed. - `pnpm -r typecheck` — passed. - `PAPERCLIP_IN_WORKTREE=false PAPERCLIP_DATABASE_RESTORE_IN_PROGRESS=false PAPERCLIP_RESTORE_IN_PROGRESS=false pnpm test:run` — all 3,528 server tests passed. The later UI/workspace phase passed 3,384 tests and hit one unrelated `CompanyEnvironments` navigation timing failure. - `pnpm --filter @paperclipai/ui exec vitest run src/pages/CompanyEnvironments.test.tsx -t "opens the edit form on a standalone page with existing values and closes after save"` — the unrelated UI test passed in isolation. - `pnpm build` — passed. Policy matrix: | Policy | Final target | Expected result | Result | | --- | --- | --- | --- | | `auto` | local | Dispatch with holder note | Passed | | `auto` | sandbox | Defer with `workspace_busy` | Passed | | `auto` | instance-forced Kubernetes | Defer with `workspace_busy` | Passed | | `serialize` | local | Defer with `workspace_busy` | Passed | | `allow` | sandbox | Dispatch with holder note | Passed | Existing serialization, retry, and stale-holder tests also pass. ## Risks - `auto` changes the post-#10699 local and SSH behavior back to concurrent dispatch. Concurrent agents can mutate the same working tree, so each dispatched run receives an explicit coordination warning. - `allow` is an operator override and can permit overlap in sandbox or plugin environments. - Unknown environment drivers serialize in `auto` mode. This keeps the fallback conservative. - There is no database migration. An absent field resolves to `auto`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 (Codex). The exact API snapshot and context-window size are not exposed to the agent. Reasoning, repository editing, terminal tool use, and local code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c09ea7112b |
feat(observability): add granular OpenTelemetry spans for sandbox startup and execution (#10758)
## Thinking Path > - Paperclip coordinates AI agent work and records execution data. > - The sandbox startup path and the host-to-sandbox execution path need clearer OpenTelemetry spans. > - The trace data now shows wait time, critical path data, and bounded labels. > - This pull request adds bounded spans and closed attribute helpers for sandbox startup and execution. > - The result is better trace data with low cardinality and no secret leakage. ## Linked Issues or Issue Description - No public GitHub issue exists for this branch. - Related PR: #10536 - This PR extends the sandbox OpenTelemetry work with exec spans, root timing, and bounded labels. ## What Changed - Added a closed span-attribute contract for sandbox startup data. - Added a bounded label helper for command and region values. - Threaded the active step context into the execution path. - Added a `sandbox.exec` span with true timestamps and bounded attributes. - Added skipped-step and root-span timing data with low-cardinality context. - Moved handshake sub-times and bridge batch data onto spans. ## Verification - `pnpm --filter @paperclipai/adapter-utils run typecheck` - `pnpm --filter @paperclipai/server run typecheck` - `pnpm --filter @paperclipai/adapter-utils exec vitest run` - `pnpm --filter @paperclipai/server exec vitest run environment-execution-target` - `git log --oneline origin/master..origin/feat/sandbox-startup-otel-spans` - `git diff origin/master...origin/feat/sandbox-startup-otel-spans --stat` ## Risks - Span attribute rules could still miss a future field if new code skips the shared helper. - Low-cardinality labels may hide some detail, but that is the intended tradeoff. - Tracing stays fail-open, so a missing tracer still hides data rather than stopping work. ## Model Used - OpenAI Codex, GPT-5, tool use enabled. The runtime did not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge The current head still has one failing required check: `e2e`. The PR stays unready for board handoff until that check passes. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
185515c97b |
fix(external-objects): refresh PR status labels (#10704)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The issue properties panel can show external objects such as GitHub pull requests. > - Those objects are resolved by external-object providers and then displayed as compact status labels. > - A GitHub pull request could remain in the fallback `unknown` state and appear as `Not yet resolved`. > - That label is confusing when the object is known but has not been refreshed yet. > - This pull request refreshes due external objects from the heartbeat scheduler and improves the unknown-status copy. > - The benefit is a properties panel that moves from pending refresh to the real pull request state without a manual refresh. ## Linked Issues or Issue Description No public GitHub issue exists for this bug. I searched for related public issues and pull requests using the terms `Not yet refreshed`, `external objects refresh`, and `external PR status`, and did not find a duplicate implementation. **What happened?** The issue properties panel could show a GitHub pull request as `Not yet resolved` even when the referenced pull request was valid. The object stayed stale unless a manual refresh path ran. **Expected behavior** A known external object should show pending-refresh copy while it waits for provider data. When the scheduler refreshes it, the properties panel should show the provider status such as open, merged, or closed. **Steps to reproduce** 1. Create or view an issue that references a GitHub pull request. 2. Open the issue properties panel. 3. Observe the external object row before a manual refresh has run. **Paperclip version or commit** Current `master` before this pull request. **Deployment mode** Local dev and self-hosted server. **Installation method** Built from source. **Agent adapter(s) involved** Not adapter-specific. **Database mode** Not database-related. **Access context** Board view. **Privacy checklist** I reviewed this description and did not include logs, credentials, private URLs, internal issue IDs, or PII. ## What Changed - Added a heartbeat scheduler tick that refreshes due external objects for active companies. - Kept manual external-object refresh behavior on the same service path. - Changed display copy so known provider objects use liveness labels such as `Not yet refreshed`, while fresh unknown provider statuses show `Status unavailable`. - Added server and UI tests for scheduled refresh and label behavior. ## Verification - `corepack pnpm install --frozen-lockfile` - `pnpm check:token-gates` - `pnpm exec vitest run server/src/__tests__/external-objects-service.test.ts server/src/__tests__/server-startup-feedback-export.test.ts ui/src/components/ExternalObjectPill.test.tsx ui/src/components/IssueProperties.test.tsx ui/src/lib/external-objects.test.ts` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/server build` - `pnpm --filter @paperclipai/ui build` - `pnpm run typecheck:build-gaps` - GitHub PR checks passed on head `0e7fcd30` - Greptile reported 5/5 on head `0e7fcd30` with no unresolved review threads Notes: - I ran recursive typecheck and build first. Both hit container resource limits with exit 137 during concurrent package work, so I reran the affected server and UI targets separately. - An unrelated workspace-runtime auto-port test fails in this container with a PID ownership mismatch. It is outside the files changed here. ## Risks Low to medium risk. The scheduler does more periodic external-object work, so the main risk is extra provider refresh load. The implementation bounds the work to active companies, due non-terminal objects, and 50 objects per company per tick. The path also stays behind the external-objects experimental setting. ## Model Used OpenAI GPT-5 Codex in the Codex execution environment, with shell and GitHub CLI tool use. The runtime did not expose a more specific internal model ID or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
717684ad8f |
Add project folder browsing to skill imports (#9930)
## Thinking Path > - Paperclip is the control plane teams use to manage AI agents and their reusable capabilities. > - Skills Manager lets operators discover and import skills from project workspaces. > - Automatic discovery only surfaces skills in conventional locations, so valid skills stored elsewhere in a project are invisible. > - Operators need a safe way to navigate project folders without exposing paths outside the selected workspace. > - This pull request adds company-scoped workspace folder browsing and selection to the project skill import flow. > - The benefit is that operators can find and import valid skill folders regardless of repository layout while preserving workspace boundaries. ## Linked Issues or Issue Description - **Subsystem affected:** Cross-cutting (`server/`, `ui/`, and `packages/shared`). - **Problem or motivation:** Project skill imports rely on conventional directory discovery, which prevents operators from selecting valid `SKILL.md` folders stored in atypical locations. - **Proposed solution:** Add a company-scoped browse endpoint and a folder browser in the import dialog. The server resolves real paths, rejects traversal outside the workspace, skips symlinks and high-noise directories, identifies skill directories/files, and caps listings at 250 entries. - **Alternatives considered:** Expanding the automatic scan to every directory would be slower and noisier, while accepting arbitrary filesystem paths would weaken project/workspace scoping. - **Roadmap alignment:** This extends the completed “Skills Manager, Skill Studio & Skills Store” capability in `ROADMAP.md` without duplicating planned core work. - **Additional context:** GitHub search found no duplicate or closely related public issues or pull requests. ## What Changed - Added shared browse request/result contracts and validation for project workspace navigation. - Added a company-scoped API route and service that safely lists local workspace folders and detects `SKILL.md` entries. - Added project workspace/folder navigation to the import dialog, including parent navigation, workspace switching, truncation feedback, and direct skill selection. - Added service and route regression tests for browsing, skill detection, company isolation, and traversal rejection. - Added shared response schemas and OpenAPI documentation for the browse endpoint. - Hardened explicit skill selections with realpath containment so symlinked directories cannot escape the project workspace. ## Verification - `pnpm exec vitest run server/src/__tests__/company-skills-service.test.ts server/src/__tests__/company-skills.test.ts` — 64 tests passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm exec vitest run server/src/__tests__/openapi-routes.test.ts` — 3 tests passed. - Focused post-review reruns: `company-skills-service.test.ts` — 45 tests passed; shared/server typechecks passed. - GitHub latest-head checks — all green after one transient e2e rerun; no pending or failing checks. - `pnpm check:token-gates` — all gates clean on the rebased head. ## Risks - Low-to-moderate risk: this adds a filesystem browsing surface. Realpath containment checks prevent workspace escape, symlinks are excluded, remote-managed workspaces are rejected, and directory listings are capped. - The browser intentionally hides `.git` and `node_modules`; skills inside those directories cannot be selected through this flow. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.5`, reasoning-enabled with terminal/tool use and code execution; runtime context-window size is not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details — the execution harness requires preserving the assigned branch name. - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes — no user-facing docs changes are needed beyond this PR description because the flow is self-explanatory UI behavior. - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0a09e4d975 |
feat(decisions): add desk workflow and retention (#10672)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Decisions desk shows work that needs a human decision > - The queue foundation can group and rank decision work > - Operators also need a focused daily view and a safe way to handle old work > - This pull request adds the desk controls, the aging shelf, and reversible retention > - It also binds bulk archive decisions to the exact reviewed item set > - The benefit is a smaller daily queue without lost or orphaned work ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: server, UI, database, and shared contracts. **Problem or motivation** Decision work can grow into one large company list. Operators need quick queue and date controls. Old items also need a safe retention path that does not delete work. **Proposed solution** Add queue and date controls to the Decisions desk. Compute the aging shelf on the server. Archive idle items after 90 days unless an operator keeps them. Keep archived items searchable and revivable. Notify origin agents in one batch per sweep. Bind bulk archive proposals to a signed, exact item manifest. **Alternatives considered** Client-only aging can drift across browsers and source kinds. Deleting old rows removes audit and recovery paths. An unsigned dynamic bulk query can archive items that the reviewer did not inspect. **Roadmap alignment** This work supports the Work Queues and decision-memory directions in `ROADMAP.md`. It extends the Decisions and attention-feed foundation from #10651. Related earlier work includes #9380, #10010, and #10474. ## What Changed - Added the queue rail, date chips, decide split, triage strip, queue page, and aging shelf UI. - Added server-owned shelf state with per-queue retention overrides. - Added reversible retention state, archive history, and an idempotent notification outbox. - Added the 90-day archive sweeper, Keep exemption, archived feed query, and revive actions. - Added one origin-agent notification per agent and sweep. - Added signed bulk archive proposals with exact-set and version checks. - Persisted queue-exclusion reasons atomically and kept cross-domain source resolution per-item until it has an exact-set transaction contract. - Added API contracts, OpenAPI entries, migration coverage, focused tests, and Storybook screens. ## Verification - `pnpm -r typecheck` - `pnpm test:run` (server: 329 files and 3,453 tests passed; UI: 408 files and 3,362 tests passed) - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts` - `pnpm build` - `pnpm check:token-gates` - Focused retention, attention, decisions, migration replay, startup, and UI API tests. - Complete queue snapshot regression with 51 items across the normal 50-item page boundary. The unmodified CLI test reports one warning assertion in this runtime because the harness injects static AWS credential variables. The isolated test passes when those two variables are removed. ## Risks - The migration adds retention and notification outbox tables. It uses idempotent table, index, and foreign-key creation. - Retention runs on the heartbeat scheduler interval. Compare-and-set version checks prevent stale archive writes. - Bulk archive acceptance fails closed when authority, activity, version, or the reviewed set changes. - Archive is reversible and does not delete source records. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with `gpt-5.6-sol`. The model used tool calls, code execution, database migration generation, and test execution. The context-window size is not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dcac49a4fd |
feat(workspaces): defer isolated setup until runtime start (#10653)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Isolated workspaces give each task a safe and reproducible checkout. > - The existing setup cloned the development database before an agent needed to run the app. > - This made worktree creation slower and heavier for tasks that never start a service. > - Runtime services already use one server start path for heartbeat, operator, and startup recovery flows. > - This pull request moves heavy setup to that start path and keeps worktree creation lean. > - The benefit is faster isolated workspace creation with the same reliable runtime setup when a service starts. ## Linked Issues or Issue Description Related pull request: #10652 covers the initial deferred database-seeding slice. This pull request supersedes it with end-to-end runtime provisioning and safe cleanup. **What existing behavior does this improve?** This improves isolated worktree creation, runtime service startup, and isolated instance cleanup. **Subsystem affected** Cross-cutting: CLI worktree setup, server runtime orchestration, shared workspace contracts, and development scripts. **Current behavior** Paperclip seeds an isolated development database during worktree creation. It can also leave an isolated instance directory after workspace teardown. This work happens even when no runtime service starts. **Proposed behavior** Paperclip creates the worktree with a lean eager setup. It runs an idempotent runtime provision command before the first managed service spawn. Concurrent starts share one provision attempt. Teardown removes the isolated instance safely. **Reason and benefit** Many agent tasks only edit and test code. They do not need a running Paperclip instance. Deferring the database seed reduces workspace startup cost while preserving automatic setup for tasks that start the app. **Breaking changes** None. The new runtime provision command is optional. Existing workspace behavior is unchanged when it is absent. ## What Changed - Split Paperclip worktree setup into a lean eager script and an idempotent runtime provision script. - Added `runtimeProvisionCommand` to project, issue, realized workspace, and persisted workspace contracts. - Added a per-workspace provision mutex before local service spawn for heartbeat, operator, and startup recovery flows. - Added a persisted `provisioning` service state and the `workspace_runtime_provision` operation phase. - Kept provision time outside the service readiness timeout and made failed attempts visible and retryable. - Reclaimed isolated instance data during safe workspace teardown. - Serialized deferred database seeding across processes and bound teardown to the instance root captured in persisted workspace metadata. - Added tests for config flow, concurrency, retry, no-op behavior, readiness timing, scripts, CLI commands, and cleanup. - Documented the eager and runtime provisioning contracts. ## Verification - `pnpm -r typecheck` - `pnpm build` - `pnpm test:run` (server: 3,201 passed; UI: 3,345 passed; the CLI phase exposed one environment-sensitive AWS doctor assertion because the agent runtime injects static AWS credentials) - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts -t 'passes AWS doctor checks when non-secret provider config is present'` - Focused runtime tests cover serialized provisioning, retry after stderr failure, absent-command no-op behavior, operation logging, persisted state order, and readiness timeout exclusion. - Focused CLI and cleanup tests cover concurrent seed serialization, stale-lock fail-closed behavior, persisted instance ownership, and rewritten sibling pointers. ## Risks - A faulty runtime provision script blocks service startup. Paperclip records stderr, marks the service failed, and retries on the next start. - Concurrent service requests share an in-process provision attempt, while the seed command uses an atomic filesystem lock across processes. A stale lock fails closed and requires an operator to verify no seed is running before removing it. - Isolated instance cleanup is destructive. The cleanup service validates ownership and path containment before removal. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, `gpt-5.6-sol`, with agentic reasoning, tool use, and code execution. The service does not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
30c49c8327 |
feat(decisions): add queues and prioritized attention feed (#10651)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source control plane for companies of AI agents. > - Operators use the attention feed to find decisions that need action. > - The feed has eleven source kinds, but it has no durable queue or triage state. > - The feed also returns every item and lacks decision deadlines, snooze state, and decision-focused ordering. > - This pull request adds secure queue sidecars and enriches the attention feed with triage data, filters, cursor pagination, and decide-now ranking. > - The benefit is a bounded feed that can show the most urgent decisions first without weakening source visibility rules. ## Linked Issues or Issue Description This pull request replaces the closed [#10634](https://github.com/paperclipai/paperclip/pull/10634). It combines that queue foundation with the dependent attention-feed change as one review unit. **Subsystem affected** Database schema, shared contracts, server authorization and REST APIs, and the UI attention client library. **Problem or motivation** The attention feed can contain hundreds of mixed decision items. Operators cannot group them into durable queues, set a decision deadline, snooze an item, or request a bounded page ordered by urgency. The current client must download the full feed on each refresh. **Proposed solution** Store queue membership and triage state by stable attention identity. Re-authorize each source during queue reads and writes. Enrich attention items with queue, deadline, snooze, expiry, rule, and origin data. Add activity and queue filters, opaque cursor pagination, decide-focused ordering, and a decide-now count. **Alternatives considered** Adding queue fields to every source would duplicate schema and authorization logic across eleven source kinds. Client-only filtering and sorting would still transfer the full feed and would make pagination unstable. **Roadmap alignment** This change improves the core decision-attention surface and operator oversight. It does not implement the separate general-purpose work queue milestone in `ROADMAP.md`. ## What Changed - Added company-scoped queue, membership, triage, and append-only event tables with actor and run provenance. - Added queue CRUD, item membership, starter-rule discovery, and decide-by and snooze endpoints. - Kept source authorization on each queue mutation, read, and count. - Added attention fields for expiry, rule, origin agent, queues, decide-by attribution, and snooze state. - Added activity date filters, queue filters, opaque cursor pagination, and configurable page limits. - Added decide-now ordering by deadline, expiry, severity, and activity. - Added `decideNowCount` and excluded actively snoozed items from the default feed. - Updated the shared and UI client contracts. - Added focused server, route, OpenAPI, and UI client tests. ## Verification - `pnpm exec vitest run server/src/__tests__/attention-service.test.ts server/src/__tests__/decision-queues-routes.test.ts server/src/__tests__/openapi-routes.test.ts ui/src/api/attention.test.ts ui/src/lib/attention.test.ts` (72 tests passed) - `pnpm --filter @paperclipai/db check:migrations` - `pnpm -r --filter @paperclipai/db --filter @paperclipai/shared --filter @paperclipai/server --filter @paperclipai/ui typecheck` - `git diff --check origin/master...HEAD` ## Risks - The migration adds four company-scoped tables and provenance foreign keys. Migration numbering and safety checks pass. - Attention reads can lazily create starter queues and memberships. Inserts are idempotent, audited, and transactional. - Cursor validity depends on the filtered feed. The API returns a clear validation error when the cursor item no longer exists in that feed. - Queue reads re-check source visibility. This favors correct authorization over fewer queries. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, model `gpt-5`. The runtime used agentic reasoning, repository tools, code execution, and test execution. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3dec88ce90 |
feat(agents): warn when an agent's escalation path routes to a paused manager (#10657)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents escalate work up the org chart (`reports_to`), and operators pause agents — notably, instance imports pause every agent by default > - A paused manager does not invalidate the chain (subordinates stay invokable), so nothing surfaces when an operator unpauses workers but leaves their manager paused > - Escalations then dead-letter silently: agent-created issues assigned to the paused manager sit in a queue nothing will ever run > - This pull request computes paused ancestors in the existing org-chain health model and surfaces a non-blocking warning on the agent read models and detail page > - The benefit is that the operator learns their escalation paths are dead before work vanishes into them ## Linked Issues or Issue Description Fixes #10647 (companion to #10648, which refuses agent-initiated assignment to paused agents at write time — this PR makes the standing hazard visible) ## What Changed - `AgentOrgChainHealth` gains two additive, optional fields: `pausedAncestors` (paused agents in the `reports_to` chain) and `escalationWarning` (human-readable, only set when the agent itself can work — a paused/terminated agent's escalation path is moot). Chain validity, invokability, and assignability are byte-identical. - No server route changes needed: the fields flow through every existing agent read model (list, detail, org chart) since they ride the same `getAgentWorkEligibility` computation. - Agent detail page shows an amber "Escalation path is paused" banner (same visual language as the invalid-chain banner, but non-blocking) with the warning text naming the paused manager and the two remedies. ## Verification - `pnpm vitest run packages/shared/src/agent-eligibility.test.ts` — 5 new cases: paused direct manager warns; paused grandparent through a healthy manager warns; the agent itself paused → no warning (but ancestors still reported); fully active chain → no warning, empty list; terminated ancestor keeps the invalid-chain classification without double-counting as paused. - Full `@paperclipai/shared` suite (392 tests) and `agent-eligibility-routes` (54) unchanged. - `tsc --noEmit` in shared, server, and ui. ## Risks - Low. Purely additive fields plus one UI banner; no behavior gates on the new data. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use. No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
27f8c8dbcf |
feat(server): cap agent review rounds and escalate exhausted reviews to the responsible human (#10650)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Execution policies let one agent implement and another review, cycling through changes-requested → addressed rounds > - Nothing bounds that cycle: no round counter, no escalation, no termination signal — two agents can ping-pong indefinitely, especially when the review's success criteria drift to something the implementer cannot satisfy > - On a real multi-agent instance this produced 6+ unattended rounds (~8 runs) that continued even after the human had merged the PR under review > - This pull request counts consecutive agent-initiated changes-requested rounds and, at a configurable cap, hands the still-pending review to the responsible human instead of bouncing back to the implementer > - The benefit is that unattended review loops terminate in a human decision instead of burning runs forever ## Linked Issues or Issue Description Fixes #10643 ## What Changed - `IssueExecutionState.changesRequestedCount` (schema + type, default 0): consecutive agent-initiated changes-requested rounds on the current stage. Carries through executor resubmissions, resets to 0 on approval, and resets when a **human** makes the changes-requested decision — the cap targets unattended agent↔agent ping-pong, never human review. - `IssueExecutionPolicy.maxReviewRounds` (optional, 1–50, default null → server default `DEFAULT_MAX_REVIEW_ROUNDS = 3`). - At the cap, the transition records the reviewer's changes-requested decision as usual but keeps the stage **pending** with the responsible human (`responsibleUserId`, falling back to `createdByUserId`) as the participant: the issue is assigned to that human and the pending review surfaces through the existing attention/review UI. The human then approves, requests changes (resetting the counter and handing back to the implementer), or re-scopes. - The escalated hold is sticky: transitions from anyone other than the escalated human no longer re-select a configured agent participant for the stage (which would have silently undone the escalation on the next unrelated PATCH). The escalated human's own decisions flow through the normal participant decision branch. - Issues with no responsible human keep today's hand-back behavior; the counter still accumulates so operators can see the churn. ## Verification - `pnpm vitest run server/src/__tests__/issue-execution-policy.test.ts` — 8 new cases: round counting on hand-back, count carried through resubmission, escalation at the default cap, sticky hold across unrelated transitions, human changes-requested resets the counter, human approval completes the stage, no-responsible-human fallback, and a `maxReviewRounds: 1` policy override. - `pnpm vitest run server/src/__tests__/issue-execution-policy-routes.test.ts` and the full `@paperclipai/shared` suite (387 tests) — schema additions are backward compatible (both fields optional with defaults; persisted states without the counter parse as 0). - `pnpm --filter @paperclipai/shared exec tsc --noEmit` and `cd server && pnpm run typecheck`. ## Risks - Behavior change: an agent-only review loop that previously ran forever now escalates to a human after 3 agent rounds by default. Instances that want longer loops can set `maxReviewRounds` per policy. Flows where a human participates are unaffected (human decisions reset the counter). - Escalation requires a `responsibleUserId`/`createdByUserId` on the issue; without one, behavior is unchanged. - Persisted execution states from before this change parse with `changesRequestedCount: 0` — no migration needed. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use. No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
c185e64b77 |
feat(ui): chat-style task view behind an experimental flag (#10606)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators spend most of their time on the issue detail page. They talk to the assigned agent there through comments. > - The current page reads as a ticket form. The thread sits below properties, the composer sits mid-page, and live agent activity renders as dense transcript logs. > - Talking to an agent is a conversation. A chat-first layout matches that mental model better than a ticket form. > - A layout change this large must not disrupt current users. It needs a safe opt-in path and full parity with the existing thread features. > - This pull request adds a chat-style task view behind a new "Chat-Style Tasks" experiment toggle. The flag is off by default and the existing page is unchanged when it is off. > - The benefit is a focused, readable conversation with the agent: live tool activity folds into compact summaries, the composer stays at the bottom, and properties, plan, and artifacts move into header tabs. ## Linked Issues or Issue Description Refs #49 (chat with agents is a much-wanted feature). Related PRs found in the dedup search: - #4489 — an earlier, closed attempt to promote the conversation to the primary surface on issue detail. This PR is a fresh, flag-gated take on the same goal. - #8837 — an open PR that proposes a two-column task layout. It restructures the same page but keeps the ticket paradigm; this PR is orthogonal because it is opt-in and chat-first. **Subsystem affected** UI (issue detail page). **Problem or motivation** The issue detail page presents agent conversations as a ticket: properties first, thread below, composer in the middle of the page, and raw transcript noise during live runs. Users who mainly converse with their agents must scroll past chrome to follow the conversation, and live activity is hard to read. **Proposed solution** An opt-in chat-style view of the issue detail page, gated by a new "Chat-Style Tasks" experiment toggle in Settings → Experimental. With the flag on, the thread fills the center pane, the composer docks to the bottom of the viewport, Properties / Plan / Artifacts become header tabs, live turns show a status pill with the current tool action and elapsed time, and settled turns collapse to a "Worked · N tools" summary that expands into per-tool rows. With the flag off, nothing changes. **Alternatives considered** Restyling the existing layout in place (rejected: too disruptive without an opt-out), and a separate chat page beside the issue page (rejected: splits the task's single source of truth). A per-request lab page (`/task-chat-lab`, dev-only) was kept for design iteration instead. **Roadmap alignment** ROADMAP.md "CEO Chat" wants lighter conversations that still resolve to real work objects. This PR keeps the core task-and-comments model — it only changes presentation, opt-in — so it does not duplicate that planned work. ## What Changed - New `enableTaskChatRedesign` instance setting, exposed as a "Chat-Style Tasks" experiment card in Settings → Experimental (shared feature catalog, validators, server instance-settings service, and UI settings page). - New `ui/src/components/task-chat/` component family: chat thread with turn grouping, agent reply bubbles, live status pill, collapsible turn summaries with per-tool rows, plan tab with a sticky CTA action bar, inline interaction cards, per-request mode chips, and a bottom-docked composer. - A shared tool taxonomy (`tool-taxonomy.ts`) maps tool names to verbs and icons; the status pill, tool rows, and the classic transcript view all use it. - A transcript adapter converts stored run logs into chat turns; it dedupes tool-call updates by `toolUseId` so tool counts match the expanded rows, and it keeps a tool row's first real name when later generic updates arrive. - Composer: posts on Cmd/Ctrl+Enter, supports image paste with object-URL thumbnail previews (revoked on clear/unmount), and uploads through the issue attachments route. - `IssueDetail.tsx`: with the flag on, pane tabs move to the header bar, the header is not sticky, and the chat fills the center; with the flag off, the previous layout renders unchanged. - Motion tokens for the new animations live in `ui/src/index.css` with a `motion-tokens.ts` catalog and a test that keeps the two in sync (the catalog now also covers the shared enter/exit/swap tokens that the decision/quicklook block declares). - A dev-only `/task-chat-lab` page with fixtures and a tweak panel for motion tuning. ## Verification - `pnpm typecheck` — clean across the workspace. - `pnpm check:token-gates` — 3/3 CLEAN. - `cd ui && pnpm vitest run` — 3,344 of 3,345 tests pass locally. The one failure is `IssueProperties.test.tsx` monitor-row time formatting, which is timezone-sensitive: it also fails on unmodified `origin/master` in a non-UTC timezone and passes with `TZ=UTC`. It is not related to this change. - `cd server && pnpm vitest run src/__tests__/instance-settings-service.test.ts` — 21/21 pass (covers the new setting). - Manual: start the dev server, open Settings → Experimental, enable "Chat-Style Tasks", and open any issue. The thread fills the page, the composer docks to the bottom, and Properties / Plan / Artifacts appear as header tabs. Assign an agent and comment to watch a live run: the status pill shows the current tool action with elapsed time, and the finished turn folds into a "Worked · N tools" summary. Disable the toggle and confirm the classic page is unchanged. - Visual snapshot baselines are intentionally not updated: per `doc/design/DECISION-SHEET.md`, "Per-change snapshot verification demoted to dormant (Jul 13 2026)". ## Risks - The flag-off path goes through the same `IssueDetail.tsx` file, so a regression there would affect current users. Mitigation: the classic markup renders through the same components as before behind explicit flag conditionals, and the full UI suite passes. - The transcript adapter interprets stored run-log formats, including legacy entries without `toolUseId`. Malformed logs degrade to generic tool rows rather than crashing. - The new view changes no server behavior other than one additive instance setting; it is additive and default-off. Overall risk with the flag off is low. ## Model Used - Claude (Anthropic), model id `claude-fable-5` (Claude Fable 5), extended thinking enabled, agentic tool use (file editing, shell, test execution) via Claude Code / Claude Agent SDK. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9c1f8e7887 |
feat(decisions): add first-class propose mode (#10010)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents can currently perform many mutations directly, while humans often need a durable review point before cross-issue or destructive actions occur > - Existing approvals and issue-thread interactions do not provide a standalone, reusable object for presenting options, collecting typed inputs, detecting stale targets, and auditing effect execution > - The control plane therefore needs a first-class propose mode that separates an agent's recommendation from the governed mutation it may cause > - This pull request adds Decisions v1 across the database, shared contracts, server execution and telemetry, agent skill guidance, and operator UI > - The benefit is that agents can propose multi-option actions safely while operators get explicit provenance, fail-closed execution, per-effect results, and a focused attention workflow ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting: `packages/db`, `packages/shared`, `server`, and `ui`. ### Problem or motivation Agents need a governed way to propose consequential work without immediately mutating issues, especially when one choice can affect several issue trees. Existing approvals and issue-thread interactions do not provide a standalone object with typed options, target snapshots, effect-level authorization, expiration, execution outcomes, and reusable attention-feed presentation. ### Proposed solution Add first-class Decisions that store options and typed inputs, surface open proposals in the operator attention feed, validate target freshness and the origin-agent/operator authorization intersection at decision time, execute a bounded set of auditable effects, and retain terminal outcomes. Decisions v1 supports comments, status and assignee changes, follow-up issue creation, blocker resolution, and issue-tree cancellation, plus bundle grouping, expiration/dismissal, rule-key telemetry, and agent-facing API guidance. ### Alternatives considered - Extend approvals with arbitrary effects: rejected because approvals represent governed yes/no actions and would become an unsafe generic mutation envelope. - Model every proposal as an issue-thread interaction: rejected because decisions can span several targets and need independent lifecycle, telemetry, idempotency, and effect results. - Let agents perform the mutation and ask for retrospective review: rejected because it removes the pre-execution governance boundary this feature is meant to provide. ### Roadmap alignment Aligns with `ROADMAP.md` sections **Agent Reviews and Approvals**, **Enforced Outcomes**, **MCP Tool Gateway & Apps (governed tool access)**, and **Activity History** by making explicit decisions, authorization gates, auditable execution, and terminal outcomes first-class control-plane objects. ### Additional context This does not replace existing approvals or issue-thread interactions, and it does not add an unrestricted generic mutation effect. ## What Changed - Added company-scoped decision, option, target, and effect-execution schema plus migration and shared TypeScript/Zod contracts. - Added decision routes and services for propose, list/get, decide, dismiss, cancel, target freshness checks, authorization intersection, idempotency, activity logging, and execution auditing. - Added rule-key decision telemetry and attention-feed metadata so open decisions are visible and measurable. - Added agent skill documentation for proposing and resolving decisions through the Paperclip API. - Added the Decisions UI: API client, query keys, inline attention resolver, bundle grouping, target-issue strip, terminal history, destructive confirmation, and per-effect result rendering. - Added server service coverage, DecisionCard state tests, and Storybook stories for the supported visual states. ## Verification - `pnpm -r typecheck` — passed. - `pnpm test:run` — 2,876 passed, 1 skipped, with one unrelated cross-suite cleanup-order failure in `heartbeat-responsible-user-invariant.test.ts`; the failing file passes in isolation (`6/6`). - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-responsible-user-invariant.test.ts` — passed. - `pnpm --filter @paperclipai/ui exec vitest run src/components/DecisionCard.test.tsx` — passed (`9/9`). - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/authz-existence-oracle-guard.test.ts src/__tests__/openapi-routes.test.ts` — passed (`5/5`). - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/decisions-service.test.ts` — passed (`16/16`). - `pnpm --filter paperclipai exec vitest run src/__tests__/company-import-export-e2e.test.ts` — passed (`1/1`). - `pnpm --filter @paperclipai/server typecheck` and `pnpm --filter paperclipai typecheck` — passed. - `pnpm build` — passed. - Rebased-head focused suite — passed (`6` files, `88` tests): shared decision contracts, Decisions service, OpenAPI routes, startup feedback export, DecisionCard states, and attention helpers. The follow-up stale-secondary-target regression passes in the DecisionCard suite (`10/10`). - Rebased-head scoped typechecks — passed for `@paperclipai/shared`, `@paperclipai/db`, `@paperclipai/server`, and `@paperclipai/ui`. - Rebased-head migration numbering and safety checks — passed after renumbering the additive migration to `0193` and making it replay-safe for environments that applied the earlier feature-branch number. - `pnpm check:token-gates` — passed with all gates clean. - GitHub PR workflow and Greptile review for `1f9f7645882d05dfdd9c99377c03a1f53f20e8be` — running after the stale-secondary-target fix and PR metadata refresh on July 27, 2026. - `pnpm --filter @paperclipai/ui build-storybook` exposes an existing Storybook version mismatch (`storybook` 10.4.6 vs `@storybook/addon-docs` 10.5.0); Decisions stories were validated with the docs addon temporarily disabled and the tracked config remains unchanged. ## Risks - **Migration:** Adds replay-safe migration `0193`; migration numbering and safety checks pass. The new tables and indexes are additive. - **Authorization:** Effect execution intersects the proposing agent's permissions with the responsible user context and fails closed; mistakes could reject a valid proposal rather than silently over-authorize it. - **Concurrency:** Target snapshots and idempotency keys protect against stale or duplicate execution, but reviewers should focus on mixed-effect partial outcomes and retry behavior. - **UI:** Decisions are integrated into the existing attention feed rather than a separate navigation surface, reducing routing risk but increasing the importance of attention-item metadata compatibility. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI using `gpt-5.6-sol` for final PR preparation, review fixes, and verification; repository tools and code execution were enabled, and context-window size is not exposed in this runtime. - Anthropic Claude Opus 4.8 with 1M context assisted with the Decisions UI implementation, as recorded in the relevant commits. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
71231dfa38 |
feat(audit): agent audit UI — company page + per-agent tab (#9744)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Operators need an audit record of agent actions across tasks, comments, documents, approvals, and runs > - The permission-gated audit read API provides that record, but operators cannot inspect it in the product > - A readable UI must preserve company boundaries, server-side permission decisions, and redaction > - Audit exports must also be safe to open in spreadsheet software and must record the export itself > - This pull request adds company and per-agent audit views plus a guarded CSV export > - The benefit is a searchable, filterable, and reviewable agent action history with direct links back to work ## Linked Issues or Issue Description **Feature.** This change adds the frontend and CSV export for the agent action audit log. Refs #9731 and #9735. - Problem: agent actions are recorded, but operators have no readable product surface to inspect or export them. - Solution: add a company audit page and a per-agent Audit tab that use the permission-gated audit API. - Alternative: build a separate plugin-only surface. This was rejected because the existing permission model already supports a unified, server-authoritative view. This pull request targets the audit epic branch, which contains the merged #9735 audit API. ## What Changed - Added a company Audit page and sidebar entry. - Added a per-agent Audit tab with a fixed agent filter. - Added filters for agent, responsible user, action domain, entity type, and date range. - Added task and run links, responsible-user context, cursor pagination, and readable action text. - Added a permission-denied Enterprise card for callers without `audit:view_agent_actions`. - Added a CSV export that is permission-gated, capped, self-audited, CSV-escaped, and protected against spreadsheet formula injection. - Preserved the merged audit API cursor validation, redaction, and sub-millisecond pagination behavior. ## Verification - `pnpm exec vitest run ui/src/pages/audit/AuditFeed.test.tsx` — 6 passed. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/agent-action-audit-routes.test.ts` — 8 passed with embedded PostgreSQL. - `pnpm -r typecheck` — passed across all workspaces. - `pnpm build` — passed across all workspaces. - `pnpm test:run` — all completed shards passed except one environment-sensitive CLI assertion caused by injected static AWS credential variables; the exact test passes 8/8 with those variables unset. - Manual Chromium QA exercised the populated feed, active filters, permission-denied card, per-agent tab, and CSV export. ## Screenshots and Manual QA - [All audit states exercised in Chromium](https://github.com/paperclipai/paperclip/pull/9744#issuecomment-4998997001) - [Detailed browser report and per-agent tab root cause](https://github.com/paperclipai/paperclip/pull/9744#issuecomment-4998771061) The per-agent redirect defect found during QA is fixed in this branch. ## Risks Low to moderate risk. The UI and export route are additive and use the existing company-scoped permission gate. The main risks are large exports and spreadsheet interpretation. The export is capped at 10,000 rows, records truncation accurately, and prefixes formula-like cells as text. There are no schema changes or migrations. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Anthropic Claude Opus 4.8, 1M context, extended thinking, tool use, and code execution produced the original implementation. - OpenAI Codex, GPT-5 (deployment ID and context window not exposed), reasoning, tool use, code execution, browser-test orchestration, and GitHub review tooling repaired and verified the pull request. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
627728bdde |
feat: add authoritative issue PATCH receipts (#10478)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents update tasks through the issue API. > - The update response did not state which values changed. > - Blocker updates also did not echo the scalar blocker IDs. > - Agents therefore used an extra GET request to confirm a successful write. > - This pull request adds an authoritative change receipt and an optional small response. > - The benefit is fewer API calls with a clear and compatible write contract. ## Linked Issues or Issue Description No public GitHub issue exists for this change. ### Subsystem affected Cross-cutting: `server/`, `packages/shared`, and the UI issue cache. ### Problem or motivation A successful issue PATCH returned the updated issue, but it did not identify the effective changes. Blocker writes returned relation summaries without the scalar IDs. Agents could not distinguish a confirmed clear operation from missing data. The response must confirm committed field and blocker changes while existing UI clients continue to receive the full issue by default. ### Proposed solution Add a `changes` receipt. Add a conditional `blockedByIssueIds` echo. Support `Prefer: return=minimal`. Keep the full response as the default. ### Alternatives considered Make the small response the default for agent tokens. This would create different response contracts by actor type, so this pull request does not use that design. ### Roadmap alignment This is a focused control-plane reliability improvement. It does not duplicate an open roadmap milestone. ## What Changed - Compute committed issue row and relation changes in the issue service. - Omit no-op fields and truncate changed long text values to 200 characters. - Echo blocker ID arrays for blocker set and clear requests. - Add the opt-in `Prefer: return=minimal` response and `Preference-Applied` header. - Keep receipt metadata out of React Query issue caches. - Add route and embedded Postgres tests for the new contract. ## Verification - `pnpm exec vitest run server/src/__tests__/issue-activity-events-routes.test.ts` - `pnpm exec vitest run server/src/__tests__/issues-service.test.ts -t "returns authoritative update receipts for row fields and blocker relations"` - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - `git diff --check` ## Risks - Low compatibility risk. The default response only adds receipt fields. - Minimal mode is opt-in. Existing clients do not receive a smaller body. - The receipt excludes `updatedAt` because the response already returns it as the freshness anchor. - Prose API and agent workflow guidance will follow after the server contract is available. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex based on GPT-5. The exact deployment ID, context window size, and reasoning mode are not exposed to the agent. The agent used repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fc5a30805e |
feat(cli): add managed install, update, and service lifecycle (#10045)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies > - Operators need a predictable installation path that survives beyond an ephemeral `npx` process > - A durable installation needs an owned per-user payload store, stable command shim, safe shell integration, and supported service lifecycle > - Updates must preserve recoverability by backing up data, installing side-by-side, verifying the new payload, and retaining rollback state > - Bootstrap scripts and privileged service operations must fail closed across download, filesystem, ownership, and consent boundaries > - This pull request integrates managed install, update, rollback, service, uninstall, doctor, bootstrap-installer, and runtime-serving support into one workflow > - The benefit is a recoverable, inspectable, and documented installation lifecycle with explicit safety boundaries across Linux, macOS, containers, WSL, npm, npx, and source checkouts ## Linked Issues or Issue Description ### Problem Paperclip lacks a first-class durable installation and lifecycle workflow. Operators currently have to assemble npm/npx installation, PATH setup, background-service management, updates, rollback, diagnostics, and uninstall behavior themselves. That makes upgrades harder to recover, creates inconsistent behavior across platforms, and leaves shell/download/service trust boundaries without one documented implementation. ### Proposed Solution Add a managed per-user install store and stable shim, a verified shell bootstrap installer, service lifecycle commands, install-mode-aware update/rollback behavior, doctor checks, and documentation. Managed updates back up the database, install and smoke-test a side-by-side payload, atomically switch `current`, and retain prior payloads. The shell installer pins registry/download trust boundaries and requires explicit consent for non-interactive privileged actions. ### Alternatives Considered - Keep recommending `npx`: simple for evaluation, but ephemeral and unsuitable for stable services, atomic updates, or rollback. - Require global npm installation only: familiar, but cannot provide the owned side-by-side payload store and retained rollback semantics. - Split the capability across multiple PRs: rejected because install, update, service, uninstall, bootstrap, and serving behavior share contracts and security boundaries that need review together. ### Related Pull Requests - Supersedes #10042 and #10044 with one integrated final diff. - Incorporates and replaces the closed preparatory work in #10032 and #10034. ## What Changed - Added `paperclipai install`, `update`/`upgrade`, rollback, uninstall, service lifecycle, onboarding integration, and managed-install doctor checks. - Added a private managed payload store, verified manifest/marker ownership, exclusive mutation locks, atomic manifest/current/shim writes, retained previous payloads, and provenance validation. - Added npm and GitHub-ref install sources with exact target resolution, registry isolation, database backup, side-by-side verification, atomic activation, service restart coordination, and failure rollback. - Made managed-update backups report actionable service-start and `--no-backup` recovery guidance for unreachable databases, while clean never-onboarded instances skip an empty backup. - Added systemd user and launchd service definitions, status/health/log commands, single-instance coordination, stale-port recovery, and explicit sudo/lingering consent handling. - Added the `scripts/install.sh` bootstrap path with checked two-stage downloads, pinned public npm registry usage, platform checks, dry-run/non-interactive controls, and Docker fixtures. - Added embedded Postgres/native bootstrap integration, hot-restart/systemd-notify serving support, passive update notices, configuration contracts, README/CLI/install documentation, and focused regression tests. - Security re-review should explicitly re-verify: (1) `addManagedPathBlock`/`removeManagedPathBlock` reject symlinked or non-regular rc files, assert current-user ownership, preserve restrictive modes, and replace atomically; (2) managed shim replacement rejects unsafe parents, foreign-owned or multiply linked files, and uses checked atomic replacement; (3) the shell installer and sudo path preserve explicit consent and checked downloads; and (4) installed service/runtime serving remains bound to the validated managed shim and instance configuration. ## Verification - `bash -n scripts/install.sh scripts/clean-install-git.sh scripts/clean-install-npm.sh scripts/test-install-sh-docker.sh` - `pnpm exec vitest run cli/src/__tests__/install-store.test.ts cli/src/__tests__/install-command.test.ts cli/src/__tests__/managed-install-check.test.ts cli/src/__tests__/onboard-service.test.ts cli/src/__tests__/service-health-check.test.ts cli/src/__tests__/service-manager.test.ts cli/src/__tests__/update-command.test.ts cli/src/__tests__/update-notice.test.ts packages/db/src/embedded-postgres-native.test.ts` — 9 files, 66 tests passed - `pnpm --dir cli typecheck` - `pnpm --dir cli build` - Follow-up verification: `pnpm exec vitest run cli/src/__tests__/update-command.test.ts` (14/14), `pnpm --dir cli typecheck`, `pnpm --dir cli build`, and `pnpm --filter @paperclipai/server typecheck`. - `pnpm -r typecheck` - `pnpm build` - Full `pnpm test:run` exercised all suites; an injected static AWS credential changed one unrelated doctor expectation, which passed when those credentials were removed. A second run cleared that case and exposed stale pre-existing adapter-utils `dist` output; rebuilding `@paperclipai/adapter-utils` made the isolated test pass. The updated PR CI is the authoritative clean-workspace full-suite run. ## Risks - Installer/update code writes executable shims, symlinks, shell rc blocks, service definitions, and managed payloads; ownership, regular-file, symlink, hard-link, marker, and path-containment checks fail closed before destructive changes. - The bootstrap installer executes downloaded tooling; downloads are staged and checked before execution, npm traffic is pinned to the public registry, and non-interactive privileged behavior requires explicit consent. - Linux lingering may invoke `sudo`; the command is surfaced and confirmed before execution, and unsupported service managers fall back to foreground-run guidance. - Database migrations remain forward-only; payload rollback does not reverse migrations, so managed updates create a backup before activation unless explicitly disabled. - Service restart and runtime serving touch process/port ownership; lifecycle locks, health/version checks, and stable-shim service definitions reduce split-brain and stale-process risk. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex coding agents using GPT-5.5 and GPT-5.6-sol, with reasoning, repository/API access, shell execution, and test tooling. The runtime did not expose a reliable context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b4a7a12985 |
feat: make recovery updates quieter (#10542)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The recovery subsystem restores work after an agent run stops or loses state > - Recovery notices currently use the same visual weight as normal work comments > - Recovery agents can also post long narratives that obscure the useful hand-off > - The server must identify recovery output because agents cannot set presentation controls > - This pull request adds compact recovery notices, structured action references, and brief recovery prompts > - The benefit is a quieter issue thread that still keeps recovery state inspectable ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: `server/`, `packages/shared`, and `packages/adapter-utils`. **Problem or motivation** Recovery notices and recovery-run comments can dominate an issue thread. Operators must scan routine recovery narration before they find the work hand-off. **Proposed solution** Give routine recovery output a compact system-notice presentation. Derive the presentation on the server so agents cannot hide arbitrary comments. Keep the successful missing-state summary fully visible because that comment is the recovery deliverable. **Alternatives considered** The UI could detect recovery text. That approach is fragile and does not provide structured action references. Agents could also set presentation directly, but that would weaken the current board-only security boundary. **Roadmap alignment** This change refines the completed “Self-healing runs & automatic recovery” and “Enforced Outcomes” roadmap areas. It does not add a competing roadmap capability. **Additional context** The scope covers shared comment validation, server recovery notices, agent-comment derivation, and recovery prompt text. No database migration is needed because presentation data already uses JSON. ## What Changed - Add the `compact` issue-comment presentation density to shared constants, types, and validation. - Give recovery escalation, waiting, and in-place notices compact titles and structured recovery-action metadata. - Use recovery-action metadata for notice deduplication, with the legacy text marker as a compatibility fallback. - Derive compact presentation for comments from recovery-scoped runs while preserving the board-only presentation boundary. - Keep successful missing-state recovery summaries fully visible. - Ask recovery participants to record outcomes in `resolutionNote` and keep source-issue comments brief. - Add shared, route, service, and prompt tests for the new behavior and exceptions. ## Verification - `pnpm -r typecheck` - Focused Vitest coverage: 320 tests passed across shared validators, adapter prompts, issue comments, recovery actions, and heartbeat recovery. - Full server phase: 292 files passed, 3,094 tests passed, and 2 tests skipped. - Full UI phase: 386 files passed and 3,182 tests passed. - `pnpm build` - Known master baseline: `cli/src/__tests__/secrets.test.ts` expects `pass`, but the current implementation returns `warn` when strict secret mode is disabled for Postgres. This branch does not change CLI secrets code. ## Risks - Low migration risk. The presentation column is JSON and needs no database migration. - Recovery-run detection depends on the persisted run context snapshot. - Structured metadata becomes the primary deduplication key. The existing body marker remains as a fallback for older comments. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5.6-sol`. The runtime did not expose the context-window size. The model used agentic reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9f7565f4ce |
feat: activate OpenTelemetry spans on the sandbox start path (#10536)
## Thinking Path > - Paperclip moves agent work through sandboxed execution and control-plane services. > - The sandbox start path now has a no-op span seam. > - This change turns that seam on when OTLP export is configured. > - It keeps the default path unchanged when export is off. > - The result is structured startup traces with low-cardinality attributes and explicit parent links. > - The benefit is better observability without changing normal behavior. ## Linked Issues or Issue Description No public GitHub issue exists for this change. ### Problem The sandbox start path has a tracer seam, but it stays a no-op unless the OTLP export path is active. ### Proposed solution Enable the server tracer on sandbox bring-up, open a root span, parent each startup boundary to that root, and keep the export path opt-in behind `OTEL_EXPORTER_OTLP_ENDPOINT`. ### Alternatives considered - Keep the start path as a no-op. I rejected that path because it leaves sandbox start opaque when OTLP export is already configured. - Add broad attributes for commands and paths. I rejected that path because the span allowlist must stay low-cardinality. ### Roadmap alignment This follows the current OTel sandbox-start work and keeps the default path unchanged. ## What Changed - Add a root sandbox startup span and child spans for each named startup boundary. - Keep concurrent bridge spans parented to the root span. - Inject the server tracer through the adapter deps without OpenTelemetry imports in the engine. - Attach host-received provider duration attributes only when the values are finite. - Keep span attributes inside the allowlist and keep command, path, id, and error text out of span data. ## Verification - The pushed branch already passed `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit`. - The pushed branch already passed `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/`. - The pushed branch already passed `pnpm --filter @paperclipai/server exec vitest run src/__tests__/environment-execution-target.test.ts src/__tests__/instrumentation.test.ts`. - The pushed branch already passed `pnpm --filter @paperclipai/server exec tsc --noEmit`. ## Risks - OTel export changes trace volume when the endpoint is set. - The allowlist limits trace detail, so new fields need care. - The change stays no-op when OTLP export is off. ## Model Used - OpenAI GPT-5, tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5ec7ce76e5 |
Upload company import packages as compressed zip uploads (fix large-company imports through Cloud) (#10531)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company Import (#10507, hardened in #10523) lets a user upload a company package on the Import page > - The page expanded the user's `.zip` into a files map and POSTed it as ONE inline JSON body — ~40MB for a real company because attachment blobs get base64-inflated > - On Paperclip Cloud that body travels browser → harness proxy → tenant, where it truncated in transit → body-parser 400 → the browser saw "Failed to fetch", and nothing imported > - Two compounding causes: the giant inline body itself, and the board async opt-in riding an `x-paperclip-cloud-*` header that the Cloud harness strips as anti-spoofing (so async never engaged and the import held one fragile synchronous connection) > - This pull request uploads the raw compressed `.zip` as a multipart request (about a third the size, already compressed) parsed server-side into the same bundle the importer consumes, and moves the async opt-in to a proxy-safe `?async=1` > - The benefit is that a large-company import actually completes through Cloud: a small compressed upload, a real async job that survives dropped connections ## Linked Issues or Issue Description - Refs #10507 / #10523 (Import/Export and its hardening). No open issue; problem described above (large-company browser import through a proxy: inline JSON body truncates → 400 → "Failed to fetch"; async opt-in header stripped by the front door → async never engages). ## What Changed - **Multipart zip transport.** The Import page uploads the raw `File` as `multipart/form-data` (field `package`, import options in a JSON `meta` field); the server unzips it into `{ rootPath, files }` and runs the exact existing preview/import logic. The `application/json` inline path is byte-identical for CLI/programmatic callers. Bare `application/zip` (meta via `?meta=`) is also accepted for programmatic use. - **Shared node zip reader.** `packages/shared/src/portability-zip.ts` (node-only subpath, not re-exported to the browser bundle — same pattern as `portability-hash.ts`); the CLI's `zip.ts` becomes a thin re-export. Identical codec (STORE + DEFLATE via `inflateRawSync`, rejects data descriptors/zip64). - **Proxy-safe async signal.** `wantsAsyncImport` = `?async=1` (board browsers, survives the harness) OR the existing `x-paperclip-cloud-async-import` header (cloud tenants, set server-side). The UI async client now uses `?async=1`. Backward compatible. - **Size + preflight.** New `PORTABLE_ZIP_UPLOAD_LIMIT_BYTES = 128MB`; the inline 56MB preflight no longer gates the zip path (it shows the compressed size instead). Async submit/poll/resume, the duplicate-guard fingerprint (now over the resolved bundle), pause-on-import, progress/error panels, and activation all apply to the multipart path. - OpenAPI documents json + multipart + zip bodies and the `async` query param. ## Verification - Full typecheck chain (shared, server, ui, cli) clean. - 152 tests across 8 files: new `portability-zip.test.ts` (STORE/DEFLATE/base64-blob byte-exact round-trip, truncation throws, data-descriptor rejection); `company-portability-routes.test.ts` +7 (multipart import+preview equals the inline bundle; async multipart 202→poll→success; board async via `?async=1` with no cloud header; cloud-tenant async via header; sync fallback with neither; truncated-zip 400, nothing imported); `CompanyImport.test.tsx` asserts the local zip sends the raw File and the inline preflight no longer blocks; `openapi-routes.test.ts` green. - NOT yet measured: the end-to-end browser upload through the live Cloud harness — verified on staging after deploy before closing out. ## Risks - Import semantics unchanged — only transport changed; the JSON inline path is byte-identical, the cloud-tenant header async path untouched. Multipart parsing is server-side (memory-bound: a ~13MB zip → ~30MB files map, fine on the server). - The bare `application/zip` path is programmatic-only and covered by content-type dispatch but not a dedicated route test (the multipart path is). ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), Claude Code CLI, extended thinking + tool use; root-caused against live logs/DB and the harness proxy source. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
276ae3a75d |
Harden company import: durable UI, async jobs, integrity guard, batched inserts (#10523)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company Import/Export (#10507) moves whole companies between instances as portability bundles > - Real-world use on a large company (1,418 issues, ~10.6k comments) surfaced a cluster of related failures: the import took hours and the browser connection died while the server kept running, a retry silently produced a second partial import, the progress/error UI gave no durable signal, and a cloud-tenant user couldn't even open the companies afterward > - Root cause of the slowness: importBundle inserted every issue, comment, and document as a separate round-trip to a network Postgres — an N+1-over-network pattern > - This pull request hardens the whole import path: durable progress/error UI, an async server-side job so imports survive dropped connections (with a duplicate-submit guard), a fail-closed guard against incomplete payloads, and batched inserts that cut a large import from hours to minutes > - The benefit is that migrating a real, large company actually completes, is legible while it runs, and can't half-import twice ## Linked Issues or Issue Description - Refs #10507 (the Import/Export feature this hardens). Supersedes #10513 (the progress/error-UI piece, folded in here). No open issue; problem described above (large-company import: slow, connection-fragile, silently duplicable, opaque UI). ## What Changed - **Batched inserts (perf):** importBundle pre-generates entity ids in JS and inserts in chunked multi-row statements, so children no longer wait on parents' generated ids. A 1,418-issue import drops from ~15,600 insert statements to **82** (190×); benchmark below. Import semantics — collision handling, pause-on-import, label/blocker/monitor/attachment/embedded-asset handling, blob sha verification — are unchanged (full portability suite green). - **Async import jobs for board sessions:** the existing cloud-tenant async job path opens to board sessions with per-actor job keys; the import page submits, polls, and resumes watching after a reload or dropped connection instead of holding one fragile request. A non-terminal job blocks a duplicate submit (409 returns the running job), preventing the double-import. - **Fail-closed completeness guard:** an optional `expectedFileCount` on inline imports; the server rejects (422 `import_payload_incomplete`) a body carrying fewer files than declared, so a re-framed/short payload fails loudly instead of half-importing. - **Durable progress/error UI (was #10513):** persistent progress panels with size-aware copy, persistent error panels with retry guidance, and inline explanation when the preview button is disabled; request-lifecycle guards so stale previews/imports can't publish or detach. ## Verification - `pnpm -r` typechecks (shared, server, ui) clean. - `company-portability.test.ts` (76) + `company-portability-routes.test.ts` (30) green — the import correctness net — plus new `CompanyImport.test.tsx` async/resume/409 coverage and a new batching regression test (a 50-issue import issues <50 issue-insert statements; rows land unchanged). - **Batching benchmark (embedded Postgres):** at 1,418 issues × 7 comments × 1 doc — 82 insert statements vs ~15,598 one-per-row (190×), ~1s wall-clock; a row-verifying run at that scale imports all 1,418 issues / 9,926 comments / 1,418 documents with unique identifiers and no warnings (no rows dropped by chunking). Over a network DB the round-trip reduction is the hours→minutes lever. - What is NOT directly measured here: wall-clock against a real network Postgres (that happens on a staging deploy); the local timing is network-free. ## Risks - Batching is the load-bearing change: it rewrites the import write path. Mitigated by the unchanged 106-test correctness suite, a new scale/row-integrity test, and per-writer transactions (a failure rolls back its table group; not a single outer transaction across writers — noted, correctness preserved). - Async jobs are in-memory (lost on server restart → pollers 404 and can resubmit); matches the pre-existing cloud-tenant job semantics. - `expectedFileCount` is optional (older callers unaffected); over-count is allowed, only under-count fails closed. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), Claude Code CLI, extended thinking + tool use; implementation across Fable 5 subagents with live diagnosis against a running instance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
fcf66f3a91 |
feat(skills): add managed skill rename API (#9688)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies. > - Company skills are reusable capabilities that operators install, edit, assign, and materialize for agents. > - Managed local skills currently lack a safe backend operation for changing their display name and canonical slug/key together. > - Treating rename as an ordinary save can leave duplicate records, stale runtime materializations, or agent assignments pointing at the old key. > - This pull request adds a company-scoped managed-skill rename contract, service operation, and REST endpoint with focused authorization and activity logging. > - The benefit is an atomic-enough, recoverable rename path that keeps disk state, database identity, and agent skill assignments synchronized. ## Linked Issues or Issue Description - Refs #2121 - Problem: managed company skills need a dedicated rename operation rather than save-time duplication behavior. - Expected behavior: renaming a managed skill updates its name, slug, key, source directory, frontmatter, runtime materialization, and assigned-agent references while preserving version pins. ## What Changed - Added shared request/result types and Zod validation for managed skill rename requests. - Added `POST /api/companies/:companyId/skills/:skillId/rename` with `skills.edit` policy checks and `company.skill_renamed` activity logging. - Restricted renames to Paperclip-managed local skills and added slug, key, and target-directory conflict handling. - Moved the managed directory, rewrote only the `SKILL.md` frontmatter name, updated the database row, and rolled filesystem changes back when persistence fails. - Rewrote assigned agents' desired-skill keys while preserving pinned version IDs and removed stale runtime materialization. - Added focused route and service coverage for success, no-op, name-only changes, conflicts, unsupported sources, assignment rewrites, rollback-sensitive behavior, and runtime cleanup. - Rejected multiline rename names before they can inject extra `SKILL.md` frontmatter fields. ## Verification - `pnpm exec vitest run server/src/__tests__/company-skills-service.test.ts server/src/__tests__/company-skills-routes.test.ts` — 106 tests passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. ## Risks - Filesystem and database updates cannot share one native transaction; the service stages filesystem changes and explicitly restores the original directory and markdown when the database transaction fails. - Renames intentionally reject catalog, remote, project-scanned, and unmanaged local skills to avoid changing identities owned by external sources. - No database migration is required; the endpoint updates existing company-skill and agent configuration fields. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex coding agent (exact underlying model ID and context-window size were not exposed to this runtime), with reasoning, repository tool use, code execution, and test execution. The rescued source commit also records assistance from Claude Opus 4.8. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
916c13501f |
Replace host-to-host Cloud Sync with full-fidelity company Import/Export (#10507)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - A company accumulates real state — issues, labels, blockers, documents, work products, monitors, attachments, agents, routines — and people need to move that state between instances: self-hosted to cloud, cloud back to self-hosted, or plain backups > - The experimental, flag-gated Cloud Sync transport (#6548) tried to solve this host-to-host: the source pushed into a receiver over HTTPS with a cross-instance consent/token handshake, which required the destination to be publicly reachable and broke for common self-hosted topologies (plain-HTTP LAN/VPN origins); the receiver half never landed upstream at all > - Meanwhile the portability bundle and the existing export/import pages already move companies offline with none of those networking constraints — but silently dropped labels, blockers, issue documents, work products, monitors, and every attachment > - This pull request removes the host-to-host transport and makes Import/Export the single data-movement path: the pages become first-class company-settings destinations, exports declare exactly what they do not carry, and bundle schemaVersion 6 now carries all of the above, with attachments as content-addressed sha256 blobs verified before a single row is written > - The benefit is a migration and backup flow that works between any two instances with no reachability requirements, no cross-instance auth, and no silent data loss ## Linked Issues or Issue Description - Refs #6548 — the original Cloud Sync sender this PR supersedes and removes. - Related, not duplicates: #1697 (goals in the portability manifest — orthogonal field addition), #954 (an earlier import/export + skill-visibility proposal predating the current portability bundle). - No open issue describes this directly, so in brief (feature-request shape): **Problem** — moving a company between instances silently lost labels (imports with label references actually hard-failed), blocker relations, issue documents, work products, monitor state, and all attachments, and the alternative Cloud Sync transport required the destination to be publicly reachable over HTTPS plus a consent handshake, which failed for typical self-hosted setups. **Desired behavior** — one Import/Export flow in company settings that produces a portable bundle carrying all of that data, tells the operator up front what it cannot carry, imports with automations paused, and offers real one-click activation afterwards. ## What Changed - New export fidelity report (`GET /api/companies/:companyId/export/fidelity`) + an "Export fidelity" panel on the Export page listing anything a bundle will not include (now only: approvals, cost history, activity history) - Imports accept `pauseAutomations`; imported agents and routines land paused, the import result reports created routines, and the Import page ends in an activation panel that actually resumes selected agents/activates routines - Export and Import pages promoted into the company-settings nav; the Cloud Upstream wizard, ux-lab page, and API client removed; the old settings route redirects to Export - Host-to-host transport removed: upstream-sync/receiver-client routes and services, CLI `cloud connect`/`cloud push` + keypair store, the shared upstream transfer contract, and the `enableCloudSync` flag; migration `0196` drops the two experimental `cloud_upstream_*` sender tables - Bundle schemaVersion 6: labels (definitions + per-task names, remapped by name on import), blocker relations (`blockedBy` slugs, cycle-tolerant), issue documents (`tasks/<slug>/documents/<key>.md`), work products (system refs nulled), monitors (notes/scheduledBy restored, imported un-armed) - Attachments travel as content-addressed `blobs/<sha256>` entries (deduped; comment-scoped attachments re-link via comment index); every blob is hash-verified **before any write**, so a corrupted bundle cannot leave a partially imported company; both zip codecs now round-trip extensionless/binary entries byte-exactly; the Import page preflights the inline body limit and offers continue-without-attachments - v5 (and older) bundles still import, with an informational warning; bundles newer than v6 are rejected cleanly - Docs: board-operator import/export guide, CLI README, README/ROADMAP updated ## Verification - `pnpm -r` typechecks (shared, db incl. migration numbering/safety checks, server, ui, cli) and `pnpm check:token-gates` — clean - Vitest: full server + shared sweep 4,888 passed / 1 skipped, with the only 3 failures being pre-existing on `master` (2× heartbeat-workspace-branch-containment, 1× workspace-runtime auto-port; reproduced identically with this change stashed); ui + cli suites green; the embedded-Postgres export-fidelity suite applies the full migration chain including the new `0196` against a fresh database - Live end-to-end on a scratch instance: seeded a company with labels, a blocker pair, an issue document, a work product, a monitor, an agent, a routine, and two binary attachments (one comment-scoped) → export → import into a fresh company → labels remapped to new ids, blocker edge and document restored, monitor un-armed with notes intact, attachments byte-identical (sha256-compared through the API), agents/routines paused → activation panel resumed them; a v5-shaped bundle imported with only the info warning; flipping one byte in a blob made the import 422 with **zero** rows created - Reviewer repro: create a company with a labeled issue + attachment → Settings → Export → download → Settings → Import on another company/instance → watch the preview, apply with "start paused", then activate ## Risks - Migration `0196` drops `cloud_upstream_connections`/`cloud_upstream_runs` — experimental tables behind a default-off flag; their connection/run history is intentionally discarded - Breaking removals are all of experimental, flag-gated surface: `/api/upstream-sync/*` + `/api/cloud-upstreams/*` routes, `paperclipai cloud connect|push`, and the `enableCloudSync` flag (stale keys in stored instance settings parse harmlessly) - Import remains non-atomic on mid-apply errors generally (pre-existing behavior); the new blob verification specifically moved ahead of all writes so tampered bundles cannot create partial state - GitHub-sourced imports do not fetch `blobs/*` and skip attachments with a warning ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), via Claude Code CLI with extended thinking, tool use, and subagent orchestration; implementation and review split across Fable 5 subagents, with live end-to-end verification against a running instance ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
15ce70dc18 |
feat(server): thread plural referenced-project workspaces through run prep (#10448)
## Thinking Path > - Paperclip coordinates work for autonomous companies. > - A run needs a workspace view before execution starts. > - That view now needs to cover one anchor project and more referenced projects. > - Those extra workspaces must stay separate and must not change the anchor path when the feature stays off. > - This PR threads plural workspace data through run prep behind a default-off kill switch. > - It also keeps extra project workspaces isolated and makes the realization contract round-trip the new shape. > - The benefit is safer run prep for referenced projects without changing the current default path. ## Linked Issues or Issue Description This PR does not link a public GitHub issue. It follows the internal run-prep task for plural referenced-project workspaces. Problem: - Run prep resolves the anchor project today, but it does not yet carry each referenced project into the run workspace view. - That gap blocks runs that need a second repo or sibling project during preparation. Proposed solution: - Thread a plural workspace result through run prep. - Keep the anchor path unchanged when the kill switch is off. - Resolve each referenced project into its own managed checkout directory when the flag is on. Alternatives considered: - Keep one shared workspace and layer the extra repos into it. Rejected because it would blur isolation and make failures harder to bound. - Upload the extra workspaces immediately. Rejected because this PR only prepares the data path. Roadmap alignment: - This change sits in the workspace and sandbox path. - It matches the roadmap work on workspace strategy and cloud or sandbox agents. ## What Changed - Added `additionalWorkspaces[]` to the run workspace result. - Split workspace resolution into an anchor path and an optional referenced-project path behind `PAPERCLIP_MULTI_PROJECT_WORKSPACE_SYNC`. - Kept per-project failure isolation so one bad clone does not stop the run. - Keyed managed workspace directories by `projectId` so sibling workspaces stay separate. - Added `additionalSources[]` to the workspace realization request and kept read and write paths backward compatible. - Added tests for the anchor-only path, the new workspace shape, and the per-project directory rule. ## Verification - `pnpm --filter @paperclipai/server run typecheck` - `pnpm --filter @paperclipai/shared run typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-project-env.test.ts` 21/21 - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/workspace-runtime.test.ts` 98/98 ## Risks Low risk. The new path stays behind a default-off kill switch, so the anchor flow does not change when the flag is off. The main risk is a bad referenced project clone. That case now drops only the affected project and keeps the run alive. The shared type change also needs every consumer to use the new array field where extra workspaces matter. ## Model Used OpenAI Codex (GPT-5; exact internal model ID not exposed in this environment; tool use enabled) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f9034ab3ca |
build(deps-dev): bump @types/node from 22.19.21 to 22.20.1 (#10304)
Bumps [@types/node](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/node) from 22.19.21 to 22.20.1. <details> <summary>Commits</summary> <ul> <li>See full diff in <a href="https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/node">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
c274f10abc |
feat(server): computed owner instance-admin elevation for cloud-managed instances, behind platform floors (#10343)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Cloud-managed instances authenticate tenant users through a trusted-header path (`resolveCloudTenantActor`) that deliberately never grants `instance_admin`, so every tenant user is company-scoped > - On a dedicated (single-owner) managed instance that leaves the paying owner unable to administer their own instance: instance settings, the environments admin surface, and the custom sandbox image flow are all instance-admin gated (the environments UI can't even show the provider/image of the platform sandbox because the restricted read view blanks `config` entirely) > - Re-granting the old blanket `instance_user_roles` row would repeat the mistake the shared-pool hardening fixed: DB rows go stale, resurrect via restores, and elevate through every auth path > - This pull request elevates only the stack `owner`, computed per request at the trusted-header boundary behind a new managed-tier feature flag, and ships that elevation together with code floors on the platform-owned surfaces an instance admin must not control on a managed instance > - The benefit is that dedicated-stack owners can administer their own instance while platform credentials, execution policy, backups, and runtime code-install stay platform-owned, and self-hosted behavior is unchanged ## Linked Issues or Issue Description No public GitHub issue exists for this change; the underlying issue is described here following the feature-request template. ### Problem or motivation - On cloud-managed instances, tenant users resolved from trusted headers are always company-scoped. For dedicated instances with a single paying owner, the owner cannot reach any instance-admin surface of their own instance (instance settings, environments administration, custom image setup), and the restricted environment read view hides even structural fields like the sandbox provider and image. - The previous hardening intentionally removed blanket elevation (and purges stale `instance_user_roles` rows on every trusted-header authentication). That protection must not regress for shared multi-tenant pools. ### Proposed solution - Owner-only, computed, flag-gated elevation plus code floors on platform-owned surfaces, in one PR so the elevation can never ship without the floors. ### Alternatives considered - Re-inserting an `instance_user_roles` row for owners (the pre-hardening model): rejected — DB rows go stale, survive restores, and elevate through every auth path; #7525 removed exactly this. - Widening only the environments read view without any elevation: rejected — it fixes one screen but still leaves a dedicated-instance owner unable to administer instance settings or custom images. - Elevating additional stack roles (`member`/`admin`/`support`): rejected — only the owner has an ownership claim over the whole instance; other roles stay company-scoped. ### Roadmap alignment - Extends the shipped "Cloud deployments" roadmap work (multi-tenant isolation, company-scoped cloud tenants, managed-instance bootstrap) without overlapping planned core items, and leaves self-hosted behavior unchanged. ## What Changed - **New feature key** `enableOwnerInstanceAdmin` (`packages/shared`): boolean flag in `instanceExperimentalSettingsSchema`, catalog tier `managed`, `cloudDefault: true`, `selfHostedDefault: false`. Inert on self-hosted instances — the elevation path only exists behind the cloud tenant trust token. - **Computed elevation** (`server/src/middleware/auth.ts`): `resolveCloudTenantActor` now returns `isInstanceAdmin: true` only when the trusted-header stack role is `owner` **and** the flag is enabled. The flag is resolved through the instance-settings service so the managed-config overlay applies (the control plane can disable elevation fleet-wide without touching tenant databases; a DB row edit or restore cannot resurrect it). Resolution fails closed on settings read errors. The `instance_user_roles` never-insert and the per-request stale-row purge are byte-identical. `member`/`admin`/`support` stack roles stay company-scoped. - **Authorization guard split** (`server/src/services/authorization.ts`): the blanket-allow now trusts the actor's *computed* `isInstanceAdmin` flag (only the attested resolver can set it for `cloud_tenant` actors) while keeping the `instance_user_roles` DB lookup excluded for `cloud_tenant` — a stale or hand-inserted role row still elevates nothing. - **Floor F1 — platform environment credentials** (`server/src/routes/environments.ts`): on cloud-managed instances, platform-provisioned environment rows (`managedByPaperclip` marker, plus the legacy managed-Kubernetes marker) use a single floored view for every reader on all environment routes (list, get, create, update, delete responses): `envVars` are never echoed and credential-shaped `config` keys (reusing the managed-config `SECRET_LIKE_CONFIG_KEY_PATTERN`) are dropped — for **all** actors including instance admins — while structural config (provider, image, template, region, …) and the managed markers stay visible. This also fixes the environments UI for managed sandboxes, which previously lost the provider/image entirely in the restricted view. The floor also covers writes: `PATCH /environments/:id` and `DELETE /environments/:id` on a platform-provisioned row are rejected (403, `environment_platform_managed`) for every actor including instance admins, and the guard binds to the persisted row's markers so a patch cannot strip the managed marker to lift the floor. The one recovery path is a metadata-only PATCH that solely clears the marker keys (null/false), for rows stamped through the old unrestricted API before the markers became reserved — and it never applies to a row whose slot markers are live platform state: the single local row (`environments_local_driver_idx`), which `ensureLocalEnvironment` adopts and stamps on cloud-managed instances from every caller (company creation, the heartbeat, run orchestration), and the single marked sandbox row (`environments_managed_sandbox_idx`) while a managed-sandbox bootstrap path is configured (managed-config `environments` section or `PAPERCLIP_EXECUTION_MODE=kubernetes`) and the provisioner therefore adopts and refreshes it on every boot. Clearing a live slot row's markers would let the next write reclassify it as tenant-managed and bypass the floor; conversely, when no sandbox provisioning path is configured the platform holds no claim on any sandbox row, so a platform marker there is stale by definition and the recovery patch applies. Every marker outside a live slot is clearable, so no legacy row is ever locked permanently. Custom-image setup and probes on the platform sandbox stay available to instance admins — those are the owner-facing flows this elevation exists for. The marker keys themselves are reserved: client create/update payloads that set `managedByPaperclip` or `managedKubernetesSandbox` are rejected (422, `environment_platform_marker_reserved`) on cloud-managed instances, so a tenant row can never be stamped platform-provisioned through the API and self-locked behind the write floor (the provisioner writes markers at the service layer, not through these routes). Tenant-created environments are otherwise unaffected. - **Floor F2 — executionMode** (`server/src/routes/instance-settings.ts`): on cloud-managed instances, `PATCH /instance/settings/general` rejects writes that would change `executionMode` (403, `execution_mode_platform_managed`). Same-value echoes pass so settings forms that submit the full general-settings object keep working. The boot-time execution-policy bootstrap path is untouched (it calls the service directly). - **Floor F3 — manual database backups** (`server/src/routes/instance-database-backups.ts`): floored off on cloud-managed instances (403, `database_backups_platform_managed`); backups are platform-owned there, and the result would also echo a server-side filesystem path. - **Floor F4 — adapter code install** (`server/src/routes/adapters.ts`): `POST /adapters/install` and `POST /adapters/:type/reinstall` are floored off on cloud-managed instances (403, `adapter_install_platform_managed`). Adapter packages execute in the server process, so a runtime install would let an instance admin read the platform trust anchors out of the process environment. This mirrors the existing bundled-only plugin install floor; adapter code on managed instances comes bundled with the platform image. ## Instance-admin surface audit Before widening who can hold `isInstanceAdmin`, every instance-admin-gated surface in `server/src` was enumerated and reviewed for whether its response or side effects could echo process environment values or platform credentials (tenant trust token, JWT signing keys, database connection strings, provider API keys): 29 distinct gate definitions covering ~90+ call sites, in four groups — sole instance-admin gates (12), instance-admin-or-company-permission gates (10), response-shaping/scope-widening sites (6), and the central `allow_instance_admin` short-circuit in the authorization service (58 `decide()` call sites). Findings and dispositions: - **Environment read/write responses** exposed platform sandbox `envVars`/credential-shaped config to instance admins → closed by floor F1. - **Manual backup trigger** echoed a server filesystem path and triggers a platform-owned operation → closed by floor F3. - **Adapter install/reinstall** loads externally fetched code into the server process (indirect, complete env exposure) → closed by floor F4. The sibling plugin-install path already had a bundled-only floor on managed instances and needed no change. - **Token-minting surfaces** (gateway tokens, custom-image terminal/connection tokens) mint credentials scoped to the instance's own resources, not platform trust anchors → acceptable for an owner-admin of a dedicated instance; unchanged. - All remaining gated surfaces return ordinary instance-scoped business data; none echo `process.env` or platform secrets directly. OAuth client secrets are referenced by env-var *name* only; SSH private keys are stored as secret refs before persistence and are not echoed. Operational note for managed platforms: this model assumes the process environment of a managed instance holds only that instance's own credentials. Platform operators should keep provider credentials per-instance (never fleet-shared) since an instance admin ultimately controls in-process code on their own instance. ## Verification - `pnpm vitest run server/src/middleware/cloud-tenant-actor.test.ts` — resolver matrix: owner × flag on/off, flag via managed overlay (on-over-DB-off and off-over-DB-on), member/admin/support × flag on, no-token self-hosted, fail-closed settings read, purge still runs and no role row is ever inserted (14 tests). - `pnpm vitest run server/src/__tests__/authorization-service.test.ts` — computed flag elevates a `cloud_tenant` actor; a stale `instance_user_roles` row still never does; `session` actors unchanged (full suite, embedded Postgres). - `pnpm vitest run server/src/__tests__/environment-routes.test.ts` — F1: no secret echo to admins on get/list, structural config visible to restricted readers, platform-row PATCH/DELETE rejected for admins (including a marker-stripping patch), marker-clear recovery allowed for stale legacy rows and for a marked sandbox row when no provisioning path is configured, but refused on the managed local row and on the sandbox slot row under a managed-config `environments` entry or the forced kubernetes execution mode, client marker-stamping creates/patches rejected, tenant rows still readable and writable, self-hosted read+write regression (60 tests). - `pnpm vitest run server/src/__tests__/environment-service.test.ts` — `ensureLocalEnvironment` adopts a pre-existing local row on cloud-managed instances (marker stamped, other metadata preserved, idempotent — no rewrite on re-ensure) and leaves self-hosted rows untouched (22 tests, embedded Postgres). - `pnpm vitest run server/src/__tests__/instance-settings-routes.test.ts server/src/__tests__/instance-database-backups-routes.test.ts` — F2 change-vs-echo matrix incl. self-hosted regression; F3 floor for both admin shapes (32 tests). - `pnpm vitest run server/src/__tests__/adapter-routes-authz.test.ts` — F4 floor; self-hosted install/reinstall behavior unchanged (existing cases). - `pnpm vitest run server/src/__tests__/first-admin-claim.test.ts server/src/__tests__/bootstrap-claim-routes.test.ts server/src/__tests__/managed-config.test.ts server/src/__tests__/health.test.ts server/src/__tests__/instance-settings-managed-overlay.test.ts server/src/services/managed-environments.test.ts server/src/services/execution-policy-bootstrap.test.ts` — first-admin bootstrap gate and managed-config behavior unchanged (91 tests). - `pnpm vitest run packages/shared/src/feature-catalog.test.ts` — catalog/schema sync tests cover the new key (selfHostedDefault must equal the schema default). - `pnpm run typecheck` — all 31 workspace projects clean. ## Risks - Self-hosted behavior is unchanged: every floor binds to `isCloudManagedInstance()` (tenant trust token present), the new flag defaults off with no elevation path, and regression tests pin the self-hosted branches. - The elevation is fail-closed and stateless: turning the flag off (managed overlay or DB) de-elevates on the next request; there is no role row to clean up and restores cannot resurrect elevation. - On a cloud-managed instance a pre-existing unmarked local row is adopted (stamped `managedByPaperclip`) by the next ensure and becomes platform-owned — the intended managed-product semantic: the platform owns the single local slot. Self-hosted instances are untouched. - F1 widens restricted readers' view of platform-provisioned rows from fully blanked `config`/`metadata` to structural-only `config` plus markers. Platform-delivered config is guaranteed secret-free by the managed-config contract (secret-shaped keys fail startup), and the floor re-drops secret-shaped keys defensively. - One extra instance-settings read per trusted-header request for owner-role actors (the resolver already performs several queries per request). ## Model Used Claude Fable 5 (Anthropic) — model id `claude-fable-5`, extended thinking enabled, agentic tool use via Claude Code; read-only explore subagents on the same model were used for the surface audit sweep. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c3bd0c5d50 |
feat(skills): add beta releases for the core Paperclip skill (#10228)
## Thinking Path > - Paperclip is the open source control plane people use to organize and operate AI-agent companies. > - Agent behavior depends partly on the bundled Paperclip core skill synchronized into each runtime. > - The existing database and runtime plumbing already supports immutable skill-version snapshots and per-agent version selections, but no product workflow exposed that capability. > - Replacing the live bundled skill globally would make champion adoption risky and difficult to compare across agents. > - This pull request adds an experimental, instance-level beta-skills gate plus a repository release registry, immutable seeded releases, enforcement, and a per-agent release picker. > - The benefit is controlled per-agent evaluation of frozen core-skill releases while the default-off path remains behaviorally unchanged. ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting: `server/`, `ui/`, `packages/db`, and `packages/shared`. ### Problem or motivation Paperclip needs a safe way to evaluate improved versions of its core operating skill without globally replacing the live default. Today the version-snapshot and per-agent pin plumbing exists, but operators cannot use it. A global replacement would make regressions difficult to contain and would prevent controlled comparisons across agents. ### Proposed solution Add a default-off instance experiment that exposes immutable, named core-skill releases. When enabled, operators can pin each agent to a seeded release; when disabled, every agent resolves the live default while saved pins remain intact. Validate pinned writes at the API boundary, gate reads at runtime, and expose the selection in the agent Skills tab. ### Alternatives considered - **Replace the bundled core skill globally:** rejected because it changes every agent at once and provides no rollback/isolation boundary. - **Ship releases as separate skills:** rejected because releases are versions of one core capability, not independently enabled skills. - **Store release snapshots only outside the repository:** rejected because repository provenance and hashes make builds reproducible and reviewable. ### Roadmap alignment This extends the Skills Manager / Skill Studio direction in `ROADMAP.md` by making core-skill versions operable per agent. It does not duplicate another open implementation PR; GitHub searches found no related `enableBetaSkills` change. ### Additional context The feature remains experimental and default off. The V7 champion was selected through a multi-model evaluation process, and the frozen release contents are verified by SHA-256 below. ## What Changed - Added the default-off instance-level `enableBetaSkills` experimental flag. - Added `skills-releases/paperclip/` with the ordered release registry and frozen `v0` plus `v7-roster` snapshots. - Added release metadata to `company_skill_versions` and idempotent release seeding. The migration was planned as `0191`, then renumbered to `0192` because current `master` claimed `0191` before final rebase. - Added read-time gating and write-time validation so disabled instances always resolve the live default and reject pinned-version writes. - Added the per-agent Release picker in the agent Skills tab, including responsive layout and beta-pin state. - Kept `EDITS.md` out of the release registry and PR diff. ### V7 Adoption Evidence - Paid roster: 6 models, 94-case suite. - Result: 553/564 pass-within-2, mean 92.17/94, versus the P2 baseline of 544/564. - Reference model improved 84→91; maximin improved 84→90. - Final report: https://pages.paperclip.ing/skills/optimization/paperclip/pap-14624-p3-final-20260721/ ### Provenance - `v7-roster` is the Phase 1 champion plus additions-only edits E107–E112. Per-edit rationale remains in the evals repository at `source/v7-roster/EDITS.md` and is deliberately excluded from this PR. - `v0` is the `skills/paperclip` tree from commit `ea66ea81`. - Champion selection was accepted on July 21, 2026 via board card `9c304fc2` (PAP-14624 G3). - This delivery mechanism was accepted on July 24, 2026 via plan revision `2367abd2` (PAP-14858). ### QA Evidence - P4 QA matrix comment `b7f40522-4e9b-4a3a-9821-28e86fe1a987`: all 6 acceptance criteria passed. - Automated QA matrix: 166 tests passed with 0 failures, including real filesystem materialization and full SHA-256 assertions. - UI QA exercised the real agent Skills tab at desktop and mobile widths with the experimental flag both on and off. ## Verification - `pnpm check:token-gates` - Focused beta-release matrix: 169 tests passed across shared validators, server services/routes/heartbeat behavior, instance settings UI, and release picker UI. - `pnpm -r typecheck` - `pnpm build` - `pnpm test:run`: server and UI partitions passed; one CLI doctor test inherited temporary AWS credentials from the agent heartbeat and expected no static credentials. The isolated rerun with `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, and `AWS_SESSION_TOKEN` unset passed 8/8. - V7 SHA-256: - `SKILL.md`: `53ab290489684cbf116fdd1406a95f6b6f53c9c36358b1bf8bfeae481e253575` - `references/cases.md`: `3b821f59064a7761091020a14819a8d787131f24029748563d6c0e1be7e6eaec` - `references/workflows.md`: `69747bd6e05f7e3673d1e67b07ff295df1869c05e1fd029804d5fa9177db92cd` - Confirmed 49 changed files, no `pnpm-lock.yaml`, no workflow changes, and no `EDITS.md`. ## Risks - **Migration:** low-to-moderate risk. Three nullable columns and one partial unique index are added idempotently; existing rows remain valid. - **Behavior:** low risk while the flag is off because read-time resolution forces the live default and saved pins are preserved but inactive. - **Frozen content:** release snapshots intentionally diverge from future live skill edits; provenance and hashes make that divergence explicit and reproducible. - **UI:** low risk. The picker only renders for the bundled core skill when the experimental flag is enabled and seeded releases exist. > This extends the existing Skills Manager / Skill Studio direction described in `ROADMAP.md`; it does not duplicate another open implementation PR. The GitHub PR search found no related `enableBetaSkills` change. ## Model Used - OpenAI Codex using `gpt-5.5` with reasoning and terminal/code-execution tools; context-window size is not exposed by this runtime. Earlier implementation commits also record Claude Opus 4.8 assistance where applicable. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [ ] I have not referenced internal/instance-local Paperclip issues or links (required governance identifiers are included above; no internal URL is included) - [ ] My branch name describes the change and contains no internal Paperclip ticket id (the approved delivery plan mandated this shared branch name) - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c111ee4cb3 |
feat(server): add per-user document stars (#9952)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies and their work. > - Artifacts and documents are first-class outputs, but users need a personal way to keep important documents easy to find. > - Existing resource memberships already model per-user starred projects and agents with company scoping and activity logging. > - Documents lacked the equivalent membership model, route, and artifact filtering behavior. > - The shared membership contract also needs to remain safe for existing UI project/agent mutation helpers when documents become a recognized resource type. > - This pull request extends the existing resource-membership system with per-user document stars and a starred artifacts view. > - The benefit is a company-scoped, idempotent server foundation for a dedicated starred-documents experience without weakening authorization or artifact filtering semantics. ## Linked Issues or Issue Description ### Problem / Motivation Board users cannot star individual documents, and the company artifacts API cannot return only the current user's starred documents. ### Proposed Solution Add company/user-scoped document memberships, a board-only document star route, document membership data in the shared contract, and a `starred=true` artifacts filter. ### Alternatives Considered A document column was rejected because stars are per-user; a separate star API was rejected because projects and agents already use resource memberships. ### Roadmap Alignment This extends the existing Artifacts & Work Products roadmap area and does not duplicate another open pull request found in the repository search. ## What Changed - Added the `document_memberships` schema and migration with company/user/document uniqueness and starred ordering. - Extended shared resource-membership and artifact-query contracts for documents and `starred=true`. - Added company-scoped document star/unstar service and board-only route behavior with activity logging. - Added starred document artifact filtering, including user-authored documents, document kinds, cursor ordering, and incompatible-kind handling. - Preserved idempotency under concurrent star requests and synchronized UI membership defaults/helpers with the expanded contract. - Added focused shared, route, service, and UI regression coverage. ## Verification - `pnpm exec vitest run packages/shared/src/resource-memberships.test.ts server/src/__tests__/company-artifacts-service.test.ts server/src/__tests__/resource-memberships-routes.test.ts` - `pnpm exec vitest run ui/src/components/SidebarAgents.test.tsx ui/src/components/SidebarProjects.test.tsx ui/src/components/SidebarStarredProjects.test.tsx` - `pnpm --filter @paperclipai/db check:migrations` - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` ## Risks - The migration adds a new membership table and non-concurrent indexes; migration safety gates pass with the repository's established policy. - The starred artifacts query intentionally returns only documents and relaxes the normal agent-authored/system-kind predicates for documents the current user explicitly starred. - Document membership mutations remain board-user-only; agent callers receive no document-star capability. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI; runtime model ID and context-window size were not exposed to this session. Reasoning, repository tool use, code execution, and test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0cf64d36a5 |
feat(secrets): write through external values and deep-link details (#10196)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies. > - Its secrets subsystem can resolve external references such as AWS Secrets Manager values without copying those values into Paperclip custody. > - Operators also need to rotate a referenced secret's value while preserving the same provider reference for consumers inside and outside Paperclip. > - Previously, external-reference rotation could only retarget metadata, and secret detail sheets were driven by local component state rather than shareable navigation state. > - This pull request adds an optional provider write capability, implements AWS Secrets Manager write-through rotation, and exposes capability-aware rotate modes in the UI. > - It also makes secret and each-user definition detail sheets URL-driven and adds a copy-link action. > - The benefit is that operators can update the canonical external value safely while keeping AWS rotation tracking intact, and they can share or navigate directly to secret details. ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting (`server/`, `ui/`, and `packages/shared`). ### Problem or motivation External-reference secrets can follow a provider-managed value, but Paperclip could not write a replacement value back to providers that support it. Operators had to leave Paperclip, update the value separately, and then return without an auditable Paperclip rotation record. Secret detail sheets also could not be shared or restored through browser history because their selection lived only in component state. ### Proposed solution Add an optional `updateExternalSecretValue` provider capability and surface it as `supportsExternalValueWrites`. Implement AWS writes with `PutSecretValue` while leaving the resolution `versionId` unset so future reads continue following `AWSCURRENT`. Add write-value and retarget modes to the rotate dialog for capable providers. Drive secret detail selection from `?secret=` / `?definition=` query parameters and provide a copy-link action. ### Alternatives considered Converting an external reference into a Paperclip-managed secret would break consumers that depend on the existing provider reference. Pinning reads to the newly written AWS version would prevent later out-of-band rotations from flowing through. Keeping sheet selection only in React state would not support browser Back or shareable links. ### Roadmap alignment This extends the completed “Secrets Manager with per-agent access” roadmap capability; it does not duplicate a separate planned roadmap item. Public GitHub searches found no duplicate or closely related issue or PR. ### Additional context The PR includes focused provider, service, and UI render coverage. Cutter also generated previews for the deep-linked detail sheet and capability-aware rotate modes. ## What Changed - Added optional external-value write support to the secret provider contract and provider descriptors. - Implemented AWS Secrets Manager write-through with `PutSecretValue`, audit material, and compensation when persistence fails after the provider write. - Allowed `secretService.rotate()` value updates for external references while rejecting ambiguous value-plus-retarget combinations and unsupported providers. - Added capability-aware “Write new value” and “Change reference” rotate modes with updated custody and action copy. - Made secret and each-user definition detail sheets source their selection from URL query parameters, compose with folder paths, close through browser history, and expose a copy-link action. - Added provider, service, and UI render coverage for write-through, rollback, capability messaging, dialog modes, and deep links. ## Verification - `pnpm vitest run server/src/__tests__/aws-secrets-manager-provider.test.ts` — 18 passed. - `pnpm vitest run server/src/__tests__/secrets-service.test.ts` — 75 passed. - `pnpm vitest run ui/src/pages/Secrets.render.test.tsx` — 31 passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm check:token-gates` — passed with all gates clean. ## Risks - External value writes affect the canonical provider secret and therefore all consumers of that AWS secret; the UI explicitly labels this custody behavior. - A provider write can succeed before Paperclip persistence fails. The service records the written version and AWS support includes compensation coverage to restore the prior value where possible; unrecoverable failures return explicit audit-safe error details. - URL-driven sheet state changes navigation behavior; render tests cover deep links, Back/close behavior, and composition with folder query state. - No database migration or breaking API requirement is introduced; providers without the optional capability retain reference-only behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using exact model ID `gpt-5.6-sol`, high reasoning mode, Codex CLI `0.142.5`, with repository, shell, Git, GitHub CLI, and code-execution tools. The runtime did not expose a context-window size. - Earlier implementation commits were assisted by Anthropic `Claude Fable 5` as recorded in their commit trailers; the exact backend model ID and context-window size were not preserved in the workspace metadata. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes, or confirmed no documentation change is required - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f6ab82d490 |
feat(interactions): add interaction withdrawal and terminal-issue expiry (#10251)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents and boards coordinate through issue-thread interactions (request_confirmation, ask_user_questions, suggest_tasks, …) that wait as `pending` cards until someone resolves them > - Two lifecycle gaps existed: an interaction's creator could not take back a card it no longer stands behind, and interactions left `pending` on issues that reached a terminal status lingered forever as live-looking approval requests > - Stale pending cards mislead humans (they look actionable), distort attention/liveness signals, and in the worst case invite acting on a proposal whose issue is already closed or cancelled > - This pull request adds an explicit withdraw route for pending interactions and automatically expires pending interactions when their issue reaches a terminal status (including a catch-up sweep for issues closed before this change) > - The benefit is that interaction cards now faithfully reflect reality: only genuinely actionable requests stay pending, and creators can retract requests that events have overtaken ## Linked Issues or Issue Description Fixes #5787 Refs #7403 Related prior PRs found while searching for duplicates (all overlap partially; none combine both lifecycle paths or the route-level authorization used here): #6709 and #7312 (creator-withdraw attempts), #8169 (terminal expiry), #8081 and #5137 (generalized cancel/expire endpoints), #6094 (stale confirmation auto-resolve). Related merged context: #9568 (agent cancel for ask_user_questions), #10119 (tolerating legacy `withdrawn_by_creator` result outcomes — the reader side of the outcome this PR writes). ## What Changed - New route `POST /issues/:id/interactions/:interactionId/withdraw` that resolves a `pending` interaction to status `withdrawn` with a structured result (`outcome: "withdrawn"`, optional trimmed `reason`), stamps `resolvedBy*`/`resolvedAt`, touches the issue, logs activity, and emits resolved-interaction telemetry - Withdrawal authorization: board users, the interaction's creator agent, or the issue's current assignee agent (assignees additionally pass the standard issue-mutation gate); task-watchdog runs are explicitly rejected, and authorization-boundary plus low-trust control-plane checks apply - Withdrawing an already-resolved interaction returns `409`; unknown/cross-issue/cross-company interaction ids return `404` - New service method `expirePendingInteractionsForTerminalIssue`: when an issue transitions to a terminal status, all of its `pending` interactions are resolved to `expired` with `outcome: "issue_closed"`, guarded by a `status = 'pending'` predicate so concurrent resolutions are not overwritten - The same expiry runs as a catch-up when interactions are listed on an already-terminal issue, so cards stranded by issues closed before this change also get cleaned up; expired request_confirmations are logged with a distinguishing source - Shared package: new `withdrawIssueThreadInteractionSchema` validator, `WithdrawIssueThreadInteraction` type, and `withdrawn` / `issue_closed` result-outcome support for all interaction kinds (kind-aware result shapes for `ask_user_questions` and `request_item_verdicts`) - UI helper `ui/src/lib/issue-thread-interactions.ts` recognizes the new outcomes for card rendering - Docs: bundled skill API reference updated with the withdraw endpoint - Review follow-up: terminal expiry moved from the HTTP route hooks into `issueService.update`'s status-transition block, so direct service callers (tree control, recovery, pipelines, status cards) expire pending cards too; the list-endpoint catch-up remains for issues closed before this change - Review follow-up: withdrawing or issue-close-expiring a `request_confirmation` also settles its linked `tool_action_requests` row (withdraw -> `cancelled`, issue closed -> `expired`), so a parked tool call cannot stay approvable after its card is gone - Review follow-up: interaction cards render dedicated copy for the new outcomes ("Withdrawn" with the reason, "Expired when issue closed") instead of falling through to superseded-by-comment / stale-target variants; withdrawn plan reviews badge as "Withdrawn" rather than "Changes requested" ## Screenshots Card states rendered from a local ux-lab harness with mocked data ([full gallery](https://pages.paperclip.ing/pr-10251-interaction-withdrawal-cards/)): | Light | Dark | | --- | --- | |  |  | ## Verification - `pnpm --filter @paperclipai/shared build` — clean tsc - `cd server && npx vitest run src/__tests__/issue-thread-interaction-routes.test.ts` — 22 tests pass, including new coverage for: creator-agent withdraw success, non-creator/non-assignee agent 403, watchdog-run 403, double-withdraw 409, and board-user withdraw - `cd server && npx vitest run src/services/issue-thread-interactions.test.ts` — 4 tests pass, including terminal-issue expiry writing `issue_closed` results and leaving already-resolved interactions untouched - `cd ui && pnpm typecheck` — clean - `cd server && npx vitest run src/__tests__/issues-service.test.ts` — includes a new embedded-Postgres test proving a direct `issueService.update` terminal transition expires pending interactions and writes the activity-log entry - `cd ui && npx vitest run src/components/IssueThreadInteractionCard.test.tsx` — 32 tests, including new coverage for withdrawn / issue-closed confirmation and question cards - `cd server && npx tsc --noEmit` — matches the pre-existing repo error baseline exactly (no new errors) - Manual: `POST /issues/:id/interactions/:interactionId/withdraw` with `{"reason":"superseded"}` as the creator agent resolves the card to `withdrawn`; closing an issue with a pending confirmation flips it to `expired` with `outcome: "issue_closed"` ## Risks - Interactions on terminal issues now auto-expire (including retroactively via the list-time catch-up), so consumers that expected to resolve a pending interaction on a closed issue will get `409`; this is the intended semantics and matches how the attention feed already wants to treat dead cards - New result outcomes (`withdrawn`, `issue_closed`) are written to stored results; readers were already made tolerant of these outcome strings in #10119, so mixed-version reads are safe - No schema/migration changes; per-row conditional updates (`status = 'pending'`) avoid clobbering concurrent resolutions - Withdrawal is a new mutation surface, but it is strictly narrower than existing resolve paths (board, creator, or assignee only; watchdog runs blocked) ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic Mythos-class tier) with extended thinking and agentic tool use (Claude Code harness); commit authored in a Paperclip-managed engineering session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4e00818574 |
fix(runtime): support in-place workspace realization (#10230)
## Thinking Path > - Paperclip is the control plane people use to coordinate AI agents and their execution environments. > - Environment realization decides where an agent runs and which filesystem and toolchain are authoritative. > - Copy-based realization is unsafe for container-anchored tasks because absolute paths such as `/app` can point outside the synchronized tree and task-specific binaries may be absent. > - That mismatch can let an agent successfully verify work in a phantom writable path while sync-back silently discards the result. > - Existing task environments already provide the authoritative filesystem and toolchain, so they should be executed in place rather than copied. > - Copy mode still needs explicit confinement rules so aliases target the synchronized workspace and unsynchronized writable paths fail visibly. > - This pull request adds typed realization metadata, propagates the authoritative root through orchestration, and teaches Codex to honor it. > - The benefit is that container-anchored tasks operate on verifier-visible state with the intended tools, while copy mode remains safe and backward compatible. ## Linked Issues or Issue Description No public GitHub issue exists for this defect. GitHub duplicate searches for in-place execution, workspace realization, and authoritative workspace roots found no related pull request to link. ### What happened? Environment-backed agent runs were always realized through a copied workspace. Tasks anchored to absolute container paths could therefore write outside the synchronized tree, and task-provided toolchains were unavailable in the copy. A run could report success even though sync-back discarded its output. ### Expected behavior Existing task environments should run against their real authoritative root and toolchain. Copy-mode runs should map declared absolute aliases into the synchronized tree and reject writable paths that cannot be restored. ### Steps to reproduce 1. Run a Codex task environment whose required files live under `/app` or `/workspace` and whose required binary exists only in the task container. 2. Observe that copy realization changes the effective filesystem/toolchain or permits writes outside the synchronized root. 3. Complete and verify the task inside the agent sandbox. 4. Observe that the verifier cannot see out-of-tree artifacts or that task-specific commands were unavailable. ### Reproduction context - Paperclip commit: `3a16b91217483d2c233926de5b7f7bc3a1077924` - Deployment: built from source in a task-container execution environment - Adapter: Codex local - Database: not database-related - Access context: agent execution ## What Changed - Added typed `copy | in_place` workspace-realization metadata, authoritative roots, confined aliases, and outbound restore paths to shared execution-target contracts. - Selected in-place realization for existing task environments and skipped archive prepare/restore when the authoritative environment is used directly. - Propagated the authoritative root into adapter context so Codex uses it for cwd and `PAPERCLIP_WORKSPACE_*` semantics, including ACP execution. - Bound copy-mode aliases such as `/app` to the synchronized workspace and rejected writable out-of-tree paths without explicit restore mappings. - Added focused regression coverage while preserving existing copy-mode archive restore behavior. ## Verification - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm exec vitest run packages/adapter-utils/src/local-process-sandbox.test.ts packages/adapters/codex-local/src/server/acp.test.ts packages/adapters/codex-local/src/server/execute.remote.test.ts server/src/__tests__/environment-run-orchestrator.test.ts` — 48 passed, 4 skipped. - `pnpm -r typecheck` — passed. - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY -u AWS_SESSION_TOKEN pnpm test:run` — passed across all general and serialized Vitest shards. - `pnpm build` — passed. - Codex `k=1` acceptance run completed July 24, 2026 at 23:54:30 UTC with 4 completed, 0 exceptions, and mean reward 1.0: `build-cython-ext`, `openssl-selfsigned-cert`, `prove-plus-comm`, and `sqlite-db-truncate` each received terminal grade 1.0 against real task-environment paths and toolchains. ## Risks - In-place mode deliberately exposes the authoritative task root to the adapter; incorrect environment metadata could point execution at the wrong root. Typed metadata and focused orchestration tests cover selection and propagation. - Copy-mode writable-path validation is stricter and may reject previously accepted unsafe configurations. The rejection is intentional and produces a visible error instead of silently losing output. - The acceptance run is focused on four Codex task-environment workloads, not a broad cross-adapter benchmark. Existing copy-mode archive tests and the full repository suite remain green. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI coding agent; exact model ID and context-window size were not exposed to this runtime. Capabilities used: extended reasoning, repository editing, shell execution, test/build execution, Git, GitHub CLI, and Paperclip API tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
30ff3d7c58 |
feat(routines): expose activity gate API (#9438)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Scheduled routines provide recurring control-plane work without manual intervention > - The new activity gate can suppress scheduled runs when no external work occurred > - The core scheduler and database support landed without a public create/update contract > - Agents, operators, and managed plugins need validated fields plus discoverable semantics to opt in safely > - This pull request exposes the activity gate through routine APIs, revisions, plugin contracts, tests, and skill documentation > - The benefit is backward-compatible control over idle scheduled work without losing activity-triggered follow-up ## Linked Issues or Issue Description - Refs #8534 ## What Changed - Added shared activity-gate policy and scope enums with create/PATCH validation. - Persisted activity-gate fields through routine creation, updates, revision snapshots, pipeline snapshots, and revision restores. - Defaulted legacy revision snapshots during restore and added regression coverage for pre-field snapshots. - Extended managed-plugin routine declarations, production reconciliation, and the SDK test harness to preserve non-default gate settings. - Added end-to-end API coverage for create/PATCH/list/detail round-trips, defaults, and invalid enum rejection. - Documented schedule-only semantics, activity windows, own-run/read-action exclusions, scopes, and an hourly quiet-night watcher example. ## Verification - `pnpm exec vitest run packages/shared/src/validators/routine.test.ts server/src/__tests__/routines-service.test.ts server/src/__tests__/routines-e2e.test.ts` - `pnpm exec vitest run packages/shared/src/validators/plugin.test.ts packages/plugins/sdk/tests/testing-actions.test.ts server/src/__tests__/plugin-managed-routines.test.ts server/src/__tests__/routines-service.test.ts -t 'activity gate|preserves declared activity gate settings|resolves routine agent and project refs'` - `pnpm exec vitest run ui/src/lib/workspace-routines.test.ts ui/src/pages/Routines.test.tsx` - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/plugin-sdk typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - GitHub CI: all final-head checks green; Storybook visual regression skipped by path rules. - Greptile: 5/5 with no unresolved review threads. ## Risks - Low risk: defaults remain `always` and `company`, preserving existing routine behavior and old revision snapshots. - Managed plugin manifests can now declare the same validated gate settings as the public routine API; omitted values retain core defaults. - Revision snapshots now include the new fields so policy changes are not lost or treated as no-ops during restore. > For core feature work, checked `ROADMAP.md`: this extends the existing Scheduled Routines roadmap item and does not duplicate a separate planned capability. ## Model Used - OpenAI GPT-5.5 via Codex CLI, with repository tool use and code execution; context-window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8f08ec5ce6 |
feat(status-cards): join summary-mentioned issues to watched set (#10205)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Status cards summarize changing company work and watch issues so later changes can produce useful deltas > - A summary can explicitly reference issues that are important to the update even when those issues do not match the card's configured queries > - Previously, those referenced issues were not retained in the watched set, so their later status, assignee, or comment changes could be missed > - The watched snapshot must avoid artificial additions or removals caused only by a summary changing which issues it references > - This pull request resolves issue references when a summary is written, persists them, and joins them to the watched snapshot with stable delta semantics > - The benefit is that status cards continue tracking the exact issues their latest update called out while keeping follow-up updates relevant and non-duplicative ## Linked Issues or Issue Description ### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am on the latest released version of Paperclip (or can reproduce on `master`). - [x] I have confirmed the error originates in Paperclip itself — not in my agent adapter, API provider, or local configuration. ### What happened? When a status-card summary explicitly referenced an issue by identifier or `/issues/<uuid>` URL, that issue was not automatically retained in the card's watched set unless it independently matched a configured query. Later status, assignee, or comment changes to an issue highlighted by the latest update could therefore be omitted. ### Expected behavior References in the latest summary should resolve only within the card's company, appear in dry runs and the watched-issues UI, count and fingerprint like query matches, and enter or leave the watched set without artificial added/removed deltas already represented by the summary change. ### Steps to reproduce 1. Create a status card whose query does not match a second issue in the same company. 2. Write a summary that references the second issue by identifier or issue URL. 3. Inspect the card's watched count or Watched issues tab. 4. Change the referenced issue's status, assignee, or comments and run the next update. 5. Before this change, the referenced issue is absent from the watched snapshot and its later change does not produce the expected delta. ### Paperclip version or commit - Reproduced on `master` before this PR (base commit `762ce5b4ef`). ### Deployment mode - Local dev (`pnpm dev`), built from source. ### Agent adapter(s) involved - Not adapter-specific (core bug). ### Database mode - Embedded Postgres test environment; the schema change uses standard PostgreSQL JSONB. ### Access context - Board (human operator). ## What Changed - Added migration `0191` and schema support for persisted `status_cards.mentioned_issue_ids`. - Resolved summary references by issue identifier or `/issues/<uuid>` URL within the status card's company when summaries are written. - Joined mentioned issues into watched counts and fingerprints so later status, assignee, and comment changes generate normal update deltas. - Suppressed artificial added/removed deltas when the latest summary starts or stops mentioning an issue. - Added `mentionedIssues` to dry-run responses and a “Mentioned in the latest update” group in the Watched issues tab. - Updated the summarizer prompt to explain that referenced issues automatically join the watched set. - Added focused server and UI coverage for reference resolution, snapshot behavior, deltas, API responses, and rendering. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/status-cards.test.ts src/__tests__/status-card-update-engine.test.ts` — 31 tests passed. - `pnpm --filter @paperclipai/ui exec vitest run src/pages/StatusCards/StatusCardTile.test.tsx` — 11 tests passed. - Earlier implementation verification also passed database/shared/server typechecks, UI `tsc -b`, the broader StatusCards UI test set, and embedded-Postgres migration application. ### Visual Verification - Greptile T-Rex ran Playwright browser checks successfully and captured the Status Card drawer Watched tab showing the new “Mentioned in the latest update” grouping: https://app.greptile.com/trex/runs/15796101/artifacts ## Risks - The migration adds a nullable JSONB column and is backward-compatible; existing cards have no mentioned issues until their next summary write. - Reference extraction is company-scoped to prevent cross-company issue association. - Watched counts and future fingerprints change for cards whose latest summaries reference issues; tests cover additions, removals, and suppression of spurious deltas. - This targeted status-card fix does not introduce a new roadmap subsystem or external integration. ## Model Used - Anthropic Claude Fable 5 (Paperclip model alias; exact underlying provider model ID and context window were not recorded in the implementation task metadata), with extended reasoning, tool use, and code execution. - OpenAI Codex coding agent (runtime model identifier and context window not exposed to this task) prepared the PR, rebased the branch, and ran focused verification with terminal tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3a16b91217 |
feat(status-cards): single-message setup drives query and update prompt (#10202)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies and their ongoing work. > - Status cards turn a standing question into recurring, agent-generated summaries on the board. > - The existing setup split intent across a watch prompt and separate update instructions, which made creation and later behavior harder to understand. > - A status card should have one durable source of truth for both deciding what to watch and telling the summarizer what each update must contain. > - This pull request makes the card prompt that source of truth, simplifies creation to one step, and lets operators choose the running agent immediately. > - The benefit is a smaller mental model, fewer configuration modes, and consistent update instructions throughout the card lifecycle. ## Linked Issues or Issue Description Status cards currently require operators to express the same intent in two places: the watch prompt and optional update instructions with append/replace/none modes. This feature simplifies the experimental status-card workflow so a single prompt defines both the watch query and every generated update. The create flow must also support selecting the responsible agent without a second setup step. Related prior status-card work: #10101. ## What Changed - Use the status card's single prompt to compile the watch query and directly instruct every summary update. - Add migration `0190_status_card_single_prompt` to remove `status_cards.instructions_mode` and `status_cards.instructions`. - Add `agentId` to `createStatusCardSchema`, validate company membership, and default new cards to the built-in Summarizer. - Replace the two-step create flow with one prompt-and-agent dialog and extract a shared `SummarizerAgentSelect` for create/settings surfaces. - Remove the extra-instructions settings section, reset incremental history when the prompt changes, and rename the board page to "Status". - Update the bundled `status-card-query` skill and board-operator documentation, then regenerate the skills catalog manifest. ## Verification - Server status-card suites: 29/29 passing. - UI `StatusCards` suites: 22/22 passing. - Skills catalog suite: 20/20 passing. - `tsc -b` passes for server, UI, shared, and database packages. - `pnpm check:migrations` passes. - Light and dark mode screenshots cover the new create dialog and settings tab. ## Risks - Migration `0190` intentionally drops existing separate instruction text. Existing card prompts remain and become the update instructions under the new model; status cards are experimental and feature-flagged. - Prompt edits now reset the incremental summary chain and trigger a full rebuild, which is intentional because the prompt is also the update contract. - Agent selection is company-scoped; invalid agent ids return a validation error rather than creating a misrouted card. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Implementation: Anthropic Claude via the `claude_local` adapter, agent label "Claude Fable 5"; extended reasoning, tool use, and code execution. The exact provider model id and context-window value were not retained in the task metadata. - PR preparation: OpenAI GPT-5.4 through Codex CLI, with reasoning, repository inspection, GitHub CLI, and Paperclip API tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
7e40ed8c43 |
feat(status-cards): add experimental status card update view (#10101)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies. > - Operators need a board-level way to monitor a changing slice of company work without repeatedly rebuilding filters or reading raw task threads. > - Existing summaries are useful snapshots, but they do not provide a dedicated query-backed card with refresh policy, change tracking, update history, and per-update cost visibility. > - The capability needs to be safe to evaluate before it becomes part of the default product surface. > - This pull request adds end-to-end experimental Status Cards, from schema and query compilation through update orchestration and operator UI. > - The entire feature is gated behind the `enableStatusCards` experimental toggle, including its route and sidebar entry. > - The benefit is a governed, inspectable way to keep focused operational rollups current while preserving explicit controls over refresh frequency and spend. ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting (`packages/db`, `packages/shared`, `server/`, `ui/`, and bundled skills/docs). ### Problem or motivation Operators cannot currently define a reusable natural-language view of company work, compile it into an inspectable query, and keep its summary current as matching issues change. Rebuilding filters and rereading task threads makes board-level monitoring repetitive and hides the relationship between source changes, refresh cost, and the resulting summary. ### Proposed solution Add experimental Status Cards that compile operator intent into a query, summarize matched work, record each update, expose manual/interval/reactive refresh policies and costs, and preserve the last good result across stale, updating, paused, and error states. The capability is off by default and fully gated behind `enableStatusCards`, including its route and navigation entry. ### Alternatives considered - Extend existing one-off summaries: rejected because status cards require persistent query provenance, refresh policy, update history, and card-specific cost controls. - Add a dashboard-only filter widget: rejected because it would not provide governed background refresh, an update ledger, or an inspectable compile pipeline. - Ship the surface by default: rejected in favor of an experimental toggle while behavior and operator value are evaluated. ### Roadmap alignment This advances Paperclip’s board-level execution visibility and output-first product goals. `ROADMAP.md` was checked and no duplicate status-card initiative was found. ### Additional context No related open PR was found in the public GitHub search for status cards. The PR-only design wireframes were removed from the repository after review; the published prototype remains external to the production source tree. ## What Changed - Added company-scoped status-card schema, CRUD APIs, compile provenance, update ledger, shared contracts, validators, and OpenAPI coverage. - Added the text-to-query compile pipeline, bundled `status-card-query` agent skill, query versioning, and authorized write-back flow. - Added the experimental board, create flow, lifecycle tiles, detail/settings/debug drawers, archived view, routing, navigation, and instance setting. - Added a change-gated update engine with manual, interval, and reactive refresh policies, trigger selection, active hours, and daily token caps. - Added per-update token/cost recording, today and lifetime rollups, and policy-derived cost previews. - Added operator documentation and agent-authoring hardening for compile and update behavior. - Added PR-prep integration coverage for settings/startup wiring and replaced raw UI values with design-system tokens. - Removed the PR-only `design/pap-15023-status-cards` wireframe artifacts so the repository contains only production feature assets. ## Verification - `pnpm -r typecheck` — passes on the PR head; includes `ui` `tsc -b` passing. The UI compile gate was also independently recorded as passing at `6d7f3cf96b` on July 23, 2026. - `pnpm build` — passes. - `pnpm check:token-gates` — passes with all three gates clean. - `pnpm test:run` — 2,880 tests passed and 1 skipped; the sole failure was an unrelated 10-second `afterAll` database-cleanup timeout in `execution-workspaces-service.test.ts`. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/execution-workspaces-service.test.ts` — passes on immediate focused rerun (25/25). - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/instance-settings-service.test.ts src/__tests__/server-startup-feedback-export.test.ts` — passes (31/31). - `pnpm --filter @paperclipai/ui exec vitest run src/pages/StatusCards/StatusCardSettingsForm.test.tsx src/pages/StatusCards/StatusCardTile.test.tsx src/pages/StatusCards/format.test.ts src/lib/status-card-state.test.ts` — passes (26/26). - Recorded pre-PR QA: compile-pipeline e2e PASS; full lifecycle and cost QA PASS; security re-review PASS after write-back hardening; UX approved. - `pnpm exec vitest run packages/db/src/status-card-migrations.test.ts` — passes; reapplies migrations `0185`–`0189` against an already-migrated embedded Postgres database. - `pnpm --filter /db check:migrations` — passes migration numbering and safety checks. - `pnpm --filter /db typecheck` — passes. - Merged current `origin/master` on July 24, 2026 with no conflicts; migrations `0185`–`0189` remain unclaimed on master. ## Risks - The feature introduces five database migrations and a new background update path; all new DDL is repeat-safe after partial application, migration numbering/safety checks pass, and update execution is company-scoped and change-gated. - Natural-language compilation can produce invalid or overly broad queries; compile provenance, query validation, debug visibility, and version history make failures inspectable and recoverable. - Reactive or interval refresh could increase spend; active hours, max refresh frequency, daily token caps, per-update cost records, and budget-paused states bound and expose that risk. - The branch name contains an internal execution identifier because it is a fixed handoff branch; it was intentionally not renamed or rebased per the release handoff instructions. - Overall rollout risk is limited because the route, navigation, services, and UI are disabled by default behind `enableStatusCards`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using GPT-5.5 with reasoning, repository tool use, shell execution, GitHub CLI, and test/build execution. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change; the fixed execution-workspace identifier is documented as an authorized handoff exception - [x] I have run tests locally and they pass, with the one cleanup timeout passing on focused rerun - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7f766526a6 |
feat(sandbox): add task-scoped egress grants (#10155)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Confinement providers protect agent runs with default-deny network policies > - Kubernetes environments currently apply only provider-level, namespace-wide egress allowances > - Tasks that legitimately need GitHub or package registries therefore cannot request narrow access, while network failures do not explain the governing policy or how to request a grant > - This pull request adds issue-scoped egress grants that become workload-owned, run-label-selected policies and carries the effective grant through lease audit metadata > - The benefit is that internet-dependent work can run without enabling broad egress for every concurrent task, and denied requests point operators to the exact grant path ## Linked Issues or Issue Description No public issue exists. Related but distinct: Refs #9944, which adds a provider-wide open-internet posture; this PR keeps provider defaults narrow and adds per-task grants. **Problem / motivation** Kubernetes sandbox egress is configured at the provider/tenant level. A task that needs to clone from GitHub or install from PyPI cannot request those destinations without changing the policy for every run in the tenant namespace. DNS/connectivity failures also surface as generic tool errors with no policy name or remediation path. **Proposed solution** Accept `executionWorkspaceSettings.networkEgress.allowFqdns` and `allowCidrs`, forward the setting through heartbeat environment acquisition, and create a workload-owned NetworkPolicy or CiliumNetworkPolicy selected by `paperclip.io/run-id`. Record the effective grant in lease activity/metadata, expose policy context through `PAPERCLIP_NETWORK_EGRESS_*`, and append the grant path to likely policy-related stderr failures. **Alternatives considered** A provider-wide open-internet switch is broader than required and is already covered by #9944. Mutating the existing namespace policy would leak each task's destinations to other concurrent runs. Standard Kubernetes NetworkPolicy cannot enforce FQDNs exactly, so standard mode uses the existing hardened public-IPv4 TCP 80/443 fallback only for the selected run; Cilium mode remains exact. **Roadmap alignment** This extends the existing cloud/sandbox agent roadmap capability with task-level control-plane policy and does not duplicate a planned roadmap item. ## What Changed - Added validated `networkEgress` grants to issue execution workspace settings and forwarded them through environment lease acquisition. - Added workload-owned, run-label-scoped NetworkPolicy/CiliumNetworkPolicy resources for task FQDN/CIDR grants. - Added lease audit metadata, sandbox policy environment variables, and actionable network-denial stderr guidance. - Added focused parser, manifest, policy creation, and denial-message tests plus Kubernetes provider documentation. ## Verification - `pnpm -C packages/shared exec vitest run src/validators/issue.test.ts` — 27 passed. - `pnpm -C packages/plugins/sandbox-providers/kubernetes test -- --run test/unit/network-policy.test.ts test/unit/cilium-network-policy.test.ts test/unit/scoped-network-egress.test.ts` — 21 passed. - `pnpm -C server exec vitest run src/__tests__/execution-workspace-policy.test.ts` — 15 passed. - `pnpm exec vitest run server/src/__tests__/heartbeat-plugin-environment.test.ts server/src/__tests__/environment-runtime.test.ts` — 26 passed. - `pnpm --dir packages/db build && pnpm --dir packages/shared build && pnpm --dir packages/plugins/sdk build` — passed, including migration safety checks. - `pnpm --dir packages/plugins/sandbox-providers/kubernetes typecheck && pnpm --dir server typecheck` — passed after refreshing the worktree's frozen offline dependencies. - End-to-end cluster validation of the `build-cython-ext` benchmark remains for CI/maintainer Kubernetes infrastructure; the focused tests assert `github.com` and `pypi.org` produce a policy selected only by the granted run. ## Risks - Standard NetworkPolicy cannot express FQDNs, so an FQDN grant allows hardened public IPv4 TCP 80/443 for that run; use Cilium mode for exact hostname enforcement. - The new field is additive and absent by default, so existing runs keep the current provider-level policy. - Workload owner references garbage-collect scoped policies with the Job/Sandbox; a cluster/controller that ignores owner references could temporarily strand a policy that still selects no future run ID. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, exact model ID `gpt-5.6-sol`, high reasoning mode, tool use and code execution. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
148a5b11f5 |
Route blocked transitions to explicit unblock owners (#10112)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent companies. > - Issue status transitions determine whether work keeps moving or silently stalls. > - A blocked issue previously could rely on prose alone, leaving the intended unblock owner unstructured and unnotified. > - Existing blocker-attention classification could identify stalled chains, but the signal was not delivered to the board attention feed. > - Blocked transitions also need rollout-safe deduplication so upgrades do not notify for historical issues and repeated processing does not create notification storms. > - This pull request adds structured unblock descriptors, prospective transition timestamps, owner delivery, and board attention routing with focused authorization controls. > - The benefit is that newly blocked work has an explicit, routable next action without weakening company boundaries or allowing agents to inject arbitrary human attention items. ## Linked Issues or Issue Description Related documentation PR: #10094. ### Subsystem affected Cross-cutting: `server/`, `packages/db`, and `packages/shared`. ### Problem or motivation An issue can enter `blocked` without a machine-readable unblock path. Prose-only ownership does not reliably wake the responsible agent or surface human-owned work, while the existing `blockerAttention` classifier is not delivered to an operator-facing attention feed. ### Proposed solution Require new transitions into `blocked` to have unresolved blockers, a pending interaction/approval, or a structured `{ owner, action }` descriptor. Notify an allowed owner once per prospective transition, route human-owned cases to board attention, and leave pre-rollout blocked issues untouched. ### Alternatives considered - Keep prose-only blockers: rejected because ownership remains unroutable. - Backfill all historical blocked issues: rejected because upgrades would create notification storms. - Let agents target arbitrary users or the board: rejected after security review because it creates an attention-injection channel. ### Roadmap alignment Aligns with `ROADMAP.md` → “Enforced Outcomes (watchdogs, recovery actions, review gates)” by making blocked work carry an explicit continuation path. ### Additional context The implementation is prospective-only and deduplicated per blocked transition. Agent-authored descriptors are limited to the acting agent; board actors retain human-owner routing. ## What Changed - Added persisted unblock descriptors and prospective blocked-transition delivery timestamps with an idempotent migration. - Added shared types and validation for board, user, and agent unblock owners. - Enforced valid blocked transitions and same-company owner validation in the issue update route. - Restricted agent-authored descriptors to the acting agent itself, preventing board/user attention injection by compromised agents. - Added one-per-transition agent wake delivery and prospective-only rollout gating. - Routed human-owned blocker attention into the board attention feed. - Added focused tests for validation, prospective delivery, flap deduplication, attention routing, route authorization, and stop-relay compatibility. ## Verification - `pnpm -r typecheck` - `pnpm exec vitest run server/src/__tests__/issue-agent-mutation-ownership-routes.test.ts server/src/__tests__/routable-blocked.test.ts server/src/__tests__/attention-service.test.ts packages/shared/src/validators/issue.test.ts` - `AWS_ACCESS_KEY_ID= AWS_SECRET_ACCESS_KEY= pnpm test:run` - `pnpm build` - `pnpm --filter @paperclipai/db check:migrations` ## Risks - Behavioral shift: new `blocked` transitions without a real blocker, pending governed action, or structured descriptor now return `422`. - Notification abuse is constrained by same-company validation, agent self-only routing, prospective rollout gating, and transition-scoped deduplication. - Migration risk is low: columns are additive, nullable, and use `IF NOT EXISTS`; historical blocked issues are not backfilled or notified. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI with GPT-5.4, reasoning-enabled tool use and code execution. The runtime did not expose a context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5a00040c4c |
feat(plugin-sdk): interactions/approvals respond + attachment read capabilities for the chat gateway (#10066)
Adds the remaining 5 plugin capabilities + 7 worker→host RPC methods (interactions read/respond, approvals read/respond, attachment read) needed by the Slack chat gateway plugin (v0.5.0) to pass manifest capability validation and load. - Security review: PASS (LOOA-642) after the viewer-role privilege-escalation blocker (LOOA-648) was fixed on this branch (requireActiveHumanMember now rejects viewer on impersonation write-paths, matching assertCompanyAccess). - CI: Build, Typecheck, all server suites (3/3 + serialized 4/4), workspaces, e2e shard 2/2, and all security scanners (Snyk/Socket/Superagent/Greptile/security-review) green. - One e2e flake (signoff-policy 'non-participant cannot advance stage') is unrelated: it exercises execution-policy stage advancement (routes/issues.ts, untouched by this PR) and failed on a heartbeat_run_events FK race + 409 checkout conflict. Unblocks LOOA-629 (Slack gateway go-live) and the interview-ask feature. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
216d3d2680 |
Managed-instance config: fail-closed PAPERCLIP_MANAGED_CONFIG parsing and read-time settings overlay (#10058)
**Builds on.** #10055 — the `catalogVersion` this config document pins is the feature-catalog artifact #10055 emits. **Summary.** Instances operated by a managed hosting control plane can now receive instance configuration through a single environment variable, `PAPERCLIP_MANAGED_CONFIG` (versioned JSON: `mode`, `catalogVersion`, `features`, `plugins.autoInstall`). When the variable is absent the instance is self-hosted and nothing changes. When present, parsing is strict and **fail-closed**: blank value, malformed JSON, unknown feature key, a feature key this build's feature catalog does not mark tier `managed`, missing required section, or unsupported version refuses startup with a precise error — a typo that silently does nothing is how a security control quietly fails. Managed feature values are overlaid **at read time** inside the instance settings service (never persisted), so a DB restore or manual row edit cannot resurrect a disabled capability; responses expose per-key `managedKeys` metadata (`managed: true`, `managedBy`) so clients can render locked state. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs both self-hosted and under managed hosting, where an operator's control plane owns instance configuration > - Today instance feature settings live only in the tenant database; a hosting control plane has no way to enforce a configuration that tenant-side writes or restores cannot undo > - Managed configuration will carry security posture, so delivery must be atomic and parsing must fail closed — a typo that silently does nothing is how a security control quietly fails > - This pull request adds strict parsing of one `PAPERCLIP_MANAGED_CONFIG` env var and overlays its feature values at read time inside the settings service, never persisting them > - The benefit is a minimal, auditable managed-hosting contract: absent var ⇒ self-hosted instances are byte-for-byte unchanged; present ⇒ deterministic, locked configuration surfaced to clients via per-key managed metadata ## Linked Issues or Issue Description Refs #966 — this PR delivers that issue's "managed config injection" hook, via a strict env-var contract rather than the config-file path it sketches; the issue's other hooks (identity header, health, usage webhook, lifecycle, external secrets, IAM auth) are out of scope, so the PR refs rather than closes it. *Mechanism differs from #966's proposal, so the `feature_request` fields are also filled in:* - **Problem or motivation:** managed hosting deployments need to centrally enable/disable instance features; DB-stored settings can be edited, restored, or migrated back to permissive values, and nothing marks a value as operator-enforced. - **Proposed solution:** one versioned JSON env var; fail-closed parse at startup; read-time overlay in the settings service (precedence: managed value over stored value over schema default); `managedKeys` metadata in settings responses so clients can render locked state. - **Alternatives considered:** per-feature env vars (non-atomic across a half-updated env set, unbounded env surface); seeding the DB at boot (persisted values can be edited or restored over, and cannot express "forced"); lenient warn-and-drop parsing (fails open — unacceptable for a security-bearing control). - **Roadmap alignment:** supports the in-progress "Cloud deployments" milestone in `ROADMAP.md`. ## What Changed - New `server/src/services/managed-config.ts` (pure parser over the env record) - Startup parse ordered before the first `instanceSettingsService` construction in `server/src/index.ts` - Read-time merge + `managedKeys` in the settings service - Shared validator updates ## Verification - 29 parser/overlay tests (fail-closed matrix incl. blank/whitespace env, missing sections, catalog-tier mismatch, empty-section happy path): `pnpm vitest run src/__tests__/managed-config.test.ts src/__tests__/instance-settings-managed-overlay.test.ts` (from `server/`) - 40 existing settings route/service tests green: `pnpm vitest run src/__tests__/instance-settings-routes.test.ts src/__tests__/instance-settings-service.test.ts` (from `server/`) - 15 shared validator tests: `pnpm vitest run src/validators/instance.test.ts` (from `packages/shared/`) - Server `tsc --noEmit` clean: `pnpm typecheck` (from `server/`) ## Risks - Self-hosted instances (no `PAPERCLIP_MANAGED_CONFIG` set) are byte-for-byte unchanged — the parser only runs when the variable is present. - For managed instances, a malformed document now refuses startup by design (fail-closed). This is an intentional behavioral guarantee, not a regression: the control plane owns the variable and a precise startup error is the contract. - Overlay values are never persisted, so no migration or data-shape risk. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use; independently peer-reviewed by a second AI agent before push. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ad74fb5450 |
Add a feature catalog build artifact derived from the experimental settings schema (#10055)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Instances expose ~23 experimental feature settings, all declared in one shared zod schema and toggled per instance > - Deployment tooling and hosting control planes have no machine-readable list of those feature keys for a given release — the schema is only reachable from code that imports the package > - Any external system that references feature keys therefore does so as free text, and typos drift silently > - This pull request derives a versioned `feature-catalog.json` build artifact from the schema, with a compiler-checked metadata map so the schema stays the single source of truth > - The benefit is a stable contract external tooling can validate feature-key references against, with zero runtime behavior change ## Linked Issues or Issue Description No public issue exists; `feature_request` template fields: **Problem or motivation:** External deployment tooling cannot enumerate or validate an instance's feature keys per release; free-text references fail silently when keys are renamed or removed. **Proposed solution:** A metadata map keyed by the settings schema's own keys (compiler flags drift) plus a build step emitting `feature-catalog.json` (keys, tiers, defaults, `catalogVersion`) as a release artifact. **Alternatives considered:** A hand-maintained catalog file (drifts from the schema); serving the schema from a runtime API (requires a running instance at validation time — a build artifact works offline and pins to a release). **Roadmap alignment:** Supports the in-progress "Cloud deployments" milestone in `ROADMAP.md`. ## What Changed Adds a metadata map (title, description, tier, cloud/self-hosted defaults) keyed by the keys of `instanceExperimentalSettingsSchema`, so the schema stays the single source of truth and the compiler flags any drift. A new build step (`build:feature-catalog --version <v>`) emits `feature-catalog.json` — all 23 feature keys, their tiers, and a `catalogVersion` — as a release artifact that managed-hosting control planes can validate feature-flag writes against. No runtime behavior changes. - New `packages/shared/src/feature-catalog.ts`: per-flag metadata map keyed by a type derived from the settings schema (adding/removing/renaming a flag without updating the map is a compile error), plus `featureCatalogArtifactSchema` and `buildFeatureCatalogArtifact`/`renderFeatureCatalogArtifact` for the artifact - New `scripts/generate-feature-catalog.ts` wired as `pnpm build:feature-catalog --version <v>` - `scripts/create-github-release.sh` generates the artifact and uploads it as a GitHub Release asset (with a dry-run preview line) - Tests in `packages/shared/src/feature-catalog.test.ts` ## Verification - `vitest run packages/shared/src/feature-catalog.test.ts` — 9 tests: schema-key coverage, drift detection, artifact shape - `pnpm --filter @paperclipai/shared typecheck` - Artifact generation run end-to-end: `pnpm build:feature-catalog --version 0.0.0-test` emits 23 keys with `catalogVersion` ## Risks Low risk — no runtime behavior changes; the change is metadata, a build script, and a release-artifact emission step only. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use; independently peer-reviewed by a second AI agent before push ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e55d702916 |
feat(plugin-sdk): human-attributed issue comments for chat gateway plugins (#10050)
Adds the `issue.comments.create_human_attributed` capability and `ctx.issues.createComment` `actorUserId` option, with host-side active-human-member verification. LOOA-627. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
0b496c9c03 |
feat(secrets): add run-bound agent secret access (#9921)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies. > - Agents already receive selected company secrets through `env.*` bindings at run launch, but environment injection is ambient, long-lived, and not suitable for every secret consumer. > - The existing binding and secret-access-event models already provide company-scoped authorization and per-resolution audit seams. > - Agents need an explicit way to discover only the secrets granted to them and fetch a value on demand without exposing the wider company catalog. > - That capability must remain run-bound, preserve low-trust token carve-outs, and make every value read visible in both security and operator audit trails. > - This pull request adds an `access.*` delivery namespace, two run-bound agent routes, dual audit logging, documentation, and an operator grants editor. > - The benefit is least-privilege, revocable, auditable secret access while preserving existing env injection behavior. ## Linked Issues or Issue Description No pre-existing public issue. Related work: - Refs #9797 — existing in-sheet agent access UI that this PR extends to distinguish env and API delivery. - Refs #9918 — complementary searchable-agent picker improvement for the same secrets sheet. - Refs #9530 — related company-wide metadata catalog proposal; this PR intentionally exposes only the authenticated run's granted aliases and values. **Problem / motivation:** Agents can currently consume secrets only through process environment injection. This keeps values resident for the run, does not support on-demand consumers, and cannot provide a discrete operator-visible activity event for each agent-initiated read. **Proposed solution:** Treat `company_secret_bindings` as the source of truth for agent secret grants. Keep `env.KEY` as env delivery and add `access.ALIAS` for API-only delivery; an env binding also implies read access because the value is already present in the agent process. Add run-bound list/fetch endpoints that derive scope from the authenticated heartbeat run and never accept caller-selected overlays. **Alternatives considered:** A company-wide agent-readable catalog was rejected for this value path because it increases reconnaissance and does not prove a per-secret grant. Reusing the ephemeral environment-probe resolver was rejected because it lacks binding enforcement. Approval-gated reads and user-scoped secrets remain deferred beyond v1. **Roadmap alignment:** This extends the completed **Secrets Manager with per-agent access** roadmap capability from launch-time env injection to explicit run-bound API delivery without duplicating a separate planned initiative. ## What Changed - Added `access.*` agent binding validation and a dedicated run-bound resolver that combines `secrets:read` authorization with binding-context enforcement. - Added `GET /api/agents/me/secrets` for minimal granted metadata and `POST /api/agents/me/secrets/:key/value` for on-demand value fetches with `Cache-Control: no-store`. - Preserved the existing denials for low-trust review agents, task-bridge credentials, and skill-test tokens; standard long-lived agent API keys cannot call the run-bound routes. - Added dual audit behavior: value attempts write `secret_access_events` and `activity_log` (`secret.value.read`), while metadata listing writes the lighter `secret.access.listed` activity event. - Kept env compatibility: `env.*` remains injected at launch and also implies API read for the same bound agent; `access.*` never becomes an environment variable. - Added the agent-settings **Secret access** editor plus delivery-mode/alias surfacing on the Secrets page, with focused UI tests and tokenized layout styles. - Updated OpenAPI, shared types, agent-facing skill documentation, and API reference documentation. ### UI Screenshots P3 produced and reviewed three screenshots using mock data; images are intentionally not committed to the repository: - `secret-access-editor.png` — agent settings grant editor. - `secret-access-light.png` — Secrets-page delivery surfacing in light mode. - `secret-access-dark.png` — Secrets-page delivery surfacing in dark mode. The source attachments are retained with the implementation task and linked in the internal handoff; the public page publisher was unavailable in the PR-prep runtime. ## Verification - `pnpm exec vitest run server/src/__tests__/agent-secrets-routes.test.ts server/src/__tests__/secrets-service.test.ts server/src/__tests__/secrets-routes.test.ts ui/src/lib/secret-delivery.test.ts ui/src/components/AgentSecretAccessEditor.test.tsx` — 5 files, 122 tests passed. - Security follow-up: `pnpm exec vitest run server/src/__tests__/agent-secrets-routes.test.ts server/src/__tests__/secrets-service.test.ts` — 2 files, 73 tests passed after active-run and version-consistency fixes. - Final-head CI: all feature, typecheck, build, e2e, security, and review gates pass; `General tests (server (1/3))` remains red after one rerun because unrelated `heartbeat-retry-scheduling.test.ts` cleanup deletes `heartbeat_runs` before referenced `activity_log` rows. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm check:token-gates` — feature-local arbitrary-value violations fixed; command still reports five unchanged `#9627` literals outside this PR. - End-to-end QA passed all eight acceptance criteria: grant/list, fetch, dual audit, env-implies-read, denial matrix, revocation, UI rendering, and env-injection regression. Evidence: https://github.com/paperclipai/paperclip/pull/9921#issuecomment-5027455492 - Security review returned PASS-with-required-changes; the implementation uses the required dedicated binding-enforcing resolver, run-bound JWT restriction, run-derived overlays, minimal metadata, and a resolver redaction-registration hook. Evidence: https://github.com/paperclipai/paperclip/pull/9921#issuecomment-5027455382 ## Risks - A compromised agent can exfiltrate any secret explicitly granted to it; explicit company-scoped/run-scoped grants, revocation, and audit reduce but cannot remove that inherent capability risk. - The resolver invokes a redaction-registration hook before returning values, but the current route has no persistent cross-request per-run redaction registry. Paperclip-owned later comments/events therefore cannot yet guarantee automatic scrubbing of a deliberately copied fetched value; QA classified this as non-blocking residual hardening. - Audit-event insertion currently fails open if the security-event insert itself fails; the operator activity event provides partial redundancy, but a future hardening change should define fail-closed behavior for value delivery. - This PR overlaps `ui/src/pages/Secrets.tsx` with #9918 and may require a straightforward rebase after that PR moves. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, `gpt-5.3-codex`, with reasoning, repository tool use, terminal execution, Paperclip API access, and GitHub CLI capabilities. Context-window size is not exposed by the runtime. - Anthropic Claude Opus 4.8 with 1M context and tool use assisted with the UI implementation commit. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
cac3c0fa1a |
feat(connections): add runtime subjects and grants (#9982)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Connections is the subsystem that lets operators connect external apps and govern which subjects may use those credentials > - #9958 established the v3 schema foundation and #9981 adds the AppDefinition catalog layer > - The runtime still needs subject-aware authorization state, scoped key handling, and API/OpenAPI routes so connected apps can actually be granted and used safely > - This pull request adds the runtime grants/authorization behavior on top of the catalog branch, while keeping unrelated dependency and workflow sync commits out of the stack > - The benefit is a reviewable runtime layer that can land after the catalog PR, then unblock the wizard and orchestrator cutover work ## Linked Issues or Issue Description Refs #9958 and #9981. Refs #9981. No public GitHub issue exists for this branch. This is the runtime layer for the Connections v3 stack and is rebased onto `master` after #9981 landed. ## What Changed - Adds the connection user authorization state migration and schema wiring. - Adds shared runtime subject/grant types and validators. - Adds runtime grant and scoped key behavior in the tool-access service. - Adds runtime route coverage and registers the routes in OpenAPI. - Replays only the Connections runtime commits on top of the catalog branch, dropping unrelated sync/dependency history from the prior closed runtime PR. ## Verification - `pnpm run preflight:workspace-links` - `pnpm exec vitest run packages/shared/src/validators/tool-access.test.ts server/src/__tests__/tool-access-service.test.ts` ## Risks - Medium: runtime grant enforcement is security-sensitive and must fail closed for unknown key scopes. - Migration ordering depends on the schema and catalog layers already merged through #9958 and #9981. - This PR is rebased and retargeted to `master` with runtime-only commits. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex coding agent with repository tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |