mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 06:15:21 +02:00
codex/plugin-task-execution
18
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
16b7db35ff |
Shorten planning skills and measure task decomposition (#15296)
## Thinking Path > - Paperclip manages work for AI agents. > - Planning guidance helps agents choose owners and dependencies. > - The runtime skill favors few tasks, but the catalog skill requires a child-task breakdown. > - Both add repeated process instructions that can distract from the requested outcome. > - This change keeps the ownership and dependency rules and removes the required matrix and repeated checklist. > - A bounded Product E2E comparison measures saved outcomes and task handoffs before qualification. ## Linked Issues or Issue Description Refs #11057. Related measurement work: #15218. **What existing behavior does this improve?** Planning and delegation through the runtime plan-to-tasks and bundled task-planning skills. **Current behavior** The two skills contain about 1,900 words and conflicting guidance on whether plans require child tasks. **Proposed behavior** Keep cohesive work with one owner. Split only for a real owner, parallel output, dependency, independent review, or follow-up lifecycle. Preserve existing authorization and planning mechanics. ## What Changed - Shorten both skills to about 400 words combined. Preserve their keys and installed-version behavior. - Remove the duplicate operational-skill pointer and regenerate affected source metadata. - Add twelve explicit Product E2E cells: four scenarios with current, short and disabled planning skills. - Use the current task composer and actual create-response ID; calibrate public skill APIs and browser creation without providers. - Eliminate an observed collision in chat-test company prefixes with a per-suite sequence. - Grade saved documents, exact author/run attribution, child count, prerequisite execution order, review boundaries and completion handoffs. - Retain current skill bytes and report source, selections, run accounting and failures. ## Verification - `pnpm test:e2e:runner:typecheck`: pass. - `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks pass. - `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve local Codex cells. - Archived current skills match master `72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly. - Corrected fixture: three real public-API/database calibrations pass with zero provider runs; all 35 evaluator checks and Product E2E typecheck pass. - Setup campaign [37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253) was canceled after source review found unsupported bundled edits and automatic core reinstallation. Its paid-cell step was skipped: zero provider runs, no behavioral grade. - The next setup [37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799) failed before task creation on the old title-field selector: zero actual runs, original FAIL retained, cleanup passed. A real browser/API calibration of the new helper passes with paused non-provider agents and zero runs. - Full local typecheck/build pass. Full local tests retain one unchanged five-minute Git streaming timeout (also fails isolated), 9,591 passes and 5,796 skips. CI's chat failure was a proven random fixture-prefix collision; five affected cases pass after the test-only repair. - Paid behavior comparison and new-head CI/review remain pending. This PR remains a draft. ## Risks - The shorter text may change delegation decisions. Live outcomes are not yet qualified. - The initial comparison uses one profile and one attempt per cell. It cannot establish cross-model reliability or cost trends. - Disabled means unassigned company-owned copies; the company library remains discoverable. This does not qualify global removal, automatic accepted-plan wiring changes, or installed-copy migration. - Skill availability does not prove a model read or cognitively used it. - No provider/tool protocol, permission, timeout or runtime lifecycle behavior changes in production. ## Model Used OpenAI Codex (GPT-6), with repository inspection, code editing and tool use. The exact backend model ID and context-window size are not exposed in this session. The declared eval model is native Codex `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cabc9146d0 |
feat(apps): add secure remote MCP and PostHog setup (#12339)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps give those agents governed access to external tools. > - Remote MCP setup needs secure endpoint validation and durable credentials. > - PostHog needs both browser sign-in and personal API key setup paths. > - This pull request adds the shared remote MCP foundation and the PostHog definition. > - The benefit is a secure and reusable base for later app connection work. ## Linked Issues or Issue Description Refs #11965 This is stack 1 of 11. It replaces the first reviewable part of #11965. ## What Changed - Add guarded remote MCP setup and credential handling. - Add PostHog OAuth and API key connection methods. - Add focused server, shared contract, and UI coverage. - Keep the migration replay-safe and idempotent. - Give the late-close security regression the same 10-second CI headroom as the adjacent real-timer handshake test. - Synchronize fake-timer handshake tests at the exact ensure-session boundary so real filesystem setup cannot race the fake deadline. - Drive PTY overflow coverage only after listener registration so scheduling cannot reorder the test fixture. ## Verification - pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts server/src/__tests__/plugin-worker-manager.test.ts (220 passed; affected cases also passed five focused stress repetitions) - `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts -t "never leaks a sandbox-provided value from a late close rejection into logs or the result"` (1 passed) - `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts -t "never promotes a late ensureSession resolution|closes a late-resolving real handle exactly once"` (2 passed) - `pnpm -r typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - `pnpm --filter @paperclipai/db check:migrations` - `pnpm build` ## Risks - Remote endpoint validation can reject configurations that previously passed without checks. - OAuth configuration errors can block setup until the operator corrects the provider settings. - The migration uses guarded statements so repeated execution is safe. - The test-only synchronization changes do not affect runtime behavior; they remove filesystem/fake-clock and listener-registration races observed under parallel CI load. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b58ce27a02 |
fix: isolate execution workspace summaries (#10790)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip gives operators a summary for each workspace. > - An execution workspace detail page used the parent project-workspace summary slot. > - Two execution workspaces under one project workspace could therefore show the same summary. > - This pull request gives each execution workspace its own summary scope. > - It also limits the summary snapshot and generated issue to that execution workspace. > - The benefit is that a new or parallel execution workspace cannot inherit unrelated status. ## Linked Issues or Issue Description **What happened?** An execution workspace detail page read and refreshed the summary slot for its parent project workspace. Parallel execution workspaces could show the same status and include issues from each other. **Expected behavior** Each execution workspace must have one isolated summary slot. Its generated snapshot must include only issues assigned to that execution workspace. **Steps to reproduce** 1. Create two execution workspaces under one project workspace. 2. Add different issues to each execution workspace. 3. Generate the summary in the first execution workspace. 4. Open the second execution workspace. 5. Observe that the old implementation could reuse the first summary. **Paperclip version or commit** The problem exists on `master` before this pull request. **Deployment mode** The issue affects both local trusted and authenticated deployments. ## What Changed - Added `execution_workspace` to the shared summary-slot scope contract. - Validated execution-workspace ownership and stored generated summary issues on the correct execution workspace. - Limited execution-workspace snapshots to issues with the matching execution workspace ID. - Updated the execution workspace page to use its own summary slot. - Updated Summarizer instructions, routine options, catalog metadata, documentation, and regression tests. ## Verification - `NODE_ENV=test pnpm exec vitest run packages/shared/src/summary-slot.test.ts server/src/__tests__/summary-slots.test.ts ui/src/pages/ExecutionWorkspaceDetail.test.tsx` — 30 focused tests passed; the embedded-Postgres server tests were run outside the process-restricted sandbox. - `pnpm check:token-gates` — passed. - `pnpm --filter @paperclipai/skills-catalog validate` — passed with 17 catalog skills. - [Latest-head GitHub Actions](https://github.com/paperclipai/paperclip/actions/runs/31491475405) — all 22 jobs passed on `beea14cbaf`, including typecheck, build, server/workspace tests, serialized suites, e2e, canary, and aggregate verification. One unrelated adapter cleanup test initially hit an `ENOTEMPTY` temp-directory race; its single permitted rerun passed. - Greptile — 5/5 confidence on `beea14cbaf`, 12 files reviewed, zero comments added, and zero unresolved threads. ## Risks - Low risk. The new scope is additive. - Existing project and project-workspace summary slots keep their current keys and behavior. - A summary generated for an execution workspace now excludes sibling workspace issues by design. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The deployment does not expose a more specific model ID or context-window value. It used agentic reasoning, repository tools, code execution, and GitHub tooling. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a71b9cf628 |
feat(skills): add MCP integration preparation skill (#11063)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Skills Store lets a company find and install reusable agent procedures > - The MCP integration preparation procedure existed outside the app catalog > - Paperclip users could not find or install that procedure from the product > - This pull request adds the procedure as an optional software development skill > - The skill keeps research, human approval, and connector delivery as separate gates > - The benefit is a repeatable and governed path from vendor research to one connector pull request ## Linked Issues or Issue Description **What existing behavior does this improve?** The app-shipped Skills Store can install optional skills, but it does not include the MCP integration preparation workflow from `paperclip-content`. **Subsystem affected** `packages/skills-catalog`. **Current behavior** An agent must know where the external workflow lives. The agent cannot find or install it from the Paperclip skills catalog. **Proposed behavior** The optional catalog includes `prepare-mcp-integration`. The installed skill directs agents through cited research, a research-only content pull request, an exact-revision human gate, and one Paperclip connector pull request per approved connection. **Reason and benefit** This change makes the existing integration and connector playbooks available as one installable Paperclip workflow. It also prevents agents from starting connector code before the research gate is approved. **Breaking changes** None. The skill is optional and markdown-only. Related source work: paperclipai/paperclip-content#13. ## What Changed - Add the optional `prepare-mcp-integration` catalog skill under software development - Add Paperclip catalog metadata for roles, requirements, tags, and trust classification - Add a worked Notion MCP research-gate example - Regenerate the checked-in catalog manifest - Add the new key to the shipped optional skill test ## Verification - `pnpm --filter @paperclipai/skills-catalog build:manifest` - `pnpm --filter @paperclipai/skills-catalog validate` - `pnpm --filter @paperclipai/skills-catalog test` - Confirm the generated catalog contains `paperclipai/optional/software-development/prepare-mcp-integration` - Confirm the trust level is `markdown_only` and compatibility is `compatible` ## Risks - Low risk. This change adds one optional markdown-only catalog entry. - The workflow can become stale if the two upstream playbooks change. The skill requires agents to read the current playbooks before each phase and to update upstream rules when reusable specifications change. ## Model Used OpenAI Codex with GPT-5.4, reasoning mode, shell tool use, and code editing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
efe8b3b707 |
feat(skills-catalog): add optional /simplified-english skill (#10410)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents draw writing behavior from installable skills in the shipped skills catalog (`packages/skills-catalog`) > - Agent-authored comments, plans, and documents are often wordy or ambiguous, which slows human readers > - A small, opt-in skill can set a clear house style for user-facing prose without touching agent code > - This pull request adds an optional `simplified-english` catalog skill that tells agents to write user-facing comments, plans, and documents in ASD-STE100 Simplified Technical English > - The benefit is shorter, unambiguous, one-meaning-per-sentence writing that readers understand on the first pass ## Linked Issues or Issue Description <!-- Feature request (no public GitHub issue). Described inline per CONTRIBUTING.md → "Link Issues or Describe Them In-PR". --> **Problem or motivation** Agent-authored user-facing text (issue comments, plans, documents) is frequently long-winded, uses inconsistent vocabulary, and packs multiple instructions into one sentence. Readers have to re-read it. There is no shared, installable house style for clear technical writing. **Proposed solution** Add an optional content skill, `simplified-english`, to the shipped catalog. It instructs agents to write user-facing comments, plans, and documents using only ASD-STE100 Simplified Technical English (short sentences, one instruction each, approved single-meaning words, active voice, present tense). Orgs opt in by installing it. **Alternatives considered** Baking the guidance into every agent's base instructions (too broad, not opt-in) or a bundled skill (would apply everywhere by default). An optional skill keeps it opt-in per org. **Roadmap alignment** Additive, opt-in catalog content only; no core behavior change. ## What Changed - Add `packages/skills-catalog/catalog/optional/content/simplified-english/SKILL.md` — a short optional skill instructing agents to write user-facing comments, plans, and documents in ASD-STE100 Simplified Technical English. Includes an "Approved words" section that identifies the controlled vocabulary (the ASD-STE100 Dictionary) and gives concrete house-choice substitutions. - Regenerate `packages/skills-catalog/generated/catalog.json` via `build:manifest` so the manifest includes the new skill. - Pin the new catalog key in `packages/skills-catalog/src/shipped-catalog.test.ts`. ## Verification - `pnpm --filter @paperclipai/skills-catalog build:manifest` → wrote manifest with the new skill. - `pnpm --filter @paperclipai/skills-catalog validate` → "Catalog manifest is valid". - `pnpm --filter @paperclipai/skills-catalog test` → the catalog set/count/key pinning tests pass with the new skill included. - Confirmed the generated entry: key `paperclipai/optional/content/simplified-english`, trustLevel `markdown_only`, description within the 300-char budget cap. ## Risks Low risk. Additive, markdown-only optional skill plus a regenerated manifest and a pinning-test update. No runtime code, migrations, or workflow changes. Not installed by default (`defaultInstall: false`). ## Model Used Claude Opus 4.8 (1M context), extended thinking, with tool use / code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3a16b91217 |
feat(status-cards): single-message setup drives query and update prompt (#10202)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies and their ongoing work. > - Status cards turn a standing question into recurring, agent-generated summaries on the board. > - The existing setup split intent across a watch prompt and separate update instructions, which made creation and later behavior harder to understand. > - A status card should have one durable source of truth for both deciding what to watch and telling the summarizer what each update must contain. > - This pull request makes the card prompt that source of truth, simplifies creation to one step, and lets operators choose the running agent immediately. > - The benefit is a smaller mental model, fewer configuration modes, and consistent update instructions throughout the card lifecycle. ## Linked Issues or Issue Description Status cards currently require operators to express the same intent in two places: the watch prompt and optional update instructions with append/replace/none modes. This feature simplifies the experimental status-card workflow so a single prompt defines both the watch query and every generated update. The create flow must also support selecting the responsible agent without a second setup step. Related prior status-card work: #10101. ## What Changed - Use the status card's single prompt to compile the watch query and directly instruct every summary update. - Add migration `0190_status_card_single_prompt` to remove `status_cards.instructions_mode` and `status_cards.instructions`. - Add `agentId` to `createStatusCardSchema`, validate company membership, and default new cards to the built-in Summarizer. - Replace the two-step create flow with one prompt-and-agent dialog and extract a shared `SummarizerAgentSelect` for create/settings surfaces. - Remove the extra-instructions settings section, reset incremental history when the prompt changes, and rename the board page to "Status". - Update the bundled `status-card-query` skill and board-operator documentation, then regenerate the skills catalog manifest. ## Verification - Server status-card suites: 29/29 passing. - UI `StatusCards` suites: 22/22 passing. - Skills catalog suite: 20/20 passing. - `tsc -b` passes for server, UI, shared, and database packages. - `pnpm check:migrations` passes. - Light and dark mode screenshots cover the new create dialog and settings tab. ## Risks - Migration `0190` intentionally drops existing separate instruction text. Existing card prompts remain and become the update instructions under the new model; status cards are experimental and feature-flagged. - Prompt edits now reset the incremental summary chain and trigger a full rebuild, which is intentional because the prompt is also the update contract. - Agent selection is company-scoped; invalid agent ids return a validation error rather than creating a misrouted card. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Implementation: Anthropic Claude via the `claude_local` adapter, agent label "Claude Fable 5"; extended reasoning, tool use, and code execution. The exact provider model id and context-window value were not retained in the task metadata. - PR preparation: OpenAI GPT-5.4 through Codex CLI, with reasoning, repository inspection, GitHub CLI, and Paperclip API tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
7e40ed8c43 |
feat(status-cards): add experimental status card update view (#10101)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies. > - Operators need a board-level way to monitor a changing slice of company work without repeatedly rebuilding filters or reading raw task threads. > - Existing summaries are useful snapshots, but they do not provide a dedicated query-backed card with refresh policy, change tracking, update history, and per-update cost visibility. > - The capability needs to be safe to evaluate before it becomes part of the default product surface. > - This pull request adds end-to-end experimental Status Cards, from schema and query compilation through update orchestration and operator UI. > - The entire feature is gated behind the `enableStatusCards` experimental toggle, including its route and sidebar entry. > - The benefit is a governed, inspectable way to keep focused operational rollups current while preserving explicit controls over refresh frequency and spend. ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting (`packages/db`, `packages/shared`, `server/`, `ui/`, and bundled skills/docs). ### Problem or motivation Operators cannot currently define a reusable natural-language view of company work, compile it into an inspectable query, and keep its summary current as matching issues change. Rebuilding filters and rereading task threads makes board-level monitoring repetitive and hides the relationship between source changes, refresh cost, and the resulting summary. ### Proposed solution Add experimental Status Cards that compile operator intent into a query, summarize matched work, record each update, expose manual/interval/reactive refresh policies and costs, and preserve the last good result across stale, updating, paused, and error states. The capability is off by default and fully gated behind `enableStatusCards`, including its route and navigation entry. ### Alternatives considered - Extend existing one-off summaries: rejected because status cards require persistent query provenance, refresh policy, update history, and card-specific cost controls. - Add a dashboard-only filter widget: rejected because it would not provide governed background refresh, an update ledger, or an inspectable compile pipeline. - Ship the surface by default: rejected in favor of an experimental toggle while behavior and operator value are evaluated. ### Roadmap alignment This advances Paperclip’s board-level execution visibility and output-first product goals. `ROADMAP.md` was checked and no duplicate status-card initiative was found. ### Additional context No related open PR was found in the public GitHub search for status cards. The PR-only design wireframes were removed from the repository after review; the published prototype remains external to the production source tree. ## What Changed - Added company-scoped status-card schema, CRUD APIs, compile provenance, update ledger, shared contracts, validators, and OpenAPI coverage. - Added the text-to-query compile pipeline, bundled `status-card-query` agent skill, query versioning, and authorized write-back flow. - Added the experimental board, create flow, lifecycle tiles, detail/settings/debug drawers, archived view, routing, navigation, and instance setting. - Added a change-gated update engine with manual, interval, and reactive refresh policies, trigger selection, active hours, and daily token caps. - Added per-update token/cost recording, today and lifetime rollups, and policy-derived cost previews. - Added operator documentation and agent-authoring hardening for compile and update behavior. - Added PR-prep integration coverage for settings/startup wiring and replaced raw UI values with design-system tokens. - Removed the PR-only `design/pap-15023-status-cards` wireframe artifacts so the repository contains only production feature assets. ## Verification - `pnpm -r typecheck` — passes on the PR head; includes `ui` `tsc -b` passing. The UI compile gate was also independently recorded as passing at `6d7f3cf96b` on July 23, 2026. - `pnpm build` — passes. - `pnpm check:token-gates` — passes with all three gates clean. - `pnpm test:run` — 2,880 tests passed and 1 skipped; the sole failure was an unrelated 10-second `afterAll` database-cleanup timeout in `execution-workspaces-service.test.ts`. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/execution-workspaces-service.test.ts` — passes on immediate focused rerun (25/25). - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/instance-settings-service.test.ts src/__tests__/server-startup-feedback-export.test.ts` — passes (31/31). - `pnpm --filter @paperclipai/ui exec vitest run src/pages/StatusCards/StatusCardSettingsForm.test.tsx src/pages/StatusCards/StatusCardTile.test.tsx src/pages/StatusCards/format.test.ts src/lib/status-card-state.test.ts` — passes (26/26). - Recorded pre-PR QA: compile-pipeline e2e PASS; full lifecycle and cost QA PASS; security re-review PASS after write-back hardening; UX approved. - `pnpm exec vitest run packages/db/src/status-card-migrations.test.ts` — passes; reapplies migrations `0185`–`0189` against an already-migrated embedded Postgres database. - `pnpm --filter /db check:migrations` — passes migration numbering and safety checks. - `pnpm --filter /db typecheck` — passes. - Merged current `origin/master` on July 24, 2026 with no conflicts; migrations `0185`–`0189` remain unclaimed on master. ## Risks - The feature introduces five database migrations and a new background update path; all new DDL is repeat-safe after partial application, migration numbering/safety checks pass, and update execution is company-scoped and change-gated. - Natural-language compilation can produce invalid or overly broad queries; compile provenance, query validation, debug visibility, and version history make failures inspectable and recoverable. - Reactive or interval refresh could increase spend; active hours, max refresh frequency, daily token caps, per-update cost records, and budget-paused states bound and expose that risk. - The branch name contains an internal execution identifier because it is a fixed handoff branch; it was intentionally not renamed or rebased per the release handoff instructions. - Overall rollout risk is limited because the route, navigation, services, and UI are disabled by default behind `enableStatusCards`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using GPT-5.5 with reasoning, repository tool use, shell execution, GitHub CLI, and test/build execution. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change; the fixed execution-workspace identifier is documented as an authorized handoff exception - [x] I have run tests locally and they pass, with the one cleanup timeout passing on focused rerun - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e3f8380e70 |
feat(skills): make summarize-status actions-first (#10117)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its built-in Summarizer keeps status slots useful for people overseeing issue trees > - Those summaries need to tell the reader what they must do now to unblock progress > - The existing skill instead imposed rigid Decide:/Review:/Recent work: sections, cost commentary, and restrictive issue-fetch guidance > - This pull request rewrites the summarize-status instructions to lead with 1–3 specific, concrete unblock actions while letting the model use its judgment for the remaining context > - The benefit is a shorter, clearer summary that is immediately actionable without changing slot writes or the streaming status protocol ## Linked Issues or Issue Description Refs #9713 The built-in summarizer currently prioritizes a fixed reporting template over the reader's immediate unblock actions. Summaries should instead open with the 1–3 specific actions the reader needs to take right now, then provide only the context needed to act. This prompt-only update preserves all summary-slot mechanics and protocols. ## What Changed - Rewrote the bundled `summarize-status` skill to open with 1–3 specific, concrete, actionable items needed right now to unblock the work. - Removed the rigid Decide:/Review:/Recent work: template, the Cost discipline section, and the restrictions against fetching issue detail. - Kept slot-write mechanics and the streaming `STATUS`/sentinel protocol unchanged. - Updated all materialized copies and tests for the same skill text: the `SKILL.md` source, regenerated catalog manifest hashes, compiled fallback string, summarizer built-in `AGENTS.md` and routine, summary generation-issue instructions, and the two tests pinning those strings. - Although the diff touches eight files, every file is either the same skill text in another materialized form or a test asserting it. No behavior outside the summarizer's prompt text changes. ## Verification - `pnpm --filter @paperclipai/skills-catalog test` — 20/20 tests pass. - `pnpm exec vitest run server/src/__tests__/summary-slots.test.ts server/src/__tests__/built-in-agents.test.ts` — 46/46 tests pass. - `git diff --check origin/master...HEAD` — clean. - `pnpm exec vitest run server/src/__tests__/summary-slots.test.ts` — 16/16 tests pass after the Greptile consistency fix. - Latest-head GitHub checks — 25 terminal checks, all successful, neutral, or skipped. ## Risks - Low risk: this intentionally changes generated summary wording and prioritization, but does not change APIs, persistence, slot-write behavior, or the streaming protocol. - The branch name contains an internal task identifier because it was pre-created and pre-pushed for this assigned change; the PR title and body do not expose the internal ticket. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using `gpt-5.6-sol`, high reasoning mode, with repository, terminal, GitHub CLI, and code-execution tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1f1f545238 |
feat: add built-in summarizer and summary slots (#9713)
## Thinking Path > - Paperclip is the open source control plane people use to organize, govern, and understand AI-agent work > - Operators need concise, current status views across projects and execution workspaces without manually reading every issue and run > - Paperclip already has auditable issues, documents, built-in agents, routines, and live run events, but no first-class summary-slot workflow connecting those systems > - A built-in Summarizer can generate status prose through ordinary governed tasks while summary slots provide stable, revisioned destinations for that output > - The UI needs to show current summaries, generation progress, failures, revisions, and streaming draft status in the places operators already work > - This pull request adds the end-to-end summary-slot data, API, agent, orchestration, and UI surfaces behind an experimental setting > - The benefit is decision-oriented status context that remains company-scoped, auditable, retryable, and inexpensive by default ## Linked Issues or Issue Description No public GitHub issue exists for this feature. **Problem** Operators currently have to reconstruct project and workspace status by reading many issues, runs, and comments. This makes it hard to identify decisions, review queues, recent work, and the next event worth watching. **Proposed capability** Add an experimental summary system with revisioned summary slots for projects and workspaces, a paused-by-default built-in Summarizer agent, governed generation tasks, live draft status, and reusable UI cards. **Expected behavior** - Summary data remains company-scoped and revisions remain auditable. - Generation runs through normal issue/agent orchestration and deduplicates active requests. - Only the linked built-in Summarizer generation task can author a slot revision. - Operators can generate, retry, inspect revisions, and follow draft progress from project and workspace views. - The feature remains opt-in and background generation remains paused by default. ## What Changed - Added summary-slot schema, idempotent migrations, shared contracts, validators, API paths, and service tests. - Added company-scoped summary-slot routes for reading revisions, requesting generation, and guarded Summarizer writes with activity logging. - Added terminal generation finalization, failure reasons, assignment wakeups, and orchestration integration. - Added the paused-by-default built-in Summarizer bundle, low-cost runtime defaults, status-summarization skill, and stale-summary routine. - Added summary cards, revision selection, retry/configuration states, live draft streaming, transcript chunk handling, and project/workspace integrations. - Updated Claude local parsing for streamed status output and expanded server, adapter, shared, database, catalog, and UI coverage. ## Verification - `pnpm -r typecheck` - `pnpm exec vitest run packages/db/src/summary-slots-schema.test.ts packages/shared/src/summary-slot.test.ts server/src/__tests__/summary-slot-routes.test.ts server/src/__tests__/summary-slots.test.ts server/src/__tests__/built-in-agents.test.ts ui/src/components/SummarySlotCard.test.tsx ui/src/components/SummarySlotCard.status.test.tsx ui/src/components/useSummaryDraftStream.test.tsx ui/src/lib/summary-draft-stream.test.ts ui/src/lib/run-log-chunks.test.ts ui/src/context/LiveUpdatesProvider.hook.test.tsx` — 113 tests passed - `pnpm test:run` — server and UI suites passed; one CLI AWS doctor test was affected by inherited `AWS_ACCESS_KEY_ID`/`AWS_SECRET_ACCESS_KEY`, and passed when those host credentials were removed - `pnpm exec vitest run cli/src/__tests__/secrets.test.ts` with inherited AWS credential variables removed — 8 tests passed - `pnpm build` - `pnpm check:token-gates` currently reports nine `#9627` comment references introduced by current `master`; none are in this PR diff ## Risks - Database risk is limited by incrementally ordered, idempotent migrations and migration safety checks. - Summary generation creates normal issues/runs, so misconfiguration can produce failed slots; the UI exposes retryable failure reasons and agent configuration entry points. - Streaming draft parsing depends on the documented `STATUS:` protocol; final persisted revisions remain the source of truth. - The feature is experimental, opt-in, and its built-in routine is paused with no background token spend by default. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5.3 Codex with reasoning, repository tool use, code execution, GitHub CLI, and Paperclip control-plane integration. Earlier branch commits also record Claude model co-authorship where applicable. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
953b315dfb |
Shorten skill frontmatter descriptions (#9353)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip agents can load repository and catalog skills, and Codex renders skill names and frontmatter descriptions into startup context. > - Long descriptions consume the fixed skill metadata budget before Codex can use the progressively disclosed skill bodies. > - The repo `.agents/skills` descriptions and a few shipped catalog descriptions had grown into operational documentation instead of short trigger metadata. > - This pull request keeps the strongest trigger language in frontmatter while leaving detailed procedures in each skill body. > - The benefit is lower prompt overhead, more reliable skill triggering, and a regression guard that prevents description drift from returning. ## Linked Issues or Issue Description No public GitHub issue found for this maintenance item. ### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am working against `master`. - [x] I have confirmed the issue originates in Paperclip's shipped skill metadata, not in a local agent adapter or provider. ### What happened? Codex startup renders discovered skill names and frontmatter descriptions into a fixed skill metadata budget. Several repository skill descriptions and one shipped catalog description had grown into long-form operational guidance, which can force Codex to truncate descriptions before the model has enough trigger signal to select the right skill. ### Expected behavior Skill frontmatter descriptions should stay short trigger summaries: one capability sentence plus a “use when” clause. Detailed procedures should stay in the skill body and load only after the skill triggers. ### Steps to reproduce 1. Inspect `.agents/skills/*/SKILL.md` and `packages/skills-catalog/catalog/**/SKILL.md` frontmatter descriptions. 2. Measure folded YAML `description` values. 3. Observe descriptions above the intended short-trigger range, including descriptions above 300 characters. 4. Run the new shipped catalog test to verify future descriptions stay capped. ### Paperclip version or commit Reproduced on `master` at `cc81eefb6047d8eaf57faf785f421c03dc97073c`. ### Deployment mode Local dev / source checkout metadata inspection. This is not database-related. ### Installation method Built from source. ### Agent adapter(s) involved Codex, because Codex startup uses the skill metadata prompt budget. The metadata source itself is core repository/catalog content. ### Database mode Not database-related. ### Access context Not applicable; this is static repository metadata. ### Node.js version `v22.22.2` in the verification environment. ### Operating system Linux container environment. ### Relevant logs or output Final measurement after this PR: 29 source `SKILL.md` files, max description length 215 chars, 5,449 total description chars, estimated 1,363 description tokens at 4 chars/token. ### Relevant config None. ### Additional context The shipped catalog manifest was regenerated so the generated package metadata matches the edited catalog `SKILL.md` sources. ### Privacy checklist - [x] I have reviewed all pasted output for PII and redacted where necessary. ## What Changed - Shortened long `.agents/skills/*/SKILL.md` frontmatter descriptions to concise capability plus use-when trigger clauses. - Shortened the over-budget shipped skills catalog descriptions for wireframe, Paperclip capsules, and reflection coach. - Regenerated `packages/skills-catalog/generated/catalog.json` so shipped metadata matches source skill frontmatter. - Added a Vitest regression guard that caps repo skill source descriptions and generated catalog descriptions at 300 characters. ## Verification - `pnpm --filter @paperclipai/skills-catalog build:manifest` - `pnpm --filter @paperclipai/skills-catalog test` — 5 files passed, 19 tests passed - `pnpm --filter @paperclipai/skills-catalog typecheck` - Final measurement: 29 source `SKILL.md` files, max description length 215 chars, 5,449 total description chars, estimated 1,363 description tokens at 4 chars/token. Note: the clean PR worktree was created from `origin/master` and contains only this commit, but it does not have `node_modules`; running `pnpm --filter @paperclipai/skills-catalog test` there failed at tool/package resolution (`vitest`, `tsc`, `@paperclipai/shared`). The dependency-equipped workspace passed the commands above before the commit was cherry-picked onto the clean branch. ## Risks Low risk. This changes skill metadata and tests only. The main risk is over-trimming a useful trigger phrase, mitigated by keeping explicit “use when” clauses and leaving detailed guidance in the skill bodies. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex coding agent via Paperclip/Codex, with shell and file-edit tool use. Exact API model ID and context window were not exposed in the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5c85ae64a0 |
Cases: experimental first-class case object (#9198)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The board currently uses issues for execution, but longer-lived content work needs a separate object that can survive beyond a single task thread. > - The Cases subsystem adds an experimental, company-scoped record for content artifacts and their supporting metadata. > - The backend needs durable storage, API routes, revision history, issue linkage, and company-boundary enforcement before the UI can depend on Cases. > - The UI needs an opt-in navigation surface, list/detail views, reference chips, and issue-page context so operators can inspect Cases without making them the default workflow. > - The agent-facing skills need a contract for creating and updating Cases so automated content workflows can dogfood the feature. > - This pull request ships that experimental end-to-end path behind the `enableCases` flag. > - The benefit is a first-class place to collect content work, references, attachments, revisions, and related execution threads without polluting the core issue model. ## Linked Issues or Issue Description No public GitHub issue exists for this experimental feature. Feature request fields: ### Problem Content-oriented work such as release notes, announcements, docs, and campaigns can span many execution issues, which makes the final artifact hard to find and reason about after the execution thread moves on. ### Proposed solution Add an experimental Cases object that is company-scoped, linked to issues, queryable through the API, inspectable in the board UI, and writable by agent workflows through documented conventions. ### Alternatives considered Continue encoding content artifacts directly in issues or documents only. That keeps the data model smaller, but it does not give operators a stable artifact-centric view or a clean way to link related execution history. ### Roadmap alignment Checked `ROADMAP.md`; this PR does not duplicate an existing planned core roadmap item. ## What Changed - Added the `cases` data model, migration, schema exports, and experimental `enableCases` instance setting. - Added company-scoped Cases API routes for list/detail/update, issue links, revisions, children, activity events, annotations, attachments, and idempotent agent-oriented upserts. - Scoped case and issue lookup helpers before access checks so inaccessible cross-company identifiers resolve as not found rather than leaking existence. - Fixed case PATCH timestamp handling so non-status updates cannot overwrite `completedAt` from a stale pre-transaction row snapshot. - Moved Cases list type/status/project filters into the server request before the server-side limit is applied, including multi-select filters and no-project filtering. - Added backend route coverage for creation, updates, idempotency, issue linking, attribution, company-boundary enforcement, OpenAPI registration, list filtering, timestamp patch behavior, and inaccessible lookup regressions. - Added the experimental Cases UI surface: sidebar entry, gated routes, list filters/grouping, detail overview, activity, revisions, children, attachments, and issue-page case rail. - Added case reference rendering and company-prefixed case href generation so case links resolve directly inside the active company route. - Added Paperclip skill documentation for agent workflows that create or update Cases. - Wired release-content skills to emit Cases for dogfooding. - Rebased onto current `master` and renumbered the Cases migrations to `0143`/`0144` after the latest upstream migration sequence. ## Verification - Current PR head: `ecc13be0d`. - Rebased on current `master` (`606aa4f266`) and pushed to the existing PR branch. - `git diff --check origin/master...HEAD` — passed before the first update push; subsequent committed diffs were also checked with `git diff --check` before commit. - Guardrails checked: no `pnpm-lock.yaml` changes, no `.github/workflows` changes, and changed-file count is below the Greptile 100-file limit. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/cases-routes.test.ts src/__tests__/instance-settings-service.test.ts src/__tests__/openapi-routes.test.ts` — passed, 3 files / 26 tests before review-fix commits. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/cases-routes.test.ts` — passed after each server-side Greptile fix, latest 1 file / 15 tests. - `pnpm --filter @paperclipai/server typecheck` — passed after the timestamp and lookup fixes. - `pnpm --filter @paperclipai/ui exec vitest run src/pages/Cases.test.tsx src/pages/CaseDetail.test.tsx src/pages/CompanySkills.test.tsx src/App.cases-routing.test.tsx` — passed, 4 files / 30 tests before review-fix commits. - `pnpm --filter @paperclipai/ui exec vitest run src/pages/Cases.test.tsx` — passed after the list-filter fix, 1 file / 12 tests. - `pnpm --filter @paperclipai/ui typecheck` — passed after the list-filter fix. - `pnpm check:token-gates` — passed after UI changes. - Remote PR checks on head `ecc13be0d` are green: Paperclip CI, build, typecheck, test matrix, e2e, Canary Dry Run, policy, commit review, Superagent Security Scan, Socket, Snyk, and Greptile passed; Storybook visual regression is skipped and security-review is neutral. - Greptile Review: 5/5 confidence, zero unresolved Greptile threads. ## Risks - Medium feature risk because this introduces a new experimental domain object across database, server, shared contracts, skills, and UI. - The feature is gated behind `enableCases`, which limits default operator exposure while the model is exercised. - Case links now prefer company-prefixed hrefs; the unprefixed redirect remains for externally entered URLs. - Cases list filtering now sends multi-select filters to the server before limiting; the UI still applies the same local filters as a second pass for ancestor/context rows. - Migrations were renumbered on top of current master; the SQL uses guarded `IF NOT EXISTS` / `ADD COLUMN IF NOT EXISTS` patterns where relevant for safer replay. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex in the Paperclip local coding environment was used for this PR curation, rebase verification, review-fix implementation, push, and PR description update. The runtime exposes tool use and shell execution; context-window size is not exposed by this Paperclip adapter. Several implementation commits also include AI co-author trailers recorded in git history. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I searched the GitHub PR list for similar PRs and confirmed this is not a duplicate - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8b6a06ee25 |
[codex] Add built-in agents and Reflection Coach bundle (#9206)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators need first-party agent capabilities for repeatable company work, not just manually created one-off agents. > - Built-in agents need to behave like normal company-scoped agents while preserving approval gates, permissions, budgets, and audit trails. > - Reflection and coaching work also needs bundled instructions, skill content, and a routine so the feature can be installed and reset predictably. > - The API, database, UI, portability, and tests all need to agree on the built-in lifecycle from not provisioned through setup, approval, ready, paused, and reset. > - This pull request adds built-in agent provisioning and the Reflection Coach bundle end-to-end. > - The benefit is a safer first-party path for Paperclip-managed agents without bypassing the same governance model used for operator-created agents. ## Linked Issues or Issue Description No public GitHub issue was found for this exact built-in agent and Reflection Coach bundle work. Problem/motivation: - Paperclip did not have a first-party built-in agent lifecycle for product-owned agents. - Bundled agent resources such as default instructions, skills, and routines needed managed ownership and reset semantics. - Approval-gated companies needed built-in setup to preserve requested adapter, budget, manager, and permission state through board approval. - The board UI needed clear built-in badges, setup affordances, readiness state, and bundle status without exposing secrets. Proposed solution: - Add a company-scoped built-in agent registry, provisioning/reset/reconcile/status APIs, and Reflection Coach bundled resources. - Track bundled managed resources in the database with idempotent migration behavior. - Reuse existing agent approval, authorization, budget, and activity-log paths instead of creating a bypass. - Add UI setup, badges, gates, bundle panels, and route coverage for built-in agents. Duplicate search: - Searched GitHub PRs for `built-in agents Reflection Coach repo:paperclipai/paperclip`; only this PR was returned. - Searched GitHub issues for the same query; no public issues were returned. ## What Changed - Added built-in agent definitions, lifecycle state derivation, provisioning, reset, reconcile, status, and routine-control routes. - Added the `built_in_managed_resources` migration and schema exports for bundled instructions, skill, and routine ownership. - Added the Reflection Coach built-in bundle with default instructions, skill catalog content, routine template, default permissions, and managed-resource drift handling. - Added approval-aware provisioning behavior that preserves requested adapter config, budgets, manager assignment, and built-in permissions through hire approval. - Added authorization and mutation gates for built-in agent and skill changes, including consented Reflection Coach change paths. - Added UI surfaces for built-in agent setup, roster/detail badges, readiness gates, bundle status, routine controls, and route filtering. - Added company import/export and validator coverage for built-in managed resources and low-trust/red-team presets. - Addressed Greptile follow-ups for pending approval reconciliation, consent-gate error propagation, config-read authorization fallback, approval-path manager preservation, and non-model adapter provisioning. ## Verification Local verification: - `git diff --check public/master..HEAD` passed. - `pnpm check:token-gates` passed with all gates clean. - `pnpm exec vitest run ui/src/components/ConfigureBuiltInAgentModal.test.tsx` passed: 1 file, 4 tests. - `pnpm exec vitest run ui/src/components/EntityRow.test.tsx ui/src/pages/Agents.test.tsx ui/src/components/BuiltInAgentGate.test.tsx ui/src/components/ConfigureBuiltInAgentModal.test.tsx ui/src/components/BuiltInBundlePanel.test.tsx ui/src/pages/InstanceExperimentalSettings.test.tsx ui/src/pages/Routines.test.tsx` passed: 7 files, 64 tests. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/built-in-agents.test.ts src/__tests__/authorization-service.test.ts src/__tests__/company-skills-routes.test.ts` passed: 3 files, 91 tests. - `pnpm --filter @paperclipai/db check:migrations` passed. - `pnpm -r typecheck` passed after the rebase; `pnpm --filter ui typecheck` passed after the final UI review fix. Remote verification on latest head `1c61f693a4ec881d739022b0e75a8ca8bf8c2cd8`: - Merge state: `CLEAN`. - Greptile: `5/5`, zero unresolved Greptile threads. - PR check rollup: all checks successful, neutral, or skipped as expected. - Passing gates include Build, Typecheck + Release Registry, all server shards, all workspace shards, all serialized server suites, e2e, Canary Dry Run, policy, review, verify, Socket, Superagent, and Snyk. ## Risks - This adds a new managed-resource table and migration; the migration uses idempotent create/add/index guards and passed migration safety checks. - Built-in agent provisioning touches approval and authorization paths; tests cover pending approval preservation, stale retry rejection, consent gates, and config-read fallback behavior. - Reflection Coach creates managed instructions, skill, and routine resources; drift/reset behavior is covered by service tests and redacted API responses. - Non-model adapter setup now provisions a `needs_setup` built-in row before command/endpoint fields are complete; this matches the server lifecycle and is covered by the setup modal regression test. ## Model Used OpenAI Codex coding agent based on GPT-5. Exact hosted model ID, context-window size, and reasoning-mode labels are not exposed in this runtime; tool use, shell execution, GitHub CLI/API access, and local code editing were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
72c42fad99 |
[codex] Add optional Ramp skill (#9157)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Skills are how operators give agents reusable, reviewable operating instructions without baking every integration into core runtime code. > - Finance setup is a sensitive workflow because account onboarding, incorporation, cards, spend controls, and data sharing can all create real-world effects. > - Ramp publishes an agent-facing setup skill and playbooks, but Paperclip needs a curated wrapper that makes those instructions subordinate to Paperclip governance. > - This pull request adds an optional Ramp catalog skill that fetches Ramp's live entrypoint while preserving Paperclip approval gates. > - The benefit is that companies can opt into Ramp setup assistance while reviewers can see the source model, allowed hosts, and fail-closed safety rules in one shipped catalog entry. ## Linked Issues or Issue Description No public GitHub issue exists for this optional catalog skill. Feature request context: - **Problem / motivation:** Paperclip companies need a safe, installable way for agents to follow Ramp's public agent setup flow without giving those fetched instructions authority over financial, legal, credential, or spend decisions. - **Proposed solution:** Ship a markdown-only optional `paperclipai:optional:finance:ramp` skill that points agents at Ramp's live get-started skill, documents the thin-wrapper source model, allowlists the Ramp host, and requires Paperclip approval for financial, incorporation, credential, connector, third-party tool, and money-movement actions. - **Alternatives considered:** Vendoring a snapshot would reduce runtime source drift but would stale quickly as Ramp updates its own onboarding flow. The wrapper instead fetches fresh instructions while explicitly failing closed on unclear provenance and keeping fetched instructions subordinate to Paperclip instructions. - **Roadmap alignment:** This fits the completed Skills Manager roadmap area by adding a focused optional catalog skill rather than expanding core workflow code. ## What Changed - Added a markdown-only optional Ramp skill under the finance catalog. - Documented the source model for live Ramp instructions, the allowed host, provenance handling, community/unclear playbook approval requirements, and safety rules. - Added mandatory Paperclip approval gates for Ramp account setup, incorporation/legal filings, CLI installers, connector/auth flows, third-party browser/MCP/CLI tooling, financial data sharing, Agent Cards, spend controls, and money movement. - Updated the Skills Store guide to document thin fetch-and-follow wrappers for curated optional skills. - Regenerated the shipped skills catalog manifest. - Added catalog tests for the Ramp entry, approval-gate wording, mixed-provenance handling, and avoiding remote-fetch execution hard-stop patterns. ## Verification - `pnpm --filter @paperclipai/skills-catalog build:manifest` - `pnpm --filter @paperclipai/skills-catalog validate` - `pnpm --filter @paperclipai/skills-catalog test -- src/shipped-catalog.test.ts` — 1 file, 8 tests passed. - `git diff --check` ## Risks - Ramp-hosted instructions can change after install. The wrapper mitigates this by keeping fetched content subordinate to Paperclip instructions, limiting the source host, failing closed on unclear provenance, and requiring scoped approvals before governed actions. - The skill is markdown-only and optional, so it does not add executable package code or install by default. - The generated catalog manifest changes hashes for the shipped catalog entry; catalog validation passed after regeneration. ## Model Used OpenAI Codex, GPT-5-based coding agent, with repository file access, shell execution, GitHub CLI/tooling, and medium-reasoning mode. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
32ef854771 |
docs: expand capsule identicon prototyper guidance (#8665)
Reviewed by CTO for PAP-12039. Documentation/catalog-only change with regenerated manifest; focused catalog verification, CI, security scans, and Greptile are green. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
1a3e398107 |
Add bundled Paperclip capsules skill
Add the bundled paperclip-capsules skill, durable generator workflow references, catalog manifest updates, and shipped catalog coverage. |
||
|
|
bee2b25f5d | Add referenced last30days catalog entry | ||
|
|
1cf3b792b5 |
Bundle the wireframe skill into the skills catalog
Adds the wireframe skill (low-fi black-and-white SVG wireframes + viewer page) as a bundled catalog skill under catalog/bundled/product/wireframe, alongside its references/ docs and assets/ templates. Regenerates generated/catalog.json (8 -> 9 skills). The skill ships static svg/html template assets, so its derived trust level is "assets" rather than "markdown_only". The server's real install-time security gate (assertCatalogSkillInstallable) blocks only "scripts_executables", and "assets" skills are installable, so the shipped-catalog markdown-only invariant is refined to gate on executable scripts instead. No skill ships executable scripts. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9eac727cf1 |
[codex] Add skills CLI and catalog management (#6782)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies through company-scoped control-plane workflows. > - Agents need reusable, inspectable skills that can be installed, reset, audited, exported, and assigned without bespoke local setup. > - The existing skill truth model needed cleanup so bundled skills, optional catalog skills, runtime skills, and adapter-provided skills have clear provenance. > - Operators also need a practical CLI and board UI for discovering and managing company skills. > - This pull request adds the skills CLI, packaged skills catalog, company skills APIs, and catalog-aware board UI. > - The benefit is a more reusable Paperclip company setup where skills are portable, auditable, and easier for operators and agents to manage. ## What Changed - Added `paperclipai skills` CLI commands and coverage for catalog listing, installing, resetting, and inspecting company skills. - Added a packaged `@paperclipai/skills-catalog` workspace with bundled and optional skill content plus validation/build tests. - Added shared company-skill types and validators used across CLI, server, and UI contracts. - Added server catalog APIs/services for company skill catalog operations, reset semantics, audit behavior, and portability provenance. - Updated adapter skill handling so runtime/catalog provenance remains explicit across local adapters. - Added board UI support for browsing and managing catalog-backed company skills. - Updated docs for the skills CLI/catalog flow and the company skills Paperclip skill reference. - Rebased the branch onto current `paperclipai/paperclip:master`; no `pnpm-lock.yaml`, `.github/workflows`, or migration files are included in the final PR diff. ## Verification - Passed: `pnpm run preflight:workspace-links && pnpm exec vitest run cli/src/__tests__/skills.test.ts packages/skills-catalog/src/catalog-builder.test.ts packages/skills-catalog/src/shipped-catalog.test.ts packages/shared/src/validators/company-skill.test.ts packages/adapter-utils/src/server-utils.test.ts packages/plugins/create-paperclip-plugin/src/entrypoints.test.ts server/src/__tests__/company-skills-catalog-service.test.ts server/src/__tests__/company-skills-routes.test.ts server/src/__tests__/company-portability.test.ts`. - Passed: `pnpm exec vitest run server/src/__tests__/workspace-runtime.test.ts -t "default branch|origin/master|symbolic-ref"`. - Attempted: full `server/src/__tests__/workspace-runtime.test.ts`. Four provisioning tests failed while seeding an isolated worktree database from the local Paperclip instance because the local plugin schema dump contains a duplicate-column foreign key (`plugin_content_machine_18a7bc327b.content_case_signals`). The default-branch tests touched by the rebase conflict passed in the focused run above. - Checked final diff: no `pnpm-lock.yaml`, no `.github/workflows`, and no migration-file changes relative to `master`. ## Risks - Medium: this is a broad skills/catalog change touching CLI, server APIs, shared contracts, adapter skill sync, and UI. - Catalog validation and reset semantics need careful reviewer attention because they affect reusable company setup and portability. - No database migrations are included in this PR, so there is no migration ordering/idempotency risk in the final diff. - No lockfile is included by design; dependency resolution will be handled by the repository lockfile workflow. ## Model Used - OpenAI Codex coding agent based on GPT-5, running in Paperclip via the `codex_local` adapter with shell, git, GitHub CLI, and code-editing tool access. Exact hosted model build/context-window metadata is not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run targeted tests locally and documented the local workspace-runtime seed failure above - [x] I have added or updated tests where applicable - [x] If this change affects the UI, screenshots were intentionally omitted per PAP-10124 instructions; UI behavior is covered by tests and reviewer inspection - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |