mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Runner gives agents tools to change company resources. > - Users need agents to save reusable skills during a task. > - A saved skill needs a visible result that users can inspect and edit. > - This pull request adds `create_skill` and a task feed card linked to Skill Studio. > - Users can open the saved skill from the task and edit the same resource. ## Linked Issues or Issue Description **Subsystem affected** Runner tools, company skill storage, task feed, and Skill Studio. **Problem or motivation** The Runner has no dedicated tool to create a company skill. A user cannot follow a creation result from the task feed to the saved skill. **Proposed solution** Add a company-scoped `create_skill` tool. Save the skill with the existing company policy. Add one creation card to the task. Open a named sidebar tab from that card. Let the user open the same skill in Skill Studio. **Alternatives considered** An agent can write a local file, but that file is not a company skill. A second document copy in the task would become stale after a Studio edit. The sidebar therefore reads the saved skill directly. **Roadmap alignment** This extends the shipped Skills Manager, Skill Studio, and Skills Store milestone. The maintainer requested and approved this scope. Search found no duplicate `create_skill` PR or issue. Related UI validation work: #8715. This PR does not change that validation display. ## What Changed - Add the real Runner tool, its contract, and its mock implementation. - Validate the complete SKILL.md and derive company, task, agent, and run identity from authentication. - Apply the existing company skill policy. Do not assign the skill to an agent. - Make keyed retries return one skill and one creation event. Reject conflicting retries. - Make concurrent file creation safe. Never replace an existing published skill during creation. - Add a creation card, a named sidebar tab, and an Open in Skill Studio action. - Show saved Studio edits when the user returns to the task. - Add storage, policy, mode, retry, UI, and Product E2E tests. Document the tool. - Fix deleted-name reuse, onboarding panel persistence, immediate feed refresh, and mock validation parity from review. - Serialize Studio file edits and renames with skill deletion and recreation. Reject stale editor requests before they can change a replacement skill. - Generate the standalone mock parser and validator from the production contract. Use portable UUIDs so the browser scenario bundle builds. ## Verification - All latest-head PR checks pass on `145dd76a5`, including all server shards, browser E2E, Runner verification, build, typecheck, and release dry run. Greptile: 5/5 with no open findings. An interrupted CI runner was retried successfully. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm check:token-gates`: passed. - Review regressions: 73 storage tests, 6 real API tests, 63 UI tests, and 61 semantic runtime tests passed. Parser synchronization passed. - CI exposed existing fire-and-forget Sentry test races. Reproduced the resumption race locally, then synchronized the related sweep and finalizer assertions on the actual report; all 27 tests across the three affected files pass. - Runner scenario browser build and strict content-security-policy check: passed. - Runner suite: 2,012 tests passed; 10 skipped. - `pnpm test:run`: the general-server batch had 12,416 passes and two failures. The old tool-count assertion was fixed; all 16 authority tests then passed. The chat webhook test had a socket error; it passed four isolated reruns. - Both workspace test groups passed. The isolated route suites completed. Two socket failures in the initial route batches passed on individual reruns; all remaining 61 files passed. - Product E2E `create-skill-studio`: passed with local Codex and local ACPX Claude. - Manual browser test: submit a task, observe the real tool call and creation card, open the sidebar, edit in Studio, save, and return. The task reached Done. The saved second revision and sidebar tab survived a server restart. - The new companion headless Runner Eval passed. Companion coverage PR: https://github.com/paperclipai/paperclip-evals/pull/23. Daytona was not run because no immutable runner image was configured. ## Risks - Database writes and local file writes cannot share one transaction. Recovery accepts only an exact file-for-file retry after a database rollback. Conflicting files remain untouched. - The sidebar displays the current skill. The feed card remains the historical creation receipt. - No database migration, dependency, or workflow change is included. - Remote Daytona behavior still needs a run with a configured immutable image. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) handled design, integration, review, and browser verification. OpenAI `gpt-5.6-luna` assisted with bounded implementation and eval work. Both used code execution and tool access. The host did not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
66 lines
3.4 KiB
Markdown
66 lines
3.4 KiB
Markdown
# Skills created by the Runner
|
|
|
|
The native Runner exposes `create_skill` in Auto and skill-test work. An agent
|
|
can save a reusable, single-file skill directly in the company's skill library.
|
|
The tool is unavailable in Ask and pre-acceptance Plan modes.
|
|
|
|
```json
|
|
{
|
|
"name": "release-review",
|
|
"description": "Review release notes.",
|
|
"markdown": "---\nname: release-review\ndescription: Review release notes.\n---\n\n# Review\nCheck each release note against its change.\n",
|
|
"idempotencyKey": "release-review-1"
|
|
}
|
|
```
|
|
|
|
The name is lowercase with hyphens. The complete `SKILL.md` must have matching
|
|
name and description frontmatter and a nonempty instruction body. An optional
|
|
`slug` must equal the name. Company, task, agent and run identity come from the
|
|
authenticated run; they are not tool inputs.
|
|
|
|
The tool uses the same company skill policy and storage as Skill Studio. Skills
|
|
are open by default unless company policy restricts creation. Creating a skill
|
|
does not assign it to an agent or introduce a new approval step. Once the skill
|
|
is saved, the task can finish normally if no other work remains.
|
|
|
|
The result contains the skill ID, name, slug, description, version ID and Studio
|
|
path. Reuse the same idempotency key and inputs after a lost response. A retry
|
|
returns the existing skill without another creation event. A key reused with
|
|
different inputs, or a name that belongs to another skill, returns a conflict.
|
|
Published files are never replaced as part of a competing creation request.
|
|
Deleting a managed skill releases its name for a later creation. Deletion uses
|
|
the same name lock as creation and retains its source files until the database
|
|
deletion commits. Imported local and project source folders are not removed.
|
|
|
|
## User interface
|
|
|
|
Successful creation adds a **Skill created** card to the originating task's
|
|
feed. The card opens a named skill tab in the task sidebar. Repeated clicks
|
|
focus that tab. The tab reads the company skill directly, so it is not a second
|
|
editable copy of the instructions.
|
|
|
|
**Open in Skill Studio** opens that same skill for editing. After saving and
|
|
returning to the task, the sidebar shows the latest saved version. Reloading the
|
|
task restores the tab. The historical card remains a creation receipt even if
|
|
the skill is later edited or removed. The sidebar reports missing skills,
|
|
denied access and retryable load errors explicitly.
|
|
|
|
## Verification
|
|
|
|
Server regressions cover authenticated creation, company isolation, policy
|
|
denials and revocation, mode restrictions, initial version/file persistence,
|
|
idempotency, conflicting retries and concurrent creation. UI tests cover the
|
|
creation card, saved tabs, frontmatter-free preview and error states.
|
|
|
|
The companion Runner Eval case `create-skill` tests provider tool use against
|
|
the seeded mock control plane, not production storage or company policy.
|
|
The standalone mock uses a generated production frontmatter parser and schema
|
|
validator. Regenerate it with
|
|
`node packages/shared/scripts/generate-runner-skill-frontmatter.ts` after changing
|
|
the shared frontmatter contract. A shared-package test checks synchronization.
|
|
Invalid inputs are covered by deterministic tests: a live model should not be
|
|
penalized for declining to send a schema-invalid request. Product E2E's
|
|
`create-skill-studio` case checks real creation, task completion, the feed card,
|
|
sidebar, Studio editing and the updated skill after returning to the task.
|
|
See [the eval guide](evals.md) for the difference between these suites.
|