mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 06:15:21 +02:00
## Thinking Path > - Paperclip is an open source control plane for AI-agent companies. > - Company skills give agents reusable work instructions. > - Skill Studio can edit skill files and save version history. > - Agents can create a skill, but they do not have a first-class update tool. > - An agent update needs a version check and safe retry behavior to prevent lost edits. > - This pull request adds `update_skill` through the existing company skill file API. > - The change keeps company policy, version history, and audit records in one path. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: the server API, shared validation, and runner tool catalog. **Problem or motivation** An agent can create a company skill but cannot update its `SKILL.md` through a first-class tool. An unguarded retry can also create duplicate versions or overwrite a newer edit. **Proposed solution** Add `update_skill` with a required current version ID and a retry key. Route it through the existing skill file API. Reject stale versions and changed-input retries. Save the version and audit event together. **Alternatives considered** A separate write endpoint would duplicate the Skill Studio mutation path and policy checks. This PR reuses that path instead. **Roadmap alignment** This work extends Skills Manager and Skill Studio, which are listed in `ROADMAP.md`. **Additional context** The tool accepts a complete `SKILL.md`, not a partial patch. Callers must read the current version before they edit it. ## What Changed - Add optional version and retry fields to the existing skill file update contract. - Add a guarded API update with a stable retry receipt and attributed audit event. - Add `update_skill` to native and semantic runner tool catalogs, with mode and policy gates. - Add unit, integration, protocol, and semantic-tool regression coverage. - Document agent use and extend the OpenAPI request contract. ## Verification - Focused tests and direct server and runner TypeScript checks passed before this PR. - `git diff --check` passed after the rebase onto `master`. - CI passed on the latest PR head, including the full test matrix, typecheck, and build. Local full typecheck and build stopped because `cargo` is not installed. The local full test run ended without a verdict. - No dedicated end-to-end eval scenario was added or run. The protocol coverage and semantic-tool test cover the new action deterministically. - Reviewers can read a skill version, call `update_skill`, repeat the same key, then try a stale version and a changed-input key. Only the first edit must create a new version. ## Risks - File writes and database transactions must stay in sync when a write fails. The integration tests cover failed writes and retry behavior, but CI must verify them on the PR head. - Existing Skill Studio callers do not send the new optional coordination fields. Their request shape remains valid. ## Model Used - OpenAI Codex CLI assisted with this change. The runner did not expose the exact model ID or context window. The agent used code execution and repository tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details; exact model ID and context window were not exposed) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused tests; full suite is pending CI) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green (53 pass, 4 skip on the latest head) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
87 lines
4.4 KiB
Markdown
87 lines
4.4 KiB
Markdown
# Skills created by the Runner
|
|
|
|
The native Runner exposes `create_skill` in Auto and skill-test work. An agent
|
|
can save a reusable, single-file skill directly in the company's skill library.
|
|
The tool is unavailable in Ask and pre-acceptance Plan modes.
|
|
|
|
The native Runner also exposes `update_skill` in the same modes. Read the
|
|
existing skill and its `currentVersionId` through the company skill API before
|
|
editing. Supply the **complete** primary file, not a fragment:
|
|
|
|
```json
|
|
{
|
|
"skillId": "<existing skill UUID>",
|
|
"expectedVersionId": "<current version UUID>",
|
|
"markdown": "---\nname: release-review\ndescription: Review release notes.\n---\n\n# Review\nCheck each release note and its tests.\n",
|
|
"idempotencyKey": "review-update-1"
|
|
}
|
|
```
|
|
|
|
The tool uses the existing Skill Studio file endpoint and company `skills.edit`
|
|
policy. It returns a skill ID, path, resulting version ID, and Skill Studio
|
|
path. Retry a lost response with **exactly** the same key and inputs; that
|
|
returns the original receipt without another version or activity record.
|
|
Reusing the key with different inputs, or updating after the version changes,
|
|
returns a conflict. After a version conflict, reread the skill and choose a
|
|
new key for the revised edit. Policy is checked again even on exact retries.
|
|
|
|
```json
|
|
{
|
|
"name": "release-review",
|
|
"description": "Review release notes.",
|
|
"markdown": "---\nname: release-review\ndescription: Review release notes.\n---\n\n# Review\nCheck each release note against its change.\n",
|
|
"idempotencyKey": "release-review-1"
|
|
}
|
|
```
|
|
|
|
The name is lowercase with hyphens. The complete `SKILL.md` must have matching
|
|
name and description frontmatter and a nonempty instruction body. An optional
|
|
`slug` must equal the name. Company, task, agent and run identity come from the
|
|
authenticated run; they are not tool inputs.
|
|
|
|
The tool uses the same company skill policy and storage as Skill Studio. Skills
|
|
are open by default unless company policy restricts creation. Creating a skill
|
|
does not assign it to an agent or introduce a new approval step. Once the skill
|
|
is saved, the task can finish normally if no other work remains.
|
|
|
|
The result contains the skill ID, name, slug, description, version ID and Studio
|
|
path. Reuse the same idempotency key and inputs after a lost response. A retry
|
|
returns the existing skill without another creation event. A key reused with
|
|
different inputs, or a name that belongs to another skill, returns a conflict.
|
|
Published files are never replaced as part of a competing creation request.
|
|
Deleting a managed skill releases its name for a later creation. Deletion uses
|
|
the same name lock as creation and retains its source files until the database
|
|
deletion commits. Imported local and project source folders are not removed.
|
|
|
|
## User interface
|
|
|
|
Successful creation adds a **Skill created** card to the originating task's
|
|
feed. The card opens a named skill tab in the task sidebar. Repeated clicks
|
|
focus that tab. The tab reads the company skill directly, so it is not a second
|
|
editable copy of the instructions.
|
|
|
|
**Open in Skill Studio** opens that same skill for editing. After saving and
|
|
returning to the task, the sidebar shows the latest saved version. Reloading the
|
|
task restores the tab. The historical card remains a creation receipt even if
|
|
the skill is later edited or removed. The sidebar reports missing skills,
|
|
denied access and retryable load errors explicitly.
|
|
|
|
## Verification
|
|
|
|
Server regressions cover authenticated creation, company isolation, policy
|
|
denials and revocation, mode restrictions, initial version/file persistence,
|
|
idempotency, conflicting retries and concurrent creation. UI tests cover the
|
|
creation card, saved tabs, frontmatter-free preview and error states.
|
|
|
|
The companion Runner Eval case `create-skill` tests provider tool use against
|
|
the seeded mock control plane, not production storage or company policy.
|
|
The standalone mock uses a generated production frontmatter parser and schema
|
|
validator. Regenerate it with
|
|
`node packages/shared/scripts/generate-runner-skill-frontmatter.ts` after changing
|
|
the shared frontmatter contract. A shared-package test checks synchronization.
|
|
Invalid inputs are covered by deterministic tests: a live model should not be
|
|
penalized for declining to send a schema-invalid request. Product E2E's
|
|
`create-skill-studio` case checks real creation, task completion, the feed card,
|
|
sidebar, Studio editing and the updated skill after returning to the task.
|
|
See [the eval guide](evals.md) for the difference between these suites.
|