Files
PaperClipAI/doc/runner-created-skills.md
DottaandPaperclip eb049aebf2 feat(skills): let agents update company skills safely (#15049)
## Thinking Path

> - Paperclip is an open source control plane for AI-agent companies.
> - Company skills give agents reusable work instructions.
> - Skill Studio can edit skill files and save version history.
> - Agents can create a skill, but they do not have a first-class update
tool.
> - An agent update needs a version check and safe retry behavior to
prevent lost edits.
> - This pull request adds `update_skill` through the existing company
skill file API.
> - The change keeps company policy, version history, and audit records
in one path.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: the server API, shared validation, and runner tool
catalog.

**Problem or motivation**

An agent can create a company skill but cannot update its `SKILL.md`
through a first-class tool. An unguarded retry can also create duplicate
versions or overwrite a newer edit.

**Proposed solution**

Add `update_skill` with a required current version ID and a retry key.
Route it through the existing skill file API. Reject stale versions and
changed-input retries. Save the version and audit event together.

**Alternatives considered**

A separate write endpoint would duplicate the Skill Studio mutation path
and policy checks. This PR reuses that path instead.

**Roadmap alignment**

This work extends Skills Manager and Skill Studio, which are listed in
`ROADMAP.md`.

**Additional context**

The tool accepts a complete `SKILL.md`, not a partial patch. Callers
must read the current version before they edit it.

## What Changed

- Add optional version and retry fields to the existing skill file
update contract.
- Add a guarded API update with a stable retry receipt and attributed
audit event.
- Add `update_skill` to native and semantic runner tool catalogs, with
mode and policy gates.
- Add unit, integration, protocol, and semantic-tool regression
coverage.
- Document agent use and extend the OpenAPI request contract.

## Verification

- Focused tests and direct server and runner TypeScript checks passed
before this PR.
- `git diff --check` passed after the rebase onto `master`.
- CI passed on the latest PR head, including the full test matrix,
typecheck, and build. Local full typecheck and build stopped because
`cargo` is not installed. The local full test run ended without a
verdict.
- No dedicated end-to-end eval scenario was added or run. The protocol
coverage and semantic-tool test cover the new action deterministically.
- Reviewers can read a skill version, call `update_skill`, repeat the
same key, then try a stale version and a changed-input key. Only the
first edit must create a new version.

## Risks

- File writes and database transactions must stay in sync when a write
fails. The integration tests cover failed writes and retry behavior, but
CI must verify them on the PR head.
- Existing Skill Studio callers do not send the new optional
coordination fields. Their request shape remains valid.

## Model Used

- OpenAI Codex CLI assisted with this change. The runner did not expose
the exact model ID or context window. The agent used code execution and
repository tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details; exact model ID and context window were not exposed)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused tests; full suite
is pending CI)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (53 pass, 4 skip on the latest
head)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 12:37:55 -05:00

4.4 KiB

Skills created by the Runner

The native Runner exposes create_skill in Auto and skill-test work. An agent can save a reusable, single-file skill directly in the company's skill library. The tool is unavailable in Ask and pre-acceptance Plan modes.

The native Runner also exposes update_skill in the same modes. Read the existing skill and its currentVersionId through the company skill API before editing. Supply the complete primary file, not a fragment:

{
  "skillId": "<existing skill UUID>",
  "expectedVersionId": "<current version UUID>",
  "markdown": "---\nname: release-review\ndescription: Review release notes.\n---\n\n# Review\nCheck each release note and its tests.\n",
  "idempotencyKey": "review-update-1"
}

The tool uses the existing Skill Studio file endpoint and company skills.edit policy. It returns a skill ID, path, resulting version ID, and Skill Studio path. Retry a lost response with exactly the same key and inputs; that returns the original receipt without another version or activity record. Reusing the key with different inputs, or updating after the version changes, returns a conflict. After a version conflict, reread the skill and choose a new key for the revised edit. Policy is checked again even on exact retries.

{
  "name": "release-review",
  "description": "Review release notes.",
  "markdown": "---\nname: release-review\ndescription: Review release notes.\n---\n\n# Review\nCheck each release note against its change.\n",
  "idempotencyKey": "release-review-1"
}

The name is lowercase with hyphens. The complete SKILL.md must have matching name and description frontmatter and a nonempty instruction body. An optional slug must equal the name. Company, task, agent and run identity come from the authenticated run; they are not tool inputs.

The tool uses the same company skill policy and storage as Skill Studio. Skills are open by default unless company policy restricts creation. Creating a skill does not assign it to an agent or introduce a new approval step. Once the skill is saved, the task can finish normally if no other work remains.

The result contains the skill ID, name, slug, description, version ID and Studio path. Reuse the same idempotency key and inputs after a lost response. A retry returns the existing skill without another creation event. A key reused with different inputs, or a name that belongs to another skill, returns a conflict. Published files are never replaced as part of a competing creation request. Deleting a managed skill releases its name for a later creation. Deletion uses the same name lock as creation and retains its source files until the database deletion commits. Imported local and project source folders are not removed.

User interface

Successful creation adds a Skill created card to the originating task's feed. The card opens a named skill tab in the task sidebar. Repeated clicks focus that tab. The tab reads the company skill directly, so it is not a second editable copy of the instructions.

Open in Skill Studio opens that same skill for editing. After saving and returning to the task, the sidebar shows the latest saved version. Reloading the task restores the tab. The historical card remains a creation receipt even if the skill is later edited or removed. The sidebar reports missing skills, denied access and retryable load errors explicitly.

Verification

Server regressions cover authenticated creation, company isolation, policy denials and revocation, mode restrictions, initial version/file persistence, idempotency, conflicting retries and concurrent creation. UI tests cover the creation card, saved tabs, frontmatter-free preview and error states.

The companion Runner Eval case create-skill tests provider tool use against the seeded mock control plane, not production storage or company policy. The standalone mock uses a generated production frontmatter parser and schema validator. Regenerate it with node packages/shared/scripts/generate-runner-skill-frontmatter.ts after changing the shared frontmatter contract. A shared-package test checks synchronization. Invalid inputs are covered by deterministic tests: a live model should not be penalized for declining to send a schema-invalid request. Product E2E's create-skill-studio case checks real creation, task completion, the feed card, sidebar, Studio editing and the updated skill after returning to the task. See the eval guide for the difference between these suites.