Files
PaperClipAI/doc/runner-created-skills.md
T
DottaandPaperclip 9fd2e50310 feat: create company skills from runner tasks (#13538)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Runner gives agents tools to change company resources.
> - Users need agents to save reusable skills during a task.
> - A saved skill needs a visible result that users can inspect and
edit.
> - This pull request adds `create_skill` and a task feed card linked to
Skill Studio.
> - Users can open the saved skill from the task and edit the same
resource.

## Linked Issues or Issue Description

**Subsystem affected**

Runner tools, company skill storage, task feed, and Skill Studio.

**Problem or motivation**

The Runner has no dedicated tool to create a company skill. A user
cannot follow a creation result from the task feed to the saved skill.

**Proposed solution**

Add a company-scoped `create_skill` tool. Save the skill with the
existing company policy. Add one creation card to the task. Open a named
sidebar tab from that card. Let the user open the same skill in Skill
Studio.

**Alternatives considered**

An agent can write a local file, but that file is not a company skill. A
second document copy in the task would become stale after a Studio edit.
The sidebar therefore reads the saved skill directly.

**Roadmap alignment**

This extends the shipped Skills Manager, Skill Studio, and Skills Store
milestone. The maintainer requested and approved this scope. Search
found no duplicate `create_skill` PR or issue. Related UI validation
work: #8715. This PR does not change that validation display.

## What Changed

- Add the real Runner tool, its contract, and its mock implementation.
- Validate the complete SKILL.md and derive company, task, agent, and
run identity from authentication.
- Apply the existing company skill policy. Do not assign the skill to an
agent.
- Make keyed retries return one skill and one creation event. Reject
conflicting retries.
- Make concurrent file creation safe. Never replace an existing
published skill during creation.
- Add a creation card, a named sidebar tab, and an Open in Skill Studio
action.
- Show saved Studio edits when the user returns to the task.
- Add storage, policy, mode, retry, UI, and Product E2E tests. Document
the tool.
- Fix deleted-name reuse, onboarding panel persistence, immediate feed
refresh, and mock validation parity from review.
- Serialize Studio file edits and renames with skill deletion and
recreation. Reject stale editor requests before they can change a
replacement skill.
- Generate the standalone mock parser and validator from the production
contract. Use portable UUIDs so the browser scenario bundle builds.

## Verification

- All latest-head PR checks pass on `145dd76a5`, including all server
shards, browser E2E, Runner verification, build, typecheck, and release
dry run. Greptile: 5/5 with no open findings. An interrupted CI runner
was retried successfully.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm check:token-gates`: passed.
- Review regressions: 73 storage tests, 6 real API tests, 63 UI tests,
and 61 semantic runtime tests passed. Parser synchronization passed.
- CI exposed existing fire-and-forget Sentry test races. Reproduced the
resumption race locally, then synchronized the related sweep and
finalizer assertions on the actual report; all 27 tests across the three
affected files pass.
- Runner scenario browser build and strict content-security-policy
check: passed.
- Runner suite: 2,012 tests passed; 10 skipped.
- `pnpm test:run`: the general-server batch had 12,416 passes and two
failures. The old tool-count assertion was fixed; all 16 authority tests
then passed. The chat webhook test had a socket error; it passed four
isolated reruns.
- Both workspace test groups passed. The isolated route suites
completed. Two socket failures in the initial route batches passed on
individual reruns; all remaining 61 files passed.
- Product E2E `create-skill-studio`: passed with local Codex and local
ACPX Claude.
- Manual browser test: submit a task, observe the real tool call and
creation card, open the sidebar, edit in Studio, save, and return. The
task reached Done. The saved second revision and sidebar tab survived a
server restart.
- The new companion headless Runner Eval passed. Companion coverage PR:
https://github.com/paperclipai/paperclip-evals/pull/23. Daytona was not
run because no immutable runner image was configured.

## Risks

- Database writes and local file writes cannot share one transaction.
Recovery accepts only an exact file-for-file retry after a database
rollback. Conflicting files remain untouched.
- The sidebar displays the current skill. The feed card remains the
historical creation receipt.
- No database migration, dependency, or workflow change is included.
- Remote Daytona behavior still needs a run with a configured immutable
image.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) handled design, integration, review, and
browser verification. OpenAI `gpt-5.6-luna` assisted with bounded
implementation and eval work. Both used code execution and tool access.
The host did not expose the context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 13:01:58 -05:00

3.4 KiB

Skills created by the Runner

The native Runner exposes create_skill in Auto and skill-test work. An agent can save a reusable, single-file skill directly in the company's skill library. The tool is unavailable in Ask and pre-acceptance Plan modes.

{
  "name": "release-review",
  "description": "Review release notes.",
  "markdown": "---\nname: release-review\ndescription: Review release notes.\n---\n\n# Review\nCheck each release note against its change.\n",
  "idempotencyKey": "release-review-1"
}

The name is lowercase with hyphens. The complete SKILL.md must have matching name and description frontmatter and a nonempty instruction body. An optional slug must equal the name. Company, task, agent and run identity come from the authenticated run; they are not tool inputs.

The tool uses the same company skill policy and storage as Skill Studio. Skills are open by default unless company policy restricts creation. Creating a skill does not assign it to an agent or introduce a new approval step. Once the skill is saved, the task can finish normally if no other work remains.

The result contains the skill ID, name, slug, description, version ID and Studio path. Reuse the same idempotency key and inputs after a lost response. A retry returns the existing skill without another creation event. A key reused with different inputs, or a name that belongs to another skill, returns a conflict. Published files are never replaced as part of a competing creation request. Deleting a managed skill releases its name for a later creation. Deletion uses the same name lock as creation and retains its source files until the database deletion commits. Imported local and project source folders are not removed.

User interface

Successful creation adds a Skill created card to the originating task's feed. The card opens a named skill tab in the task sidebar. Repeated clicks focus that tab. The tab reads the company skill directly, so it is not a second editable copy of the instructions.

Open in Skill Studio opens that same skill for editing. After saving and returning to the task, the sidebar shows the latest saved version. Reloading the task restores the tab. The historical card remains a creation receipt even if the skill is later edited or removed. The sidebar reports missing skills, denied access and retryable load errors explicitly.

Verification

Server regressions cover authenticated creation, company isolation, policy denials and revocation, mode restrictions, initial version/file persistence, idempotency, conflicting retries and concurrent creation. UI tests cover the creation card, saved tabs, frontmatter-free preview and error states.

The companion Runner Eval case create-skill tests provider tool use against the seeded mock control plane, not production storage or company policy. The standalone mock uses a generated production frontmatter parser and schema validator. Regenerate it with node packages/shared/scripts/generate-runner-skill-frontmatter.ts after changing the shared frontmatter contract. A shared-package test checks synchronization. Invalid inputs are covered by deterministic tests: a live model should not be penalized for declining to send a schema-invalid request. Product E2E's create-skill-studio case checks real creation, task completion, the feed card, sidebar, Studio editing and the updated skill after returning to the task. See the eval guide for the difference between these suites.