Files
PaperClipAI/packages/paperclip-runner/docs/capability-semantic-tools.md
DottaandPaperclip eb049aebf2 feat(skills): let agents update company skills safely (#15049)
## Thinking Path

> - Paperclip is an open source control plane for AI-agent companies.
> - Company skills give agents reusable work instructions.
> - Skill Studio can edit skill files and save version history.
> - Agents can create a skill, but they do not have a first-class update
tool.
> - An agent update needs a version check and safe retry behavior to
prevent lost edits.
> - This pull request adds `update_skill` through the existing company
skill file API.
> - The change keeps company policy, version history, and audit records
in one path.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: the server API, shared validation, and runner tool
catalog.

**Problem or motivation**

An agent can create a company skill but cannot update its `SKILL.md`
through a first-class tool. An unguarded retry can also create duplicate
versions or overwrite a newer edit.

**Proposed solution**

Add `update_skill` with a required current version ID and a retry key.
Route it through the existing skill file API. Reject stale versions and
changed-input retries. Save the version and audit event together.

**Alternatives considered**

A separate write endpoint would duplicate the Skill Studio mutation path
and policy checks. This PR reuses that path instead.

**Roadmap alignment**

This work extends Skills Manager and Skill Studio, which are listed in
`ROADMAP.md`.

**Additional context**

The tool accepts a complete `SKILL.md`, not a partial patch. Callers
must read the current version before they edit it.

## What Changed

- Add optional version and retry fields to the existing skill file
update contract.
- Add a guarded API update with a stable retry receipt and attributed
audit event.
- Add `update_skill` to native and semantic runner tool catalogs, with
mode and policy gates.
- Add unit, integration, protocol, and semantic-tool regression
coverage.
- Document agent use and extend the OpenAPI request contract.

## Verification

- Focused tests and direct server and runner TypeScript checks passed
before this PR.
- `git diff --check` passed after the rebase onto `master`.
- CI passed on the latest PR head, including the full test matrix,
typecheck, and build. Local full typecheck and build stopped because
`cargo` is not installed. The local full test run ended without a
verdict.
- No dedicated end-to-end eval scenario was added or run. The protocol
coverage and semantic-tool test cover the new action deterministically.
- Reviewers can read a skill version, call `update_skill`, repeat the
same key, then try a stale version and a changed-input key. Only the
first edit must create a new version.

## Risks

- File writes and database transactions must stay in sync when a write
fails. The integration tests cover failed writes and retry behavior, but
CI must verify them on the PR head.
- Existing Skill Studio callers do not send the new optional
coordination fields. Their request shape remains valid.

## Model Used

- OpenAI Codex CLI assisted with this change. The runner did not expose
the exact model ID or context window. The agent used code execution and
repository tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details; exact model ID and context window were not exposed)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused tests; full suite
is pending CI)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (53 pass, 4 skip on the latest
head)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 12:37:55 -05:00

5.7 KiB

Capability Semantic Tool Catalog

The semantic tool catalog is the transport-neutral surface an agent may call. It is a frozen set of JSON-schema descriptors — no HTTP method, path, or provider detail. Provider-specific shapes are produced only by bindings, so every provider sees the same operation set.

Each protocol action is single-sourced in its own module under src/protocol-actions/. That module owns the action's policy metadata, documentation, examples, and its live and scenario JSON-schema presentations. src/catalog/canonical-operations.ts, src/semantic-tools/catalog.ts, and src/tools/capability-semantic-tool-catalog.ts are compatibility projections; do not add action definitions to them. See catalog reconciliation.

Sources: src/protocol-actions/, src/catalog/canonical-operations.ts, src/tools/capability-semantic-tool-catalog.ts, src/tools/capability-semantic-tool-types.ts, src/tools/capability-tool-bindings.ts, barrel src/tools/index.ts.

Two dispositions, plus a separate control-plane list

A tool descriptor's disposition is either always_agent_tool or optional_agent_tool. Control-plane-owned operations are a separate frozen list, not a disposition — they are never exposed as tools. See capability disposition and authorization and exposure.

The catalog holds 40 tools: 14 always-agent tools and 26 optional tools across 10 groups.

Always-agent tools (14)

get_task_context, get_task_history, list_documents, read_document, list_document_revisions, report_progress, answer_status_question, finish_task, block_task, request_review, write_document, request_human_input, register_deliverable, inspect_operation_result.

Optional tools (26), by group

Group Tools
discovery search_tasks, list_agents, list_projects, list_goals
delegation_dependencies create_task, reassign_task, set_dependencies
governance list_approvals, request_approval, decide_approval, comment_on_approval
cases list_cases, upsert_case
workspace_runtime get_workspace_runtime, control_workspace_service
routines list_routines, manage_routine
company_skills create_skill, update_skill, list_company_skills, sync_company_skills
secrets list_secret_metadata, read_secret_value
portability_admin export_company, administer_company
test_escape_hatch generic_api_request

Descriptor shape

Each descriptor carries operationId, version (1), title, description, input/output JSON schemas, disposition, an optional optionalGroup, requiredClaims, and optionally allowedRoles, taskModes, a sideEffectClass (read, task_write, company_write, governance, workspace_control, secret_read, admin, test_escape_hatch), an idempotency level, redaction rules, and an abstract mockCommandMapping.

The mockCommandMapping is one of context_read, snapshot_read, semantic_command, operation_result, or mock_extension — describing what the tool does against the mock, never how a transport would carry it.

Transport neutrality

Bindings, not descriptors, produce provider shapes, both derived from the same visibleTools.tools array:

  • CapabilityFakeAgentToolBinding emits {operationId, description, inputSchema, outputSchema}.
  • CapabilityCodexToolBinding emits {type: "function", name, description, strict: true, parameters}.

Because both bindings derive from one array, the fake-agent and Codex operation surfaces are byte-identical. The conformance suite asserts this parity (the 18/18 fake-agent/Codex operation matrix).

Running the tests

pnpm --filter @paperclipai/paperclip-runner exec vitest run \
  src/tools/capability-semantic-tools.test.ts

The "catalog" describe asserts a unique, versioned, provider-neutral catalog and that every optional group is present.

Reassigning existing tasks

reassign_task moves another existing company task to an eligible agent. Read search_tasks first and send the observed expectedAssigneeActorId (nullable) and expectedStatusVersion, the new assigneeActorId, a handoff reason, and a stable idempotencyKey. It requires delegation:tasks:assign and production tasks:assign authorization in standard mode. The same task retains its documents, dependencies, scope, priority, and history.

The old execution and any live goal must stop before ownership changes. Active work returns to todo; backlog and blocked work retain their status and are not started. An audited reassignment stop does not create a recovery blocker or retry the outgoing run when it ends without a semantic completion result. Assignment changes advance the observed version, including ordinary API changes, to reject stale handoffs. A durable audit receipt makes retries idempotent across caller runs; retrying after failed dispatch repairs the wake with a stable key and an owner/status/version guard.

The current caller task, conversations, human-assigned tasks, completed or cancelled tasks, and pending reviews cannot be reassigned with this tool. Use create_task to delegate new work from the current task. A stop failure or concurrent ownership change fails without applying the handoff.