Files
PaperClipAI/packages/paperclip-runner/docs/capability-semantic-tools.md
DottaandPaperclip eb049aebf2 feat(skills): let agents update company skills safely (#15049)
## Thinking Path

> - Paperclip is an open source control plane for AI-agent companies.
> - Company skills give agents reusable work instructions.
> - Skill Studio can edit skill files and save version history.
> - Agents can create a skill, but they do not have a first-class update
tool.
> - An agent update needs a version check and safe retry behavior to
prevent lost edits.
> - This pull request adds `update_skill` through the existing company
skill file API.
> - The change keeps company policy, version history, and audit records
in one path.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: the server API, shared validation, and runner tool
catalog.

**Problem or motivation**

An agent can create a company skill but cannot update its `SKILL.md`
through a first-class tool. An unguarded retry can also create duplicate
versions or overwrite a newer edit.

**Proposed solution**

Add `update_skill` with a required current version ID and a retry key.
Route it through the existing skill file API. Reject stale versions and
changed-input retries. Save the version and audit event together.

**Alternatives considered**

A separate write endpoint would duplicate the Skill Studio mutation path
and policy checks. This PR reuses that path instead.

**Roadmap alignment**

This work extends Skills Manager and Skill Studio, which are listed in
`ROADMAP.md`.

**Additional context**

The tool accepts a complete `SKILL.md`, not a partial patch. Callers
must read the current version before they edit it.

## What Changed

- Add optional version and retry fields to the existing skill file
update contract.
- Add a guarded API update with a stable retry receipt and attributed
audit event.
- Add `update_skill` to native and semantic runner tool catalogs, with
mode and policy gates.
- Add unit, integration, protocol, and semantic-tool regression
coverage.
- Document agent use and extend the OpenAPI request contract.

## Verification

- Focused tests and direct server and runner TypeScript checks passed
before this PR.
- `git diff --check` passed after the rebase onto `master`.
- CI passed on the latest PR head, including the full test matrix,
typecheck, and build. Local full typecheck and build stopped because
`cargo` is not installed. The local full test run ended without a
verdict.
- No dedicated end-to-end eval scenario was added or run. The protocol
coverage and semantic-tool test cover the new action deterministically.
- Reviewers can read a skill version, call `update_skill`, repeat the
same key, then try a stale version and a changed-input key. Only the
first edit must create a new version.

## Risks

- File writes and database transactions must stay in sync when a write
fails. The integration tests cover failed writes and retry behavior, but
CI must verify them on the PR head.
- Existing Skill Studio callers do not send the new optional
coordination fields. Their request shape remains valid.

## Model Used

- OpenAI Codex CLI assisted with this change. The runner did not expose
the exact model ID or context window. The agent used code execution and
repository tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details; exact model ID and context window were not exposed)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused tests; full suite
is pending CI)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (53 pass, 4 skip on the latest
head)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 12:37:55 -05:00

122 lines
5.7 KiB
Markdown

# Capability Semantic Tool Catalog
The semantic tool catalog is the transport-neutral surface an agent may call. It
is a frozen set of JSON-schema descriptors — no HTTP method, path, or provider
detail. Provider-specific shapes are produced only by *bindings*, so every
provider sees the same operation set.
Each protocol action is single-sourced in its own module under
`src/protocol-actions/`. That module owns the action's policy metadata,
documentation, examples, and its live and scenario JSON-schema presentations.
`src/catalog/canonical-operations.ts`, `src/semantic-tools/catalog.ts`, and
`src/tools/capability-semantic-tool-catalog.ts` are compatibility projections;
do not add action definitions to them. See
[catalog reconciliation](../spec/capability/catalog-reconciliation.md).
Sources: `src/protocol-actions/`, `src/catalog/canonical-operations.ts`,
`src/tools/capability-semantic-tool-catalog.ts`,
`src/tools/capability-semantic-tool-types.ts`, `src/tools/capability-tool-bindings.ts`,
barrel `src/tools/index.ts`.
## Two dispositions, plus a separate control-plane list
A tool descriptor's `disposition` is either **`always_agent_tool`** or
**`optional_agent_tool`**. Control-plane-owned operations are a separate frozen
list, not a disposition — they are never exposed as tools. See
[capability disposition](capability-disposition.md) and
[authorization and exposure](capability-authorization-and-exposure.md).
The catalog holds **40 tools**: 14 always-agent tools and 26 optional tools
across 10 groups.
<!-- These counts are drift-checked against src/tools/capability-semantic-tool-catalog.ts
by src/catalog/catalog-docs.test.ts; update the catalog, not the numbers. -->
### Always-agent tools (14)
`get_task_context`, `get_task_history`, `list_documents`, `read_document`,
`list_document_revisions`, `report_progress`, `answer_status_question`,
`finish_task`, `block_task`, `request_review`, `write_document`,
`request_human_input`, `register_deliverable`, `inspect_operation_result`.
### Optional tools (26), by group
| Group | Tools |
| --- | --- |
| discovery | `search_tasks`, `list_agents`, `list_projects`, `list_goals` |
| delegation_dependencies | `create_task`, `reassign_task`, `set_dependencies` |
| governance | `list_approvals`, `request_approval`, `decide_approval`, `comment_on_approval` |
| cases | `list_cases`, `upsert_case` |
| workspace_runtime | `get_workspace_runtime`, `control_workspace_service` |
| routines | `list_routines`, `manage_routine` |
| company_skills | `create_skill`, `update_skill`, `list_company_skills`, `sync_company_skills` |
| secrets | `list_secret_metadata`, `read_secret_value` |
| portability_admin | `export_company`, `administer_company` |
| test_escape_hatch | `generic_api_request` |
## Descriptor shape
Each descriptor carries `operationId`, `version` (1), `title`, `description`,
input/output JSON schemas, `disposition`, an optional `optionalGroup`,
`requiredClaims`, and optionally `allowedRoles`, `taskModes`, a `sideEffectClass`
(`read`, `task_write`, `company_write`, `governance`, `workspace_control`,
`secret_read`, `admin`, `test_escape_hatch`), an `idempotency` level, `redaction`
rules, and an abstract `mockCommandMapping`.
The `mockCommandMapping` is one of `context_read`, `snapshot_read`,
`semantic_command`, `operation_result`, or `mock_extension` — describing *what*
the tool does against the mock, never *how* a transport would carry it.
## Transport neutrality
Bindings, not descriptors, produce provider shapes, both derived from the same
`visibleTools.tools` array:
- `CapabilityFakeAgentToolBinding` emits `{operationId, description, inputSchema,
outputSchema}`.
- `CapabilityCodexToolBinding` emits `{type: "function", name, description, strict:
true, parameters}`.
Because both bindings derive from one array, the fake-agent and Codex operation
surfaces are byte-identical. The [conformance suite](capability-eval-conformance.md)
asserts this parity (the 18/18 fake-agent/Codex operation matrix).
## Running the tests
```sh
pnpm --filter @paperclipai/paperclip-runner exec vitest run \
src/tools/capability-semantic-tools.test.ts
```
The "catalog" describe asserts a unique, versioned, provider-neutral catalog and
that every optional group is present.
## Related
- [Authorization and exposure](capability-authorization-and-exposure.md)
- [Mock ControlPlanePort](capability-mock-control-plane-port.md)
- [Eval conformance](capability-eval-conformance.md)
## Reassigning existing tasks
`reassign_task` moves another existing company task to an eligible agent. Read
`search_tasks` first and send the observed `expectedAssigneeActorId` (nullable)
and `expectedStatusVersion`, the new `assigneeActorId`, a handoff `reason`, and a
stable `idempotencyKey`. It requires `delegation:tasks:assign` and production
`tasks:assign` authorization in standard mode. The same task retains its
documents, dependencies, scope, priority, and history.
The old execution and any live goal must stop before ownership changes. Active
work returns to todo; backlog and blocked work retain their status and are not
started. An audited reassignment stop does not create a recovery blocker or
retry the outgoing run when it ends without a semantic completion result.
Assignment changes advance the observed version, including ordinary
API changes, to reject stale handoffs. A durable audit receipt makes retries
idempotent across caller runs; retrying after failed dispatch repairs the wake
with a stable key and an owner/status/version guard.
The current caller task, conversations, human-assigned tasks, completed or
cancelled tasks, and pending reviews cannot be reassigned with this tool. Use
`create_task` to delegate new work from the current task. A stop failure or
concurrent ownership change fails without applying the handoff.