mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 16:35:27 +02:00
## Thinking Path > - Paperclip is an open source control plane for AI-agent companies. > - Company skills give agents reusable work instructions. > - Skill Studio can edit skill files and save version history. > - Agents can create a skill, but they do not have a first-class update tool. > - An agent update needs a version check and safe retry behavior to prevent lost edits. > - This pull request adds `update_skill` through the existing company skill file API. > - The change keeps company policy, version history, and audit records in one path. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: the server API, shared validation, and runner tool catalog. **Problem or motivation** An agent can create a company skill but cannot update its `SKILL.md` through a first-class tool. An unguarded retry can also create duplicate versions or overwrite a newer edit. **Proposed solution** Add `update_skill` with a required current version ID and a retry key. Route it through the existing skill file API. Reject stale versions and changed-input retries. Save the version and audit event together. **Alternatives considered** A separate write endpoint would duplicate the Skill Studio mutation path and policy checks. This PR reuses that path instead. **Roadmap alignment** This work extends Skills Manager and Skill Studio, which are listed in `ROADMAP.md`. **Additional context** The tool accepts a complete `SKILL.md`, not a partial patch. Callers must read the current version before they edit it. ## What Changed - Add optional version and retry fields to the existing skill file update contract. - Add a guarded API update with a stable retry receipt and attributed audit event. - Add `update_skill` to native and semantic runner tool catalogs, with mode and policy gates. - Add unit, integration, protocol, and semantic-tool regression coverage. - Document agent use and extend the OpenAPI request contract. ## Verification - Focused tests and direct server and runner TypeScript checks passed before this PR. - `git diff --check` passed after the rebase onto `master`. - CI passed on the latest PR head, including the full test matrix, typecheck, and build. Local full typecheck and build stopped because `cargo` is not installed. The local full test run ended without a verdict. - No dedicated end-to-end eval scenario was added or run. The protocol coverage and semantic-tool test cover the new action deterministically. - Reviewers can read a skill version, call `update_skill`, repeat the same key, then try a stale version and a changed-input key. Only the first edit must create a new version. ## Risks - File writes and database transactions must stay in sync when a write fails. The integration tests cover failed writes and retry behavior, but CI must verify them on the PR head. - Existing Skill Studio callers do not send the new optional coordination fields. Their request shape remains valid. ## Model Used - OpenAI Codex CLI assisted with this change. The runner did not expose the exact model ID or context window. The agent used code execution and repository tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details; exact model ID and context window were not exposed) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused tests; full suite is pending CI) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green (53 pass, 4 skip on the latest head) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
122 lines
5.7 KiB
Markdown
122 lines
5.7 KiB
Markdown
# Capability Semantic Tool Catalog
|
|
|
|
The semantic tool catalog is the transport-neutral surface an agent may call. It
|
|
is a frozen set of JSON-schema descriptors — no HTTP method, path, or provider
|
|
detail. Provider-specific shapes are produced only by *bindings*, so every
|
|
provider sees the same operation set.
|
|
|
|
Each protocol action is single-sourced in its own module under
|
|
`src/protocol-actions/`. That module owns the action's policy metadata,
|
|
documentation, examples, and its live and scenario JSON-schema presentations.
|
|
`src/catalog/canonical-operations.ts`, `src/semantic-tools/catalog.ts`, and
|
|
`src/tools/capability-semantic-tool-catalog.ts` are compatibility projections;
|
|
do not add action definitions to them. See
|
|
[catalog reconciliation](../spec/capability/catalog-reconciliation.md).
|
|
|
|
Sources: `src/protocol-actions/`, `src/catalog/canonical-operations.ts`,
|
|
`src/tools/capability-semantic-tool-catalog.ts`,
|
|
`src/tools/capability-semantic-tool-types.ts`, `src/tools/capability-tool-bindings.ts`,
|
|
barrel `src/tools/index.ts`.
|
|
|
|
## Two dispositions, plus a separate control-plane list
|
|
|
|
A tool descriptor's `disposition` is either **`always_agent_tool`** or
|
|
**`optional_agent_tool`**. Control-plane-owned operations are a separate frozen
|
|
list, not a disposition — they are never exposed as tools. See
|
|
[capability disposition](capability-disposition.md) and
|
|
[authorization and exposure](capability-authorization-and-exposure.md).
|
|
|
|
The catalog holds **40 tools**: 14 always-agent tools and 26 optional tools
|
|
across 10 groups.
|
|
|
|
<!-- These counts are drift-checked against src/tools/capability-semantic-tool-catalog.ts
|
|
by src/catalog/catalog-docs.test.ts; update the catalog, not the numbers. -->
|
|
|
|
### Always-agent tools (14)
|
|
|
|
`get_task_context`, `get_task_history`, `list_documents`, `read_document`,
|
|
`list_document_revisions`, `report_progress`, `answer_status_question`,
|
|
`finish_task`, `block_task`, `request_review`, `write_document`,
|
|
`request_human_input`, `register_deliverable`, `inspect_operation_result`.
|
|
|
|
### Optional tools (26), by group
|
|
|
|
| Group | Tools |
|
|
| --- | --- |
|
|
| discovery | `search_tasks`, `list_agents`, `list_projects`, `list_goals` |
|
|
| delegation_dependencies | `create_task`, `reassign_task`, `set_dependencies` |
|
|
| governance | `list_approvals`, `request_approval`, `decide_approval`, `comment_on_approval` |
|
|
| cases | `list_cases`, `upsert_case` |
|
|
| workspace_runtime | `get_workspace_runtime`, `control_workspace_service` |
|
|
| routines | `list_routines`, `manage_routine` |
|
|
| company_skills | `create_skill`, `update_skill`, `list_company_skills`, `sync_company_skills` |
|
|
| secrets | `list_secret_metadata`, `read_secret_value` |
|
|
| portability_admin | `export_company`, `administer_company` |
|
|
| test_escape_hatch | `generic_api_request` |
|
|
|
|
## Descriptor shape
|
|
|
|
Each descriptor carries `operationId`, `version` (1), `title`, `description`,
|
|
input/output JSON schemas, `disposition`, an optional `optionalGroup`,
|
|
`requiredClaims`, and optionally `allowedRoles`, `taskModes`, a `sideEffectClass`
|
|
(`read`, `task_write`, `company_write`, `governance`, `workspace_control`,
|
|
`secret_read`, `admin`, `test_escape_hatch`), an `idempotency` level, `redaction`
|
|
rules, and an abstract `mockCommandMapping`.
|
|
|
|
The `mockCommandMapping` is one of `context_read`, `snapshot_read`,
|
|
`semantic_command`, `operation_result`, or `mock_extension` — describing *what*
|
|
the tool does against the mock, never *how* a transport would carry it.
|
|
|
|
## Transport neutrality
|
|
|
|
Bindings, not descriptors, produce provider shapes, both derived from the same
|
|
`visibleTools.tools` array:
|
|
|
|
- `CapabilityFakeAgentToolBinding` emits `{operationId, description, inputSchema,
|
|
outputSchema}`.
|
|
- `CapabilityCodexToolBinding` emits `{type: "function", name, description, strict:
|
|
true, parameters}`.
|
|
|
|
Because both bindings derive from one array, the fake-agent and Codex operation
|
|
surfaces are byte-identical. The [conformance suite](capability-eval-conformance.md)
|
|
asserts this parity (the 18/18 fake-agent/Codex operation matrix).
|
|
|
|
## Running the tests
|
|
|
|
```sh
|
|
pnpm --filter @paperclipai/paperclip-runner exec vitest run \
|
|
src/tools/capability-semantic-tools.test.ts
|
|
```
|
|
|
|
The "catalog" describe asserts a unique, versioned, provider-neutral catalog and
|
|
that every optional group is present.
|
|
|
|
## Related
|
|
|
|
- [Authorization and exposure](capability-authorization-and-exposure.md)
|
|
- [Mock ControlPlanePort](capability-mock-control-plane-port.md)
|
|
- [Eval conformance](capability-eval-conformance.md)
|
|
|
|
## Reassigning existing tasks
|
|
|
|
`reassign_task` moves another existing company task to an eligible agent. Read
|
|
`search_tasks` first and send the observed `expectedAssigneeActorId` (nullable)
|
|
and `expectedStatusVersion`, the new `assigneeActorId`, a handoff `reason`, and a
|
|
stable `idempotencyKey`. It requires `delegation:tasks:assign` and production
|
|
`tasks:assign` authorization in standard mode. The same task retains its
|
|
documents, dependencies, scope, priority, and history.
|
|
|
|
The old execution and any live goal must stop before ownership changes. Active
|
|
work returns to todo; backlog and blocked work retain their status and are not
|
|
started. An audited reassignment stop does not create a recovery blocker or
|
|
retry the outgoing run when it ends without a semantic completion result.
|
|
Assignment changes advance the observed version, including ordinary
|
|
API changes, to reject stale handoffs. A durable audit receipt makes retries
|
|
idempotent across caller runs; retrying after failed dispatch repairs the wake
|
|
with a stable key and an owner/status/version guard.
|
|
|
|
The current caller task, conversations, human-assigned tasks, completed or
|
|
cancelled tasks, and pending reviews cannot be reassigned with this tool. Use
|
|
`create_task` to delegate new work from the current task. A stop failure or
|
|
concurrent ownership change fails without applying the handoff.
|