## Thinking Path > - Paperclip manages AI agents and their task execution. > - Agents need a durable way to return to work after a delayed check. > - The issue monitor scheduler already provides a one-shot wake for an assignee. > - Native runners reject generic execution-policy writes and had no bound monitor tool. > - Scheduling alone is insufficient because native completion also needs to accept a timed wait. > - This pull request adds an authorized monitor tool and connects it to completion and the existing scheduler. > - An agent can now schedule its next check, end the run, and resume on the same task. ## Linked Issues or Issue Description **What happened?** A native runner could not set its own task monitor. `call_api` correctly rejected execution-policy writes, while `schedule_wake` had no production binding. `paperclip_finish` also rejected monitor waits. **Expected behavior** A standard native run can set a one-shot monitor on its current task or another accessible task assigned to the same agent. After a confirmed schedule on the current task, it can yield. The scheduler later delivers `issue_monitor_due`. **Steps to reproduce** 1. Start a standard native task. 2. Ask the agent to check the task again later and end its current run. 3. Inspect available tools and try the generic issue execution-policy update. 4. Observe the missing native tool and the lifecycle-write denial. Related PRs: #14680 concerns monitor notes in the shared wake prompt. #11919 changes attempt-limit scope. This PR adds native scheduling and completion authority and retains the existing cumulative attempt bounds. It does not depend on either PR. ## What Changed - Add provider-neutral `set_task_monitor` with a default current-task target, future timestamp, required notes, existing bounds, and explicit clearing. - Check company, task visibility, ownership, runtime permissions, work mode, and active-run authority. Preserve review-only restrictions. Reject the reserved server-owned quota-recovery name before saving or accepting a native wait. - Commit the monitor, audit event, and retry receipt together. Retry receipts survive a successor run without re-arming cleared or consumed timers. - Permit `paperclip_finish` to yield to a persisted monitor. Recheck ownership and the schedule when committing final disposition. Release execution without an immediate continuation. - Preserve due monitors during native execution. Fence wake admission and consumption against replacement, clearing, reassignment, and completion. Preserve unrelated review policy. - Expose scheduled and consumed monitor instructions in task context. Update provider schemas, Rust validation, generated contracts, and execution documentation. - Add an opt-in live Codex smoke script with isolated data and explicit run/session/runner/process evidence. ## Verification - Repository `pnpm -r typecheck` and `pnpm build` passed after rebase. Server typecheck passed again after review fixes. All CI test shards pass on `3def77b1b`, including runner TypeScript/Rust, server, serialized server, workspace, and browser tests. All CI gates are green, including the canary dry run. Greptile is 5/5 on the same commit with zero unresolved threads. - The local monolithic `pnpm test:run`, started before the rebase, was interrupted after current-head CI test coverage passed. It is not counted as a standalone full-suite pass; the focused local regression suites passed. - Targeted server tests cover scheduling, replacement, clearing, policy preservation, cumulative bounds, cross-run retries, permissions, provider-neutral discovery, review restrictions, completion authority, and scheduler/finalizer races. - Runner contract/catalog/semantic tests and Rust terminal-tool tests cover the new operation and monitor completion. - Live Codex test passed twice (latest live run on `e0bcd63e6`) in a temporary database and workspace, with a 300,000 ms warm window. First run `67bd7709-c089-4d4a-9d2b-0d6b618a34b0` yielded at `2026-10-07T13:15:03.274Z`. Second run `97c323f9-595a-4cc5-a007-db5a2fbb937c` started at `13:15:30.952Z`, received `issue_monitor_due`, and completed the same task. Exactly one monitor wake was recorded. - Both live runs used native session `7b1dd753-1c9b-4e7a-b22f-a125dbc3748c`, runner `e90d9a1b-3502-4ee3-b15e-edc024c555d4`, provider session `01a11680-6d01-70c0-9a55-db7246ed66c3`, and PID `64218` with the same process start time. This proves warm reuse for that local Codex test, not only successful scheduling. - Reproduce the paid live test with `node --import ./server/node_modules/tsx/dist/loader.mjs server/scripts/smoke-native-task-monitor.ts --run`, with the installed Codex binary on `PATH` and a valid local login. ## Risks - The scheduler now defers monitor dispatch while the task has an active native run. A stuck run still depends on the existing recovery lifecycle. - Idempotency uses the existing run ledger; no table or migration is added. - Other providers share the tested tool and completion contracts. Only Codex received a live model test. - Existing `call_api` lifecycle restrictions remain enforced. Monitor waits do not bypass task blockers, reviews, or approvals. ## Model Used OpenAI Codex, GPT-6 family, with tool use, code execution, and TypeScript/Rust editing. The session does not expose the exact deployed model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
90 KiB
Paperclip agent operation groups
Status: canonical explanatory contract for the Paperclip runner V1 surface.
This document keeps three independent meanings of group separate. PRP families describe wire evidence and controller commands; capability placement decides who owns an operation; behavioral eval groups organize the 106 scenario corpus. None of the three axes can be used as a substitute for another.
The generated totals are 105 PRP events in 31 event families, 18 controller commands in 7 command families, 10 control-plane operations, 57 reconciled semantic operations (18 always, 39 optional), and 106 scenarios in 16 behavior groups.
Axis 1: PRP v1 event and command families
PRP records ordered, replayable execution evidence. It is not the model's Paperclip tool API. Events are ordered per sourceInstanceId; duplicate sourceEventId values are idempotent; source-sequence gaps remain evidence; replay is side-effect free; unknown required versions fail closed.
Event families
| Family | Purpose | Events | Count |
|---|---|---|---|
runner |
Runner connection, drain, and diagnostic lifecycle. | runner.connectedrunner.reconnectedrunner.reconciledrunner.disconnectedrunner.drainingrunner.backpressurerunner.suspendingrunner.suspendedrunner.stoppedrunner.diagnostic |
10 |
runtime |
Runner phase transitions. | runtime.phase.changed |
1 |
sandbox |
Sandbox resource measurements. | sandbox.metric |
1 |
workspace |
Workspace readiness. | workspace.readyworkspace.change.updatedworkspace.diff.recordedworkspace.file.referenced |
4 |
harness |
Provider harness startup, readiness, exit, and diagnostics. | harness.startingharness.readyharness.exitedharness.diagnostic |
4 |
plan |
Complete provider-authored within-turn checklist snapshots, separate from durable Paperclip Plan documents. | plan.updated |
1 |
tool |
Provider-neutral process, MCP, dynamic, and built-in execution activity. | tool.execution.startedtool.execution.progressedtool.execution.completed |
3 |
research |
Provider-reported search, page-open, and in-page research activity. | research.startedresearch.progressedresearch.completed |
3 |
delegation |
Child-agent delegation lifecycle and aggregate status. | delegation.starteddelegation.updateddelegation.completed |
3 |
model |
Requested/effective model routing and verification state. | model.route.changedmodel.verification.updated |
2 |
context |
Context-window compaction markers without hidden summaries. | context.compacted |
1 |
artifact |
Authorized artifact viewing and structured generated outputs. | artifact.viewedartifact.generated |
2 |
review |
Provider review-mode state, separate from Paperclip authority. | review.mode.changed |
1 |
hook |
Bounded provider hook lifecycle and blocking outcomes. | hook.startedhook.completed |
2 |
memory |
Authorized or unavailable memory citation references. | memory.citation.referenced |
1 |
safety |
Provider safety review state attached to governed work. | safety.review.startedsafety.review.completed |
2 |
terminal |
Content-free terminal input activity metadata. | terminal.input.sent |
1 |
wait |
Intentional provider waits, distinct from warm idle and human input. | wait.startedwait.completed |
2 |
provider |
Redacted provider notices and actionable warnings. | provider.notice.recorded |
1 |
session |
Provider-neutral session open, resume, reconciliation, close, and failure. | session.startingsession.startedsession.resumingsession.resumedsession.reconciledsession.updatedsession.closedsession.failed |
8 |
turn |
Model turn submission through terminal turn disposition. | turn.submittedturn.acceptedturn.startedturn.completedturn.failedturn.interruptedturn.cancelled |
7 |
item |
Provider-neutral model/tool item lifecycle. | item.starteditem.deltaitem.completeditem.failed |
4 |
usage |
Provider/model-attributed usage and accounting boundaries. | usage.reported |
1 |
semantic_tool |
Canonical authorized Paperclip tool input and result evidence. | semantic_tool.inputsemantic_tool.resultsemantic_tool.reconciled |
3 |
mcp_app |
MCP App discovery, initialization, tool, action, host-context, and teardown evidence. | mcp_app.discoveredmcp_app.resource.resolvedmcp_app.initializingmcp_app.readymcp_app.tool_inputmcp_app.tool_resultmcp_app.action.requestedmcp_app.action.resolvedmcp_app.host_context.changedmcp_app.failedmcp_app.teardown |
11 |
runtime_request |
Runtime permission/input request lifecycle. | runtime_request.createdruntime_request.resolvedruntime_request.expiredruntime_request.cancelled |
4 |
interaction |
Issue-thread interaction proposal, materialization, response, delivery, and rejection. | interaction.request.proposedinteraction.request.materializedinteraction.request.rejectedinteraction.response.progressedinteraction.response.resolvedinteraction.response.delivered |
6 |
run |
Structured result negotiation and terminal run outcome. | run.attachedrun.detachedrun.result.proposedrun.result.acceptedrun.result.rejectedrun.terminal |
6 |
attention |
Continuation and attention routing lifecycle. | attention.request.proposedattention.request.routedattention.request.resolvedattention.request.expiredattention.request.superseded |
5 |
work |
Recorded work-assessment evidence. | work.assessment.recorded |
1 |
issue |
Issue-status decision proposal outcome and application evidence. | issue.status.decision.recordedissue.status.decision.appliedissue.status.decision.rejectedissue.status.decision.superseded |
4 |
Controller-command families
| Family | Purpose | Commands | Count |
|---|---|---|---|
run |
Prepare or cancel a run. | run.preparerun.attachrun.cancel |
3 |
session |
Open, snapshot, or close a normalized provider session. | session.opensession.snapshotsession.closesession.budget.increasesession.destroy |
5 |
turn |
Start, steer, interrupt, or stop a model turn. | turn.startturn.steerturn.interruptturn.stop |
4 |
request |
Resolve a pending runtime request. | request.resolve |
1 |
interaction |
Acknowledge delivery of an interaction response. | interaction.receipt |
1 |
semantic_tool |
Returns an authorized, correlated Paperclip tool result to the provider through runnerd. | semantic_tool.result |
1 |
runner |
Drain or shut down the runner process. | runner.drainrunner.suspendrunner.shutdown |
3 |
Axis 2: capability placement
Placement has exactly three outcomes:
control_plane_owned: Paperclip or the runner performs the operation; it is absent from the model tool catalog.always_agent_tool: every eligible active-task run receives the operation after task-mode and actor checks.optional_agent_tool: the operation is exposed only when every declared claim, task-mode, role, and policy condition passes.
Control-plane-owned operations
| Operation | Why it is not a model tool |
|---|---|
checkout_task |
Atomic checkout and execution-lock ownership. |
release_task |
Run cleanup and checkout release. |
select_work |
Scoped wake/inbox work selection. |
route_wake |
Attention, blocker, interaction, approval, and continuation routing. |
enforce_budget |
Budget hard stop, pause, and run-stop reason. |
append_audit_record |
Immutable mutation evidence. |
persist_run |
Durable run/event/checkpoint persistence. |
replay_run |
Side-effect-free replay and reconstruction. |
schedule_blocker_wake |
Dependency-resolution wake scheduling. |
reconcile_run |
Work assessment, terminal status arbitration, and recovery reconciliation. |
Always-agent operations (18)
answer_status_question, block_task, finish_task, get_agent_instruction_history, get_task_context, get_task_history, inspect_operation_result, list_document_revisions, list_documents, read_agent_instructions, read_document, register_deliverable, report_progress, request_human_input, request_review, restore_agent_instructions, update_agent_instructions, write_document.
Optional operations (39) and grant groups (14)
Grant groups are documentation/exposure bundles, not additional authority. The operation descriptor's exact requiredClaims remains decisive.
| Grant group | Operations | Required claims represented | Purpose |
|---|---|---|---|
feedback |
submit_complaintsubmit_suggestion |
none | Internally attributed complaints and improvement suggestions; no additional grant required beyond active run authority. |
discovery |
search_taskslist_agentsget_agentlist_projectslist_goalsset_task_title |
discovery:agents:readdiscovery:goals:readdiscovery:projects:readdiscovery:tasks:read |
Company-visible task, agent, project, and goal discovery. |
projects |
create_projectlist_project_repositories |
none | Project creation and authorized repository discovery through the live company/run authority. |
delegation_dependencies |
create_taskreassign_taskset_dependencieshire_agent |
delegation:agents:createdelegation:tasks:assigndelegation:tasks:createdependencies:write |
Create or reassign delegated work and maintain dependency edges. |
governance |
list_approvalsget_approvalget_approval_contextrequest_approvaldecide_approvalcomment_on_approval |
governance:approvals:commentgovernance:approvals:decidegovernance:approvals:readgovernance:approvals:request |
Read, request, comment on, and decide approvals under governed-action checks. |
cases |
list_casesupsert_case |
cases:readcases:write |
Read and update case summaries without reusing issue-document authority. |
workspace_runtime |
get_workspace_runtimecontrol_workspace_service |
workspace:controlworkspace:read |
Inspect and control the active issue workspace runtime. |
wake_scheduling |
schedule_wakeset_task_monitor |
control_plane:wakes |
Schedule a bounded continuation wake when the current task owns the future check. |
routines |
list_routinesmanage_routine |
routines:readroutines:write |
Inspect or manage company routines. |
company_skills |
create_skillupdate_skilllist_company_skillssync_company_skills |
company_skills:readcompany_skills:write |
Create, update, or inspect company skills and synchronize the current agent's skills. |
secrets |
list_secret_metadataread_secret_value |
secrets:metadata:readsecrets:values:read |
Inspect secret metadata or use a brokered secret value without exposing plaintext evidence. |
portability_admin |
export_companyadminister_company |
company:adminportability:export |
Export portable company state; broad administration remains deferred until split into governed operations. |
test_escape_hatch |
generic_api_request |
test:generic_api_request |
Controlled skill-test transport only; never product coverage. |
api_fallback |
search_apicall_api |
api:callapi:discover |
Discover and invoke HTTP API operations unsupported by available dedicated tools. Production authority and route authorization remain required. |
Reconciled semantic-operation ledger
live_codex means a live provider dispatcher exists. Production binding is tracked separately: bound rows are advertised by Paperclip's run-scoped authority over the shared PRP route, while audit_pending rows remain unavailable to production agents. generic_api_request is test-only and cannot satisfy product coverage.
| Operation | Placement | Claims | Modes / roles | Side effect | Idempotency | Redacts | Mock | Catalogs / current runner | Production / PRP evidence |
|---|---|---|---|---|---|---|---|---|---|
administer_company |
optional_agent_tool |
company:admin |
standardskill_testroles: boardceoadmin |
admin |
required |
no | mock_extension:company.admin |
scenarioscenario_mock |
unboundcompany admin/portability item event plus audit record catalog PRP status: audit_pending |
answer_status_question |
always_agent_tool |
none | standardaskplanningskill_test |
task_write |
required |
no | semantic_command:report_progress |
scenario + livelive_codex |
unboundsemantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events catalog PRP status: audit_pending |
block_task |
always_agent_tool |
none | standardskill_test |
task_write |
required |
no | semantic_command:block_task |
scenario + livelive_codex |
unboundsemantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events catalog PRP status: audit_pending |
call_api |
optional_agent_tool |
api:call |
standardaskplanningskill_test |
company_write |
none |
no | inline/no mapping | livelive_codex |
PaperclipRunnerToolAuthorityAuthenticated PRP tool input/result and existing HTTP route authorization/activity records. catalog PRP status: bound |
comment_on_approval |
optional_agent_tool |
governance:approvals:comment |
standardaskplanningskill_test |
governance |
required |
no | semantic_command:comment_on_approval |
scenario + livelive_codex |
unboundapproval lifecycle plus governed-wait continuation and audit events catalog PRP status: audit_pending |
control_workspace_service |
optional_agent_tool |
workspace:control |
standardskill_test |
workspace_control |
required |
no | semantic_command:control_workspace_service |
scenario + livelive_codex |
unboundworkspace service lifecycle event catalog PRP status: audit_pending |
create_project |
optional_agent_tool |
none | standardskill_test |
company_write |
required |
no | inline/no mapping | livelive_codex |
PaperclipRunnerToolAuthorityAuthenticated project tools, persisted projects and repository workspaces, and run-bound activity. catalog PRP status: bound |
create_skill |
optional_agent_tool |
none | standardskill_test |
company_write |
required |
no | semantic_command:create_skill |
scenario + livelive_codex |
PaperclipRunnerToolAuthorityAuthenticated company skill API, saved skill and initial version, and task-bound creation activity. catalog PRP status: bound |
create_task |
optional_agent_tool |
delegation:tasks:create |
standardskill_test |
company_write |
required |
no | semantic_command:create_task |
scenario + livelive_codex |
issues.create / issues.createChildsemantic-operation item event plus company-entity state diff and audit record catalog PRP status: bound |
decide_approval |
optional_agent_tool |
governance:approvals:decide |
standardskill_testroles: boardapproversecurity |
governance |
required |
no | semantic_command:decide_approval |
scenario + livelive_codex |
unboundapproval lifecycle plus governed-wait continuation and audit events catalog PRP status: audit_pending |
export_company |
optional_agent_tool |
portability:export |
standardskill_test |
admin |
required |
no | mock_extension:portability.export |
scenarioscenario_mock |
unboundcompany admin/portability item event plus audit record catalog PRP status: audit_pending |
finish_task |
always_agent_tool |
none | standardskill_test |
task_write |
required |
no | semantic_command:finish_task |
scenario + livelive_codex |
unboundsemantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events catalog PRP status: audit_pending |
generic_api_request |
optional_agent_tool |
test:generic_api_request |
skill_test |
test_escape_hatch |
required |
yes | mock_extension:test.generic_api |
scenario + livetest_only |
unboundtest-only; excluded from product PRP evidence catalog PRP status: audit_pending |
get_agent |
optional_agent_tool |
discovery:agents:read |
standardaskplanningskill_test |
read |
none |
no | inline/no mapping | livelive_codex |
unboundread projection surfaced via a tool-result item event; no control-plane state diff catalog PRP status: audit_pending |
get_agent_instruction_history |
always_agent_tool |
none | standardaskplanningskill_test |
read |
none |
no | inline/no mapping | livelive_codex |
agentInstructionRevisionServicetool-result item event plus canonical revision receipt for writes; no scenario mock coverage catalog PRP status: bound |
get_approval |
optional_agent_tool |
governance:approvals:read |
standardaskplanningskill_test |
read |
none |
no | inline/no mapping | livelive_codex |
unboundread projection surfaced via a tool-result item event; no control-plane state diff catalog PRP status: audit_pending |
get_approval_context |
optional_agent_tool |
governance:approvals:read |
standardaskplanningskill_test |
read |
none |
no | inline/no mapping | livelive_codex |
unboundread projection surfaced via a tool-result item event; no control-plane state diff catalog PRP status: audit_pending |
get_task_context |
always_agent_tool |
none | standardaskplanningskill_test |
read |
none |
no | context_read:active_task |
scenario + livelive_codex |
PaperclipRunnerToolAuthority active issue/run + accepted plan revisionbound company/assignment query plus exact accepted document revision projection catalog PRP status: bound |
get_task_history |
always_agent_tool |
none | standardaskplanningskill_test |
read |
none |
no | snapshot_read:active_task_history |
scenario + livelive_codex |
unboundread projection surfaced via a tool-result item event; no control-plane state diff catalog PRP status: audit_pending |
get_workspace_runtime |
optional_agent_tool |
workspace:read |
standardaskplanningskill_test |
read |
none |
no | snapshot_read:active_task_workspace |
scenario + livelive_codex |
unboundread projection surfaced via a tool-result item event; no control-plane state diff catalog PRP status: audit_pending |
hire_agent |
optional_agent_tool |
delegation:agents:create |
standardskill_test |
company_write |
required |
no | inline/no mapping | livelive_codex |
PaperclipRunnerToolAuthorityAuthenticated PRP tool input/result and agent-hire state diff with activity record. catalog PRP status: bound |
inspect_operation_result |
always_agent_tool |
none | standardaskplanningskill_test |
read |
none |
no | operation_result |
scenarioscenario_mock |
unboundread projection surfaced via a tool-result item event; no control-plane state diff catalog PRP status: audit_pending |
list_agents |
optional_agent_tool |
discovery:agents:read |
standardaskplanningskill_test |
read |
none |
no | snapshot_read:company_actors |
scenario + livelive_codex |
unboundread projection surfaced via a tool-result item event; no control-plane state diff catalog PRP status: audit_pending |
list_approvals |
optional_agent_tool |
governance:approvals:read |
standardaskplanningskill_test |
read |
none |
no | snapshot_read:company_approvals |
scenario + livelive_codex |
unboundread projection surfaced via a tool-result item event; no control-plane state diff catalog PRP status: audit_pending |
list_cases |
optional_agent_tool |
cases:read |
standardskill_test |
read |
none |
no | mock_extension:cases.list |
scenarioscenario_mock |
unboundread projection surfaced via a tool-result item event; no control-plane state diff catalog PRP status: audit_pending |
list_company_skills |
optional_agent_tool |
company_skills:read |
standardskill_test |
read |
none |
no | mock_extension:company_skills.list |
scenarioscenario_mock |
unboundread projection surfaced via a tool-result item event; no control-plane state diff catalog PRP status: audit_pending |
list_document_revisions |
always_agent_tool |
none | standardaskplanningskill_test |
read |
none |
no | snapshot_read:active_task_document_revisions |
scenario + livelive_codex |
unboundread projection surfaced via a tool-result item event; no control-plane state diff catalog PRP status: audit_pending |
list_documents |
always_agent_tool |
none | standardaskplanningskill_test |
read |
none |
no | snapshot_read:active_task_documents |
scenario + livelive_codex |
unboundread projection surfaced via a tool-result item event; no control-plane state diff catalog PRP status: audit_pending |
list_goals |
optional_agent_tool |
discovery:goals:read |
standardskill_test |
read |
none |
no | mock_extension:discovery.goals |
scenarioscenario_mock |
unboundread projection surfaced via a tool-result item event; no control-plane state diff catalog PRP status: audit_pending |
list_project_repositories |
optional_agent_tool |
none | standardaskplanningskill_test |
read |
none |
no | inline/no mapping | livelive_codex |
PaperclipRunnerToolAuthorityAuthenticated project tools, persisted projects and repository workspaces, and run-bound activity. catalog PRP status: bound |
list_projects |
optional_agent_tool |
discovery:projects:read |
standardaskplanningskill_test |
read |
none |
no | mock_extension:discovery.projects |
scenario + livelive_codex |
PaperclipRunnerToolAuthorityAuthenticated project tools, persisted projects and repository workspaces, and run-bound activity. catalog PRP status: bound |
list_routines |
optional_agent_tool |
routines:read |
standardskill_test |
read |
none |
no | mock_extension:routines.list |
scenarioscenario_mock |
unboundread projection surfaced via a tool-result item event; no control-plane state diff catalog PRP status: audit_pending |
list_secret_metadata |
optional_agent_tool |
secrets:metadata:read |
standardskill_test |
read |
none |
no | mock_extension:secrets.metadata |
scenarioscenario_mock |
unboundread projection surfaced via a tool-result item event; no control-plane state diff catalog PRP status: audit_pending |
manage_routine |
optional_agent_tool |
routines:write |
standardskill_test |
admin |
required |
no | mock_extension:routines.manage |
scenarioscenario_mock |
unboundcompany admin/portability item event plus audit record catalog PRP status: audit_pending |
read_agent_instructions |
always_agent_tool |
none | standardaskplanningskill_test |
read |
none |
no | inline/no mapping | livelive_codex |
agentInstructionRevisionServicetool-result item event plus canonical revision receipt for writes; no scenario mock coverage catalog PRP status: bound |
read_document |
always_agent_tool |
none | standardaskplanningskill_test |
read |
none |
no | snapshot_read:active_task_document |
scenario + livelive_codex |
unboundread projection surfaced via a tool-result item event; no control-plane state diff catalog PRP status: audit_pending |
read_secret_value |
optional_agent_tool |
secrets:values:read |
standardskill_test |
secret_read |
none |
yes | mock_extension:secrets.value |
scenarioscenario_mock |
unboundredacted tool-result item event; secret value never reaches the wire catalog PRP status: audit_pending |
reassign_task |
optional_agent_tool |
delegation:tasks:assign |
standard |
company_write |
required |
no | semantic_command:reassign_task |
scenario + livelive_codex |
PaperclipRunnerToolAuthority.reassign_tasksemantic-operation item event plus company-entity state diff and audit record catalog PRP status: bound |
register_deliverable |
always_agent_tool |
none | standardplanningskill_test |
task_write |
required |
no | semantic_command:register_deliverable |
scenario + livelive_codex |
unboundsemantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events catalog PRP status: audit_pending |
report_progress |
always_agent_tool |
none | standardaskplanningskill_test |
task_write |
required |
no | semantic_command:report_progress |
scenario + livelive_codex |
unboundsemantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events catalog PRP status: audit_pending |
request_approval |
optional_agent_tool |
governance:approvals:request |
standardskill_test |
governance |
required |
no | semantic_command:request_approval |
scenario + livelive_codex |
unboundapproval lifecycle plus governed-wait continuation and audit events catalog PRP status: audit_pending |
request_human_input |
always_agent_tool |
none | standardplanningaskskill_test |
task_write |
required |
no | semantic_command:request_human_input |
scenario + livelive_codex |
unboundsemantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events catalog PRP status: audit_pending |
request_review |
always_agent_tool |
none | standardskill_test |
task_write |
required |
no | semantic_command:request_review |
scenario + livelive_codex |
unboundsemantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events catalog PRP status: audit_pending |
restore_agent_instructions |
always_agent_tool |
none | standardplanning |
company_write |
recommended |
no | inline/no mapping | livelive_codex |
agentInstructionRevisionServicetool-result item event plus canonical revision receipt for writes; no scenario mock coverage catalog PRP status: bound |
schedule_wake |
optional_agent_tool |
control_plane:wakes |
standardskill_test |
task_write |
required |
no | inline/no mapping | livelive_codex |
unboundsemantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events catalog PRP status: audit_pending |
search_api |
optional_agent_tool |
api:discover |
standardaskplanningskill_test |
read |
none |
no | inline/no mapping | livelive_codex |
PaperclipRunnerToolAuthorityAuthenticated PRP tool input/result and existing HTTP route authorization/activity records. catalog PRP status: bound |
search_tasks |
optional_agent_tool |
discovery:tasks:read |
standardaskplanningskill_test |
read |
none |
no | snapshot_read:company_tasks |
scenario + livelive_codex |
unboundread projection surfaced via a tool-result item event; no control-plane state diff catalog PRP status: audit_pending |
set_dependencies |
optional_agent_tool |
dependencies:write |
standardskill_test |
company_write |
required |
no | semantic_command:set_dependencies |
scenario + livelive_codex |
issues.update.blockedByIssueIdssemantic-operation item event plus company-entity state diff and audit record catalog PRP status: bound |
set_task_monitor |
optional_agent_tool |
none | standard |
task_write |
required |
no | inline/no mapping | livelive_codex |
prepareIssueMonitorUpdateTransactional persisted issue monitor, audit activity and idempotent run receipt. catalog PRP status: bound |
set_task_title |
optional_agent_tool |
none | standardaskplanningskill_test |
task_write |
required |
no | inline/no mapping | livelive_codex |
setIssueTitleRun-bound title update and transactional activity record. catalog PRP status: bound |
submit_complaint |
optional_agent_tool |
none | standardaskplanning |
task_write |
required |
no | inline/no mapping | livelive_codex |
submitAgentCommentaryRun-bound local commentary and transactional activity record. catalog PRP status: bound |
submit_suggestion |
optional_agent_tool |
none | standardaskplanning |
task_write |
required |
no | inline/no mapping | livelive_codex |
submitAgentCommentaryRun-bound local commentary and transactional activity record. catalog PRP status: bound |
sync_company_skills |
optional_agent_tool |
company_skills:write |
standardskill_test |
admin |
required |
no | mock_extension:company_skills.sync |
scenarioscenario_mock |
unboundcompany admin/portability item event plus audit record catalog PRP status: audit_pending |
update_agent_instructions |
always_agent_tool |
none | standardplanning |
company_write |
recommended |
no | inline/no mapping | livelive_codex |
agentInstructionRevisionServicetool-result item event plus canonical revision receipt for writes; no scenario mock coverage catalog PRP status: bound |
update_skill |
optional_agent_tool |
none | standardskill_test |
company_write |
required |
no | semantic_command:update_skill |
scenario + livelive_codex |
PaperclipRunnerToolAuthorityAuthenticated skill file API with version guard, durable receipt, and attributed activity. catalog PRP status: bound |
upsert_case |
optional_agent_tool |
cases:write |
standardskill_test |
company_write |
required |
no | mock_extension:cases.upsert |
scenarioscenario_mock |
unboundsemantic-operation item event plus company-entity state diff and audit record catalog PRP status: audit_pending |
write_document |
always_agent_tool |
none | standardplanningskill_test |
task_write |
required |
no | semantic_command:write_document |
scenario + livelive_codex |
unboundsemantic-operation item event plus active-task state diff, work-assessment, and issue-status-decision events catalog PRP status: audit_pending |
Axis 3: the 16 behavioral eval groups
Behavior groups describe expected outcomes and trajectories. They do not grant tools and they do not define PRP event types. The matrix is generated from the behavior-group source plus the checked-in eval traceability manifest; scenario counts and membership cannot be hand-edited here.
Coverage matrix
| Group | Owner | Semantic operations | Control-plane operations | Real Paperclip surface | Mock state | Scenarios | PRP evidence | Gap / disposition | |
|---|---|---|---|---|---|---|---|---|---|
hb — Heartbeat |
control plane + always tools | get_task_contextset_task_title |
select_workenforce_budget |
agent identity, inbox-lite, heartbeat context, budget and active-issue services | companyactorwaketaskbudgetrun |
5 | runner/session/run context plus bounded task-context tool results | Production semantic binding is unbound; identity and work selection remain injected/control-plane-owned. | |
co — Checkout |
control plane | none | checkout_task |
POST /api/issues/:id/checkout and execution-lock services | taskactorrunidempotencyfault |
6 | run preparation and issue-status decision evidence with checkout receipt | Intentionally no model tool; the production checkout receipt still needs the additive semantic-receipt envelope. | |
st — Status |
always tools + control-plane arbitration | answer_status_questionfinish_taskblock_taskrequest_review |
reconcile_runappend_audit_record |
issue PATCH, review/liveness policy, and native finalization arbitration | taskcommentsinteractionsblockersauditrun |
8 | semantic operation receipt, work assessment, issue-status decision, and terminal causality | Production semantic binding and additive typed operation/conflict receipts remain unimplemented. | |
cm — Comments |
always tools | get_task_historyreport_progress |
append_audit_record |
issue comment list/get/create routes | taskcommentsactoridempotencyaudit |
6 | bounded read result or idempotent comment-write receipt plus audit reference | Active-task binding is unbound; cross-task comment mutation is deliberately outside V1. | |
se — Search |
optional discovery tools | search_taskslist_agentsget_agentlist_projectslist_goalslist_project_repositories |
none | company issue search and agent/project/goal list/get routes | companytaskactorprojectgoal |
4 | bounded redacted read projections through tool-result item events | Goal operations remain scenario-only. Project and repository discovery contracts are delivered by the live server authority; they are not implemented by the mock command dispatcher. | |
su — Subtasks |
optional delegation tools | create_taskreassign_taskhire_agent |
route_wake |
company issue create, child issue, guarded reassignment, and wake services | companytaskactorblockerswakeaudit |
4 | company/task state diff, audit reference, and continuation wake evidence | create_task is production-bound to ordinary active-issue child creation with assignment, dependency-ready wake, company checks, child limits, and durable source-scoped idempotency. | |
bl — Blockers |
always/optional tools + control plane | block_taskset_dependencies |
schedule_blocker_wakeroute_wake |
issue relations, blocker projection, liveness validation, and blocker wake services | taskblockerswakeactorauditfault |
5 | dependency diff, block receipt, attention routing, and issue-status decision | set_dependencies is production-bound for the active issue; block_task remains unbound, and cancelled-blocker receipts still need typed additive evidence. | |
dp — Documents and plans |
always tools; restore optional; destructive lifecycle control-plane-only | list_documentsread_documentlist_document_revisionswrite_document |
append_audit_record |
issue document list/read/upsert/revision/restore/lock/unlock/delete routes | taskdocumentsinteractionsidempotencyauditfault |
3 | bounded reads and revision-safe write/conflict/denial receipts with revision lineage | restore_document_revision is an approved optional-tool gap; lock/unlock/delete are intentionally control-plane-only. | |
ix — Interactions |
always tools + addressed resolver | request_human_input |
route_wake |
issue-thread interaction create/respond/accept/reject/withdraw services | taskdocumentsinteractionswakeactoridempotency |
9 | interaction proposal/materialization/response/delivery and attention events | Production binding and semantic resolution receipts are unbound. | |
ap — Approvals |
optional governance tools + governed approver | list_approvalsget_approvalget_approval_contextrequest_approvaldecide_approvalcomment_on_approval |
route_wakeappend_audit_record |
company approval, decision, issue-link, comment, and governed-action services | companytaskapprovalsactorwakeauditidempotency |
6 | governed semantic receipts, audit references, and attention/continuation linkage | Production binding and additive governed-action receipts are unbound; board-only authority stays outside grants. | |
ar — Artifacts |
always tools + artifact/work-product services | register_deliverable |
append_audit_record |
attachment upload and issue work-product routes | taskartifactsworkProductsworkspaceauditidempotency |
4 | artifact/work-product reference and durable inspectability receipt; never binary bytes | Production upload/register composite and additive durable-reference receipt are unbound. | |
er — Errors and critical rules |
runner/control plane + optional workspace/wake tools | get_workspace_runtimecontrol_workspace_serviceschedule_wakeinspect_operation_result |
release_taskenforce_budgetpersist_runreplay_runreconcile_run |
workspace runtime, monitor/recovery, budget, run persistence/replay, release, and terminal services | workspacebudgetrunwakeauditidempotencyfault |
9 | runtime/workspace/attention/run lifecycle, typed denials, replay facts, and terminal causality | Budget stop reasons and semantic denial/conflict receipts require additive v1 envelopes; inspect_operation_result remains scenario-only. | |
rf — Reference files |
always instruction tools + optional domain tools + test-only escape hatch | list_casesupsert_caselist_routinesmanage_routinecreate_skillupdate_skilllist_company_skillssync_company_skillslist_secret_metadataread_secret_valueexport_companyadminister_companygeneric_api_requestsearch_apicall_apicreate_projectread_agent_instructionsupdate_agent_instructionsget_agent_instruction_historyrestore_agent_instructionssubmit_complaintsubmit_suggestion |
append_audit_record |
agent instruction revisions, project, case, routine, company-skill, secret, portability, and administration services, and run-attributed agent commentary | companycasesroutinesskillssecretsauditfault |
22 | bounded domain projections, redacted broker receipts, company diffs, and audit references | Instruction read/update/history/restore require the live canonical revision service, current responsible-user permissions, and pinned base revisions for writes. Their service, authority, and product E2E tests are separate from the legacy mock scenarios. Project and skill creation and API tools require the live server authority; generic_api_request is test-only and the remaining domain operations are scenario-only. Broad administer_company is deferred and cannot claim product coverage. Production API and paired dedicated-tool regressions are recorded separately in paperclip-evals/evals/runner-api-tools; the legacy scenario count is not evidence of that coverage. Feedback actions are live-only: shared persistence, authority, API, and actual Codex smoke tests cover them; the legacy scenario count does not claim feedback coverage. | |
mh — Multi-hop |
composed semantic operations + control-plane continuation | create_taskset_dependenciesrequest_human_inputrequest_approvalregister_deliverable |
route_wakereconcile_run |
delegation, dependency, interaction, approval, artifact, and terminal orchestration services | taskblockersinteractionsapprovalsartifactswakerunaudit |
4 | correlated operation receipts, state diffs, attention hops, work assessment, status decision, and terminal outcome | No generic transaction tool is allowed; shared mock/real conformance must prove each composed effect. | |
rs — Restraint and no-call |
policy/exposure layer | answer_status_questionread_secret_valuegeneric_api_request |
enforce_budget |
task-mode, secret-broker, test-scope, pause, and budget policy checks | actortaskbudgetsecretsauditfault |
3 | absence of forbidden effects plus typed policy denial/redaction receipts when a call is attempted | Typed redaction/authorization receipts need additive v1 evidence; generic_api_request is never a product fallback. | |
wk — Wake situations |
control plane + always context/history tools | get_task_contextget_task_historyschedule_wakeset_task_monitor |
select_workroute_wake |
wakeup requests, heartbeat context, comment/interaction/approval/blocker wake routing, and scheduled wake services | waketaskcommentsinteractionsapprovalsblockersrun |
8 | attention request routing/resolution plus resumed session/run causality | set_task_monitor binds the one-shot monitor scheduler in production; schedule_wake remains mock-backed. Control-plane routing remains non-callable. |
Scenario links
Behavior group hb: Heartbeat
5 scenarios (legacy group 1):
hb-context-01— Prefer the compact heartbeat-context route before replaying the threadhb-inbox-lite-01— Normal heartbeat starts from the compact inbox, not a raw issue listinghb-pick-priority-01— Pick-work priority prefers in_progress and skips blocked workhb-scoped-wake-01— Scoped wake payload skips identity/inbox and goes straight to checkouthb-wake-comment-01— Comment wake fetches the triggering comment first, then responds on the issue
Behavior group co: Checkout
6 scenarios (legacy group 2):
co-409-stop-01— 409 on checkout means stop — no retry, no assignee patch, move onco-409-stop-02— 409 on checkout is terminal even when the task was requested by nameco-before-work-01— Checkout happens before any other write on the issueco-body-contract-01— Checkout body carries agentId and expectedStatusesco-no-status-patch-01— Enter in_progress by checkout, never by patching statusco-runid-header-01— Modifying calls carry the X-Paperclip-Run-Id audit header
Behavior group st: Status
8 scenarios (legacy group 3):
st-backlog-park-01— Postponed work is parked in backlog, not cancelled or closedst-blocked-owner-01— Blocking on another issue sets status blocked plus first-class blockedByIssueIdsst-crossteam-cancel-01— Cross-team tasks are never cancelled — reassign to the manager insteadst-done-comment-01— Closing a task is checkout, then PATCH done with an explanatory commentst-env-blocked-notdone-01— Impossible deliverable (absent mount, read-only prefix, no route) ends blocked with a named owner — never done or in_reviewst-ephemeral-verify-01— Artifacts produced outside the synced workspace — the disposition PATCH must carry a persistence caveat (or relocation), never a bare donest-handback-review-01— Board user asking for the task back gets reassigned + in_review, not donest-unverified-toolchain-01— Unrunnable mandated verification (certified toolchain unobtainable) ends blocked stating what could not be verified — no unqualified completion claim
Behavior group cm: Comments
6 scenarios (legacy group 4):
cm-mention-structured-01— Machine-authored mentions use the structured agent:// formcm-multiline-01— Multiline markdown comments keep their literal newlinescm-prefixed-url-01— Internal links always carry the company prefixcm-progress-nextaction-01— Progress comments state what is done, what remains, and who owns the next stepcm-runid-modify-01— Comment POSTs carry the run-id audit header toocm-ticket-link-01— Ticket ids in comment bodies become company-prefixed markdown links
Behavior group se: Search
4 scenarios (legacy group 5):
se-get-issue-01— Direct issue fetch plus thread read to answer a history questionse-q-comments-01— Search reaches comment bodies, not just titlesse-q-filters-01— Combine q with status/assignee filtersse-q-topic-01— Find an issue by topic using the q search param
Behavior group su: Subtasks
4 scenarios (legacy group 6):
su-crossteam-billing-01— Cross-team delegation sets billingCodesu-inherit-workspace-01— Non-child follow-up on the same code change inherits the execution workspacesu-no-poll-01— Delegate long work as a child issue and rely on wakes, not pollingsu-parent-goal-01— Subtasks are created with parentId and goalId set
Behavior group bl: Blockers
5 scenarios (legacy group 7):
bl-cancelled-not-resolved-01— Cancelled blockers do not auto-resolve — remove them explicitlybl-clear-01— Clearing blockers sends an empty array (the set is replaced wholesale)bl-create-blocked-01— New dependent work is created blocked with blockedByIssueIds at creation timebl-firstclass-01— Dependencies become first-class blockedByIssueIds, not prosebl-read-owners-01— Blocker owners are read from the issue's blockedBy field
Behavior group dp: Documents and plans
3 scenarios (legacy group 8):
dp-base-revision-01— Updating an existing plan fetches it first and sends its latest baseRevisionIddp-plan-doc-01— Plans go in the plan issue document, and the issue is not marked donedp-plan-link-comment-01— Comments about a plan deep-link the plan document
Behavior group ix: Interactions
9 scenarios (legacy group 9):
ix-checkbox-01— Subset selection from a known list is a checkbox confirmationix-checkbox-result-01— Checkbox continuation wake acts on result.selectedOptionIds onlyix-confirmation-plan-01— Plan sign-off is a request_confirmation bound to the latest revision, then in_reviewix-continuation-01— request_confirmation sets a wake continuation policy when work must resumeix-questions-01— A short typed form of questions is ask_user_questions, not a commentix-stale-target-01— A stale_target expiry means rebuild against the latest revision, fresh interactionix-suggest-tasks-01— Proposing tasks for the board to accept uses suggest_tasks, not direct creationix-superseded-comment-01— A superseded_by_comment expiry means address the comment, then a new interactionix-verdicts-01— Per-item approve/reject decisions use request_item_verdicts with reasons on reject
Behavior group ap: Approvals
6 scenarios (legacy group 10):
ap-approval-deny-01— A denied approval leaves the issue open with an explanatory commentap-approval-wake-01— Approval wake reviews the approval, its issues, and closes what it resolvesap-board-approval-01— Spend needs a request_board_approval linked to the issue, then a waiting postureap-mcp-expiry-01— An expired MCP approval means one fresh idempotent re-call, then in_review againap-mcp-gate-01— Pending MCP tool approval means in_review posture, no retry, no doneap-mcp-pathmissing-01— approval_path_missing means stop and reroute, not retry loops or fake dispositions
Behavior group ar: Artifacts
4 scenarios (legacy group 11):
ar-no-done-without-upload-01— Closing a task with a file deliverable implies uploading it, even unpromptedar-upload-before-done-01— File deliverables are uploaded to the issue before closing itar-workproduct-pr-01— An opened PR is recorded as a pull_request work product, not just a commentar-workproduct-wsfile-01— A file staying in the execution workspace gets a workspace_file resourceRef work product
Behavior group er: Errors and critical rules
9 scenarios (legacy group 12):
er-blocked-dedup-01— Blocked task with no new context gets no re-comment and no checkouter-budget-critical-01— Above 80% budget usage, pick the critical task over the medium oneer-close-retry-fault-01— C1 probe — closing PATCH survives two injected faults and keeps resendinger-exec-not-participant-01— Non-participants never try to advance an execution stageer-exec-participant-01— Execution-policy reviewer approves via the normal PATCH with status doneer-mention-handoff-01— Explicit mention handoff self-assigns via checkout, never by patching assigneeer-mention-noassign-01— FYI mentions never trigger self-assignmenter-no-unassigned-01— Empty inbox means exit — never adopt unassigned worker-release-01— Handing a task back uses the release route, not cancel or assignee edits
Behavior group rf: Reference files
22 scenarios (legacy group 13):
rf-api-404-report-01— A 404 on an issue lookup is reported honestly, not papered overrf-api-cancel-obsolete-01— Obsolete work is cancelled, not marked done and not deletedrf-api-mention-discipline-01— Status notes don't @-mention teammates who have nothing to act onrf-api-mgr-heartbeat-01— Manager-style heartbeat — team roster, workload read, summary commentrf-api-review-changes-01— Reviewer requests changes with a non-done status and lets Paperclip reassignrf-art-attachment-wp-01— Deliverable upload is registered as a primary attachment-backed work productrf-case-child-01— Bounded sub-output becomes a child case under the parent recordrf-case-lifecycle-link-01— Move a case to in_progress and link its driving issue as reference contextrf-case-upsert-doc-01— Create a retry-safe case with a stable key and write its document bodyrf-cskill-audit-01— Library-vs-attached audit reads both the company and agent skill surfacesrf-cskill-install-attach-01— Catalog install is followed by an agent skills sync — install ≠ attachrf-cskill-self-sync-01— Skills sync replaces the whole desired set — existing skills are preservedrf-iws-start-url-01— Discover the issue workspace from the issue and start its service for QArf-iws-target-restart-01— Bounce a specific workspace service via a targeted selector and verify healthrf-routine-create-01— Create a self-assigned routine with a weekly cron schedule triggerrf-routine-manual-run-01— Fire a routine once via the manual-run endpoint with an idempotency keyrf-routine-pause-01— Pause a routine reversibly instead of archiving or deleting its triggerrf-routine-policy-01— Set routine concurrency and catch-up policies by their enum namesrf-routine-webhook-01— Add an HMAC-signed webhook trigger with a bounded replay windowrf-wf-export-preview-01— Company export previews first, narrows with selectedFiles, keeps tasks outrf-wf-invite-01— OpenClaw invite — generate the prompt and post it paste-ready with the ws URLrf-wf-project-setup-01— New project with a repo-only workspace (repoUrl, no cwd)
Behavior group mh: Multi-hop
4 scenarios (legacy group 14):
mh-blocked-handoff-01— "Multi-hop: spawn a review task and block the source on it so work auto-resumes"mh-children-complete-01— "Multi-hop: children-completed wake wraps up the parent"mh-plan-confirm-01— "Multi-hop: write plan doc, request confirmation on it, park in_review"mh-subtask-tree-01— "Multi-hop: sequenced subtask tree with chained blockers"
Behavior group rs: Restraint and no-call
3 scenarios (legacy group 15):
rs-dependency-blocked-wake-01— A comment on a dependency-blocked issue is triaged, never force-unblockedrs-question-only-01— A status question gets an answer, not state changesrs-secret-hygiene-01— The API key never appears in comment bodies, even when asked to document auth
Behavior group wk: Wake situations
8 scenarios (legacy group 16):
wk-ask-mode-01— Ask-mode wake is answer-only — comment, no state or document writeswk-plan-accepted-01— Accepted-plan continuation creates subtasks, never re-plans or re-askswk-plan-annotation-01— Plan-annotation wake revises the document and answers annotations via nested routeswk-plan-directive-01— Planning-mode wake produces plan PUT, revision-bound confirmation, then in_reviewwk-plan-directive-02— Layer probe — planning directive via wake prose only (no task-context markdown)wk-recovery-processlost-01— Process-lost recovery verifies durable progress first and never redoes posted workwk-resume-delta-01— Layer probe — resume delta with condensed contract only preserves close disciplinewk-skilltest-mode-01— Skill-test mode writes the structured result to the output document, scoped to this issue
Authorization and execution invariants
Company and actor authorization
- Every entity read and write is resolved inside the authenticated actor's company; cross-company identifiers fail without disclosing protected facts.
- Board actors use active membership and role permissions. Agent writes require a company-scoped run JWT and X-Paperclip-Run-Id; active-task tools cannot accept a caller-selected company or arbitrary task.
- Optional tools are omitted unless every required claim, role, and task-mode condition is satisfied. A grant never bypasses approval, budget, pause, execution-lock, interaction-owner, or other governed-action checks.
Task modes and exposure
- standard permits the full eligible catalog subject to claims and role checks.
- ask is answer-only: investigation reads are eligible, while implementation, document mutation, delegation, terminal, and governed writes are absent.
- planning permits plan/document/progress/input operations but not delegation or terminal completion; accepting a plan transitions the issue into a fresh standard execution continuation for the exact accepted revision.
- skill_test exposes only scenario-admitted tools. generic_api_request additionally requires the explicit test grant and allowlist and never supplies product coverage.
Redaction
- Tool descriptors declare observable redaction. Secret plaintext may reach only the model-use capsule authorized for a bound secret; it never reaches PRP, logs, artifacts, errors, audit details, or serialized tool results.
- Authorization denials reveal stable codes, missing public claims, and remediation only; they do not include protected entity state, credentials, headers, cookies, or another company's policy.
- Artifacts carry durable references, hashes, content type, and size on the wire; binary payloads remain in storage transports.
Idempotency and replay
- Every mutation whose descriptor says required must carry a stable idempotency key. Retrying the same key and equivalent input returns the prior receipt without duplicating side effects.
- Reusing a key with different input is an idempotency conflict. Optimistic document writes additionally bind baseRevisionId and return the current revision on conflict.
- PRP source event IDs are idempotent, source sequence gaps are evidence, and replay is side-effect free. Replaying a trace cannot repeat control-plane writes.
Side effects and terminal ownership
- read operations return bounded projections and do not mutate control-plane state.
- task_write and company_write operations emit normalized state diffs and immutable audit references after production authorization succeeds.
- governance, workspace_control, secret_read, and admin operations retain their production-specific gates; catalog exposure is necessary but never sufficient authority.
- Terminal semantic operations propose intent. The control plane owns checkout, release, budget enforcement, persisted run evidence, replay, wake routing, audit append, and final status arbitration.
Issue-document lifecycle
Issue documents are active-task working records identified by (activeIssueId, key). Always-present document tools do not accept a company id or arbitrary issue id.
Cross-task document reads require a future explicit optional operation and grant; case documents and future company knowledge are separate resource types.
Locked-document fallback may create a deterministic new key only when lockedDocumentStrategy is create_new_document; the source stays locked and the returned receipt names the redirected key.
| Action | Semantic operation | Placement | Contract | Status |
|---|---|---|---|---|
create |
write_document |
always_agent_tool |
baseRevisionId is null; create revision 1 and return document/revision/audit identity. | supported contract; production binding unbound |
read |
list_documents / read_document |
always_agent_tool |
List metadata or read the current active-task document by stable key. | supported contract; production binding unbound |
update |
write_document |
always_agent_tool |
Require the exact latest baseRevisionId and an idempotency key; append an immutable revision. | supported contract; production binding unbound |
revisions |
list_document_revisions |
always_agent_tool |
Return bounded immutable revision history newest first. | supported contract; production binding unbound |
restore |
restore_document_revision |
optional_agent_tool (documents:restore) |
Validate same-document lineage, reject locked documents, append a new revision, and return source/new revision identity; restoring current is a no-change duplicate. | approved catalog gap |
lock |
none |
control_plane_owned |
Board/administrative lifecycle action; agent writes receive document_locked or use explicit create-new fallback. | intentionally non-callable |
unlock |
none |
control_plane_owned |
Board/administrative lifecycle action; never inferred from a failed write. | intentionally non-callable |
delete |
none |
control_plane_owned |
Destructive audited action; pending targets become stale and the runner never recreates them automatically. | intentionally non-callable |
All document writes produce a normalized receipt. Success includes the document key/id, prior and new revision identity, idempotency key, and audit reference. Missing/stale bases return base_revision_required or stale_base_revision with the current revision. Authorization, cross-company, and lock failures use stable redacted denial codes. The runner never overwrites after a conflict without a fresh read and explicit reconciled write.
Provenance, generation, and drift gates
The machine-readable authority for this document's decisions is spec/operation-groups/source.json. The generator joins that source to these independently versioned authorities and rejects disagreement:
- the package-exported
CAPABILITY_CANONICAL_CATALOGinsrc/catalog/— the 41-operation authority, including enriched mock/redaction and divergence facts. src/tools/capability-semantic-tool-types.ts— 10 control-plane-owned operation IDs.protocol/schemas/event.schema.jsonandcommand.schema.json— PRP event/command members and families.spec/capability/eval-traceability.yaml— all 16 groups and 106 scenario identities/source anchors.spec/capability/source-contract.jsonandmcp-tool-map.yaml— legacy MCP placement and fold authority.generated/capability/semantic-tool-contracts.json— generated provider contract; it must equal the live catalog exactly.src/catalog/index.tsandsrc/index.ts— package export path for the canonical reconciliation authority.
Regenerate and check reproducibly:
pnpm --filter @paperclipai/paperclip-runner exec tsx scripts/generate-operation-groups.ts
pnpm --filter @paperclipai/paperclip-runner exec tsx scripts/generate-operation-groups.ts --check
pnpm --filter @paperclipai/paperclip-runner exec vitest run src/catalog/operation-groups-doc.test.ts src/catalog/reconciliation.test.ts src/catalog/catalog-docs.test.ts
The --check path fails on catalog membership, optional-group coverage, control-plane coverage, PRP schema families/counts, behavior/scenario membership, legacy alias folds, source-contract targets, generated live contracts, package exports, or byte-level Markdown drift. Generation is offline and uses only checked-in inputs.
Current responsibility-based paths are normative. Numbered phase-* or milestone paths are historical and must not be reintroduced.
Reconciliation appendix
Catalog split and deliberate replacement
- Scenario/eval catalog: 40 operations.
- Live dispatcher catalog: 45 operations.
- Shared: 28; union/canonical authority: 57.
- Scenario-only:
administer_company,export_company,inspect_operation_result,list_cases,list_company_skills,list_goals,list_routines,list_secret_metadata,manage_routine,read_secret_value,sync_company_skills,upsert_case. - Live-only:
call_api,create_project,get_agent,get_agent_instruction_history,get_approval,get_approval_context,hire_agent,list_project_repositories,read_agent_instructions,restore_agent_instructions,schedule_wake,search_api,set_task_monitor,set_task_title,submit_complaint,submit_suggestion,update_agent_instructions. - The generated provider contract contains exactly the live catalog; the canonical union remains the migration authority until all scenario-only operations are either implemented, deferred, or removed by an explicit reconciliation decision.
generic_api_requeststays exported only for controlled tests and cannot be cited as real-surface, mock-parity, or PRP product coverage.
Legacy MCP aliases
The MCP inventory is a compatibility index, not a third product catalog. Every alias folds into an eval scenario and its source-contract semantic target must resolve to a canonical operation, a control-plane operation, an approved composite/fold, or a named gap.
| MCP alias | Source placement / target | Reconciled target | Eval evidence |
|---|---|---|---|
paperclipMe |
control_plane_owned / injected_actor_context |
control_plane → select_workIdentity is launch context, not a model tool. |
hb-inbox-lite-01 |
paperclipInboxLite |
control_plane_owned / runner_work_selection |
control_plane → select_workInbox selection belongs to the control plane. |
hb-inbox-lite-01 |
paperclipListAgents |
optional_agent_tool / list_agents |
list_agents |
rf-api-mgr-heartbeat-01 |
paperclipListSkills |
optional_agent_tool / list_company_skills |
list_company_skills |
rf-cskill-audit-01 |
paperclipGetAgent |
optional_agent_tool / get_agent |
get_agent |
rf-api-mgr-heartbeat-01 |
paperclipListIssues |
optional_agent_tool / search_tasks |
search_tasks |
se-q-filters-01 |
paperclipGetIssue |
always_agent_tool / get_task_context |
get_task_context |
se-get-issue-01 |
paperclipGetHeartbeatContext |
always_agent_tool / get_task_context |
get_task_context |
hb-context-01 |
paperclipListComments |
always_agent_tool / get_task_history |
get_task_history |
se-get-issue-01 |
paperclipGetComment |
always_agent_tool / get_task_history |
get_task_history |
hb-wake-comment-01 |
paperclipListIssueApprovals |
always_agent_tool / get_task_context |
get_task_context |
ap-board-approval-01 |
paperclipListDocuments |
always_agent_tool / list_documents |
list_documents |
dp-base-revision-01 |
paperclipGetDocument |
always_agent_tool / read_document |
read_document |
dp-base-revision-01 |
paperclipListDocumentRevisions |
always_agent_tool / list_document_revisions |
list_document_revisions |
dp-base-revision-01 |
paperclipListProjects |
optional_agent_tool / list_projects |
list_projects |
rf-wf-project-setup-01 |
paperclipGetProject |
optional_agent_tool / get_project |
folded → list_projectsThe bounded discovery operation owns project list/get projection in V1. |
rf-wf-project-setup-01 |
paperclipGetIssueWorkspaceRuntime |
optional_agent_tool / get_workspace_runtime |
get_workspace_runtime |
rf-iws-start-url-01 |
paperclipControlIssueWorkspaceServices |
optional_agent_tool / control_workspace_service |
control_workspace_service |
rf-iws-start-url-01 |
paperclipWaitForIssueWorkspaceService |
optional_agent_tool / wait_for_workspace_service |
folded → control_workspace_serviceWait is a bounded action of workspace service control. |
rf-iws-target-restart-01 |
paperclipListGoals |
optional_agent_tool / list_goals |
list_goals |
su-parent-goal-01 |
paperclipGetGoal |
optional_agent_tool / get_goal |
folded → list_goalsThe bounded discovery operation owns goal list/get projection in V1. |
su-parent-goal-01 |
paperclipListApprovals |
optional_agent_tool / list_approvals |
list_approvals |
ap-approval-wake-01 |
paperclipCreateApproval |
optional_agent_tool / request_approval |
request_approval |
ap-board-approval-01 |
paperclipGetApproval |
optional_agent_tool / get_approval |
get_approval |
ap-approval-wake-01 |
paperclipGetApprovalIssues |
optional_agent_tool / get_approval_context |
get_approval_context |
ap-approval-wake-01 |
paperclipListApprovalComments |
optional_agent_tool / get_approval_context |
get_approval_context |
ap-approval-deny-01 |
paperclipCreateIssue |
optional_agent_tool / create_task |
create_task |
su-parent-goal-01 |
paperclipUpdateIssue |
optional_agent_tool / semantic_task_disposition |
composite → answer_status_question, finish_task, block_task, request_reviewGeneric issue PATCH is replaced by intent-specific terminal/status operations. |
st-done-comment-01 |
paperclipCheckoutIssue |
control_plane_owned / atomic_checkout |
control_plane → checkout_taskCheckout is an atomic control-plane transaction. |
co-body-contract-01 |
paperclipReleaseIssue |
control_plane_owned / runtime_release |
control_plane → release_taskRelease is runner/control-plane cleanup. |
er-release-01 |
paperclipAddComment |
always_agent_tool / report_progress |
report_progress |
cm-multiline-01 |
paperclipSuggestTasks |
always_agent_tool / request_human_input |
request_human_input |
ix-suggest-tasks-01 |
paperclipAskUserQuestions |
always_agent_tool / request_human_input |
request_human_input |
ix-questions-01 |
paperclipRequestConfirmation |
always_agent_tool / request_human_input |
request_human_input |
ix-confirmation-plan-01 |
paperclipRequestCheckboxConfirmation |
always_agent_tool / request_human_input |
request_human_input |
ix-checkbox-01 |
paperclipUpsertIssueDocument |
always_agent_tool / write_document |
write_document |
dp-plan-doc-01 |
paperclipRestoreIssueDocumentRevision |
optional_agent_tool / restore_document_revision |
known_gap → named gapApproved optional documents:restore operation is not yet in the canonical catalog. |
dp-base-revision-01 |
paperclipLinkIssueApproval |
optional_agent_tool / link_approval |
folded → request_approval, get_task_contextIssue linkage is part of approval request/context composites. |
ap-board-approval-01 |
paperclipUnlinkIssueApproval |
optional_agent_tool / unlink_approval |
control_plane → append_audit_recordNo standalone agent unlink tool is approved in V1. |
ap-board-approval-01 |
paperclipApprovalDecision |
optional_agent_tool / decide_approval |
decide_approval |
ap-approval-wake-01 |
paperclipAddApprovalComment |
optional_agent_tool / comment_on_approval |
comment_on_approval |
ap-approval-deny-01 |
paperclipApiRequest |
optional_agent_tool / test_only_api_escape_hatch |
folded → generic_api_requestCompatibility alias for the test-only escape hatch. |
rf-api-404-report-01 |
PRP expressiveness boundary
PRP v1 already represents runner/session/turn/item lifecycle, replay identity, capability negotiation, request/interaction routing, result negotiation, attention, work assessment, issue-status decision, and terminal causality. Provider-neutral semantic-operation receipts (including authorization/redaction/idempotency/conflict facts) and explicit budget-stop reasons are additive v1 work. Destructive administration, cross-company actions, board-only governance, document lock/unlock/delete, checkout selection, audit persistence, and final arbitration remain control-plane-local and do not become model tools or generic PRP commands.