Files
PaperClipAI/doc/LOW-TRUST-PRESETS.md
T
DottaandPaperclip 71a6535792 docs(execution-semantics): define routable blocking and watchdog restoration (#10094)
## Thinking Path

> - Paperclip is the open source control plane people use to manage
AI-agent companies.
> - Execution semantics define when work may stop, how child work
reports upward, and how watchdog recovery proves liveness was restored.
> - The existing docs did not require a routable waiting path before
entering `blocked`, leaving permission-denial and review workflows
vulnerable to dead ends.
> - Child-to-parent reporting and low-trust review delegation also
needed explicit channels and ownership boundaries.
> - Watchdog recovery needed bounded restoration verification rather
than treating a recovery write as proof of restored progress.
> - The security-sensitive defaults were reviewed and approved before
publication.
> - This pull request documents those contracts and synchronizes the
bundled Paperclip skills that operationalize them.
> - The benefit is clearer stop conditions, safer reporting defaults,
and bounded watchdog recovery without sacrificing liveness.

## Linked Issues or Issue Description

- No public GitHub issue exists for this execution-semantics
documentation phase, so the problem and solution are described here in
full.
- Problem: `blocked` could be entered without a routable owner/action
path, review findings could be reported on the wrong issue or treated as
blockers, and watchdog recovery lacked bounded verification of restored
liveness.

## What Changed

- Require a routable waiting path for `blocked` and clarify that
permission denial alone is not a blocker.
- Define canonical child-to-parent completion/report channels,
review-delegation ownership, and the sanctioned courier pattern.
- Specify atomic watchdog recovery batches, fingerprint validation,
bounded restoration attempts, and human escalation.
- Document the low-trust preset default for report comments and align
both bundled Paperclip skills.

## Scope and Sequencing

- This is the contract-first P1 documentation PR. Runtime enforcement is
intentionally excluded and follows in the separately owned
implementation phases; these docs define the acceptance contract those
phases must satisfy.
- The five-file patch is approved and frozen for this PR, so review
findings about absent runtime support are tracked as
implementation-phase requirements rather than edits to this docs-only
change.

## Verification

- `git diff --check origin/master...HEAD`
- Confirmed the PR diff is exactly 5 approved docs/skill files with 91
insertions and 1 deletion.
- Confirmed the patch preserves the approved execution-semantics
contract after replay on latest `origin/master`.

## Risks

- Low implementation risk: documentation and skill guidance only; no
runtime code, schema, migration, dependency, lockfile, or workflow
changes.
- Semantic risk is limited to readers or agents applying the clarified
contracts; the security-sensitive R1a default was approved before
publication.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex coding agent (exact underlying model ID and
context-window size are not exposed by this runtime), with reasoning,
terminal tool use, Git/GitHub operations, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced private URLs or internal issue identifiers
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [x] I have run the applicable docs-only checks locally and they pass
- [x] Tests are not applicable because this is a
docs/skill-guidance-only change
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-27 19:10:28 -05:00

71 lines
3.3 KiB
Markdown

# Low-Trust Presets
Paperclip ships core trust preset names so containment decisions are enforced in
Community Edition even when EE policy editing is unavailable.
## Presets
- `standard`: the default V1 company-visible collaboration model. This preserves
existing behavior for normal agents.
- `low_trust_review`: an opt-in containment preset for automated work that may
consume hostile or prompt-injected input, such as untrusted pull requests,
external tickets, dependency diffs, or generated review output.
## Boundary Model
`low_trust_review` is resolved from existing JSON policy fields:
- agent permissions: `permissions.trustPreset` and
`permissions.authorizationPolicy.trustBoundary`
- project policy:
`executionWorkspacePolicy.authorizationPolicy.trustBoundary`
- issue/run policy: `executionPolicy.authorizationPolicy.trustBoundary`
The resolver intersects those sources. Narrower wins. A low-trust preset must
resolve to a concrete company-local project, root issue, or issue-id scope. If a
policy source names another company, uses an unsupported preset, or lacks that
scope for risky access, Paperclip fails closed.
## Containment, Not Privacy
This is containment for hostile automated work. It is not a general project,
issue, or human privacy system.
V1 standard work remains company-visible by default: board users and in-company
actors can inspect company work objects unless a separate access-control feature
changes that behavior. Low-trust containment instead limits what the low-trust
agent can read or mutate through the Paperclip API and prevents raw untrusted
output from being automatically promoted into higher-trust agent context.
Low-trust agents cannot read or mutate agent configuration, instruction bundles,
or company skill configuration through direct grants. Configuration changes from
low-trust work must go through higher-trust review and promotion paths instead.
## Child→Parent Reporting Under Containment
The direct-parent report comment (`doc/execution-semantics.md` §6, "Child→Parent
Reporting") is **off by default** for `low_trust_review`: a contained run reads
untrusted input, so a free-prose comment into the higher-trust parent thread is
a prompt-injection promotion path. Contained reviewers report by completing
their own review issue (`done` — the verdict is the deliverable; the
`issue_blockers_resolved` wake carries it upward) and by the platform's
system-attributed stop-only relay when they enter `blocked` or `cancelled`.
Never instruct a contained delegate to comment on its parent issue.
## Runtime Containment
Managed `low_trust_review` runs fail closed unless Paperclip can enforce the
runtime boundary:
- the selected execution environment must use the `sandbox` driver
- the effective execution workspace mode must be `isolated_workspace`
- the issue being run must be inside the resolved low-trust boundary
- secret references must use binding ids explicitly allowed by the boundary
- inline sensitive environment values such as API keys and tokens are rejected
- workspace runtime-service mutations are denied unless the boundary explicitly
grants the `runtime.manage` tool class
The Docker workflow in `doc/UNTRUSTED-PR-REVIEW.md` remains useful for manual
local review, but Paperclip-managed low-trust execution requires a sandboxed
environment instead of a host-local adapter process.