Files
PaperClipAI/doc/LOW-TRUST-PRESETS.md
T
DottaandPaperclip 71a6535792 docs(execution-semantics): define routable blocking and watchdog restoration (#10094)
## Thinking Path

> - Paperclip is the open source control plane people use to manage
AI-agent companies.
> - Execution semantics define when work may stop, how child work
reports upward, and how watchdog recovery proves liveness was restored.
> - The existing docs did not require a routable waiting path before
entering `blocked`, leaving permission-denial and review workflows
vulnerable to dead ends.
> - Child-to-parent reporting and low-trust review delegation also
needed explicit channels and ownership boundaries.
> - Watchdog recovery needed bounded restoration verification rather
than treating a recovery write as proof of restored progress.
> - The security-sensitive defaults were reviewed and approved before
publication.
> - This pull request documents those contracts and synchronizes the
bundled Paperclip skills that operationalize them.
> - The benefit is clearer stop conditions, safer reporting defaults,
and bounded watchdog recovery without sacrificing liveness.

## Linked Issues or Issue Description

- No public GitHub issue exists for this execution-semantics
documentation phase, so the problem and solution are described here in
full.
- Problem: `blocked` could be entered without a routable owner/action
path, review findings could be reported on the wrong issue or treated as
blockers, and watchdog recovery lacked bounded verification of restored
liveness.

## What Changed

- Require a routable waiting path for `blocked` and clarify that
permission denial alone is not a blocker.
- Define canonical child-to-parent completion/report channels,
review-delegation ownership, and the sanctioned courier pattern.
- Specify atomic watchdog recovery batches, fingerprint validation,
bounded restoration attempts, and human escalation.
- Document the low-trust preset default for report comments and align
both bundled Paperclip skills.

## Scope and Sequencing

- This is the contract-first P1 documentation PR. Runtime enforcement is
intentionally excluded and follows in the separately owned
implementation phases; these docs define the acceptance contract those
phases must satisfy.
- The five-file patch is approved and frozen for this PR, so review
findings about absent runtime support are tracked as
implementation-phase requirements rather than edits to this docs-only
change.

## Verification

- `git diff --check origin/master...HEAD`
- Confirmed the PR diff is exactly 5 approved docs/skill files with 91
insertions and 1 deletion.
- Confirmed the patch preserves the approved execution-semantics
contract after replay on latest `origin/master`.

## Risks

- Low implementation risk: documentation and skill guidance only; no
runtime code, schema, migration, dependency, lockfile, or workflow
changes.
- Semantic risk is limited to readers or agents applying the clarified
contracts; the security-sensitive R1a default was approved before
publication.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex coding agent (exact underlying model ID and
context-window size are not exposed by this runtime), with reasoning,
terminal tool use, Git/GitHub operations, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced private URLs or internal issue identifiers
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [x] I have run the applicable docs-only checks locally and they pass
- [x] Tests are not applicable because this is a
docs/skill-guidance-only change
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-27 19:10:28 -05:00

3.3 KiB

Low-Trust Presets

Paperclip ships core trust preset names so containment decisions are enforced in Community Edition even when EE policy editing is unavailable.

Presets

  • standard: the default V1 company-visible collaboration model. This preserves existing behavior for normal agents.
  • low_trust_review: an opt-in containment preset for automated work that may consume hostile or prompt-injected input, such as untrusted pull requests, external tickets, dependency diffs, or generated review output.

Boundary Model

low_trust_review is resolved from existing JSON policy fields:

  • agent permissions: permissions.trustPreset and permissions.authorizationPolicy.trustBoundary
  • project policy: executionWorkspacePolicy.authorizationPolicy.trustBoundary
  • issue/run policy: executionPolicy.authorizationPolicy.trustBoundary

The resolver intersects those sources. Narrower wins. A low-trust preset must resolve to a concrete company-local project, root issue, or issue-id scope. If a policy source names another company, uses an unsupported preset, or lacks that scope for risky access, Paperclip fails closed.

Containment, Not Privacy

This is containment for hostile automated work. It is not a general project, issue, or human privacy system.

V1 standard work remains company-visible by default: board users and in-company actors can inspect company work objects unless a separate access-control feature changes that behavior. Low-trust containment instead limits what the low-trust agent can read or mutate through the Paperclip API and prevents raw untrusted output from being automatically promoted into higher-trust agent context.

Low-trust agents cannot read or mutate agent configuration, instruction bundles, or company skill configuration through direct grants. Configuration changes from low-trust work must go through higher-trust review and promotion paths instead.

Child→Parent Reporting Under Containment

The direct-parent report comment (doc/execution-semantics.md §6, "Child→Parent Reporting") is off by default for low_trust_review: a contained run reads untrusted input, so a free-prose comment into the higher-trust parent thread is a prompt-injection promotion path. Contained reviewers report by completing their own review issue (done — the verdict is the deliverable; the issue_blockers_resolved wake carries it upward) and by the platform's system-attributed stop-only relay when they enter blocked or cancelled. Never instruct a contained delegate to comment on its parent issue.

Runtime Containment

Managed low_trust_review runs fail closed unless Paperclip can enforce the runtime boundary:

  • the selected execution environment must use the sandbox driver
  • the effective execution workspace mode must be isolated_workspace
  • the issue being run must be inside the resolved low-trust boundary
  • secret references must use binding ids explicitly allowed by the boundary
  • inline sensitive environment values such as API keys and tokens are rejected
  • workspace runtime-service mutations are denied unless the boundary explicitly grants the runtime.manage tool class

The Docker workflow in doc/UNTRUSTED-PR-REVIEW.md remains useful for manual local review, but Paperclip-managed low-trust execution requires a sandboxed environment instead of a host-local adapter process.