mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 06:15:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Low-trust agents must work within their assigned scope. > - Task creation currently rejects these agents before checking assignment permission or scope. > - Persistent instruction saves also reject direct requests from authorized chat owners. > - This PR checks the requested action and its recorded authority instead of denying all such work. > - Agents can organize permitted work and follow their owner's instruction-edit requests while outside work stays restricted. ## Linked Issues or Issue Description **What happened?** Low-trust agents cannot create self-assigned tasks or subtasks, even within their allowed scope. An authorized user also cannot ask an agent in their own Agent Chat to update its managed `AGENTS.md`. Agent-folder collection can hide the permission rejection behind a generic save failure. **Expected behavior** Allow task creation when assignment permissions and project or root-task scope permit it. Allow instruction self-edits during authenticated owner-chat execution, subject to the user's current edit permission. Outside tasks, subtasks, connector messages, and peer agents do not inherit that instruction authority. Explain the actual denial when a save fails. **Steps to reproduce** 1. Configure an active agent with `low_trust_review` and a project or root-task boundary. 2. Ask it to create an in-scope task assigned to itself, or a subtask of its own task. 3. As a user with permission to configure that agent, ask it in your Agent Chat to update its managed `AGENTS.md`. 4. Observe blanket permission denials rather than action-specific checks. **Paperclip version or commit** Rebased onto master at `8ec4b84e1`. This is a core authorization change, independent of adapter choice. **Deployment mode** Authenticated server. Regression coverage uses the server services, HTTP routes, native tool authority, and embedded PostgreSQL. Related work: #14775 adds human-directed task execution. #13599 concerns instruction-path configuration; this PR leaves that configuration restricted. #11988 proposes separate active-review instruction protection. #10693 reports unclear authorization denials on a different API surface. ## What Changed - Apply task-assignment checks to both HTTP creation routes and native task creation, including unassigned work. Preserve low-trust policy and source attribution on the created task and its initial plan. - Allow self-assigned decomposition within the permitted project or root-task tree. Resolve workspace-derived project scope before authorization, and reauthorize existing tasks before duplicate detection returns them. Keep cross-project and peer-assignment checks. - Derive instruction self-edit authority from the accepted run identity and authenticated owner-message wake. Recheck current permissions at save time. Bind retries to the same request and chat session. - Reject inherited instruction authority from outside tasks, subtasks, plugins, connectors, stale sessions, cancelled runs, and peer edits. - Surface permission errors in instruction and agent-folder save receipts. Tell chat agents to explain the rejected action and the specific restriction. - Update the low-trust policy and implementation documentation. ## Verification - All 297 tests in 11 focused server suites pass after the rebase. These cover owner-chat saves, private copies, warm agent directories, reset and retry boundaries, permission revocation, task creation routes, and native tool authority. - After review fixes, all 126 tests in the four affected authorization/chat suites pass. Workspace scope regressions and 146 existing creation/ownership/workspace-route tests also pass. - The final duplicate-task and CI fixes pass all 39 tests across chat-project tools, duplicate creation, environment-selection guards, and assignee-invokability routes. The duplicate-task test reproduced an unauthorized response before the fix and verifies denial plus permitted reuse afterward. - `pnpm --filter @paperclipai/server typecheck` passes after rebasing; `pnpm --filter @paperclipai/server exec tsc --noEmit` also passes after the review fixes. - `git diff --check origin/master...HEAD` passes. - Final head `7e73270b86748792649e4ae6fbc6879f73b42b73`: all 54 checks passed, with two expected skips and no pending or failed checks. This includes builds, typechecking, the full test matrix, end-to-end tests, runner verification, the canary dry run, and security scans. - Greptile is 5/5 on that exact head, with no unresolved review threads. This change has not been deployed to staging. ## Risks This changes authorization behavior. The instruction exception must not become an inherited task permission. The check uses server-owned execution records, requires the agent's own chat and instructions, and keeps normal protected-change and responsible-user checks. Saves fail closed when current provenance or permission is missing. Owner chat grants a turn-scoped capability; the server does not classify the message intent or require approval of the exact new file bytes. Prompt injection within an authorized owner-chat turn remains a model-level risk. This is the requested owner-chat trust boundary, without a new per-edit confirmation flow. No database migration or broad trust-preset change is required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code editing, shell tools, and test execution. The exact runtime model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
154 lines
8.6 KiB
Markdown
154 lines
8.6 KiB
Markdown
# Low-Trust Presets
|
|
|
|
Paperclip ships core trust preset names so containment decisions are enforced in
|
|
Community Edition even when EE policy editing is unavailable.
|
|
|
|
## Presets
|
|
|
|
- `standard`: the default V1 company-visible collaboration model. This preserves
|
|
existing behavior for normal agents.
|
|
- `low_trust_review`: an opt-in containment preset for automated work that may
|
|
consume hostile or prompt-injected input, such as untrusted pull requests,
|
|
external tickets, dependency diffs, or generated review output.
|
|
|
|
## Boundary Model
|
|
|
|
`low_trust_review` is resolved from existing JSON policy fields:
|
|
|
|
- agent permissions: `permissions.trustPreset` and
|
|
`permissions.authorizationPolicy.trustBoundary`
|
|
- project policy:
|
|
`executionWorkspacePolicy.authorizationPolicy.trustBoundary`
|
|
- issue/run policy: `executionPolicy.authorizationPolicy.trustBoundary`
|
|
|
|
The resolver intersects those sources. Narrower wins. A low-trust preset must
|
|
resolve to a concrete company-local project, root issue, or issue-id scope. If a
|
|
policy source names another company, uses an unsupported preset, or lacks that
|
|
scope for risky access, Paperclip fails closed.
|
|
|
|
## Containment, Not Privacy
|
|
|
|
This is containment for hostile automated work. It is not a general project,
|
|
issue, or human privacy system.
|
|
|
|
V1 standard work remains company-visible by default: board users and in-company
|
|
actors can inspect company work objects unless a separate access-control feature
|
|
changes that behavior. Low-trust containment instead limits what the low-trust
|
|
agent can read or mutate through the Paperclip API and prevents raw untrusted
|
|
output from being automatically promoted into higher-trust agent context.
|
|
|
|
Low-trust agents cannot read or mutate privileged agent configuration or company
|
|
skill configuration through direct grants. Persistent instruction edits are also
|
|
restricted, except for the direct-chat self-edit described below. Configuration
|
|
changes from outside-triggered work must go through higher-trust review and
|
|
promotion paths instead.
|
|
|
|
Low trust does not prohibit task creation. Agents may create tasks assigned to
|
|
themselves, and subtasks of their own tasks, within their permitted project or
|
|
root-task scope. Creation uses the normal task-assignment permission checks,
|
|
including explicit assignment restrictions and the responsible user's authority,
|
|
even for an unassigned task. Other assignees must be explicitly within the trust
|
|
boundary; board-user assignments remain forbidden. New tasks retain the creator's
|
|
effective containment policy and quarantined source attribution. An exact issue-id
|
|
scope does not implicitly authorize new descendants, and a permitted parent does
|
|
not grant access to an unrelated project. An explicit root-task scope includes
|
|
its descendants. Self-assigned decomposition is not a delegation cycle.
|
|
|
|
Permission denials must identify the rejected action and specific restriction.
|
|
Agents must surface failed instruction-save receipts and explain when the current
|
|
run lacks direct-chat authority or the user lacks editing permission. The user
|
|
can request the edit in their authorized chat or apply it directly. Creating
|
|
another task does not authorize the instruction change.
|
|
|
|
## Direct-chat instruction edits
|
|
|
|
When an authorized user directly asks an agent in their own Agent Chat to edit
|
|
its persistent instructions, the agent may save its own `AGENTS.md` even under
|
|
`low_trust_review`. The exception requires the current accepted execution identity
|
|
and a recorded authenticated board-message wake for that chat, user, and agent.
|
|
User attribution alone, outside connector/plugin messages, ordinary tasks,
|
|
subtasks, and messages from another user do not qualify. An earlier owner message
|
|
cannot authorize a later external instruction. External content read during the
|
|
chat is not a request from the user to change instructions.
|
|
|
|
The authenticated owner-chat turn is the permission boundary. The server checks
|
|
message provenance and current edit access; it does not classify the message's
|
|
natural-language intent or require a separate approval for each edit. The agent
|
|
must follow the owner's request and treat tool output as data. This means prompt
|
|
injection during an authorized owner-chat turn remains a model-level risk; the
|
|
exception does not provide content-bound approval of the resulting file bytes.
|
|
|
|
Normal instruction editing permissions and explicit protected-change rules still
|
|
apply. The user's current permission is rechecked at every save, including private
|
|
working-copy collection after a successful, failed, or timed-out execution.
|
|
Cancellation, chat reset, reassignment, message deletion, and revoked access
|
|
invalidate the exception. A server-linked retry of the same chat request may use
|
|
its original authenticated message; unrelated continuations may not. This
|
|
authority is never written into the inherited trust policy and does not grant
|
|
peer instruction edits, general configuration access, or other privileged APIs.
|
|
|
|
## Human-directed work
|
|
|
|
An authenticated board user can talk to a low-trust agent in their own Agent
|
|
Chat or assign it a task outside its default intake boundary. Existing conversation
|
|
identity authorizes owner chat. Ordinary tasks use the human requester already
|
|
recorded on the run's wakeup requests, including coalesced requests. Each request
|
|
retains its server-owned origin; plugin-attributed users do not count as board
|
|
instructions. Legacy board assignment requests remain valid through their existing
|
|
assignment source, reason, and human requester, without trusting old plugin human
|
|
attribution. A board backlog assignment is retained as a completed assignment
|
|
request with no run, so it authorizes the later system launch without starting
|
|
work prematurely. No separate permission table or client-supplied human identity is needed. Responsible-user
|
|
attribution, external connector sender attribution, and agent claims do not qualify.
|
|
|
|
At dispatch and API authorization, Paperclip checks the live run and current
|
|
assignment in one database snapshot. The exception permits reading, commenting
|
|
on, and updating only that run's exact task. It does not extend to another task,
|
|
a child, a whole project, configuration, instructions, secrets, or runtime
|
|
management. Normal responsible-user checks still apply. The exception is never
|
|
stored in the inherited trust boundary.
|
|
|
|
Automatic retries and continuations follow the existing `retryOfRunId` database
|
|
links, checking the same company, agent, and task at every step. Cancelled runs
|
|
cannot authorize a retry. The common assignment transaction cancels prior human
|
|
requests, including plugin and service reassignments. Assigning the task back
|
|
cannot revive those requests, and late run settlement cannot undo cancellation. Sandbox and
|
|
isolated workspace requirements still apply, as do malformed-policy and
|
|
company-boundary checks.
|
|
|
|
## Child→Parent Reporting Under Containment
|
|
|
|
The direct-parent report comment (`doc/execution-semantics.md` §6, "Child→Parent
|
|
Reporting") is **off by default** for `low_trust_review`: a contained run reads
|
|
untrusted input, so a free-prose comment into the higher-trust parent thread is
|
|
a prompt-injection promotion path. Contained reviewers report by completing
|
|
their own review issue (`done` — the verdict is the deliverable; the
|
|
`issue_blockers_resolved` wake carries it upward) and by the platform's
|
|
system-attributed stop-only relay when they enter `blocked` or `cancelled`.
|
|
Never instruct a contained delegate to comment on its parent issue.
|
|
|
|
## Runtime Containment
|
|
|
|
Managed `low_trust_review` runs fail closed unless Paperclip can enforce the
|
|
runtime boundary:
|
|
|
|
- the selected execution environment must use the `sandbox` driver
|
|
- the effective execution workspace mode must be `isolated_workspace`
|
|
- the issue being run must be inside the resolved low-trust boundary or be the
|
|
exact human-directed task described above
|
|
- secret references must use binding ids explicitly allowed by the boundary
|
|
- inline sensitive environment values such as API keys and tokens are rejected
|
|
- workspace runtime-service mutations are denied unless the boundary explicitly
|
|
grants the `runtime.manage` tool class
|
|
|
|
When the task's project has no configured workspace and no layer specifies a
|
|
workspace strategy, sandbox execution uses a private directory for that company
|
|
and task. The directory persists across turns and reassignment and never imports the shared
|
|
project directory or agent home. No Git repository is required for this case.
|
|
Configured workspaces and explicit Git strategies keep their existing validation;
|
|
a missing or broken checkout does not fall back to an empty directory.
|
|
|
|
The Docker workflow in `doc/UNTRUSTED-PR-REVIEW.md` remains useful for manual
|
|
local review, but Paperclip-managed low-trust execution requires a sandboxed
|
|
environment instead of a host-local adapter process.
|