mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-11 14:10:50 +02:00
4e8cd757eb696b2d61a0dce75f0b7c0312bfdf3a
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5ed0b74b34 |
fix(runtime): scope PAPERCLIP_ env-binding strip to reserved keys (#9974)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent runs get their environment from user/adapter/project/routine env bindings resolved by the server heartbeat, plus `PAPERCLIP_*` runtime vars (identity, wake, workspace, API access) injected by the harness > - The heartbeat stripped **every** `PAPERCLIP_`-prefixed binding before resolution, so legitimately user-named keys (e.g. cloud provider token bindings like `PAPERCLIP_CLOUD_PROD_PROVIDER_RAILWAY_*`) were silently dropped and never reached the run env > - At the same time, several adapters honored an explicitly configured `PAPERCLIP_API_KEY` over the harness-minted run token, which is exactly the one key config must never control > - This pull request replaces the blanket prefix strip with a precise three-rule policy: never accept `PAPERCLIP_API_KEY` from config, always let harness-assigned runtime vars win, and let every other `PAPERCLIP_*`-named user binding flow through > - The benefit is that user secrets with a `PAPERCLIP_`-style name work like any other binding, while runtime identity and API credentials stay fully harness-controlled ## Linked Issues or Issue Description **Bug description** (no public issue exists): - **What happened:** Env bindings whose key starts with `PAPERCLIP_` (e.g. a cloud provider token a user deliberately named `PAPERCLIP_CLOUD_PROD_PROVIDER_RAILWAY_TOKEN`) were silently stripped by the server before secret resolution, so the spawned agent never received them. No error, no access event — the variable just never appeared. - **Expected behavior:** A user-named `PAPERCLIP_*` binding should reach the run env unless the harness itself uses that key. Only `PAPERCLIP_API_KEY` should be categorically rejected, and harness-assigned runtime vars (`PAPERCLIP_RUN_ID`, `PAPERCLIP_AGENT_ID`, wake/workspace vars, …) should always win over config. - **Steps to reproduce:** Configure an agent/project env binding named `PAPERCLIP_<ANYTHING>` (plain or secret_ref), run a heartbeat, and inspect the spawned process env — the key is absent. - **Deployment mode:** local server, any local adapter. Related prior PRs (different, save-time/API-layer blanket-ban approach; this PR supersedes that direction with a runtime allow-except-reserved policy): Refs #8239, Refs #8439. ## What Changed - `server/src/services/heartbeat.ts`: the pre-resolution strip now removes only `PAPERCLIP_API_KEY` (hard denylist) instead of every `PAPERCLIP_`-prefixed binding; other `PAPERCLIP_*` keys flow into binding resolution. Low-trust inline-sensitive-env checks now also cover those keys. - `packages/adapter-utils/src/server-utils.ts`: new `isForbiddenConfigEnvKey()` helper; the shared `refreshPaperclipWorkspaceEnvForExecution` merge drops `PAPERCLIP_API_KEY` from config and keeps harness-assigned `PAPERCLIP_*` keys authoritative. - `packages/adapter-utils/src/acpx-engine/execute.ts`: removed the explicit-`PAPERCLIP_API_KEY`-from-config allowance; the run token (`authToken`) is now always applied; config `PAPERCLIP_API_KEY` is ignored. - All local adapters (`claude-local`, `codex-local`, `cursor-local`, `gemini-local`, `grok-local`, `opencode-local`, `pi-local`) plus `cursor-cloud`, `hermes`, and the server `process` adapter: removed `hasExplicitApiKey`-style allowances so the harness token always wins, and guarded the remaining unguarded env-merge loops (claude-local inline loop, process adapter) with the same policy. - Tests updated/added: heartbeat binding-strip test now asserts the three-rule policy; adapter-utils merge tests assert the `PAPERCLIP_API_KEY` ban and `PAPERCLIP_*` pass-through; acpx engine tests moved credential fixtures to `authToken` and assert config `PAPERCLIP_API_KEY` is ignored while other `PAPERCLIP_*` config keys forward and still bust the session fingerprint on rotation. ## Verification - `pnpm vitest run packages/adapter-utils/src/server-utils.test.ts packages/adapter-utils/src/acpx-engine/execute.test.ts` — 127 passed - `pnpm vitest run server/src/__tests__/heartbeat-project-env.test.ts server/src/__tests__/heartbeat-local-environment.test.ts server/src/__tests__/claude-local-execute.test.ts server/src/__tests__/codex-local-execute.test.ts server/src/__tests__/cursor-local-execute.test.ts server/src/__tests__/gemini-local-execute.test.ts` — 68 passed - Adapter package execute suites and the server tests touching API-key fixtures (`heartbeat-run-log`, `redaction`, `effective-run-config-fingerprints`, `agent-permissions-routes`) — green. Three pre-existing sandbox/SSH fixture failures reproduce identically on clean `master` on this host and are unrelated. - `pnpm --filter <pkg> typecheck` for server, adapter-utils, and all nine touched adapter packages — all pass. ## Risks - Behavioral change: a deployment that relied on configuring a static `PAPERCLIP_API_KEY` in adapter config env loses that override — by design; the harness-minted run token is now the only source. When no run token exists, no API key is injected at all. - `PAPERCLIP_*`-named user bindings now reach binding resolution and run envs; a key that collides with a harness runtime var is still discarded at merge time, so runtime identity/wake/workspace vars cannot be spoofed. - Low risk otherwise: no migrations, no API surface changes. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic Claude 5 family, Mythos-class tier), extended thinking enabled, agentic tool use (file edits, shell, test runner) via Claude Agent SDK. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d32ed88443 |
fix(recovery): route recovery by failure cause (#9634)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent companies and their work. > - Its recovery subsystem detects stranded issue execution and decides whether to retry, escalate, or request operator intervention. > - The existing recovery path used a mostly generic owner ladder and generic execution contract, so transient failures could wake a manager who then performed the deliverable instead of repairing and returning the task. > - Provider quota failures also entered the same takeover path even when the correct action was to wait for capacity and retry the original assignee. > - Recovery actions already retain the source owner and evidence needed to choose a cause-specific route, render a scoped contract, and measure whether work was handed back. > - This pull request adds a cause-keyed recovery playbook, propagates its contract through every built-in adapter, and makes resolved recovery actions return work to the original owner by default. > - The benefit is bounded self-recovery that preserves task ownership, avoids needless management takeover, and makes recovery outcomes observable. ## Linked Issues or Issue Description No matching public GitHub issue was found. Related recovery work was reviewed but is not duplicated here: #9630 restores bounded recovery continuations, #8807 changes one assignee-ranking case, and #9404 records runtime-failure transition evidence. This change instead introduces cause-specific routing and recovery contracts across the recovery lifecycle. ### What happened? When an issue became stranded, recovery generally selected an owner through the same fallback ladder and rendered the normal execution contract. That made the recovery wake look like ordinary deliverable work, even when the correct action was to retry the original agent, repair its runtime, or wait for a provider quota reset. ### Expected behavior Recovery should select a response by failure cause, tell the recipient to recover rather than complete the deliverable, suppress takeover wakes for provider quota waits, and return repaired work to its original assignee unless the recovery owner explicitly completes it. ### Actual behavior Recovery could escalate transient failures to management, omit the cause-specific next action from the wake, and leave the recovery owner assigned after the runtime problem was resolved. ### Impact The generic path creates avoidable management work, ownership churn, and budget consumption while obscuring whether recovery successfully returned work to the responsible agent. ## What Changed - Added cause-keyed routing for process loss, missing disposition, provider quota limits, Codex output inactivity, workspace validation failures, and fallback recovery causes. - Added recovery-scoped wake rendering that replaces the generic execution contract with the failure summary, original assignee, attempt count, next action, and cause-specific playbook instruction. - Propagated the structured recovery contract through all built-in adapter execution paths, including Hermes local and gateway adapters. - Added provider-quota wait monitoring so capacity failures schedule the original assignee instead of enqueueing a takeover wake. - Added hand-back behavior and `handed_back` / `owner_completed` outcome accounting when recovery actions are resolved. - Added focused routing, renderer, quota-monitor, and hand-back regression coverage plus implementation-spec documentation. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/server-utils.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts server/src/__tests__/heartbeat-workspace-branch-containment.test.ts server/src/__tests__/issue-recovery-actions.test.ts` - 4 test files passed; 194 tests passed. - Targeted `pnpm --filter ... typecheck` across `@paperclipai/adapter-utils`, `@paperclipai/shared`, `@paperclipai/server`, `@paperclipai/ui`, and all nine changed adapter packages. - 13 affected workspace packages passed typecheck. - `pnpm check:token-gates` - All UI token gates passed. ## Risks - Recovery routing behavior changes for stranded work, so an incorrectly classified cause could select a different recipient than before; fallback causes retain the existing management ladder. - Provider quota detection depends on structured failure evidence and conservative text matching; unmatched failures continue through fallback recovery. - Adapter prompt plumbing changes across built-ins, covered by shared renderer tests and compile-time call signatures. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with exact model ID `gpt-5.6-sol`, using reasoning, tool use, and code execution. The runtime does not expose its configured context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
534aee66ae |
Add cursor_cloud adapter for Cursor SDK + Cloud Agents API v1 (#5664)
## Thinking Path
> - Paperclip orchestrates AI agents for zero-human companies
> - There are many adapter types, one per agent-runtime product (Claude,
Codex, OpenCode, Cursor local CLI, etc.)
> - Cursor shipped a public TypeScript SDK on 2026-04-29 that exposes
Cursor's full hosted-agent platform (cloud VMs, harness, MCP, skills,
hooks)
> - Paperclip had no first-class adapter for this — agents that wanted
to use Cursor's managed cloud runtime had to fall back to the local CLI
adapter, which loses the cloud session, streaming, and durable run model
> - This PR adds a new `cursor_cloud` adapter built directly on
`@cursor/sdk`, with Paperclip's heartbeat mapped to Cursor's
durable-agent + per-run model
> - The benefit is that any Paperclip agent can now drive a Cursor cloud
agent across heartbeats with native session reuse, streaming, and
cancellation, while Paperclip remains the source of truth for issue/task
state
## What Changed
- New built-in adapter package `packages/adapters/cursor-cloud` (15
files, ~1.7k LOC) backed by `@cursor/sdk` ^1.0.12
- `src/server/execute.ts` — SDK-first lifecycle: `Agent.create` /
`Agent.resume` / `Agent.getRun` / `agent.send` / `run.stream` /
`run.wait`, with session reuse keyed on the (runtime env type, env name,
repo set) tuple
- `src/server/session.ts` — codec for `cursorAgentId` + `latestRunId` +
repo metadata, persisted in `runtime.sessionParams`
- `src/server/test.ts` — environment probe via `Cursor.me()` and
optional model validation via `Cursor.models.list()`
- `src/ui/parse-stdout.ts` + `src/cli/format-event.ts` — normalize
Cursor SDK message types (`status`, `thinking`, `assistant`, `user`,
`tool_call`, `tool_result`, `result`) into Paperclip transcript events
for the UI and CLI
- Registrations: `packages/shared/src/constants.ts`,
`packages/adapter-utils/src/session-compaction.ts`,
`server/src/adapters/{registry,builtin-adapter-types}.ts`,
`ui/src/adapters/{registry,adapter-display-registry}.ts` +
`ui/src/adapters/cursor-cloud/index.ts`, `cli/src/adapters/registry.ts`,
plus workspace deps in `cli`/`server`/`ui` `package.json`
- `ui/src/components/AgentConfigForm.tsx` — hide local-Cursor
`mode`/thinking-effort field for `cursor_cloud` (different config
surface)
- 11 vitest tests covering execute paths (fresh create, matching-resume,
active-run reattach, non-finished result), session codec round-trip,
transcript parsing, and config building
## Verification
Reviewer steps:
```bash
pnpm install
pnpm --filter @paperclipai/adapter-cursor-cloud typecheck # → clean
pnpm vitest run packages/adapters/cursor-cloud # → 11/11 passing
```
End-to-end check against a real Cursor cloud agent (requires
`CURSOR_API_KEY` and Cursor GitHub-app install on the target repo):
1. Create a `cursor_cloud` agent in Paperclip with `repoUrl` set to the
test repo, `repoStartingRef: main`, and `env.CURSOR_API_KEY` set
2. Trigger a heartbeat → adapter calls `Agent.create({ cloud: { env: {
type: "cloud" }, repos: [...] } })`, streams events, terminates on
`finished`
3. Trigger a second heartbeat → adapter calls `Agent.resume` or
`agent.send` follow-up depending on prior-run state, reusing
`cursorAgentId`
4. The Paperclip UI/CLI transcript reflects Cursor `status` / `thinking`
/ `assistant` events as they stream
5. Cancellation from Paperclip maps to `run.cancel()` or Cloud API v1
`cancelRun` for cross-heartbeat cancellation
A direct-SDK smoke run against a real repo (devinfoley/my_test_project @
main) confirmed: `Cursor.me()` ok → `Agent.create` → `agent.send` →
`run.stream()` (30 events) → terminal status `finished` in ~11s.
## Risks
- **New adapter, additive only.** No existing adapter or registry is
replaced; current `cursor` local-CLI adapter is untouched. Default
behavior of any existing agent is unchanged.
- **External dependency on `@cursor/sdk`.** Cursor's SDK is v1.0.x and
may evolve. Mocked unit tests cover the public surface used here; if the
SDK breaks compatibility we update the adapter independently.
- **Cost/budget.** `cursor_cloud` runs on Cursor's billed cloud VMs;
operators must understand they are spending money outside Paperclip's
budget controls when they enable this adapter. Same shape as other
API-billed adapters.
- **No webhook support in V1.** The SDK already provides
stream/wait/cancel/reattach, so V1 does not require a public callback
URL. If a future use case needs out-of-band wakes, we add a Cloud API v1
webhook bridge as a separate change. This is called out in the issue
plan document.
- **Lockfile.** Per repo policy, `pnpm-lock.yaml` is intentionally not
in this PR — CI's lockfile workflow will update it on merge given the
manifest changes.
## Model Used
- Provider: Anthropic Claude (via Claude Code / Paperclip `claude_local`
adapter)
- Model: `claude-opus-4-7` (Claude Opus 4.7), knowledge cutoff January
2026
- Mode: standard tool-use with extended reasoning
- Context: ~200k token window
- Capabilities used: code generation, multi-file edits, shell/test
execution, GitHub PR workflow
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have run tests locally and they pass (11/11 in
`packages/adapters/cursor-cloud`)
- [x] I have added or updated tests where applicable (4 new test files,
11 cases)
- [ ] If this change affects the UI, I have included before/after
screenshots (the only UI change is hiding the local-Cursor mode field on
the `cursor_cloud` adapter — happy to attach a screenshot if the
reviewer wants one)
- [x] I have updated relevant documentation to reflect my changes (issue
plan document supersedes the pre-SDK design; tracked in PAPA-203)
- [x] I have considered and documented any risks above
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
|