Commit Graph
18 Commits
Author SHA1 Message Date
DottaandPaperclip 2ec82c5774 fix(runner): preserve task context when tool connections change (#14963)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use tools through company-scoped connections and provider
sessions.
> - Resolving a tool connection currently forces a fresh session even
when the provider can load new tools into the existing conversation.
> - A fresh provider conversation can receive too little history to
continue the task.
> - This pull request adds explicit tool-refresh capabilities and uses
them in both runner paths.
> - Fresh attempts receive bounded task history with source IDs and
retrieval instructions.
> - The benefit is that agents can continue the same task after a
connection changes.

## Linked Issues or Issue Description

**What happened?**

A resolved tool connection forced a fresh provider conversation. The new
conversation could lose the original goal and prior answers. Claude also
rejected resume when only the MCP server set changed.

**Expected behavior**

Resume the provider conversation when its harness can refresh tools.
When a fresh session is required, supply enough bounded history to
continue the task. Preserve company, agent, task, workspace, model,
instruction, and skill checks.

**Steps to reproduce**

1. Start a conversation and agree on a task and its constraints.
2. Request and connect a tool needed for the task.
3. Continue the conversation after the connection resolves.
4. Check that the agent remembers the task and can use the new tool.

**Paperclip version or commit**

The bug was reproduced on master at `c46e41e81`. This branch is rebased
on current master.

**Deployment mode**

Self-hosted server. Both legacy adapters and the native runner are
affected.

Related public work: Refs #13282 for task-backed conversations. Refs
#13057 for the broader session-compaction proposal. Refs #14659 for
another report about local CLI session continuity. This change fixes
tool-connection continuation. Provider authentication repairs keep their
existing recovery behavior.

## What Changed

- Expose tool-refresh support in native harness descriptors and legacy
adapter metadata.
- Request tool refresh after connection resolution. Keep provider
authentication repair as a fresh-session wake.
- Reload current tools and credentials while retaining supported Claude,
Codex, Grok, and other provider conversations.
- Allow MCP-only changes during qualified native recovery. Keep all
other compatibility checks.
- Refresh managed-provider and ACPX tool bindings when attaching a new
run.
- Add a fresh-session handoff for both runner paths. Bound database
reads, excerpts, and the final packet to 24,000 bytes.
- Include the original request, recent messages, decisions, plans, prior
answers, and source IDs. Mark omitted content. Apply reset boundaries,
wake cutoffs, quarantine, and secret redaction.
- Add regression tests and document the capabilities and handoff
behavior.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed. Rust formatting passed.
- Final review fixes passed 314 server tests, 333 adapter utility tests,
139 native-session runtime tests, and 12 managed-provider Rust tests.
They verify historical quarantine, raised budgets across attachment, no
history reads on successful resume, and handoff delivery on fresh retry.
- Broader branch verification also passed 1,401 adapter utility tests,
1,047 runner TypeScript tests, 43 Grok adapter tests, and 311 Rust core
tests.
- Live Claude CLI and Grok ACP probes preserved the provider session ID,
recalled a prior task constraint, and called a newly added read-only MCP
tool.
- GitHub CI passed on `b21486d18084a7aa4cafbe8e012f7cad6585d9cc`: 55
successful checks and 4 skipped checks. This includes all test shards,
all eight browser shards, runner checks, and the Grok clean public npm
install canary. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37057514976).
- A full local test attempt encountered a separate Git snapshot timeout.
All affected local suites passed after the final edits, and the full CI
test gates passed.
- Review the capability matrix in `packages/paperclip-runner/README.md`.
Repeat the four reproduction steps with a supported provider and with an
unsupported harness.

## Risks

- Provider tool refresh can fail. Existing recovery falls back to a
fresh conversation where policy permits it.
- A new transport can replace an old process while preserving the
provider conversation. Tests cover current credentials and unchanged
identity.
- Long history can omit older context. Explicit markers and source IDs
let the agent retrieve needed context within task scope.
- Unknown and unqualified harnesses use the fresh-session path. No
database migration is required.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and live
provider testing. The exact model ID and context-window size are not
exposed in this session. Claude and Grok also ran as test subjects.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 15:11:44 -05:00
da887ea3e9 fix(runner): honor Codex effort selected in composer (#14568)
## Thinking Path

> - Paperclip manages AI agents that work on assigned tasks.
> - The task composer lets a person choose an assignee, model, and
effort for the next run.
> - A Paperclip Runner agent can use Codex as its provider.
> - The composer hid Codex effort for that agent because it checked only
the older Codex adapter.
> - The native Runner input also did not carry an effort choice to
Codex.
> - This pull request carries the chosen effort from the composer to
each Codex turn.
> - People can now select a supported effort and get the effort they
selected.

## Linked Issues or Issue Description

Refs #14322

**What happened?**

The composer showed a model but no effort slider when the assignee used
Paperclip Runner with the Codex provider. A task-level model override
also did not reach the native Runner input.

**Expected behavior**

The composer shows effort choices for a known Codex model. The next
native Codex turn uses the selected model and effort.

**Steps to reproduce**

1. Open a task composer.
2. Select an agent that uses Paperclip Runner with the Codex provider.
3. Select a known Codex model such as `gpt-6-astra`.
4. Open the assignee and model picker. The effort slider is missing
before this change.

## What Changed

- Show known Codex effort levels for Paperclip Runner Codex assignees.
- Save the task effort override in the native run input and send it to
Codex on each turn.
- Apply the task's merged model and effort overrides when the native run
starts.
- Apply a task model override for OpenCode Runner without changing the
agent's provider.
- Add Runner effort tests and desktop and mobile Storybook cases.

## Verification

- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- `pnpm build-storybook` passed.
- `pnpm check:token-gates` passed.
- Focused UI, server, Runner contract, and Codex driver tests passed.
- The full CI test matrix, build, typecheck, and canary dry run passed
on the latest head.

## Risks

- Native Runner inputs add an optional Codex effort field to the current
v5 input. Older inputs keep their previous behavior.
- A known model rejects an effort that its catalog does not support.
Unknown models do not show a slider.

> This fixes an existing composer bug. I checked `ROADMAP.md`; it does
not describe this bug as planned work.

## Model Used

OpenAI Codex, GPT-6. The exact deployment ID and context window are not
exposed in this session. The model used reasoning, code execution, and
repository tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI GPT-6 Astra <noreply@openai.com>
2026-09-29 09:41:05 -05:00
DottaandPaperclip 992f720262 fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English
(ASD-STE100). -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task descriptions, comments, continuation data, skills, and
execution rules enter several agent adapters.
> - The same source can be rendered by more than one automatic input
carrier.
> - Failed resumes can also rebuild input from stale or compact context.
> - This pull request gives each Paperclip-owned source one delivery
owner and preserves the required transport boundaries.
> - It adds deterministic adapter, interaction, runner, and browser
tests for these boundaries.
> - The benefit is more predictable context delivery with explicit
evidence for later live qualification.

## Linked Issues or Issue Description

Related: #13144 removes a duplicate environment payload and bounds wake
lists. Related: #11360 addresses Hermes resume behavior. This pull
request preserves compatible active-session formats while repairing
context ownership and stale question creation.

**What happened?**

Task descriptions and comments could enter more than one automatic
context block. Native transports could wrap a complete model input in a
second task envelope. Some legacy and gateway adapters could omit the
owned assignment on ordinary tasks or rebuild a failed resume with stale
compact context. A continuation could also request a question after
newer human comments had arrived.

**Expected behavior**

Each task or comment source has one automatic model-facing owner.
Distinct comment IDs and repeated wording remain distinct. Fresh
fallback attempts rebuild the required full context. A question request
is rejected when newer queued human direction makes it stale. Harness
access policy remains owned by execution configuration.

**Steps to reproduce**

1. Build a task with a description and current comments.
2. Capture the actual adapter or runner input.
3. Compare source ownership and task-envelope nesting.
4. Queue a human comment before a continuation requests a question.
5. Trigger a failed resume and inspect the fresh retry input.
6. Run the focused adapter, interaction, runner, and browser checks.

## What Changed

- Add shared prompt-section selection at the provider-attempt boundary.
- Deliver owned assignment context through native, legacy CLI, ACP,
gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and
Hermes paths.
- Rebuild full or compact context after resume recovery changes the
attempt. Add native and Claude ACP tests of actual recovery requests.
- Preserve custom templates, loaded instruction files, execution
policies, and older active-session formats.
- Record continuation source metadata and reject stale question creation
under the issue-row lock.
- Add explicit Product E2E context-integrity profiles, prerequisite
gates, credential-isolation checks, and report fixtures.
- Bypass service-worker forwarding for same-origin Vite development
modules. A real Chromium test fails with resource exhaustion before the
repair and passes after it. Production asset caching keeps its existing
policy.
- Add browser diagnostics and service-worker module-loading regressions.
- Add an explicit zero-retry eval option. The default retry behavior
remains unchanged. Each campaign records its effective policy.
- Remove the model-facing working-directory sentence from four prompt
builders. Existing workspace, sandbox, permission, and custom-template
configuration remains unchanged.
- Align the everyday workflow assertion with the current 47-entry
catalog.

Compared with current upstream master, the branch carries the
context-ownership implementation and its tests, the explicit
context-integrity catalog and evidence harness, and the focused browser
regression checks.

## Verification

**Merge assessment:** focused regression evidence supports merge. This
is not full completion of the original broad qualification matrix. The
maintainer has authorized merge after fresh verification of the master
integration.

- Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This
integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`.
All 14 conflicts are resolved. Cancellation checks, workspace
finalization, native Grok support, and both sets of tests are retained.
- Current-head Greptile: **5/5**, with no blocking findings. The review
names this exact commit. All **59 reported checks are terminal: 55
successful, 4 skipped, zero pending or failing**. This includes the full
root general and serialized suites, separate runner checks, typecheck,
build, canary, browser E2E, Docker, and security checks. The successful
legacy security status is included in that total.
- After integration: workspace typecheck and full build passed. Separate
runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust
tests, and 39 preparation checks**. Other passing checks include 621
Product E2E harness units, 376 focused shared/adapter tests, 160
real-database/API tests, 86 Hermes tests, 18 browser-support checks, and
Product E2E typechecking. The complete root suite passed in CI. The
duplicate local monolithic root run was stopped after that CI result; it
is not counted as a completed local pass.
- New native recovery coverage retains full assignment, completion
contract, and explicit skill selection after safe replacement, for old
and prepared input formats. Full native session test file: **136/136
passed**.
- New Claude ACP coverage captures actual fresh, resumed, and
missing-session fallback requests. It verifies one assignment copy,
comment order, identical text under distinct comment IDs, and full
fallback context. Full file: **33/33 passed**. Both affected TypeScript
checks passed.
- Existing deterministic tests cover source revisions, approval and
trust boundaries, completion validation, custom templates, compatible
sessions, standalone driver wrapping, and maintained adapter transport
requests.
- Provider-free browser support: **17/17 passed** after the master
merge. Service-worker unit tests: **33/33 passed**. The module-overload
regression failed before the repair and passed after it in real
Chromium.

### Fresh live comparisons

The new batch ran exactly four Product E2E attempts. **All four passed
on the first attempt; no retries.** Each has six terminal matchers plus
the existing browser lifecycle and invariant checks.

| Exact case ID | Control | Candidate |
|---|---|---|
| `core-compatibility.runner-codex.local.plan-revise-accept` | Passed |
Passed |
|
`local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume`
| Passed | Passed |

The plan case checks a revised canonical plan and revision-bound
approval before completion. The question case restarts the server before
submitting the answer, then verifies the continuation completes.

Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate
source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical
frozen definitions and provider versions: Codex `0.156.0` with
`gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with
`claude-sonnet-5`. The September 24 head added master browser recovery
and test-only changes. The September 28 head also integrates newer
master changes, including cancellation, workspace finalization, and
native Grok. These are frozen-source live results, not exact-head live
runs.

The candidate received one description copy where the control initially
received three. The submitted initial plan envelopes were 7,969 versus
19,097 characters. Question envelopes were 7,592 versus 18,919. These
are structural measurements, not whole-provider token or dollar savings.

### Earlier evidence and failed attempts

- The preceding fresh batch has four effective passing pairs: OpenCode
comment continuation and assigned skill, native Codex comment
continuation, and native Claude comment continuation. It retains **11
attempts: eight passed and three failed**.
- Original failures remain recorded: missing local PostgreSQL library
links before task creation; host-sleep cleanup after task/page checks
passed; and a Claude **control** session-open rejection before a model
turn. Setup was repaired identically on both worktrees. The permitted
unchanged infrastructure retries passed. The underlying Claude provider
startup error was not retained and remains unknown.
- Older R2 retains **17 passes and one failure** across 18 attempts,
including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode
blank-page failure led to the service-worker repair. R2 is historical
evidence: master changed the native fixed prompt and removed duplicate
wake environment data afterward.
- The September 24 CI run initially failed one unrelated preview
readiness test (`ECONNREFUSED` on its local fixture). Its test and
production code match master. Isolated local verification passed **28
tests, 3 skipped**. One unchanged CI retry passed the full shard: **831
passed, 1 skipped**, including all **31 preview-exposure tests**. The
aggregate CI gate passed afterward. The precise startup cause remains
unknown; a port race is a hypothesis, not a proved cause.

### Limits

The original wider profile/workflow matrix, repeated trials, and remote
Daytona qualification are incomplete. These results support a focused
merge recommendation, not statistical equivalence or universal harness
qualification. Some usage receipts are missing in both variants, so no
token or dollar savings are claimed. The $500 ceiling was preserved
using conservative allowances; failed attempts and unknown charges
remain in the ledger.

Reproduce the focused additions with `pnpm exec vitest run
packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm
--filter @paperclipai/paperclip-runner exec vitest run
src/native-session-runtime.test.ts`. Full checks use `pnpm -r
typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner
checks. Paid evals require the frozen definitions, profiles, and
credentials; do not use `--all` as a substitute for the selected cases.

## Risks

- Context placement changes can affect model behavior. Deterministic
checks cover the selected paths, but live qualification remains
incomplete.
- The stale-question guard can reject a request when queued human
comments arrived during the run. This is intended.
- New stored inputs and model envelopes retain compatibility readers for
older active sessions.
- Custom templates may intentionally repeat content.
- Removing a model-facing working-directory sentence does not change
filesystem, command, sandbox, or permission configuration.
- The worker bypass applies only to same-origin development module
paths. Cache-policy tests preserve private-response handling and
production asset caching. Mounted HTTP fixture changes remain test-only.
- This PR does not claim measured token savings or statistical
equivalence across every harness.

## Model Used

OpenAI Codex, exact model gpt-6-astra, with repository tools and code
execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The
serving context-window size is not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR using the required issue fields
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run the focused local checks and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect these changes
- [x] I have considered and documented risks above
- [x] All current-head Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
for the current head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:49:14 -05:00
DottaandPaperclip d9b3a5653e feat(chat): add initial Slack communication guidance and connection menus (#13760)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors let people use the same tasks and agent tools from
external conversations.
> - Agents need communication guidance that fits the conversation
medium.
> - That guidance belongs in the original task context, without repeated
instructions on each turn.
> - Connection owners also need clear settings and a consistent way to
remove a connection.
> - This pull request adds initial Slack guidance, optional connection
instructions, and chat connection menus.
> - The benefit is clearer Slack replies with the existing Paperclip
workflow and permissions.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Agent replies in Slack and chat connection management in the Apps
catalog.

**Current behavior**

Slack tasks do not carry a saved communication profile. The catalog
shows a separate Manage button and does not offer removal on every chat
connection row.

**Proposed behavior**

Save Slack guidance when a new conversation creates a task. Restore that
original guidance when a model session is rebuilt. Do not append it to
ordinary follow-ups. Expose optional additional instructions in Slack
Settings. Put Manage and Remove connection in a three-dot menu for all
chat providers. Keep Finish setup visible for drafts.

**Reason and benefit**

Small answers fit in Slack. Substantial deliverables use ordinary
document or artifact tools with a useful Slack summary. Connection
settings apply to new tasks and cannot change permissions. Users can
remove both active and unfinished chat connections from the catalog.

**Breaking changes**

Two additive database columns store endpoint preferences and the initial
conversation snapshot. Existing endpoints default to empty preferences.
Existing conversations keep their original behavior. Non-Slack guidance
is unchanged.

Related public context:
https://github.com/paperclipai/paperclip/pull/13741 improves native chat
recovery. This change adds communication context to those existing
execution paths. A search found no duplicate communication-guidance PR.

## What Changed

- Add a provider-guidance registry, enabled for Slack first.
- Persist optional endpoint communication instructions and capture an
immutable snapshot when a conversation creates a task.
- Resolve guidance from the verified company-scoped connection. Restore
it for fresh native and legacy sessions without per-turn reminders,
extra model calls, or extra context queries.
- Add the Slack Settings field, validation, audit coverage, and
Storybook save/error states.
- Add Manage and Remove connection menus for all seven chat providers.
Keep the draft setup button. Require removal confirmation and allow
retry after failure.
- Add regression coverage, an active/draft menu story, and connector
documentation.

## Verification

All CI checks are green for 5f48df4e0. Greptile scored this head 5/5
with no actionable findings. No review threads remain unresolved.

- Passed `pnpm -r typecheck` and `pnpm build` on PR head 5f48df4e0.
- Passed design-token checks, UI typecheck, and all 20 catalog tests
after rebase. Tests cover all seven providers, active/draft removal,
confirmation, cache refresh, errors, and cancellation.
- Verified the active/draft menu in Storybook. The interaction test runs
without browser console errors.
- Passed focused guidance, endpoint persistence/isolation, heartbeat
trust, native context, ACPX, adapter utility, and CLI recovery tests.
Full UI and CLI groups passed (6,512 and 502 tests).
- Tested real Slack conversations on staging: concise updates with
public links, a planning question with buttons, a saved plan, a saved
report, task creation and assignment, and explicit detailed output. Old
tasks retained original preferences after an edit; a new task used the
changed preferences. Restored the staging setting afterward.
- Existing safe progress remained visible without duplicate final
replies or private reasoning.
- Broad local tests found resource/time-sensitive failures that passed
targeted reruns. One Cursor archive-download fixture failed on both this
branch and the unchanged main checkout. The full local suite is not
claimed clean. All PR-head CI test shards passed, including general,
serialized, Runner, and browser suites. The redundant local full-suite
rerun was stopped after CI completed successfully.
- Live delegation was not tested because the staging company has only
one agent. Live testing also found separate latency and runner
task-editing capability gaps; this PR does not add connector-specific
workflow behavior to hide them.

## Risks

- Prompt guidance changes the form of new Slack replies. Explicit
requests for detail still take precedence.
- The additive migration is idempotent. Conversation snapshots remain
fixed when connection settings change.
- Native and legacy recovery must preserve the initial context without
duplicates; targeted tests cover these paths.
- Removing a connection stops new work through the existing lifecycle
action. It retains Paperclip task history and does not delete the
external app or bot.

## Model Used

OpenAI GPT-6 through Codex, with repository editing, shell tools, and
browser testing. The host does not expose a more specific model ID or
context-window size. No separate model calls were added to the product.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted suites; broad
local limitations are listed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 14:14:25 -05:00
DottaandPaperclip 9f30eb10dd fix: reduce chat latency and preserve managed session reuse (#13710)
## Thinking Path

> - Paperclip manages AI agents and keeps their work attached to tasks.
> - Chat connectors carry user messages and agent replies between a
provider and those tasks.
> - Each extra startup and context reset delays a reply.
> - Managed account metadata was lost during adapter decoding, so
compatible follow-ups started fresh.
> - This branch fixes the reset and measures the remaining preparation,
execution, and delivery costs.
> - The changes must preserve account isolation, authorization, durable
output, and recovery ownership.

## Linked Issues or Issue Description

Refs #13699. The related service lifecycle work in #13410 and #13408 is
separate; this branch focuses on task-bound chat response latency.

**What happened?**

Managed AI follow-ups started new provider sessions even after their
configuration fingerprint stayed stable. The Codex codec removes unknown
fields. The resume check then read the removed credential identity and
treated it as a credential change.

**Expected behavior**

Compatible follow-ups resume the correct provider session. Changes to
credentials, responsible users, permissions, or task configuration
retain their reset behavior.

**Steps to reproduce**

1. Use a Slack connector with a managed AI connection.
2. Send a message, then send a same-thread follow-up.
3. Inspect the configuration reset reason and the provider session
identity.

**Paperclip version or commit**

Reproduced on `2a99de80ec52db01eead901f28323926ceaf3c1d`.

**Deployment mode**

Cloud staging with a native Codex runner.

## What Changed

- Read saved credential identity before adapter decoding discards it.
- Remove the internal credential identity from adapter-facing session
params.
- Test the real Codex codec and missing, changed, or unmanaged identity
cases.
- Preserve configured warm Codex runners and flush refreshed credentials
after every turn.
- Fence detached or closing session handles from successor credential
ownership.
- Stage current Codex launch credentials after restoring durable session
history, uploading launch assets only once.
- Reuse a runner binary already in the retained sandbox only when its
SHA-256 matches the controller-owned artifact; still verify required
capabilities before launch.
- Lock the task before the run when saving results, preventing deadlocks
with task updates.
- Scope reusable projectless sandboxes to the company, environment,
task, agent, and runtime configuration; verify Daytona sentinels for
that scope.
- Admit a new authorized chat message after a fully committed failed run
and verified process cleanup.
- Send compact deltas for verified plain-text Slack continuations. Match
the actual prior run and current comment identity/body; exclude edited
historical comments and prior agent output, preserve genuine brief edits
and the full bootstrap fallback.
- Keep attachments, omitted input, questions, approvals, recovery, and
other providers on their existing framing.
- Document managed session compatibility, credential lifecycle, and
compact continuation boundaries.

## Verification

- Workspace/session coverage: 156 tests passed.
- Native session and credential ownership coverage: 390 tests passed,
including exact artifact reuse, mismatches, failed probes, timeouts, and
explicit artifact overrides.
- Explicit continuation and durable chat authorization coverage: 172
tests passed.
- Session resume and launch preparation coverage: 416 tests passed.
- Result persistence coverage: 15 tests passed. The new concurrency test
reproduced a PostgreSQL deadlock before the lock-order fix.
- Environment lifecycle coverage: 92 tests passed, including projectless
reuse and task/agent isolation at both selection and atomic handoff.
- Daytona plugin coverage: 237 tests passed; 6 gated tests skipped.
Standalone plugin build passed.
- Compact Slack continuation and native resume coverage: 69 tests
passed, including full-bootstrap retention, matching message
authors/bodies, current-delivery selection, rejection of duplicate
identities and historical comments, brief edits, and attachment/recovery
fallbacks.
- Final frozen-head `pnpm test:run` on repository-supported Node 26: 668
suites passed, 3 skipped, 1 failed; 12,797 tests passed and 82 skipped.
The sole failure was a local `socket hang up` in
`issue-recovery-actions.test.ts`, not an authorization assertion
mismatch. All 57 tests in that suite passed three fresh reruns, and the
suite passed latest-head CI. The full local invocation is therefore not
claimed green.
- An earlier Node 24 full run exposed an unrelated macOS symlink-cleanup
failure; that 11-test catalog suite passes on Node 26 and in CI. No test
behavior or timeout was relaxed.
- Full local typecheck and build passed. Latest-head CI is green;
Greptile is 5/5 with no unresolved review threads.
- Two real Slack baseline replies took 25.1 and 24.6 seconds
(24.9-second mean). Three same-thread signed probes on this head took
23.8, 23.9, and 22.7 seconds (23.5-second mean). This is a small sample
and a modest wall-clock improvement, not a large or statistically
established speedup.
- In that same thread, uncached provider input fell from 8,514 tokens
before compact input to 694–765 tokens afterward. The current delivery
uses a 362-character delta; the full 19–21k-character bootstrap remains
available for failed resume. Verified runner artifact preparation fell
from about 1.2 seconds to 0.6 seconds.
- A fresh thread created a separate task, sandbox, and provider session
with full bootstrap (24.4 seconds). Its follow-up reused its own
sandbox/session and compact input (28.3 seconds, including 16 seconds of
model execution). Model variability and process startup remain
substantial.
- A signed duplicate webhook produced exactly one user comment, one
successful run, and one final Slack reply. Slack's API independently
confirmed the actual replies and a public task URL without an internal
or pool hostname.
- Earlier signed probes verified recovery after a failed run and reuse
across a server deployment. The final idle test observed Daytona report
the sandbox as stopped, then delivered a new reply in 19.9 seconds using
the same sandbox/provider-session identity and compact input. Slack’s
API confirmed that reply.
- Live probes use signed synthetic inbound webhooks and real outbound
Slack delivery, read back through Slack’s API. The final browser recheck
found the Mac locked and the Slack tab blocked by another extension, so
this is not claimed as full UI E2E proof.
- This is a review branch. Do not merge until the maintainer reviews it.

## Risks

- Incorrect session reuse could mix account or task context. Missing or
changed identities continue to reset, and existing authorization checks
remain in place.
- Warm mode remains opt-in. Remote warm mode requires a reusable sandbox
lease. Retained processes keep credentials until they close, so idle
expiry and ownership fences are required.
- A fresh user message may continue after a committed provider failure.
Approval, current authorization, process termination, and prior-result
checks remain required.
- Projectless sandbox reuse is task- and agent-scoped. Missing or
mismatched ownership cannot replace an existing lease; existing
workspace-scoped leases keep their scope. Opt-in reuse retains a sandbox
per task/agent, so provider auto-stop and deletion policies still
determine idle compute and storage costs. Fleet defaults are unchanged.
- Compact prompts apply only after proven resume and a matching
prior-run delta. Missing or specialized context falls back to full
input; fresh sessions always receive the full bootstrap.
- No schema or migration changes.

## Model Used

OpenAI GPT-6 through Codex, with code editing, tool use, and test
execution. The exact serving model ID and context-window size are not
exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-20 13:12:44 -05:00
DottaandPaperclip 7bc03e0acd feat(runner): default harnesses to full auto and support task reassignment (#13686)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agent Chat uses native runners to save plans and coordinate tasks.
> - Provider defaults differed across harnesses and could stop
unattended work at a second permission gate.
> - Agents also lacked a dedicated tool to move existing work to another
agent safely.
> - This change defaults native providers to full automatic permission
for provider tools and connected tools.
> - A guarded reassignment tool preserves task identity, stops the
previous run, and schedules the new owner once.
> - Codex and Claude chat acceptance tests now use production permission
defaults.

## Linked Issues or Issue Description

**Subsystem affected**

Native runner, ACPX Claude permission policy, task authority, and Agent
Chat acceptance tests.

**Problem or motivation**

A user can authorize an agent to save a plan or create a task, but
Claude's default provider gate can still stop that action. Reassignment
needs a dedicated operation that preserves context and avoids concurrent
owners or unintended recovery runs.

**Proposed solution**

Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to
`never`. Apply the defaults at configuration, execution, fresh-session,
resume, driver, and proxy boundaries. Keep explicit permission settings
and server-side company, claim, task-mode, and approval checks. Add
`reassign_task` with version checks, durable idempotency, audited
cancellation, and guarded successor scheduling.

**Alternatives considered**

A Paperclip-only allowlist still blocks provider tools and other
connections during unattended work. Full automatic permission is the
requested product default. Recreating a task discards its identity and
history. Updating assignment without stopping the previous run can leave
two agents working on the same task.

**Roadmap alignment**

This extends the existing planning, delegated work, governed tool
access, and recovery features. It adds no new service or schema
migration. Recent related tasks and open PRs were checked for duplicate
work.

**Additional context**

Related: #13678 (Agent Chat tools and recovery), #13677 (remote runner
startup). The stacked legacy-adapter companion is #13693. This also
fixes the deployed-server artifact fallback needed to stage the current
runner binary.

## What Changed

- Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex
to `never`, including missing settings at direct driver and proxy entry
points. These defaults cover provider tools and connected tools.
Preserve explicitly configured restrictive modes.
- Include assigned approval reads using canonical side-effect
classifications, so verifying a recorded approval does not trigger
another provider gate. Paperclip approval decisions still enforce
controller authority.
- Carry the new permission mode through server configuration, execution
contracts, recovery identity, TypeScript, and Rust. Keep
`approve-paperclip` as an optional restricted mode, with exact SDK rules
and closed unknown requests. It is not a default.
- Add `reassign_task` to the semantic catalog, controller, mock
authority, and generated contracts.
- Guard reassignment with company authorization, expected owner and
version, protected-state checks, and durable retry receipts.
- Honor explicit backlog task creation atomically with the initial plan,
without scheduling a wake. Preserve backlog holds regardless of
dependency readiness.
- Stop active work before changing ownership. Restore the prior owner
through a guarded, idempotent wake if final handoff validation fails.
Keep intentional reassignment stops out of failure recovery. Preserve
backlog and blocked states without waking them early.
- Add authorization, concurrency, replay, stop, and permission boundary
regressions. Add Codex and Claude chat reassignment cases and run native
chat cases with production defaults.
- Clarify shared runner guidance: save plans and Paperclip documents
directly with `write_document`; create and register a local file only
when a downloadable file is requested.
- Document provider defaults and the operator choices for existing
agents.

## Verification

- Current head `d82fbb0f03546d27cecf072250e4172e0b1ee662`: **55 checks
passed**, with two intentional skips. [PR
checks](https://github.com/paperclipai/paperclip/pull/13686/checks).
- Greptile reviewed that exact head at **5/5**. The security reviewer
acknowledged the intended full-auto default, and the acknowledged
discussions are resolved.
- Full workspace `pnpm -r typecheck` and `pnpm build` passed locally
after rebasing onto current master. Targeted adapter/server, runner,
API, default/resume, and heartbeat configuration tests passed.
- **All six real-provider acceptance cases passed on their first
attempt, with cleanup passing:** plan handoff, task reassignment, and
backlog creation/status, each on native Claude and Codex. Evidence
records Claude's effective `approve-all` mode. [Campaign and
downloadable
evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548).
- The live campaign tested combined revision
`a37881c824dcd7170380fc4b788732fc743e5da7`. The final PR heads add only
a heartbeat test expectation correction; application code is unchanged
from that live-tested revision.
- The campaign's result-enforcement job passed. Its separate report
publisher failed because the trusted workflow's `patchedDependencies`
configuration differs from its frozen lockfile. All six results and
screenshots remain available as GitHub artifacts. The overall manual
workflow is red for this publishing failure.
- Full-suite coverage is supplied by the passing CI partitions. The
separate unsharded local run was stopped after the corresponding CI
partitions passed; it is not counted as a completed local run.
- Reassignment tests cover stale state, cross-company access, denied
authority, cancellation failure, compensating wake, and idempotent
retries. Backlog tests verify the original creation audit, saved plan,
exact task count, and absence of task-bound runs.

## Risks

- Agents with no explicit permission mode now receive full provider tool
permission, including connected tools. This is a deliberate broad
default. Existing explicit restrictive modes still apply. Controller
authorization, company isolation, workspace boundaries, and Paperclip
governance remain in force.
- Reassignment crosses run cancellation and task ownership transactions.
Durable stop intent, revalidation, audit receipts, and guarded queue
dispatch cover interruptions and retries.
- The new permission enum requires a current runner artifact. The remote
artifact fallback uses the same resolved controller binary for upload
and execution.
- Live provider behavior remains subject to the selected model. Targeted
live results do not qualify the full catalog.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The exact deployment model ID and context-window size are not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 16:49:18 -05:00
DottaandPaperclip 43acbcc398 fix(runner): preserve sessions and complete question and approval continuations (#13655)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner connects task state to provider sessions.
> - Follow-up turns must retain provider memory and carry new user
direction.
> - Lost session IDs caused repeated context and extra input tokens.
> - Native question answers and approval races could leave valid work
blocked.
> - This pull request repairs those paths and adds regression coverage.
> - Agents can continue accepted work without repeating the conversation
or losing the user's answer.

## Linked Issues or Issue Description

Refs #13574. That merged PR shortened continuation prompts and moved
question instructions into tool documentation. This change preserves
sessions and fixes failures exposed by broader testing. Related runtime
work: #13408 and #13410.

**What happened?**

Native follow-up turns could lose the provider session ID. Completion
guidance could replace the original task with its latest comment. Claude
native questions could remain pending after the user answered. Approval
during a running tool call could suspend the run before the tool
response arrived. Onboarding and chat handoff instructions also caused
repeated planning or missing plan documents.

**Expected behavior**

Reuse a valid provider session. Send only new events when that session
already has the history. Preserve the task requirements and apply later
user direction. Store the question answer and deliver it to the waiting
run. Finish governed tool responses before suspending. Execute the
accepted plan without asking for the same approval again.

**Steps to reproduce**

Run the continuation, local-session-integrity, first-task, and
agent-chat suites with native Codex and Claude. Include
provider-question-bridge, accept-while-running, and plan-handoff.

**Paperclip version or commit**

This branch is based on master d54b75011. The active full catalog run
tests 4e75881db. Later review fixes have separate regression coverage.

**Deployment mode**

Isolated local instances and Daytona sandboxes in the existing Runner
full-stack E2E harness.

## What Changed

- Retain provider session identity across turns and late usage
snapshots. Send new continuation events on session reuse, with full
context available for a fresh session.
- Preserve task requirements and later direction in completion guidance.
Return the current contract revision after a stale completion
submission.
- Bridge native Claude questions to saved Paperclip cards. Submit
answers through the saved card and resume the same run.
- Delay governed suspension until tool results settle. Add a
deterministic test barrier for approval during an active run.
- Clarify free-text question examples, explicit onboarding plans, and
execution of accepted chat plans.
- Fix continuation readiness, verified output evidence, and declared
screenshot collection.
- Qualify the legacy Claude test CLI at 2.1.277. The old 2.1.19 CLI did
not discover mounted skills. Update the existing workflow pin and
isolated launcher together.
- Refresh the Daytona image lockfile integrity pin after reviewing
master patch updates.
- Carry continuation mode as runtime metadata instead of inferring it
from user-visible text. Install the test Claude CLI without lifecycle
scripts.

## Verification

- Targeted paid verification: 20/20 cases passed across
local-environment campaigns before the rebase.
[Report](https://pages.paperclip.ing/runner-e2e-seven-fixes-35397904249/).
- Harness checks: 379 unit tests passed; harness typecheck passed.
- Latest-head PR checks: 55 passed, two intentionally skipped. Greptile
is 5/5; the security scan passes.
- Review regressions: 350 executor tests and 204 session/driver tests
passed. A script-free Claude install was verified with the actual CLI.
- Full catalog, including the explicit-only everyday suite: [run
35417932353](https://github.com/paperclipai/paperclip/actions/runs/35417932353).
Completed: **164/205 passed; 41 failed**. [Full dashboard and failure
investigation](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/).
Includes 204 case artifacts and one pre-case GitHub authorization
timeout; missing evidence is not scored as a pass. The full run tested
`4e75881db`; Final-head metadata/CLI smoke cases both passed. In the
separate [six infrastructure
retries](https://github.com/paperclipai/paperclip/actions/runs/35419769343),
the GitHub timeout case passed and all five Docker preflight failures
repeated. [Follow-up
dashboard](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/follow-up/).
- Full local typecheck and build passed on the rebased branch. The full
local unit run completed with 657 passing files, two test timeouts and
one suite setup timeout. All three affected files passed when rerun in
isolation (84 tests). The first full local run was not clean.
- Focused regression coverage includes the live question bridge,
same-run response delivery, UI routing, stale revisions, approval
overlap, and session reuse.

## Full-catalog follow-ups

- Test infrastructure: 14 Claude everyday cells probe an absent host
CLI; six cells failed pre-task GitHub/Docker qualification (GitHub
passes on retry; all five Docker cases repeat; the workflow preflight
allowlist omits their case IDs); five ACPX Codex cells cannot create
sandbox namespaces.
- Runtime: four OpenCode completion-criteria mismatches masked by
shutdown errors, one service-approval suspension failure; three Daytona
recovery failures encounter existing skill files; one duplicate
completion wake.
- Confirmed test defects: question pagination and a noncanonical plan
document key.
- Product/behavior: mismatched visible/required question sets, an
attachment instead of the requested task document, one lone-option
onboarding question, early completion instead of review, and a Codex
Mini completion-schema failure.
- The report job itself fails on trusted master’s stale patch/lock
configuration. The linked report is rebuilt with the shared renderer
from original cell results and public fixture screenshots; it excludes
private snapshots, logs and traces.

These are investigated follow-ups, not silently regraded passes.
First-task passed 51/52. The PR checks are green independently of the
broader catalog’s behavioral/infrastructure failures.

## Risks

- Session reuse depends on a valid provider identity and context
coverage. Fresh-session fallback and reset tests cover this boundary.
- Native question delivery spans saved interaction state and a live
provider run. Tests cover duplicate events, closed runs, and same-run
answers.
- Provider behavior varies. The full paid catalog may expose failures
beyond these targeted fixes; those results will be reported without
relaxing valid approval or output checks.
- The legacy Claude version update is limited to test infrastructure. No
database migration is included.

## Model Used

OpenAI Codex, GPT-6 family, with repository inspection, code execution,
and browser/E2E tools. The exact runtime model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted checks and all
three timeout-file reruns pass; full-run timeout caveat above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 07:42:57 -05:00
DottaandPaperclip e26d787928 Shorten continuation prompts and verify question tool guidance (#13574)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents must continue tasks using user answers without losing earlier
requirements or approval gates.
> - The wake prompt mixed human decisions with prior tool evidence and
repeated detailed question instructions.
> - Those instructions belong with the question tool, with a short
routing hint in the wake.
> - The Runner evals need to prove that answers, approvals, and
completed work survive later turns.
> - This PR shortens the prompts, separates authenticated answers, and
adds continuation tests with useful screenshots.

## Linked Issues or Issue Description

Refs #13517. This is a follow-up to the merged onboarding skill and
Runner E2E work. Related #13539 covers responses received while a run is
active; this PR preserves its cases and adds continuation coverage.
Existing continuation/recovery and question PRs were searched; none
covers this prompt/documentation and eval change.

**What existing behavior does this improve?**

The instructions sent when an agent continues a task, the native
human-input tool documentation, and the evidence captured by Runner
full-stack E2E.

**Current behavior**

The wake repeats a long question-tool guide. Human answers appear
alongside untrusted prior results. Screenshot capture can finish at DOM
load while the task still shows a spinner, even when backend behavior
checks pass.

**Proposed behavior**

Keep earlier requirements unless the user changes them. Treat
clarification as distinct from approval. Give authenticated human
responses a scoped field. Keep tool and agent results as evidence. Put
detailed question behavior in the tool descriptor and retain one routing
sentence in the native wake. Wait for the correct task and loaded
conversation before taking screenshots.

**Reason and benefit**

Reduce repeated prompt text and make authority boundaries clear. Test
that real question cards, later answers, approval gates, and completed
child tasks still work. Make screenshots useful for human review.

## What Changed

- Shorten shared continuation instructions for legacy and native
runners. Separate authenticated user responses from tool results and
agent summaries.
- Remove the detailed question guide from native wake prompts. Keep its
behavior in the canonical `request_human_input` descriptor and existing
payload schema. Regenerate semantic contracts and fixture hashes.
- Add five continuation cases across four local profiles. Add a
dedicated choice-then-text case for native Codex and native Claude. All
22 cells join the shared full E2E campaign.
- Cover revised scope, clarification without approval, hostile
instructions in a handoff file, and reuse of a completed child after
restart. Keep production instructions and fixed user facts.
- Capture continuation screenshots only when the intended task and
conversation have rendered. Add provider-free browser regressions for
loaders and wrong-task capture.
- Preserve current master’s extra tool and onboarding cases. The default
campaign now contains 166 cells; 35 manual everyday cells remain
separate.

## Verification

- `pnpm -r typecheck`: passed after replay on current master.
- `pnpm test:e2e:runner:unit`: 340 passed. Harness typecheck passed.
- `PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm
test:e2e:runner:browser-support`: 4 passed. These tests failed against
immediate screenshot capture and passed after the fix.
- Focused continuation and native-input tests: 36 passed locally. The
tool-authority suite could not initialize embedded PostgreSQL locally,
including one isolated retry; its 17 assertions did not run locally. The
full remote server shards passed on this PR commit.
- `pnpm build`: passed after replay on current master. `pnpm test:run`
was attempted locally but hit the same embedded PostgreSQL
initialization failure; the remaining local run was stopped after
complete remote CI passed. This is not claimed as a full local test
pass.
- [Full PR
CI](https://github.com/paperclipai/paperclip/actions/runs/35232755685):
passed on `6a22128c14f4552d0613a6d9a25955db4a1ed02f`. All
server/chat/workspace/serialized shards, browser shards, Runner checks,
typecheck, build, canary and policy checks passed. The isolated native
Runner build and security checks also passed: 57 successful checks, with
two expected Storybook skips.
- Greptile reviewed the exact PR head at 5/5, with no findings or
unresolved review threads. The PR has no merge conflicts.
- [Live question-docs
report](https://pages.paperclip.ing/runner-e2e-question-docs-35227647794/):
3/3 passed at source `83dd132f2` before replay on master. Native Codex
and Claude each asked a choice, waited, asked a text question, and saved
both answers. Claude also passed a completed-child restart case. All
three native turns are checked for absence of the old question block.
- [Earlier continuation
report](https://pages.paperclip.ing/runner-e2e-continuation-35154943615/):
all five continuation cases passed on native Claude. The report retains
campaign and revision provenance and separately shows two unresolved
onboarding behavior failures.
- [Before/after prompt
report](https://pages.paperclip.ing/runner-prompt-comparison-20260917/):
full text, current recorded Claude inputs, and reproducible
reference-token counts. The controlled wake comparison removes 401
reference tokens; the net counted input reduction is 339 after charging
the larger tool description. These are text-size estimates, not measured
billing savings.

## Risks

- Prompt wording affects model behavior. Live results cover the stated
cases, not every provider or conversation. Legacy profiles are
registered but were not rerun for this change.
- The optional continuation field changes prompt data only; there is no
database migration or new production API.
- Authenticated answer projection excludes generated summaries and
agent-resolved interactions. It preserves the answer’s question or
approval scope.
- The screenshot guard can expose UI loading failures that earlier runs
hid. Backend grading alone no longer makes those captures valid.
- The two prior onboarding failures remain separate product issues: work
before acceptance and a missing saved plan. This PR does not claim the
entire onboarding suite passes.

## Model Used

OpenAI Codex, GPT-6, with reasoning, repository tools, code execution,
and browser verification. The exact deployed model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — targeted tests above; the
full local database-startup limit is documented
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-17 09:32:54 -05:00
DottaandPaperclip 11921075a4 Add first-task onboarding skill and Runner E2E coverage (#13517)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The first task helps a new user define and approve useful work.
> - That workflow needs reusable instructions and tests against the
production experience.
> - Native Codex and Claude must load the assigned skill, including
after resume.
> - Maintainers need recorded conversations and precise failed checks to
judge regressions.
> - This pull request adds the first-task skill and a suite in the
shared Runner E2E harness.
> - It keeps behavior results separate from informational quality scores
and incomplete recordings.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The first onboarding task and the Runner E2E report used to review it.

**Current behavior**

Onboarding embeds its policy in a hidden brief. Native Codex drops the
skill-instructions setting at the Rust boundary. The shared E2E harness
has no onboarding suite or full conversation view.

**Proposed behavior**

Assign and invoke `/first-task` for the onboarding task. Send selected
Codex skills as structured protocol inputs. Run twelve scenarios across
legacy Codex, legacy Claude, native Codex, and native ACPX Claude.
Include all 48 cells in full campaigns. Show recorded chat, question and
approval cards, exact checks, instructions, and billing in the shared
dashboard.

**Reason and benefit**

Measure the real onboarding experience before changing prompts.
Distinguish infrastructure failures, behavior failures, and unexercised
journey steps.

**Breaking changes**

No database migration or production API change. First-task instructions
now live in an assigned skill. The user-edited persona is preserved; the
skill includes the maintainer-approved proposal-mode mapping and
saved-plan requirement.

Related: #11043 is earlier onboarding work. #13422 already fixes native
Claude model pinning, context delivery, and read permissions on master;
this branch includes those fixes through its base. The new Claude
recovery test supplements them.

## What Changed

- Extract and assign the first-task skill while retaining the production
greeting and opening question.
- Carry the Codex skill-instructions flag through thread start and
resume. Resolve explicit task skill references only against assigned
skills and send native skill inputs.
- Invoke an unambiguously selected assigned skill through Claude ACPX’s
native slash-command parser on initial and resumed turns, retaining the
entire task/wake envelope as its argument. Do not carry that invocation
into ordinary tasks.
- Restore the saved single-task proposal modes: confirmation card, or
saved plan with revision-targeted checkbox approval. Explicit plan
requests also require a saved plan.
- Add first-response and complete-journey cases with fixed user facts,
acceptance checkpoints, durable outcome checks, and accounting for child
runs.
- Fail the eval when choice questions have fewer than two real options.
Recognize planning documents without treating them as completed work.
- Add optional, bounded quality judging as explicit post-processing.
- Render full conversations and static interaction cards in the shared
report. Conversations start folded. Show original and regraded results
and incomplete journeys distinctly.
- Keep credential-persistence scanning outside the first-task behavioral
suite; retain public evidence redaction.
- Refresh generated capability references after the API-reference edits.
- Correct shared native question guidance and tool schemas: choices need
at least two meaningful options; open-ended questions use canonical text
fields with the required compatibility payload. Verify both formats
through real tool-authority persistence.
- Disable announcements automatically for every isolated Runner E2E
process and label the gallery environment/provider/target explicitly.
- Remove CI races in the GitHub connection browser test and native
session recovery test by waiting for the actual async work before
asserting its results.

## Verification

- `pnpm exec vitest run
server/src/services/onboarding-first-task-assets.test.ts
server/src/__tests__/issue-onboarding-first-task-routes.test.ts`: 19
passed.
- `pnpm --dir packages/paperclip-runner exec vitest run
src/drivers/acpx/runtime-host.test.ts
src/drivers/acpx/native-skill-prompt.test.ts
src/cli/acpx-runtime-sidecar.test.ts`: 70 passed. Native command
forwarding and the 1 MiB input boundary both failed before their fixes
and passed afterward. Coverage includes changed skills on reopen,
approval context, and an ordinary subsequent task.
- Runner E2E unit suite: 306 passed. Harness typecheck passed. The 64
first-task fixture and grader tests also pass.
- Full repository typecheck and build passed locally. Server typecheck
and Runner build passed again after the native-command change.
- Full GitHub Actions CI passed on `23e56447b`: all
server/workspace/browser shards, Runner verification, typecheck/release
registry, build, canary, policy, and Docker checks. Greptile reviewed
this exact head at 5/5 with no unresolved threads. The earlier broad
local run had database startup/timing failures that passed isolated
retries; the complete remote suite is green.
- Merge verification against current master: 312 harness tests and 13
native recovery tests passed. Regenerated semantic contracts and fixture
hashes pass their consistency check. Full local typecheck and build also
passed on the stacked queue branch. After merging the latest master and
preserving the GitHub setup timing regression in the split browser
suite, both focused GitHub browser tests passed. Three CI timing/startup
flakes passed local verification and one remote retry; all latest-head
checks are green.
- Real pinned Claude SDK and Claude ACP JSON-RPC probes against a local
mock API confirmed that `/skill-name` expands the assigned skill body
before the model request and retains the task arguments. A prose mention
does not. The probes made no paid model calls. The ACP probe used the
current first-task skill body and retained the wake arguments.
- [Full 48-case campaign and
report](https://pages.paperclip.ing/runner-e2e-first-task-35053063880/):
44 passed after three interrupted Codex cases completed in targeted
reruns. Original results, regrades, and all 51 executions remain in the
report provenance.
- [Claude campaign after the shared-question
fix](https://pages.paperclip.ing/runner-e2e-first-task-claude-35099525201/):
10/12 passed with zero single-option failures. All 12 recorded the
current assigned skill and corrected guidance. The failures exposed
skipped skill invocation and a missing saved plan. This PR adds native
command invocation and explicit saved-plan instructions; the subsequent
report below still shows behavior failures.
- [Fresh 12-case Claude
report](https://pages.paperclip.ing/runner-e2e-first-task-claude-35102737804/)
at `78452129e`: 10/12 pass after correcting two false proposal-matcher
failures. The recordings said “Here is the task I will create and
run/complete” in approval cards; the old matcher missed that word order.
Regression tests failed before the fix and pass after it. Original
results and offline regrade provenance remain linked. No agent rerun was
needed. Zero single-option-question failures; two behavior failures
remain: direct work before acceptance on a plain first message, and an
explicit plan request without a saved plan. Neither check was relaxed.
The follow-up `82087ac7e` fixes command-prefix size accounting;
`94aefb1f3` fixes only that proposal matcher.
- Report browser checks confirm folded conversations, rendered cards,
explicit Local/Daytona labels, and no page errors. The published-object
audit scanned 1,306 text files across 2,154 objects with no
credential-format findings or prohibited files. Image pixels and unknown
token formats are outside that scan.

## Risks

- Model behavior is nondeterministic. One campaign is evidence, not a
guarantee. The two remaining Claude behavior failures are visible in the
report and require further product work; this PR does not claim all
onboarding scenarios pass.
- The suite checks persisted Paperclip effects. It cannot prove the
absence of arbitrary external effects.
- Historical recordings can miss later journey steps. These remain
incomplete, never passes.
- Native profiles switch runtime after the production onboarding wizard
because it does not yet expose a native option.
- Quality scores are informational and cannot override behavioral
failures.

## Model Used

OpenAI Codex, GPT-6, with reasoning, repository tools, and code
execution. The exact deployed model identifier and context-window size
are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 14:24:32 -05:00
DottaandPaperclip 422287eecd fix: preserve runner recovery, warm sessions, and task outcomes (#13338)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner connects task messages, provider execution, and
task outcomes.
> - First-time user tests exposed gaps in recovery, completion
permissions, message delivery, and Stop behavior.
> - These gaps left usable output hidden, completed work waiting for
bookkeeping, or safe work unable to continue.
> - This pull request fixes the shared lifecycle and receipt paths while
preserving process ownership and action checks.
> - Users can continue work with accurate task state and durable
messages.

## Linked Issues or Issue Description

**What happened?**

A stopped local Codex execution could remain blocked even after its
processes had stopped and its complete transcript proved that no
external action needed replay. Claude under Conservative permissions
could fail to call task completion tools. Recovery could reuse an
assistant item ID and overwrite prior output. A delivered comment could
remain marked uncertain after navigation. Stop could look like Pause or
a new recovery incident. Workspace contention could look like
cancellation. A direct reply reopening Done could enter a clarification
loop.

**Expected behavior**

Recover automatically only with verified termination and complete action
receipts. Preserve answers and messages. Keep task completion available
under Conservative permissions without broad tool access. Show crashes
as Blocked, actual human decisions as In Review, and ordinary workspace
contention as waiting. Stop the current response and allow a new
direction.

**Steps to reproduce**

1. Create ordinary response tasks with local Codex and Claude Code, then
send follow-up messages through the task composer.
2. Interrupt a disposable local Codex runner during text-only work.
Verify automatic continuation and retained output.
3. Stop a response, send a new request, answer a clarification, and
reopen completed work with another message.
4. Navigate or reload while a comment submission is pending. Confirm the
exact persisted request receipt settles it without removing newer draft
text.
5. Run two tasks in a shared Daytona workspace. Confirm waiting does not
appear as failure.

**Paperclip version or commit**

Initial acceptance baseline: `c9021c6721f91e2c74bd9fee9d3fd41c999d17b7`.
Current integration base: `6cef9743c`. Both operator-interruption and
workspace-waiting guards are preserved; native restart and legacy
permission rules remain documented.

**Deployment mode**

An isolated source-built test-drive instance, with real local Codex and
Claude Code providers and disposable Daytona environments.

Related work: #13314, #13316, #13327, #13344, #13239, #13254, #13163.
This PR addresses additional failures from ordinary task journeys,
including controller restart handoff and repeated warm sandbox setup.
Historical task status reconciliation is excluded.

## What Changed

- Persist runner ownership immediately at spawn and resume an explicitly
adopted runner even when the controller crashed before the first driver
checkpoint. Detach the controller safely across graceful restarts,
including session startup. Prevent an old finalizer from suspending or
signaling an adopted runner. Checkpoint idle warm sessions before
shutdown. Preserve the same run and queued follow-up messages.
- Scope saved legacy queue successor checks to the queue owner while
preserving ordinary task locks, operator identity, assignment gates, and
exactly-once delivery.
- Preserve managed Codex credential files when an old session is
detached for restart; normal owned cleanup still copies refreshed auth
back and removes the scoped copy.
- Reuse the bound warm shared sandbox and fully verify an existing
staged provider pack before using it. This avoids repeated uploads when
the pack is already valid.
- Add a narrow local Codex replacement path with stopped-process proof,
a closed transcript inventory, exact completion receipts, and
fresh-session lineage. Preserve no-replay holds when evidence is
incomplete. Recovery may clear only the same run's recorded Blocked
status version; manual re-blocking and dependency changes invalidate
that receipt, while queued comments do not. Later blocks stop scheduled,
queued, and final dispatch; queued/final checks re-read dependencies
even when the task status stays In Progress.
- Permit only task delivery and human-input tools through the isolated
Claude runner's exact task bridge.
- Scope assistant item identity to the provider turn and ignore only
authority-free Codex skill-change notifications during startup.
- Reconcile composer submissions by client request ID across response
loss, navigation, and reload. Retain text typed during delivery.
- Keep acknowledged run-only Stop neutral and show workspace contention
as waiting. Project exhausted native failures as Blocked.
- Restore the guarded task-page retry action for failed legacy runs,
including the server-supported explicit new-attempt path for stopped
conversation adapters. Preserve native/process recovery holds and avoid
promising Retry while a decision or execution gate hides it.
- Refresh delivered artifacts and handle direct user replies that reopen
completed work without a clarification loop.
- Check the embedded PostgreSQL PID, data directory, and actual port
before connecting or migrating.
- Document accepted behavior and add focused regressions at lifecycle,
route, transcript, and UI boundaries.

## Verification

- Final head `fece606ac2` passes the complete GitHub CI matrix: **34
green checks, two expected Storybook skips, no failures or pending
checks**, including `ci / verify`, `ci / e2e`, full runner verification,
typecheck, build, every server/workspace shard, and all browser shards.
[CI
run](https://github.com/paperclipai/paperclip/actions/runs/34727183287).
Greptile is **5/5 with no open findings**. The final two commits only
refine test fixtures; both affected suites pass 24/24 locally and in CI,
with server typecheck green.
- Complete local Vitest coverage uses the canonical groups/shards: all
635 general server suites, all 145 serialized suites, and all workspace
packages. The aggregate began on `0a8001c18` while the final queue fix
arrived: 23,903 passed, five failed, 87 skipped. The five
port/socket/timing failures passed unchanged in follow-ups (60 tests in
the exposure/file suites and 412 tests covering the serialized failures
and unrun tails). The final queue/operator-identity suites separately
passed 52/52. This is aggregate coverage plus explicit reruns, not a
pristine single-command final-head run.
- After integration with current master,
queue/operator-identity/continuation suites passed 162/162 and affected
UI suites passed 140/140. ACP Stop/continuation and legacy
task/Inbox/message browser suites passed 9/9, including both task
recovery Retry and thread Try again, automatic saved-message delivery,
exactly one new run, Done, and retained output after reload. The default
process Stop/Pause/Resume browser case passed (the native-provider case
is opt-in and skipped by default). The complete Board attachment/receipt
browser suite passed 11/11 on a disposable instance, covering both
composers, exact receipts after lost responses, no replay, bound
attachments, and newer drafts after reload.
- Blocking-intent regressions cover pre-existing Blocked, a mismatched
run/cause, an explicit manual re-block, changed dependencies, a queued
comment after failure, and a block arriving between scheduling and
provider dispatch. The negative cases reproduced before the fix. All 478
affected executor/recovery/dispatch tests passed; both database suites
ran separately after availability-probe skips in the first combined
command. The final late-dependency check passed all 143 affected
recovery/dispatch tests (zero skips) after two new negative cases
reproduced the bug.
- Focused runtime regressions cover awaited runner ownership
publication, authenticated adoption before the first checkpoint,
old-finalizer detachment, idle and busy warm-session shutdown, rejected
checkpoint propagation, provider-pack verification, and managed-Codex
credential preservation. Four managed credential detachment cases
reproduced the bug before the fix; normal owned cleanup still succeeds
exactly once.
- Live local Claude: SIGKILL 2.6 seconds into startup recovered the same
run automatically in 53 seconds, then a normal follow-up completed in 24
seconds. SIGTERM 2.5 seconds into startup preserved the same run (54
seconds) and its queued follow-up (21 seconds). Answers remained visible
and the task reached Done.
- Live Claude Daytona: a warm follow-up retained its sandbox and fell
from 121 seconds to 44 seconds. A separate cold turn took 127 seconds;
after controller shutdown and checkpointing, its follow-up completed in
33 seconds with the same sandbox, workspace, native session, and runner.
Both answers remained visible and the task was Done.
- Other live journeys covered task completion and follow-up with local
and Daytona Codex, local Codex crash recovery, Stop then new direction,
clarification response, live artifact refresh, and shared-workspace
waiting.
- Validation limits: the opt-in native composer Stop/Pause→subtree
Resume fixture exposes terminal/result ordering and subtree-cancellation
attribution bugs that can leave a child task blocked; that new finding
is assigned to a separate follow-up and is not claimed fixed here.
Default CI skips this optional native-provider fixture. Managed-Codex
credential handoff and the queue-agent integration use automated
regression evidence. Cold custom provider-pack uploads still add startup
latency.

## Risks

- Automatic replacement remains deliberately narrow: local Codex,
verified stopped identities, unchanged retained state, and a complete
text/completion-only turn. Unknown actions, partial history, or changed
ownership remain blocked.
- Claude completion permission handling changes an upstream package
patch. The exact isolated task bridge must remain pinned; unrelated
tools keep their existing permissions.
- New task failure projection changes user-visible status. No historical
status backfill or database migration is included.
- This is a broad lifecycle fix across server and UI. Live proof covers
graceful local Claude restart during startup and idle Claude Daytona
session recovery across controller shutdown. Live abrupt SIGKILL during
local Claude startup also recovered the same run. Unknown ownership or
missing action evidence still blocks reuse. Cold custom provider-pack
uploads still add startup latency; this change avoids unnecessary repeat
uploads.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, code execution, browser
automation, and tool use. The exact hosted model ID and context window
are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 19:41:15 -05:00
DottaandPaperclip ab15aff390 feat: add experimental persistent agent chat (#13284)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Conversations must use the same tasks, controls, and execution
history.
> - Users need an ongoing chat with an agent without managing task
properties.
> - Agents should clarify and plan work, then hand execution to assigned
project tasks.
> - This pull request combines the reviewed Agent Chat stack for one
squash merge.
> - The benefit is persistent conversation with normal task governance
and shared UI.

## Linked Issues or Issue Description

**Subsystem affected**

Task lifecycle, agent runtime tools, shared task UI, and browser/paid
runner tests.

**Problem or motivation**

Users need one persistent conversation with each agent. A separate chat
store or renderer would duplicate task behavior and bypass existing
controls.

**Proposed solution**

Use a task-backed chat per company, user, and agent. Reuse the task
composer and transcript. Clarify and plan in chat, then create assigned
project tasks with the relevant plan. Keep Agent Chat behind its own
disabled-by-default experimental setting.

**Roadmap alignment**

This implements the task-backed direction in [CEO
Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat).
Related proposals: #2504 and #9693. Related request: #7981. The
maintainer requested one squash merge of the complete stack.

Consolidates the reviewed runtime
[#13281](https://github.com/paperclipai/paperclip/pull/13281), backend
[#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI
[#13283](https://github.com/paperclipai/paperclip/pull/13283) layers
with this PR's E2E coverage. All four layers passed CI and received
Greptile 5/5 before consolidation. This PR targets master and includes
the complete feature.

## What Changed

- Add personal canonical chat tasks with ordinary company visibility,
immutable identity, idempotent first sends, and an idle waiting state.
- Process `/new` in queue order. Preserve history, release a chat pause,
and fence old provider context and delayed writes.
- Keep chat lifecycle rules across recovery, finalization, assignment,
task lists, and rollups.
- Support research and plan revision in chat. Hand plans to ordinary
assigned project tasks before execution starts. Reject new chat
subtasks.
- Add repository-aware project creation and discovery tools, including
multiple repository IDs and GitHub URLs, authorization, idempotency, and
durable project-created cards.
- Reuse task UI components for chat, with starred/recent agent
navigation and a separate `enableAgentChat` experimental flag.
- Add deterministic browser tests and 24 paid chat cells across four
Codex/Claude profiles, with validated reports and screenshots.
- Integrate current master recovery, controller lease, queued-message,
and task UI changes. Gate chat interruption and deferred promotion on
ownership/feature policy. Guarantee lease renewal and active controls
are stopped even if teardown fails.
- Preserve master's migration 0273 and generate chat migration 0274 with
idempotent replay for development databases.

## Verification

- Prior exact heads of all four PRs passed Linux CI, including build,
typecheck, general/serialized tests, and browser E2E. Each had Greptile
5/5 and no unresolved findings.
- Integrated local verification passed: full repository typecheck and
production build, Storybook build, token gates, 340 focused UI tests,
all 20 deterministic chat browser tests, two migration replay tests, 88
focused chat/queue/native/controller tests, and provider/session
regressions including real lease expiry. These include the three
lifecycle regressions for the final admission/teardown fixes; server
typecheck also passes. Current head
`1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no
unresolved findings and passing security scans. All final-head CI gates
passed: build, full Runner verification, typecheck/release registry,
canary, all general/serialized test shards, and all browser E2E shards
([CI
run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)).
Local PostgreSQL startup contention required serialized retries; skipped
fixtures do not count as passing coverage.
- The earlier paid campaign passed all 24 chat cells and retained 32
screenshots:
[report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat).
It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior
evidence, not a paid run of this integrated head.
- Manual check: enable Agent Chat in Experimental settings, open an
agent, clarify and revise a plan, then hand off to an assigned project
task. Stop a reply, send `/new`, and verify fresh context with retained
history. Disable the setting and verify agent shortcuts/new chat turns
are blocked.

## Risks

- Queue/session integration can affect retries and delayed writes. Tests
cover ownership, cancellation, reset boundaries, idle recovery, and
ordinary task behavior.
- Migration 0274 adds conversation fields and constraints. Replay is
idempotent and preserves existing development chat history.
- This combines the previously reviewed stack at the maintainer's
request. Agent Chat remains off by default and is separate from
Conference Room.

## Model Used

OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code
execution, browser tools, and parallel review. The exact context-window
size is not exposed in this session. Codex and Claude also ran as test
subjects in the linked paid campaign.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 08:56:04 -05:00
DottaandPaperclip 889947c238 feat: add experimental native chat connectors (#13038)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - People also ask agents for work in their existing chat tools.
> - Each external conversation needs one task and a current authorized
source.
> - Retries, Stop, and provider failures must not duplicate work or
expose private data.
> - The first chat PR establishes the opt-in provider and data
contracts.
> - This PR adds experimental channel integration and its durable
control plane.
> - Users can request work from connected channels and inspect delivery
in Paperclip.

## Linked Issues or Issue Description

Refs #13100 and #13092. This is the second of exactly two chat PRs.
Foundation #13100 is merged and changed 143 files. Runner prerequisite
#13092 is also merged. This PR changes 400 files against master, below
the 500-file review limit. It contains no wireframe images or HTML
galleries.

## What Changed

- Add native Slack, GitHub, Microsoft Teams, Telegram, and Discord chat
connections. Keep chat disabled unless the operator enables experimental
chat connectors. Preserve the production GitHub tool connection and its
normal setup path.
- Bind each provider bot identity to one immutable Paperclip agent. Bind
each admitted external conversation to one task. Paperclip owns tasks,
runs, permissions, and audit records.
- Add durable admission, per-conversation queues, questions, task
controls, progress, final replies, images, files, and delivery receipts.
Board comments remain internal unless explicitly sent to the channel.
- Check current identity, provider reach, resource access, credentials,
runtime generation, and exact source before provider effects. Keep
private responses private. Never send raw reasoning, private logs,
credentials, or tool arguments.
- Hold uncertain sends for explicit audited resolution. Make Board
Send-to-channel atomic and idempotent. Keep reconnect and setup
credentials in Paperclip secret storage.
- Preserve current native-runner authority across retries, lost
acknowledgements, and recovery. Keep immutable input and completion
contracts separate from newer user input. Receipt reconciliation cannot
launch a provider.
- Reconcile chat close/new ordering and provider-effect lock order.
Audit resource access changes in the same transaction. Submit only the
selected resource from each UI toggle so stale pages cannot undo
unrelated access changes.
- Drain Codex stdout before certifying process exit. Bound the drain
with the existing shutdown grace. Preserve observed terminal authority
without treating an undrained process as successful or reusable.
- Incorporate master `018ca5da` with its ACP Stop, mobile task layout,
runner packaging, and official lock changes. Preserve dedicated
chat-answer continuations in both directions when ordinary queued
comments are adopted after Stop.
- Fence late adapter readiness behind an earlier Stop for the same run.
Preserve verified cleanup for registered adapters. Handle single Stop,
agent pause, duplicate Stops, and failure release without creating a
false cancellation receipt.
- Incorporate master's `6dd48cad4` wake-queue extraction. Preserve exact
failed-chat retry authorization and lineage, retired question-source
suppression, and the block on generic recovery that would discard the
admitted source. Fresh deferred input retains its separate promotion
path.
- Incorporate master `2a05b5ed3` and its queue-admission extraction,
simplified transaction ports, and separate runner CI job. Preserve exact
durable receipts, actor separation, and dedicated-answer isolation
through the new module. A failed receipt insert rolls back the
accompanying deferred-wake merge.

## Verification

Current head: `afe19299d06253cb628eb398e91d1200ea9f412a`, incorporating
master `2a05b5ed3457ea33efd6895520447d1d97fe98d8`. The conflicts are
resolved. This successor fixes two test-harness boundaries exposed by
CI: per-case route-module preparation and actual durable-save completion
before intentional runner termination. Production code and all existing
test/turn deadlines are unchanged. [Exact-head Greptile
review](https://github.com/paperclipai/paperclip/pull/13038#issuecomment-5587250594)
is **5/5**, completed September 10 at 13:20:55 UTC, with no actionable
findings or open review threads. [Fresh exact-head
CI](https://github.com/paperclipai/paperclip/actions/runs/34481724341)
passes **all 24 jobs**, including Build and both required aggregates.
Normal exact-head guarded merge was attempted and rejected by the
remaining branch approval policy: CODEOWNER review is required and no
human approval is present. Normal **squash auto-merge is enabled** as of
September 10 at 13:36:26 UTC. Requested CODEOWNERS have been notified;
no approval bypass or self-approval was used. Earlier-head results below
remain historical evidence, not qualification of this successor.

- Final exact-head Linux evidence: 995/995 chat integration cases; 36/36
agent-skills routes; 35/35 runner live-session cases, including real
process kill/resume; 1948 runner Vitest cases with three existing
benchmark/platform guards; 870/870 API-authority cases; and 104 browser
cases with four existing optional skips. Rust, conformance/replay, full
repository build, typecheck, canary, all server/workspace shards, and
both required aggregates pass with normal CI concurrency. Earlier failed
attempts remain recorded below.

- Latest test-only qualification: 141/141
route/permissions/authentication cases pass in separate cold forks, with
plain server types and independent review clear. The real-runner suite
passes 35/35, with plain runner types and independent review clear. A
controlled premature-save acknowledgement fails as expected; matching
ownership/effect/process evidence, rejected saves, real turn outcome,
test abort, and pre-kill liveness are covered. No local reproduction of
the original CI scheduling failure is claimed. The preceding [CI
run](https://github.com/paperclipai/paperclip/actions/runs/34479680858)
passes 21/24 jobs, including all 995 Linux chat cases and browser
aggregate (104 passed, four existing optional skips); only Build, the
skills serialized shard, and the required verification aggregate fail.
Its exact-head Greptile review was 5/5. Both failed job logs are
retained.

- Final fixture qualification: all eight focused Discord cases and all
995 chat integration cases pass. The exact modal statement/PID is
observed before taking the real connection lock; the test then proves
its actual blocking relationship before mutation. Original SQL
execution, provider behavior, negative assertions, and 1s/15s timeouts
remain unchanged. Independent review is clear and test/production hashes
remain frozen. The preceding [CI
attempt](https://github.com/paperclipai/paperclip/actions/runs/34477184777)
passed 22 jobs, including Build/runner, typecheck, canary, all other
test shards, and browser aggregate (104 passed, four existing optional
skips); the two fixture failures and failed verification aggregate
remain recorded, not relabeled as a pass.

- Current queue-module composition: 308/308 recovery/batching/queue/Stop
tests; 995/995 full chat integration; 89/89 module tests, including real
PostgreSQL receipt-insert rollback; 24/24 workflow/module-boundary
tests; plain server and UI types. All four actual local process/ACP
browser paths pass in 1.4 minutes. Fresh databases, no skips or retries,
stable reviewed source hashes. The initial boundary failure is retained;
its no-op service wrapper was removed without changing recovery context
or weakening the check. An exploratory standalone test-directory
typecheck fails because its new upstream transformation config is not a
standalone typechecking project; standard CI/build does not invoke it,
and no configuration was weakened to suppress those diagnostics.

- The preceding head `e02a63d462ce5d47433b0aeb632bb6fd20aab1ba` passed
[all 24 CI
jobs](https://github.com/paperclipai/paperclip/actions/runs/34436462958)
and exact-head Greptile review at 5/5. Required CODEOWNER review
prevented its normal merge before master advanced again.

- Final extracted-module composition: 307/307 recovery, batching, queue
and Stop-control tests; 995/995 full chat integration; 49/49 module
tests including eight PostgreSQL adapter cases; and 19/19 issue-update
tests. Plain server types pass. All four actual local process/ACP
browser paths pass in 1.3 minutes. Fresh databases, no skips or retries
in these cohorts, frozen source hashes, and independent review clear.

- The preceding head `3e4e1c1c` passes [all PR CI
jobs](https://github.com/paperclipai/paperclip/actions/runs/34415826820),
including Build and required `ci / verify` and `ci / e2e`. Both the
original Rust failure and the previously load-sensitive lineage fixture
pass with unchanged Linux concurrency. Master advanced afterward and
required this reconciliation.
- Final master composition: 448/448 focused UI tests, 186/186 adapter
tests, 24/24 queue/control tests, and 11/11 packaging tests. Plain UI,
server, shared, and adapter types pass. Token gates and diff checks
pass. Independent server and UI reviews are clear.
- Stop-registration regression: both real-service cases fail against
exact `a95` source and pass with the fix. The full corrected
recovery/control suite passes 265/265. Duplicate-owner and failed-Stop
controls also pass. Plain server types pass. The readiness barrier
prevents provider startup without adding an acknowledgment to an already
terminal run.
- Final qualification strengthens terminal-field equality and repeats
both affected cases successfully on a fresh database. All four actual
local process/ACP browser paths pass again in 1.3 minutes, without skips
or retries. The final screenshot shows Cancelled, a paused subtree,
retained input, and no error toast.
- Two new actual-service regressions fail before the merge fix. They
prove that queued-comment adoption could consume a dedicated chat answer
or add unrelated input to that answer. The fixed four-case cohort
passes, including ordinary upstream continuation and adapter Stop
controls. Full recovery passes 257/257. All four actual local
process/ACP Stop browser flows pass in 1.4 minutes, without skips or
retries, on a fresh database.
- The unchanged runner artifact was qualified with 171/171 transport
tests, 870/870 API-authority tests, conformance 1/1, and replay 11/11.
Six controlled reader tests prove the exit/drain repair. Its local
serial Rust workspace passed 546 top-level cases plus two invoked
helpers; the later passing Linux CI supplies default-concurrency
evidence.
- Prior exact-source full chat integration passes 995/995. Settings
regressions cover concurrent stale pages, 501 destinations, pending
state, rejected updates, and explicit retry. These deterministic tests
do not prove live provider behavior.
- Retained failed attempts and their causes are in the [qualification
log](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-chat-queue-and-webhook-repair.md).
The first merge adapter run timed out while macOS slept for 290 seconds.
Its unchanged repeat passed with a temporary sleep guard. No assertion,
deadline, or CI gate was weakened.

Review commands include `pnpm --filter @paperclipai/server exec vitest
run src/__tests__/heartbeat-process-recovery.test.ts
src/__tests__/issue-queued-comments-routes.test.ts` and `pnpm exec
playwright test --config tests/e2e/playwright.config.ts
tests/e2e/acp-stop-continuation.spec.ts`. Database suites require fresh
disposable databases. See the [browser
runbook](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-04-chat-adapters-browser-e2e-runbook.md)
for provider setup and separate live acceptance steps.

## Risks

- This remains experimental. Deterministic tests and bounded live
evidence do not establish every provider feature, tenant, permission
layout, or media shape. Teams work-tenant qualification is still open.
- Failed and uncertain provider effects remain visible and can require
operator action. A transport receipt does not prove recipient
visibility.
- Native controller and runner artifacts must remain compatible.
Preserve lease ownership, terminal authority, source binding, and
quarantine during future changes.
- Access and audit rows commit together, but activity notifications
remain best-effort. This is not a new durable event outbox.
- The PR operation does not deploy a live server, replace its runner, or
change provider permissions. Remaining live qualification is documented
in the [temporary
handoff](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-open-qualification-followups.md).

## Model Used

OpenAI Codex assisted with implementation, tool execution, testing, and
review. The work records `gpt-6-astra` assistance. The environment does
not report a context-window size. No private reasoning traces are
included.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-10 10:06:45 -05:00
Dotta 0f94521017 fix(runner): restore local session and task integrity (#12721)
## Thinking Path

> - Paperclip is the control plane for agents that perform work.
> - Paperclip Runner connects durable provider sessions to individual
task runs through PRP.
> - Provider continuity and per-run authority are different lifetimes.
> - The existing implementation mixed those lifetimes and lost event
metadata between provider frames, runnerd, persistence, API
sanitization, and the task thread.
> - That caused failed continuation, missing progress and Plans,
duplicate replies, hidden failures, and unsafe recovery.
> - This repair gives every heartbeat fresh authority, preserves
qualified provider-session continuity, and restores one lossless
presentation path without changing direct adapters.

## Linked Issues or Issue Description

**What happened?**

A second native heartbeat could reuse tickets, leases, command receipts,
sequence state, and run identity from the first heartbeat. Provider
phase and item identity could be lost before the UI read them. Redaction
could corrupt protocol discriminators while still missing malformed
credential tails. The task thread could fold progress into the final
response, hide failures, or show more than one final answer. Native
Codex also exposed approval modes that do not yet have a durable
approval bridge.

**Expected behavior**

Each heartbeat uses a new PRP authority epoch. Codex and OpenCode
preserve exact qualified provider sessions; ACPX emits an explicit
continuity event when its qualified process-replacement policy is used.
Every accepted provider event is presented, classified as internal, or
surfaced as unsupported. The task page shows chronological progress,
reasoning summaries, activity, Plans, interactions, terminal failures,
and exactly one final reply. Direct adapters retain their existing path.

**Steps to reproduce**

1. Enable the unified experimental Paperclip Runner setting.
2. Create a local native Codex, OpenCode, ACPX Claude, or ACPX Codex
agent.
3. Run response, Plan, structured-question/resume, restart,
cancellation, and failure scenarios.
4. Reload the task while active, waiting, failed, and settled.
5. On the old implementation, observe stale run authority, missing
classifications, incomplete output, or duplicated/folded replies.

**Paperclip version or commit**

The repair is based directly on `master` at
`87d05e194b643810d16d20612115acd01d735d43`.

**Deployment mode**

Local development with the embedded database.

Related work: Refs #12616, #12646, #12666, #12685, and #12700.

## What Changed

- Rotates PRP control-plane, outbox, ticket, lease, command, receipt,
and sequence authority for each heartbeat while carrying forward only a
validated provider-session identity.
- Reads `control-plane-state.json`, validates both durable schemas and
lifecycle values, resumes coherent current runs, archives qualified
settled authority, and quarantines malformed or mismatched scoped state
without moving ambiguous live legacy state.
- Preserves Codex provider phase and stable item identities so
commentary remains progress and only `final_answer` becomes final.
- Adds raw OpenCode HTTP/SSE boundary coverage and canonical reasoning
lifecycle mapping.
- Makes ACPX normalization lossless for visible reasoning, tool
lifecycle metadata, stable bounded identities, Plan revisions,
structured requests, failures, and qualified process replacement. Only
the compatible terminal assistant message is promoted as final.
- Applies schema-aware redaction before generic JWT-shaped detection and
scans every diagnostic string leaf. Malformed raw/escaped quoted
credential tails are redacted in both server and durable Rust state.
- Restores snapshot-style chronological task presentation, expandable
tool activity, inline Plan cards, visible waiting/resume/cancel/failure
states, and exactly one final answer.
- Makes `never` the only qualified native Codex permission mode and
rejects unsupported persisted native modes with remediation. OpenCode
and ACPX policies remain intact.
- Keeps the unified experimental Runner setting as the only enablement
flag. Onboarding and direct Codex, Claude, and OpenCode stay on their
legacy execution/finalization paths.
- Adds cross-language goldens, authority/recovery/fault coverage, exact
response/count assertions, and native plus legacy acceptance scenarios.

## Verification

- Pull-request GitHub Actions run Rust formatting/tests, TypeScript
checks, server/UI tests, builds, protocol drift checks, browser E2E, and
security scans.
- A separate workflow-only validation ref is pinned directly on this PR
head and runs the 35-cell paid local matrix: three core scenarios plus
structured-question resume and restart/resume for native Codex, native
OpenCode, ACPX Claude, ACPX Codex, and direct Codex/Claude/OpenCode.
Run: https://github.com/paperclipai/paperclip/actions/runs/33682434315
- Acceptance requires exact single visible replies, monotonic sequences,
matching envelope discriminators, one semantic terminal, one run
terminal, no unresolved interaction, no duplicate mutation, no secret
leakage, provider continuity, and zero native rows for direct adapters.
- Per maintainer direction, tests are running in GitHub Actions rather
than on the slower local host. Only formatters and static diff checks
were run locally.

## Risks

- Recovery from old or partial filesystem state is sensitive. The repair
fails closed, preserves active or unverifiable authority, and
quarantines only state whose scoped ownership is safe to move.
- Provider event formats can change. Closed validators and boundary
goldens turn new or malformed events into visible diagnostics instead of
silent drops.
- Shared task presentation could affect direct adapters. Runtime-fact
gating plus the direct-adapter matrix protect the existing path.
- Managed and remote providers are not qualified here. Shared code
continues to compile and fail safely, but live qualification is
deferred.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex based on GPT-5. The exact deployed snapshot and
context-window size are not exposed to this task. It used agentic
reasoning, repository inspection, code editing, Git, parallel subagents,
and GitHub Actions.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass (intentionally deferred to
GitHub Actions)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [ ] All Paperclip CI gates are green
- [ ] The paid local-provider matrix is green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-02 16:11:26 -05:00
Dotta fdf8c8464d feat(runner): add managed provider backends (#12699)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Paperclip Runner provides durable, provider-neutral agent
execution.
> - The current stack supports qualified local providers but omits the
managed provider paths from the integration branch.
> - Claude Managed Agents and AWS AgentCore need explicit profile
qualification, durable recovery, usage accounting, and cleanup controls.
> - This pull request adds those managed backends as the third part of
the Runner parity stack.
> - The benefit is managed execution without weakening the default-off
Runner rollout gate.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: Runner, server orchestration, database profiles, CLI, and
adapter configuration UI.

**Problem or motivation**

The current Runner stack cannot select or execute the managed Claude
Agents API or AWS Bedrock AgentCore Harness backends. It also lacks
qualified profile storage and recovery checks for those remote
resources.

**Proposed solution**

Add qualified managed and remote profiles, API and CLI management, exact
provider selection, durable lifecycle handling, cumulative usage
accounting, bounded cleanup, and retention acknowledgement. Keep
`enableNativeRunner` default-off.

**Alternatives considered**

A direct copy of the old integration branch was rejected because its
provider contracts, model values, credential flow, and migration history
no longer match the current base. A single large parity pull request was
also rejected because stacked review keeps each subsystem bounded.

**Roadmap alignment**

This continues the existing Runner architecture and rollout work. It
does not introduce a separate execution system.

**Additional context**

This pull request is based on the merged #12691 and #12685 stack. It
also closes the delayed security-review findings reported on #12691 by
binding qualified ACPX and OpenCode launch artifacts to the bytes
actually executed. A GitHub search for managed agent, AgentCore, and
Claude managed work found no duplicate public issue or pull request.

## What Changed

- Add Claude Managed Agents and AWS AgentCore provider executors to
runnerd.
- Add qualified managed and remote profile storage, routes, OpenAPI
contracts, CLI commands, and migration 0237.
- Validate profile ownership, enabled state, exact qualified revision,
model, agent version, and secret binding before persistence and
recovery.
- Persist durable provider session and owned skill state for
restart-safe cleanup.
- Reconcile uncertain create responses and delete remote sessions before
owned skills.
- Track cumulative provider usage and enforce positive session spend
caps.
- Recover interrupted AgentCore usage at the next turn boundary by
charging the prior invocation ceiling exactly once; keep the session
gated until an explicit monotonic budget raise.
- Isolate AgentCore AWS configuration from host profiles and
credential-process/SSO configuration while preserving workload identity.
- Require OpenCode 1.18.17 and fixed build-owned provider-pack artifact
paths; remove the ambient executable override.
- Snapshot and content-verify ACPX and OpenCode commands, scripts, and
provider executables before launch. Linux executes sealed inherited
descriptors; macOS uses authenticated private snapshots with retry-safe
rematerialization at the spawn boundary.
- Persist canonical ACPX and OpenCode launch-profile digests, reject
drift across fresh recovery, and make recovery failures sticky.
- Close and journal unsafe ACPX active-turn recovery before any provider
bootstrap or reconnect.
- Add managed provider fields to the Runner configuration UI and
permission projection.
- Preserve the default-off `enableNativeRunner` experimental flag.

## Verification

- `pnpm -r typecheck`
- `pnpm build`
- Focused managed server, database, CLI, Runner TypeScript, Rust,
Claude, AgentCore, ACPX, OpenCode, process-supervisor, and
durable-recovery tests passed.
- `cargo test -p paperclip-runner-core --lib --locked` (160 tests)
- `cargo check --workspace --all-targets --locked`
- Native Codex integration tests passed (60 tests); native provider
tests passed (7 tests); server native-runtime tests passed (87 tests).
- Verified-launch replacement, nested-spawn retry, exact-version,
profile-drift, sticky-failure, and no-bootstrap active-recovery tests
passed.
- `git diff --check`
- The PR changes 91 files. `pnpm-lock.yaml` is unchanged. The Rust
workspace lockfile adds the approved `rustix` dependency used for safe
descriptor handling while `#![forbid(unsafe_code)]` remains enabled.

## Risks

- The provider APIs can change while they are in beta. Exact
qualification and fail-closed recovery checks limit drift.
- Remote cleanup can fail after a partial create. Durable ownership
inventories and retry-safe deletion preserve recovery state.
- Migration 0237 adds profile tables. The generated migration and
snapshot pass the repository migration checks.
- Managed execution can incur provider cost. Positive default spend caps
and explicit retention acknowledgement limit accidental use.
- An interrupted AgentCore invocation without final metadata is
conservatively charged to its active session ceiling. This can overstate
cost, but cannot undercount it; later work requires an explicit budget
increase.
- Linux qualified launches use sealed memory descriptors. macOS lacks
executable-descriptor APIs, so the runner uses owner-only private
snapshots and minimizes linked-path lifetime; hostile same-UID processes
remain outside the documented local-host trust boundary.
- The global Runner feature remains default-off.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex, GPT-5, with tool use, code execution, and subagent review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-02 00:48:30 -05:00
Dotta 1ed29abaa6 fix(runner): harden dormant provider boundaries (#12654)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Runner currently enables only the Codex production path.
> - The package also contains dormant OpenCode and ACPX provider
boundaries.
> - Dormant boundaries must still fail safe before later activation
work.
> - Provider children must not inherit unrelated server secrets or host
homes.
> - Permission defaults must require interaction instead of broad
automatic approval.
> - This pull request hardens those boundaries without activating them.
> - The benefit is a safer base for later provider-specific runnerd
work.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

This improves the inactive OpenCode and ACPX provider boundary in
Paperclip Runner.

**Subsystem affected**

The adapter permission contract, Runner provider environment, and native
execution input builder.

**Current behavior**

Dormant OpenCode code can inherit the full server environment. Its
default permission mode allows operations. ACPX also defaults to broad
approval. The provider guard can accept inherited object property names.

**Proposed behavior**

Use exact provider identifiers. Use interactive defaults. Allow only
required OpenCode environment keys. Reject invalid proxy permission
modes.

**Reason and benefit**

This reduces accidental authority and secret exposure before future
provider activation.

**Breaking changes**

No production provider is activated. Codex runtime selection and Codex
credential-home discovery do not change. Dormant OpenCode and ACPX
callers that omit permission modes now receive safer defaults.

## What Changed

- Change dormant OpenCode and ACPX permission defaults to interactive
modes.
- Reject prototype property names as provider identifiers.
- Default dormant ACPX input to the qualified Codex agent profile.
- Add an explicit OpenCode runner environment allowlist.
- Exclude host homes, server credentials, database values, and Node
injection options.
- Add a fail-closed OpenCode proxy permission parser.
- Add focused tests for defaults, filtering, and invalid values.

## Verification

GitHub Actions must run:

- Adapter utility tests.
- Paperclip Runner tests, type checks, and build.
- Server native runtime tests.
- Repository test, type-check, build, policy, and security gates.

No local test command was run. The repository owner requested
GitHub-only verification.

## Risks

Future OpenCode credential providers must add required variables to the
allowlist through review. The safer defaults can pause dormant internal
scenarios that relied on implicit broad approval. Production Codex
behavior is unchanged.

## Model Used

OpenAI Codex with the GPT-5 agent model. The work used high reasoning,
repository inspection, tool use, and parallel security review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-01 04:46:36 -05:00
Dotta 131f5c4065 feat(runner): add administration and observability (#12641)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Administrators need bounded controls for experimental native
execution.
> - The lower stack adds remote Codex execution and the task workspace.
> - Operators need to configure Codex safely and inspect provider
traces.
> - Unsupported providers must not appear as runnable choices.
> - This pull request adds Codex-only administration and observability.
> - The benefit is a default-off operational surface for production
diagnosis.

## Linked Issues or Issue Description

Refs #12640.
Refs #12616.
Refs #12352.

**Subsystem affected**

Agent configuration, instance experimental settings, run ledger,
provider trace inspector, and administrator actions.

**Problem or motivation**

The native runner lacks one safe operator surface for Codex permissions,
lifecycle, raw trace capture, and run inspection. The integration branch
also contains provider choices that the production backend cannot
execute yet.

**Proposed solution**

Expose only the qualified Codex controls. Keep Paperclip Developer Mode
and runner preview ingress off by default. Gate raw trace actions by
administrator access and existing trace authorization.

**Alternatives considered**

Exposing unfinished providers would create configurations that fail at
runtime. Always-on tracing would increase sensitive data and storage
risk.

**Roadmap alignment**

This work supports governed Cloud and Sandbox agents and production
diagnostics.

## Stack

- Base PR: #12640.
- Lower PRs: #12639 and #12638.
- This PR contains only its 54-file administration and observability
delta.
- This is the final feature PR in the Codex production stack.

## What Changed

- Added Codex-only Paperclip Runner permission and lifecycle controls.
- Added bounded warm idle configuration.
- Kept the provider field fixed to Codex.
- Added administrator-only one-run raw trace requests.
- Added a persistent future-run raw trace toggle.
- Added trace status, metadata, ledger, and canonical runner inspection.
- Added JSON-RPC request-origin grouping and finalization lineage.
- Restored the stateful PRP transcript parser and focused projection
tests required by trace inspection.
- Added default-off Paperclip Developer Mode.
- Added Honeycomb run links for authorized developer mode.
- Disabled the legacy operational skill for `paperclip_runner`.
- Did not expose OpenCode, ACPX, Pi, Claude Managed, or AWS runner
choices.
- Did not change migrations, workflows, dependencies, or
`pnpm-lock.yaml`.

## Verification

- GitHub Actions will run UI tests, server tests, repository typecheck,
build, browser tests, security, and policy gates.
- Tests cover Codex configuration defaults and bounds, administrator
trace actions, persistent settings, ledger inspection, trace lineage,
and Honeycomb links.
- Existing server trace authorization and retention tests remain the
backend authority.
- Local tests were not run. The requested verification policy uses
GitHub Actions for this series.
- `git diff --check runner/task-workspace-experience...HEAD` passes.
- The delta contains 54 files.

## Risks

- Raw provider traces can contain sensitive provider data.
- Existing server authorization controls access, reveal, download,
retention, and deletion.
- The UI gates trace actions by administrator access and developer mode.
- All new instance settings remain off by default.
- Fresh Paperclip Runner configuration remains Codex-only.
- Direct adapters and legacy task behavior do not change in this PR.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex with GPT-5.6. The work used high-reasoning agent mode,
repository tools, GitHub tools, and parallel code-audit agents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with Fixes: / Closes /
Refs OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-01 03:41:23 -05:00
Dotta 0a422fda52 feat(runner): add remote execution substrate (#12638)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Runner gives native runs a durable and governed execution
path.
> - The current native path runs on the control-plane host.
> - Remote environments need an authenticated execution-target contract.
> - The contract must not change direct adapters or enable new runtimes
by default.
> - This pull request adds the remote execution substrate and Daytona
ingress.
> - The benefit is a bounded base for later remote runner transport
work.

## Linked Issues or Issue Description

Refs #12616.
Refs #12352.

**Subsystem affected**

Cross-cutting. This change touches runner transport, server
orchestration, plugin contracts, and shared settings.

**Problem or motivation**

Native execution cannot resolve an authenticated runner ingress through
a remote environment. The server also lacks one provider-neutral
contract for remote execution targets.

**Proposed solution**

Add a default-off runner preview ingress capability. Add
transport-neutral runner connectivity. Add remote execution target and
lifecycle handling. Add a Daytona ingress implementation with redacted
credentials.

**Alternatives considered**

A provider-specific server path would duplicate orchestration and
authorization. A public endpoint without an environment contract would
weaken the trust boundary.

**Roadmap alignment**

This work supports the Cloud and Sandbox agents milestone. It also
supports self-healing runs and governed tool access.

## What Changed

- Added execution-target traits for local, SSH, and sandbox
environments.
- Added plugin RPC contracts for runner ingress endpoints.
- Added authenticated Daytona preview ingress.
- Added transport-neutral PRP outbound connections.
- Added remote runner artifact verification and fail-closed provider
selection.
- Added bounded native session resume, cancellation, and lifecycle
recovery.
- Preserved Codex-only selection for fresh experimental runner starts.
- Preserved all direct adapter execution and finalization paths.
- Removed stale Pi provider-pack requirements that security review
rejected.
- Kept the rollout controls off by default.
- Did not change pnpm-lock.yaml, Cargo, database migrations, or GitHub
workflows.

## Verification

- GitHub Actions will run the repository test, typecheck, build,
security, and policy gates.
- Focused tests cover ingress validation, redaction, execution targets,
remote lifecycle, cancellation, resume, and legacy adapter selection.
- Local tests were not run. The requested verification policy uses
GitHub Actions for this series.
- `git diff --check origin/master...HEAD` passes.
- The diff contains 52 files.

## Risks

- Remote execution crosses a trust boundary.
- The implementation validates target capabilities, artifact digests,
provider-pack pins, and connection metadata.
- The feature remains default-off.
- Fresh native selection remains Codex-only.
- Existing direct adapters remain on the legacy path.
- This PR does not yet make remote Codex runnable. The next PR adds the
Rust WSS and TLS transport.

## Model Used

OpenAI Codex with GPT-5.6. The work used high-reasoning agent mode,
repository tools, GitHub tools, and parallel code-audit agents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with Fixes: / Closes /
Refs OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-01 01:29:06 -05:00
Dotta 51ad751e0b feat(runner): integrate Codex native execution (#12616)
## Thinking Path

> - Paperclip is the open source control plane for teams of AI agents.
> - Agent runs currently use direct adapters and their established
finalization paths.
> - The new runner package needs one production integration before it
can execute a real provider through the server.
> - That integration must not change direct adapters or expose
unsupported providers.
> - The rollout must also preserve native runs that were already
recorded when the feature flag changes.
> - This pull request adds a default-off, Codex-only native execution
path and its authority boundary.
> - The benefit is a recoverable production vertical slice with explicit
compatibility guards.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting server orchestration and adapter selection.

**Problem or motivation**

The runner package exists, but the server cannot yet start and recover a
governed Codex run through it. A careless integration could also route
existing direct adapters into the native runtime or lose cancellation
and finalization state.

**Proposed solution**

Add a hidden `paperclip_runner` adapter for Codex. Keep it behind the
default-off instance flag. Bind native execution, resume, cancellation,
semantic tool authority, and finalization to the recorded company,
issue, run, and coordinator identities. Leave every direct adapter on
its existing path.

**Alternatives considered**

A multi-provider launch was rejected because only Codex has the complete
production bridge in this series. Replacing direct adapter execution was
rejected because the runner remains experimental.

**Roadmap alignment**

This work supports governed tool access, action attribution, and
self-healing runs. It keeps the integration narrow and default-off.

## What Changed

- Add the Codex-only native session executor and persisted resumption
path.
- Add run-scoped semantic tool projection, authorization, receipts, and
idempotency.
- Add audited native cancellation with durable issue and coordinator
binding.
- Add result fencing so a recorded result cannot reacquire the provider
and run twice.
- Reject fresh runner starts when the rollout flag is off while
preserving recorded native recovery.
- Keep direct adapters outside native status, cancellation, record
creation, and finalization.
- Add focused conformance, recovery, cancellation, status, portability,
and compatibility coverage.

## Verification

- GitHub Actions is the authoritative test environment for this large
stack.
- The PR policy and lightweight stack checks run while this is a middle
PR.
- The full required suite runs when this PR becomes the lowest unmerged
or top PR.
- Greptile will review this exact delta after the branch is pushed.

## Risks

- The main risk is routing a legacy adapter into native execution.
Runtime selection and heartbeat tests cover that boundary.
- The next risk is stale or cross-company cancellation. Durable binding
checks and transactional audit persistence cover it.
- The adapter remains hidden and default-off. Only Codex is admitted.
- There are no database migration, lockfile, or GitHub workflow changes
in this PR.

## Stack

1. [Runner package, SDK, and developer
tools](https://github.com/paperclipai/paperclip/pull/12608)
2. This PR: Codex production server integration
3. [Provider-neutral task-thread
UI](https://github.com/paperclipai/paperclip/pull/12617)

## Model Used

OpenAI Codex with GPT-5, extended reasoning, repository tools, and
parallel review agents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-31 22:51:17 -05:00