Commit Graph
25 Commits
Author SHA1 Message Date
DottaandPaperclip d66acb7ac1 feat: automate Slack bot app setup and installation (#15413)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors give each agent a customer-owned bot and task-backed
conversations.
> - Manual Slack setup requires app creation and copying durable
credentials.
> - Operators need a shorter setup that an assisting agent can use
safely.
> - This pull request creates the app through Slack's Manifest API and
installs it through OAuth.
> - Durable registration state supports recovery without creating
another app.
> - A four-screen wizard, automatic avatar upload, and OAuth account
linking reduce setup work.
> - Connector settings and per-turn tool guidance support daily use
after installation.

## Linked Issues or Issue Description

**Subsystem affected**

Native Slack bot setup, company secret storage, chat connector
management, and agent tool guidance.

**Problem or motivation**

New Slack bots require manual app creation and copying a signing secret
and bot token. Interrupted setup can create duplicate apps. The setup
and management screens contain unnecessary controls. Agents also need
guidance for native questions, files, thread replies, and governed Slack
actions.

**Proposed solution**

Use a temporary app-configuration access token to create a
customer-owned app. Save durable secrets in the vault. Bind OAuth to the
initiating actor, company, endpoint, registration revision, scopes, and
configured origins. Preserve manual and existing-app recovery. Link the
installing user's account, send a welcome DM, and advance from saved
server evidence. Keep request-URL recovery instructions available if
automatic connection detection waits.

**Alternatives considered**

The Slack CLI adds installation requirements. Socket Mode changes
transport. A shared Paperclip-owned app changes app ownership. These
alternatives are outside this change.

**Roadmap alignment**

This extends existing chat connectors and secrets capabilities. Related
public work: #14037 and #13954 cover Slack MCP prerequisites and user
OAuth. No duplicate bot-registration PR was found.

## What Changed

- Share one reviewed manifest builder between automatic registration and
manual setup.
- Add replay-safe migration 0318 and company-bound registration state
with vault references and uncertain-creation recovery.
- Add registration, installation, callback, and resume APIs with
short-lived, single-use OAuth state.
- Save installation credentials before downstream checks and preserve
bot identity constraints.
- Reduce automatic setup to four screens. Keep advanced app details,
manual recovery, and existing-app setup.
- Upload the agent avatar with the Paperclip dark background. Link the
OAuth installer's account and send setup DMs.
- Show agent and connector-owner avatars. Simplify settings, access, and
conversation screens.
- Discover joined Slack channels and enable them by default. Start a
task from a bare mention and admit same-thread follow-ups.
- Refresh Slack tool guidance each turn. Add native-form, file,
approval, and delivery regressions plus manual model probe definitions
and sanitized acceptance records.
- Update deployment/database docs, OpenAPI, redaction, removal cleanup,
production Storybook stories, and provider browser tests.
- Merge current master and move the registration migration after its
latest migration without rewriting published commits.

The completed Slack success view intentionally has a single centered
**Done** action and no **Save & exit**, as explicitly requested by the
product owner. `DESIGN.md` records this exception; unfinished setup
steps retain the aligned wizard footer.

## Verification

- Passed after the master merge: repository typecheck, full build,
Storybook build, design-token gates, module-boundary gates, and
migration generation.
- Passed: all 352 focused Slack deterministic tests and all 14 affected
provider browser tests. Browser tests use controlled provider fixtures
and a separate throwaway instance.
- Passed on current head `c5d01e0e2`: the complete GitHub test matrix
(general server, chat, all workspaces, serialized server, and Runner),
all eight browser shards, typecheck/release registry, build, canary dry
run, security checks, and policy gates. There are 52 passing checks and
no pending or failing checks.
- Greptile completed on the exact current head with 5/5 and no
actionable findings or open review threads.
- Local repair verification passed 93 focused tests, including same-app
reinstall after revocation and rejection of consent started before
revocation, the AgentMail browser journey, and repository typecheck.
Local build and Storybook build also passed. The redundant local
full-suite rerun was stopped after the complete current-head CI matrix
passed.
- Real Slack setup and agent replies were exercised in the authorized
isolated test drive during the setup iteration.
- The ten additional model probes were attempted with legacy
`codex_local`, `gpt-5.6-sol`: five passed, two failed, and three were
partly verified. Native runtime is not qualified. See
`server/src/services/connectors/slack/evals/2026-10-08-acceptance.md`
for evidence and limits.
- Passing model probes cover native forms, downloaded file bytes, bare
mentions with thread replies, explicit posts/reactions, and saved
approval denial.
- The controlled uncertain-write probe found wrong delivery-check IDs.
The canvas fallback attempt used an invented tool name. Search
pagination/native search, a private-source denied-tool receipt, and
distinct board/webhook origins remain unqualified.

Reviewer path: enable Chat connectors, start Slack chat setup, select an
agent, enter an app-configuration access token, and approve Slack
installation. Send a message to the bot and confirm that setup advances
to success. Inspect settings and allowed channels. See
`doc/connections/SLACK-AUTOMATIC-SETUP.md` for deployment and recovery.

## Risks

- Slack app creation has no provider idempotency guarantee. A timeout
after dispatch stays uncertain until the operator checks Slack.
- OAuth needs a stable public HTTPS board origin. Webhook ingress may
use a separate configured HTTPS origin. Workspace policy can delay
installation.
- Migration 0318 can replay safely on instances that applied the earlier
development migration.
- OAuth installation now links the installer to the initiating Paperclip
user. Identity checks and company access rules still apply.
- Joined channels now enable bot responses by default. Linked-user
authorization and per-action approval rules still apply.
- Model behavior has the documented delivery-check and canvas fallback
failures. A passing CI run does not establish that every model probe
passed.
- Removing the connection does not delete the customer's Slack app. No
new first-party telemetry is added.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, code
execution, and browser verification. The runtime does not expose a more
specific authoring model ID or context-window size. The live bot probes
used OpenAI `gpt-5.6-sol` through `codex_local` in legacy mode.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 06:33:31 -05:00
DottaandPaperclip caf120105c test: prepare neutral native connection guidance evals (#15407)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents need to discover connections, obtain consent, and continue
from saved decisions.
> - We want to reduce repeated instructions only when measured behavior
supports the change.
> - The existing decline tasks tell the model not to retry. One
provider-decline check can pass without an explanation or an observed
service counter.
> - This PR adds neutral tasks and stricter saved-evidence checks before
any connection instruction reduction.
> - Production instructions remain unchanged. The new cells are
configured, not live-qualified.

## Linked Issues or Issue Description

Refs #15218. Refs #15389.

**What existing behavior does this improve?**

The Product E2E connection workflow evaluation and its instruction
measurement provenance.

**Current behavior**

Some decline prompts supply the policy they intend to test. The
provider-decline workflow does not require a saved post-decision
explanation. Its old no-call check can use a missing fixture counter as
zero. OpenCode has no connection cases in the original Everyday matrix.

**Proposed behavior**

Add an explicit-only suite with five connection stories on native Codex,
ACPX Claude, and OpenCode. Require an explanation attributed by exact
run ID after a saved decline. Observe the provider fixture counter.
Preserve the original cases and grades.

## What Changed

- Add fifteen configured cells with one attempt, twelve-minute
deadlines, and verified 1,000-cent company and agent budget stops.
- Remove procedure hints from the three new decline prompts. Keep a
user-permitted explanation fallback and the existing positive controls.
- Require saved decline state, one decision, unchanged connections,
observed zero service calls where applicable, and a post-decision
explanation from a successful run on the same task.
- Add negative grader calibration and test the actual fixture budget
payloads. Exclude the suite from default and generic selection.
- Extend the existing full-catalog measurement source manifest with
connection descriptions and schemas. Add an audit of fixed text, tool
descriptions, returned instructions, and unqualified behavior.
- Rebase on master `a6306ba606eb87c89b9ef0344e9fe8e0025580f9` and
preserve its new Cursor suites. No production, credential, workflow, or
lockfile change.

## Verification

- Before rebase: Product E2E support passed 1,424 TypeScript tests and
128 Node checks. Six catalog measurement tests, repository
typecheck/build, Product E2E typecheck, and exact fifteen-cell discovery
passed.
- The full pre-rebase repository test run was stopped when master
advanced. Its partial result is not a pass.
- After rebase and the review correction: repository build/typecheck,
Product E2E typecheck, 1,799 TypeScript support tests (one skipped), 128
Node checks, six measurement tests, and exact fifteen-cell discovery
pass. The duplicate local full-suite run was stopped incomplete after
about 20 minutes once complete CI passed; no local full-suite pass is
claimed.
- Review found that the initial grader read `runId` instead of public
`createdByRunId`. A regression calibration reproduced both rejection of
valid public comments and acceptance of the wrong alias. The fix uses
the actual field and binds the evidence type to the shared
`IssueComment` contract. A subsequent type-only import path correction
passes Product E2E typecheck.
- Final source `0de306b9664bfbdebb6709ddb54c95152740d1ad` passes
[complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37560250545):
51 successful checks and two intentional Storybook skips, plus separate
Snyk success. Fresh Greptile review is 5/5 with the single review thread
resolved and no new findings. The PR is clean and mergeable.
- Local commands: `pnpm build`, `pnpm -r typecheck`, `pnpm
test:e2e:runner:unit`, `pnpm test:e2e:runner:typecheck`, and `pnpm
test:e2e:runner -- --list --suite native-connection-guidance`. The
measurement uses
`PAPERCLIP_NATIVE_PROCEDURE_MEASUREMENT=/tmp/connection-measurement.json
pnpm exec vitest run --project @paperclipai/server
server/src/__tests__/native-procedure-measurement.test.ts`. Validation
used pinned pnpm 9.15.4.
- No paid provider campaign was started. There is no baseline/candidate
behavior result for these new cells.
- The audit records 654 UTF-8 bytes of fixed connection guidance. A
clean capture at `a04b8c6a452315625014888335d45670a2094fb6` confirms 41
supplied tools, 53,341 normalized bytes at start/resume, 50,949 at
compact continuation, and a 48,195-byte authenticated OpenCode MCP
catalog. These are byte counts, not tokens, bills, vendor-private prompt
sizes, or savings from this PR.

## Risks

- This is eval preparation. Passing support tests do not establish live
model behavior or qualify an instruction reduction.
- The explanation oracle checks attributed saved output. It does not
prove cognition or arbitrary prose truthfulness. One saved interaction
also does not prove the absence of repeated idempotent tool calls.
- Successful new authentication and tool refresh, existing-connection
agent grants, independent work while waiting, explicit retry after
decline, and blocking when mandatory work remains still need separate
coverage.
- Notion setup decline does not execute a real Notion service. Positive
service approval uses an already installed deterministic service; it
does not qualify new connection creation.
- The original historical failures remain unchanged. Future comparisons
must freeze source, fixture, model, input, and grading controls and
retain every actual attempt.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, and code
execution. The exact serving model ID and context-window size were not
exposed in this session; they are not inferred. No model provider was
invoked by the eval suite in this PR preparation.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 21:28:16 -05:00
DottaandPaperclip e38d6d16b6 feat(connections): add advanced provider setup and live browser qualification (#15341)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users connect accounts and choose an agent harness and model.
> - The runtime change in #14970 supports custom providers on those
connections.
> - Normal setup must stay simple while advanced users can choose a
compatible gateway.
> - Shared connector rows and access controls keep these choices
consistent.
> - This pull request refines the agent setup UI and adds review stories
and repeatable browser qualification.
> - The qualification checks real tools and downloaded outputs, not only
a successful run status.

## Linked Issues or Issue Description

Refs #14970, #37, #13083, #14104, #14565, #12692.

The core implementation in #14970 is merged. This branch incorporates
its squash commit and targets `master`. Both PRs contain our
implementation. #14016 is a reference only and is not a dependency. This
PR has 96 changed files.

## What Changed

- Complete model-provider connector presentation beside other
connectors. Each row uses the existing Connect action and connection
list. Tags are stored without category UI. The base PR includes the
provider forms and routes.
- Show persistent Subscription, API Key, and Advanced choices. Label
Advanced as Custom Gateway. Reuse provider logos, connection lists, and
permissions controls. Default access to the organization and all agents
when permitted; keep narrowing controls under Advanced.
- Keep Configure reachable before subscription sign-in, so users can
select a supported environment when the default cannot sign in. Testing
and saving still require a connection. Show the execution environment in
Configure. Preserve the confirmed Connect choice. Editing a method,
credential, saved account, or advanced choice requires that current
choice to connect before testing or saving. Use matching model and
thinking-effort dropdowns and retain connection icons in selected
values.
- Preserve the new harness model default when switching an existing
OpenCode agent to Codex or Claude, and resolve user-selected model names
with the effective harness.
- Load popular OpenRouter models through the shared connection-model
discovery path. Keep explicit model lists and manual model entry
available.
- Group onboarding, connection setup, agent runtime, management,
recovery, and production-component stories under AI Connections /
Provider routing.
- Add an explicit-only provider-connections browser suite for managed
local or existing local/staging targets. Use private browser profiles
and credential handoffs. Support human-assisted subscription sign-in
without sharing passwords or tokens in reports.
- Verify persisted connection identity, runtime probes, tool execution,
exact artifact bytes, completion, and context-dependent follow-up.
Retain source/model provenance, cost bounds, closed error diagnostics,
original failures, and cleanup evidence.
- Add Gemini startup-model and skill-root fixes, Grok private-history
detection, ACP filesystem regression fixtures, selected-workspace
handling for local Hermes, and artifact-helper workspace fallback.
- Keep managed Grok runtime homes disposable. Remove host-side
transcript retention/restoration because private file modes do not
isolate same-user agent processes. Ignore earlier development archives
and use a fresh task handoff when history is unavailable. Verify the
absence of restored transcripts with a separate same-user process.
- Capture stopped-run diagnostics before deleting an attached-company
fixture agent. Track creation and owned sign-in receipts; revoke only
this attempt's accounts and never adopt a concurrent campaign's newly
created account. Preserve failure signals and final status through
cleanup.
- Require the requested environment in the saved agent and every run,
including follow-ups. Reject a forced incompatible target. Keep one
cancellation state through startup, every cell, reporting, and teardown
for SIGINT, SIGTERM, and SIGHUP. Stop further paid cells after
interruption. Document qualification limits.

## Verification

- Current head `b3bb3e94d577d43d9965a6b9daba039f599b2e49` includes
master `d9f600043`. The security fix in `a758fde31` passes full
workspace typecheck, production build, and 119 connection/Grok
regressions. The unchanged UI passes all 126
configuration/model-discovery tests and token gates. The final
published-guide correction passes Grok adapter typecheck. Earlier head
`eebd8225c` passed the complete deterministic runner suite (1,404 Vitest
tests and 128 Node tests) and all CI jobs. Current-head CI run
`37520147514` passed all 47 jobs, including the full sharded Vitest and
browser matrix, production build, and canary dry run. All 55 checks
completed: 53 successes and two expected skips. The current-head
security scan passed, Greptile is 5/5, and no review threads remain
open.
- A separate same-user process reproduced reading a restored Grok
transcript before the security fix. The regression now finds no
transcript. Existing fresh-session fallback and ordinary session
metadata behavior pass.
- The final account-choice and cleanup fixes pass 85 setup tests and 26
qualification-harness tests. Regressions verify that editing a
connection invalidates confirmation, Configure remains reachable before
sign-in, diagnostics are captured before fixture deletion, and
concurrent campaigns cannot adopt or revoke each other's accounts. UI
and E2E typechecks pass.
- The Storybook build and actual Chromium production-component stories
passed during this change. Review the neighboring AI Connections /
Provider routing stories, regular connector rows, three connection
modes, model discovery, and the single execution-environment control in
Configure.
- Cancellation smoke verified authenticated cleanup before browser close
for SIGINT, SIGTERM, and SIGHUP. Regressions cover interruption during
startup and reporting, missing-file ACP resource errors, and preserved
permission denials. Both ACP runtime versions and 54 ACPX/Grok
regressions passed. The deterministic connection-intent browser suite
passed two tests.
- Historical local qualification retained 43 passing API/gateway cells
out of 46, with downloaded outputs and follow-up receipts. These
attempts span earlier builds; they do not qualify this exact commit or
staging. Subscription combinations, Gemini overloads, and the unresolved
follow-up failure remain recorded rather than counted as passing.
- Use `pnpm test:e2e:runner -- --list --suite provider-connections` to
inspect the matrix. Follow `tests/runner-e2e/PROVIDER-CONNECTIONS.md`
for credentials, target URL, sign-in assistance, budget, evidence, and
cleanup. Paid live tests remain opt-in.

## Risks

- The core implementation in #14970 is merged. This PR adds no database
migration of its own.
- Subscription login needs an interactive provider session. Dedicated
accounts and staging qualification remain follow-up work; this PR does
not certify every login combination for production.
- Managed Grok transcript resume is deferred until provider history has
an OS isolation or authorized broker solution. Follow-ups start fresh
with Paperclip task context; earlier live Grok results do not qualify
this behavior.
- Gemini CLI 0.58.0 has an upstream ACP new-file error conversion
defect. Live overloads and one unresolved follow-up timeout remain
recorded. The stock CLI is unchanged, and those cases are not marked as
passing.
- Real-provider tests spend credits and use private credential/evidence
directories. The launcher requires explicit selection and checks target
ownership. It must not attach to a developer's database by accident.
- OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore,
Process, HTTP, and legacy ACPX local remain outside custom provider
setup.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact deployment model ID and
context window size were not exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 15:21:22 -05:00
DottaandPaperclip e34abee670 feat(mcp): connect assistants to a team with user OAuth (#14846)
## Thinking Path

> - Paperclip gives teams durable tasks, agent execution, budgets, and
approvals.
> - People also use assistants in Codex, Claude, and other MCP clients.
> - Those assistants need a scoped connection that preserves the
person’s permissions and attribution.
> - Delegating a task must not turn the assistant into the assigned
agent.
> - This PR adds opt-in user OAuth, ten first-party tools, browser
consent, and workflow packages.
> - Paid product evals verify the resulting tasks, documents,
attribution, retries, and access boundaries.
> - The team keeps working after the assistant conversation ends.

## Linked Issues or Issue Description

**Problem or motivation**

A person cannot connect an external assistant to an existing team
through browser consent and safely delegate durable work as themselves.

**Proposed solution**

Expose an opt-in `/mcp/paperclip` endpoint with individually described
first-party operations. Bind every connection to a person, client,
company, resource, and scopes. Reuse domain authorization and
scheduling. Package shared team-review, delegation, and follow-up
workflows for OpenAI/Codex and Claude.

**Alternatives considered**

Related PRs #9393 and #12549 cover earlier remote MCP and board-operator
approaches. This change uses user OAuth and a bounded public catalog. It
does not expose a generic executor, operator administration, static
shared board credentials, or external agent execution. Registry listing
work in #9851 is a separate distribution step.

**Roadmap alignment**

This maintainer-requested implementation extends the governed MCP
gateway, activity attribution, durable work products, and hosted
deployment direction in `ROADMAP.md`. It implements the first release of
the saved design plan; external agent participation and granted
third-party tools remain later releases.

## What Changed

- Add MCP 2.0 discovery and task status/comment/document Events on the
same authenticated endpoint. Persist subscriptions and delivery
receipts, verify HTTPS callbacks, sign Standard Webhooks, encrypt
callback material, recheck permissions/Cloud membership, and bound
retries/expiry. Older MCP clients keep their existing tools.
- Add discovery, dynamic client registration, S256 PKCE, resource
validation, rotating refresh tokens, revocation, and company consent.
Store credentials as hashes and recheck membership at execution.
- Add tools for connection identity, agents/projects, task
search/read/create, human comments, documents/deliverables, and
pending-approval links. Preserve current domain permissions and
scheduling.
- Add durable mutation receipts across reconnects. Matching retries
replay results; uncertain outcomes keep the same request ID and require
inspection.
- Add consent and connection-management pages, OAuth log redaction,
shared plugin workflows, and separate OpenAI/Codex and Claude package
outputs.
- Add eight paid Product E2E cases across three models, independent
durable-state grading, usage evidence, cleanup, and report integration.
Add task-document guidance and regenerate the runner capability
inventories.
- Add migrations 0301 and 0302, the dated implementation plan, result
notes, and direct-client setup instructions in `doc/public-mcp.md`.

## Verification

- Merge integration `e180b1948`: resolved conflicts with current master,
preserved both eval registries, regenerated capability catalogs, and
regenerated migrations as 0301/0302 while keeping the original
replay-safe SQL byte-identical. Local migration safety/snapshot tests
(26), MCP/OAuth tests (38), redaction/OpenAPI tests (71), and eval
catalog/grading tests (198) pass. Token and capability gates pass. Full
recursive typecheck passed. Fresh Greptile review is 5/5 with no
unresolved findings. CI is green on this exact head (55 successes, two
intentional skips, one neutral result): one unchanged Cursor sandbox
test timed out at 10 seconds, then passed locally in 856 ms. A single
retry of that failed shard and the aggregate workflow passed. Merge
remains blocked on the repository code-owner approval rule.

Earlier checks passed at `6aa0962d4fb715f2190bb7bb22efacab2e58495d`: 55
successes, two intentional skips and one neutral result. [The earlier CI
run](https://github.com/paperclipai/paperclip/actions/runs/36901592350)
includes all test shards, browser tests, typecheck, build and canary dry
run. Greptile was 5/5 on that commit with no unresolved review threads.
GitHub still requires code-owner review under the repository merge
rules; passing checks do not bypass that approval. Paid source
fingerprints remain separate below and in the dated result note.

- Paid Events qualification passes **3/3**: GPT-5.4 Mini, Claude Haiku
4.5 and Claude Sonnet 4.6. Each uses a real public HTTPS callback,
signature verification and report retrieval in a fresh conversation. A
final Mini regression passes after the quota/status fixes. All evidence
validates. Bounded tunnel startup retries occur before provider calls
and remain visible; failed earlier attempts retain their original
grades.
- The earlier complete seven-case matrix passes **21/21**, with a
separate **3/3** delegation regression. Two preceding matrices also
passed 21/21 each. A complete 24-cell matrix including Events has not
been run. [The dated
results](doc/plans/2026-10-01-public-mcp-paid-eval-results.md) retain
exact source fingerprints, failures, model IDs and partial costs.
- Node 24: repository-wide `pnpm -r typecheck` and `pnpm build` pass
after merging master. Server typecheck passes after the final
quota/status changes. Eval typecheck and all 892 eval-support tests
pass.
- All 33 real MCP/OAuth tests pass. The preceding combined MCP,
redaction, private-address and DNS-rebinding run passed 129 tests; two
later MCP regressions cover quota reuse and unchanged-status
suppression. All 28 adjacent issue-tree/stale-lock route tests pass. CI
then found a null checkout result in the existing concurrent-workspace
path; logging now uses optional status access. All 12 closed-workspace
tests and all 33 MCP tests pass after that correction. The exact-start
event calibration exposed a timestamp gap; scanning now includes the
subscription start, with all 33 MCP tests and server typecheck passing.
These two narrow corrections follow the paid regression.
- A real Core → Cloud → Core authority round trip passes OAuth, MCP 2.0
subscription/delivery, current membership loss, unsubscribe, legacy SDK
tools, refresh and revocation. Its callback transport is a fixture with
independent HMAC verification. The paid Events campaigns separately
prove public HTTPS delivery.
- Earlier component qualification passed UI 7,117 tests, CLI 502, shared
832, skills catalog 20, database 160 and OpenAPI 10. Token gates, module
boundaries, migration order and plugin regeneration passed. CI covers
general/serialized suites, eight browser shards, runner checks,
typecheck, build and canary dry run.
- **Local full-suite limitation:** the earlier monolithic run was not
clean. It encountered overlapping schema rebuilding, Mac database
shared-memory limits and isolated CLI/fixture failures. Targeted reruns
passed. The existing >32 MiB Git filename stress test still hit its
300-second Mac timeout. The additional serialized sweep stopped after 62
passing suites once CI passed. Original failures and partial logs
remain; this PR does not claim a wholly green local monolithic run.
- Local Codex CLI and Claude Code OAuth login and MCP SDK
interoperability were verified. Public-store installation, actual
ChatGPT Work Cloud Events UI, staging HTTPS client behavior and hosted
newcomer provisioning remain release gates.

Enablement is moving to **Settings → Experimental → Assistant
connections (MCP)** in the stacked follow-up
[#14933](https://github.com/paperclipai/paperclip/pull/14933). Merge
both for the intended setup experience. This foundation branch alone
still uses `PAPERCLIP_PUBLIC_MCP_ENABLED=true`. After deployment, set
`PAPERCLIP_PUBLIC_URL` to the authenticated instance's HTTPS origin, and
connect to `/mcp/paperclip`. Select a team and allow writes in browser
consent. Configure an available agent and budget, then delegate and
retrieve results later. For Events, rescan the deployed plugin catalog
in ChatGPT Work Cloud; the host supplies its webhook credentials when
the user asks to watch a task. See [the setup
runbook](doc/public-mcp.md).

## Risks

- Events are at-least-once and may arrive out of order. No replay cursor
is advertised. Clients must refresh finite subscriptions, read current
state and avoid comment feedback loops. Callback material uses the
instance secrets master key; hosted subscriptions require the updated
Cloud broker and are bounded to five minutes/the access proof expiry.
- ChatGPT Work Cloud/dot event UI, plugin rescan and a hosted staging
subscription remain deployment gates. Local signed-webhook and paid
model evidence does not claim those surfaces have been exercised.
- Disabled by default. Merging adds schema and opt-in code; it does not
deploy a public endpoint, publish a store listing, create a team, or
start paid agents.
- Migrations 0301 and 0302 are additive and idempotent. Their SQL is
unchanged from the earlier preview numbers, so hash-aware upgrade
reconciliation preserves prior staging applications. Normal instance
upgrades must apply it before enabling MCP.
- Task creation and comments can schedule paid agent work. Consent and
tool descriptions disclose that effect. Revocation blocks future calls
but does not undo delegated work.
- Public deployments need edge rate limits and credential-safe logging.
Internal dispatch is restricted to the closed catalog and carries a
request-local verified actor.
- Hosted onboarding requires the companion Cloud broker, encryption-key
configuration, and tenant rollout. Self-hosted direct connections can
use this PR alone.
- Store acceptance and agent-mode participation are not claimed.
Checked-in plugin endpoints are development defaults; rebuild packages
for a real deployment before installation.

## Model Used

OpenAI GPT-6 in Codex, with reasoning, tool use, and code execution. A
more specific serving version and context-window size were not exposed
by the session. Paid eval models: `gpt-5.4-mini-2026-03-17`,
`claude-haiku-4-5-20251001`, and `claude-sonnet-4-6`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted/component checks;
full local-run limitations are recorded above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 11:48:53 -05:00
DottaandPaperclip 16b7db35ff Shorten planning skills and measure task decomposition (#15296)
## Thinking Path

> - Paperclip manages work for AI agents.
> - Planning guidance helps agents choose owners and dependencies.
> - The runtime skill favors few tasks, but the catalog skill requires a
child-task breakdown.
> - Both add repeated process instructions that can distract from the
requested outcome.
> - This change keeps the ownership and dependency rules and removes the
required matrix and repeated checklist.
> - A bounded Product E2E comparison measures saved outcomes and task
handoffs before qualification.

## Linked Issues or Issue Description

Refs #11057. Related measurement work: #15218.

**What existing behavior does this improve?**
Planning and delegation through the runtime plan-to-tasks and bundled
task-planning skills.

**Current behavior**
The two skills contain about 1,900 words and conflicting guidance on
whether plans require child tasks.

**Proposed behavior**
Keep cohesive work with one owner. Split only for a real owner, parallel
output, dependency, independent review, or follow-up lifecycle. Preserve
existing authorization and planning mechanics.

## What Changed

- Shorten both skills to about 400 words combined. Preserve their keys
and installed-version behavior.
- Remove the duplicate operational-skill pointer and regenerate affected
source metadata.
- Add twelve explicit Product E2E cells: four scenarios with current,
short and disabled planning skills.
- Use the current task composer and actual create-response ID; calibrate
public skill APIs and browser creation without providers.
- Eliminate an observed collision in chat-test company prefixes with a
per-suite sequence.
- Grade saved documents, exact author/run attribution, child count,
prerequisite execution order, review boundaries and completion handoffs.
- Retain current skill bytes and report source, selections, run
accounting and failures.

## Verification

- `pnpm test:e2e:runner:typecheck`: pass.
- `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks
pass.
- `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve
local Codex cells.
- Archived current skills match master
`72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly.
- Corrected fixture: three real public-API/database calibrations pass
with zero provider runs; all 35 evaluator checks and Product E2E
typecheck pass.
- Setup campaign
[37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253)
was canceled after source review found unsupported bundled edits and
automatic core reinstallation. Its paid-cell step was skipped: zero
provider runs, no behavioral grade.
- The next setup
[37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799)
failed before task creation on the old title-field selector: zero actual
runs, original FAIL retained, cleanup passed. A real browser/API
calibration of the new helper passes with paused non-provider agents and
zero runs.
- Full local typecheck/build pass. Full local tests retain one unchanged
five-minute Git streaming timeout (also fails isolated), 9,591 passes
and 5,796 skips. CI's chat failure was a proven random fixture-prefix
collision; five affected cases pass after the test-only repair.
- Paid behavior comparison and new-head CI/review remain pending. This
PR remains a draft.

## Risks

- The shorter text may change delegation decisions. Live outcomes are
not yet qualified.
- The initial comparison uses one profile and one attempt per cell. It
cannot establish cross-model reliability or cost trends.
- Disabled means unassigned company-owned copies; the company library
remains discoverable. This does not qualify global removal, automatic
accepted-plan wiring changes, or installed-copy migration.
- Skill availability does not prove a model read or cognitively used it.
- No provider/tool protocol, permission, timeout or runtime lifecycle
behavior changes in production.

## Model Used

OpenAI Codex (GPT-6), with repository inspection, code editing and tool
use. The exact backend model ID and context-window size are not exposed
in this session. The declared eval model is native Codex `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 09:36:01 -05:00
DottaandPaperclip a386a59998 Reduce repeated native completion guidance and preserve final replies (#15151)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native agents receive task constraints and completion tools from
Paperclip.
> - Completion tools already define the procedure for reporting a
result.
> - Repeated procedure text adds instructions to each full task turn.
> - The final reply must still explain a blocker and link a saved
document.
> - This pull request removes repeated procedure text and keeps these
visible outcome requirements explicit.
> - A document receipt supplies the exact link, and stricter evals check
the persisted reply and browser navigation.

## Linked Issues or Issue Description

Refs: #14961. Related: #14948 and #15007.

**What happened?**

Native task envelopes repeat completion procedure text. A reduced
envelope needs explicit final-reply requirements. The `write_document`
receipt also lacks a canonical document link.

**Expected behavior**

Keep the completion tools as the source of procedure details. Require
one accepted completion result before the final reply. A blocked reply
must explain the reason, owner and unblock action. A document reply must
contain a working link to the saved document.

**Steps to reproduce**

1. Run the native assigned-skill document case and native blocker case.
2. Inspect the run-attributed provider final and its persisted comment.
3. Check the blocker explanation or open the final reply's document
link.

## What Changed

- Remove repeated completion procedure text from the native task
constraints and backend instructions.
- Keep explicit blocker and document-link requirements in full task
turns.
- Return a company/task-scoped `documentHref` from `write_document`.
Preserve the link in the idempotent mutation receipt.
- Repeat canonical links for this run's current saved revisions in
accepted completion feedback. Give blocked providers final-response
guidance for the cause, owner and unblock action.
- Keep internal document/comment anchors when Markdown issue links load
cached issue details.
- Add a manual six-cell comparison suite with strict source, build,
default-instruction and budget admission.
- Capture eighteen shared runnerd RPC projections and six direct
OpenCode HTTP projections across start, resume and continuation phases,
using scripted local transports and no provider execution.
- Apply v3 checks only to the manual instruction comparison; preserve v2
checks for the existing native completion suite. Check the actual
persisted blocker reason and exact saved-document link. Click the
rendered document link and check the original content marker in the
classic document card or the new document tab.
- Forward exact OpenCode finishing calls through the controller. Wait
for acceptance, keep accepted feedback and concrete rejection text, and
reject malformed responses. Preserve ordinary dynamic-tool response
handling.
- Settle the completion decision and tool response before mapping a
racing idle/error/abort event or handling explicit close/interruption.
Reject a concurrent finishing call before controller admission.
- Add a provider-free regression through real runnerd, the OpenCode
proxy and a fake provider. Reject the first completion, accept the
corrected report in the same turn, and propose one result.
- Keep all original verdicts unchanged. Treat replay under new checks as
separate diagnostics.

## Verification

- `pnpm -r typecheck` and `pnpm build` pass locally.
- Native document-authority tests pass, including company/run
authorization and idempotent replay.
- Native runtime-context, backend and measurement tests pass.
- Final-answer calibration, protocol scoring, source-admission and
catalog tests pass. Wrong reasons, absent links and wrong link targets
fail.
- `pnpm test:e2e:runner:typecheck` passes. Discovery lists exactly six
single-attempt local cells with the declared models.
- Exported `prepareNativeInstructionPreflight` then
`verifyNativeInstructionPreflight` pass on this clean committed source.
They build locally and make zero provider calls.
- Corrective live confirmation is incomplete. Source 3a7349d passed both
Claude and both Codex cases. OpenCode saved the correct document but
omitted its final link; its blocker case was canceled before paid
execution. Preserve this failure. The e171282 confirmation was stopped
during build after fresh review found a completion-settlement race; it
executed zero providers. Source c3e0cb303 fixes that race. Two affected
OpenCode cases await fresh review and one bounded confirmation; earlier
results remain attributed to their original source.
- OpenCode proxy parsing, driver, factory and input tests: 81 pass
across retained focused runs, including six settlement races.
Evaluator/scoring/admission checks: 122 pass. The real proxy regression
passes. Fresh local prepare then verify passes with 18 shared and 6
direct scripted captures, fresh SDK/Rust builds and zero providers.
- The full local suite recorded two failures: a webhook timeout and a
Git-scan load count mismatch. Both files pass in isolation with
unchanged assertions/time budgets; preserve the original failure log.
Fresh c3e0cb303 CI and review are pending. This PR remains draft.

## Risks

- Final-answer wording can vary by provider. The checks cover the
declared release-access blocker and saved document fixture, not general
answer quality.
- A single trial does not establish general equivalence, cause, speed,
cost or live resume behavior.
- `documentHref` is an additive receipt field. It points to the current
saved document, not an immutable historical revision. Replaying an older
receipt does not fabricate a new link.
- The correction adds four production paths for document receipts,
accepted completion feedback and UI navigation, plus four OpenCode
controller/proxy paths, beyond the original three instruction paths.
Completion rejection must remain repairable; the production-boundary
regression covers it.
- Preserve the frozen comparison context for live measurement. A
merge-tree check against current master is clean. Do not relabel earlier
live results as results from a later source tree.

## Model Used

- OpenAI Codex, GPT-6 family. The exact serving model ID and context
window are unavailable in this session. Capabilities used: reasoning,
code editing, shell execution, test authoring and evidence review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 08:10:14 -05:00
DottaandPaperclip b17019e14d fix(agents): reduce default instructions and qualify stock harnesses (#14948)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its adapters supply task context and access to Paperclip skills and
tools.
> - The default hire manual and shared prompts also repeat general work
procedures.
> - Those procedures overlap with stock provider instructions and the
Paperclip skill.
> - Existing E2E fixtures supply a QA manual, so they do not qualify the
production default.
> - This pull request reduces the generic instructions and adds real
default-hire coverage.
> - The benefit is less competing guidance, with inspectable evidence
for preserved skills and task context.

## Linked Issues or Issue Description

Refs: #14920. That merged change preserves native Codex base
instructions. This PR covers the default manual, shared legacy prompts,
operational skill guidance, and the narrowly approved ACP
skill-discovery/session-environment repair for measured delivery and
credential-persistence failures.

**What existing behavior does this improve?**

New non-CEO hires without a custom bundle and legacy task/chat startup
and continuation prompts.

**Current behavior**

The shipped default manual contains 602 words. Generic task/chat prompts
and ordinary resume deltas repeat work procedures already available
through the harness and Paperclip skill.

**Proposed behavior**

The default manual contains only the eight-word company identity. Shared
startup prompts retain identity and connection guidance. Ordinary resume
deltas retain current work context without the generic execution
contract.

**Reason and benefit**

Let the stock harness guide general work. Keep Paperclip-specific
capabilities and independently test default hires, skills, ordered
comments, and chat restart.

**Breaking changes**

New default hires receive less guidance. Existing saved manuals,
explicit custom bundles, CEO templates, and specialized wake contracts
retain their behavior. The obsolete includeExecutionContract option
remains accepted for source compatibility.

## What Changed

- Reduce the default hire manual to one sentence.
- Reduce shared task/chat defaults and remove the generic
ordinary-resume contract.
- Keep connection guidance, auth, skills, custom prompts, and
specialized wake context.
- Add credential-free instruction-boundary gates and 26 explicit Product
E2E cells across eight legacy/native profiles, including two focused
Paperclip-storage cases.
- Capture public hire receipts before providers run, then grade
delivered prompts and independent task/chat outcomes.
- Add an early legacy skill API recipe for saving a task document,
checking the saved revision receipt and linking the document. Improve
stock task/heartbeat skill-selection metadata and show a clickable
Markdown UI-link example. Keep native tool completion separate.
- Advertise bounded routing descriptions and exact successfully staged
SKILL.md paths in legacy ACP Claude; keep full bodies on demand and
preserve remote path rebasing.
- Remove only the provider environment from copied persisted ACP session
records, while loading current run credentials and preserving all other
options/conversation state.
- Regenerate both capability metadata inventories and reject stale
manifests/inventories before provider admission.
- Publish the original reduction and focused skill-repair comparisons,
preserving all failures, automatic recovery, cost coverage and
limitations.

## Verification

**Behavioral qualification remains pending.** Original legacy ACP Claude
loses the issue document only in the reduced cohort beneath an unchanged
credential failure. A source-backed diagnosis finds that neither
ordinary assignment reads the staged operational skill, while the
runtime persists provider environment in session state. The new common
repairs expose skill metadata/path and omit persisted env; strict
document and credential guards stay intact. [Inspectable diagnosis and
retained
hashes](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-readiness.md).

Current repair head `de0965984ff3edf611ae6d0e7ca5c7d5ae3947bb`
incorporates master `569c7203aa24b95440682983ce7940ba1d4247bd` (merged
#14961/#15007). All 222 affected adapter tests, adapter-utils/E2E
typechecks, and final 96 variant/grader/retry calibrations pass. The
frozen historical comparator is
`c25697f4260b6f3adfea143c3ae9932e2f42986d`: 8,280 of 8,291 paths
identical, exactly two production instruction paths plus nine declared
unit expectations differ. The operational skill/discovery/environment
repairs, selected model/profile/task/core grader/auth/permissions/retry
policy are identical. Both actual launcher prepare→verify admissions
pass with zero providers. [Immutable manifest and exact
receipts](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-evidence/manifest.json).

One original legacy ACP Claude cell per variant is authorized, with
enforced single campaign attempts, 12-minute deadlines and company/agent
1,000-cent hard stops; every product recovery run/cost is counted.
Actual live outcomes are pending. Current normal CI has one failed
server shard and failed aggregate verify under diagnosis; other normal
gates including typecheck/build/Rust/all eight browser shards pass.
Fresh review completed successfully; the valid historical startup/resume
masking finding was fixed with per-invocation task/chat checks and
strict complete-snapshot capture, calibrated and resolved. Prior heads,
failures and campaigns below remain historical evidence, not checks on
this repair head.

- Prior head `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` is replayed on
merged hiring master `862a5758ba0e88a33232c1f1fa645e85c38a3113`. All 52
current-head checks pass with two intentional Storybook skips, including
repository typecheck/test/build and the browser shard. Fresh Greptile is
5/5 with zero unresolved review threads. Exact-head stock prerequisites
pass 599 assertions (598 TypeScript + 1 Rust), all six gates and
retained receipt verification, zero providers/source errors. Fingerprint
`a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`.
Combined catalog/hiring calibrations pass 67 assertions, E2E typecheck
and 26-cell stock discovery pass. Canonical contract/inventory checks
and the later issue-derived reference calibration are retained; that
reference-only follow-up is not live-qualified by earlier frozen runs.
- Prior full repository typecheck/build passed. The complete local
Vitest run executed 14,956 tests: 14,870 passed, 83 skipped, three
timing failures. All three affected files passed unchanged narrow
reruns; original failures remain retained. Current-head CI now passes
the full general checks; the original local failures remain retained.
- The original 24-pair default-manual/shared-prompt comparison has two
new overall classic Claude/OpenCode document-delivery failures plus an
additional legacy ACP Claude document loss beneath an unchanged
credential-guard failure (not closed by later runs), two newly passing
OpenCode ordered cases, seven unchanged failures and 13 unchanged
passes. Equal 15/24 totals do not establish behavioral equivalence.
[Complete original
report](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-stock-harness-live-comparison.md).
- The skill-only repair holds the eight-word manual/shared prompts and
merged #14920 fixed. All four matched profile configurations and 203
fixture/behavior files match. Candidate
`abd0b628ca642c09a54a4edc56a5227402f6686e` varies only the two skill
sources against baseline `bc83fe030234439ac51279502a28803958963e2e`.
[Candidate
workflow](https://github.com/paperclipai/paperclip/actions/runs/37060885547)
and [baseline
workflow](https://github.com/paperclipai/paperclip/actions/runs/37060888047)
each pass 571 exact-source prerequisites before providers; all eight
cells clean up successfully. Failed campaigns publish successfully and
remain failed.
- Repair pairs: Claude original Fail → Pass; Claude explicit Pass →
Pass; both OpenCode cases Fail → Fail. Explicit OpenCode's handoff
worsens beneath the unchanged failing UI-link grade: baseline gives a
clickable API URL, candidate gives a code-formatted path without an
anchor. The request's usable-link wording is narrower in the UI-only
oracle. [Complete repair report and safe
projection](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-legacy-document-skill-repair.md).
- The subsequent narrow stock metadata/link correction has two matched
Pass → Pass cases, zero new machine failures/passes and no pending
pairs. Both original-case handoff links remain deficient: candidate uses
a wrong PAP prefix, baseline supplies a bare prefix-less slug path; the
preserved original oracle only requires a durable document. Both
explicit clickable UI-link cases pass revision/content/link grading. All
four exact-source 587-check gates, single assignment runs and cleanup
pass. This does not establish fix causality because baseline also
succeeds. [Candidate
workflow](https://github.com/paperclipai/paperclip/actions/runs/37069547401)
freezes `fe9dc1e3c518825242ed889ab9c8352986f8c2ed`; [matched
baseline](https://github.com/paperclipai/paperclip/actions/runs/37069552374)
freezes `0d7ecfa96d72fba79b7f0a25052b42c0686c0488`. This is a skill-only
comparison with reduced manuals/shared prompts held constant, not a
repeat of the historical-manual comparison. Only original and clarified
explicit classic OpenCode cases are selected, two per variant/four
expected turns. 8,242 other tracked files and both profile hashes match;
protected workflows admit each exact source before credentials.
[Complete qualification
report](https://github.com/paperclipai/paperclip/blob/74d0d3d945f4c52d0814b5a845ab5bd09f33cd6b/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.md).
Candidate original loads Paperclip/reference before saving publicly;
baseline original loads it after writing locally, then saves publicly
within the same assignment. Reported cost totals are $0.0107824490
candidate / $0.0107909015 baseline, with unmetered runtime. The later
reference-only issue-derived link correction is provider-free calibrated
and **not live-qualified** by these frozen runs; no further paid runs.
- Retained tool calls show the repaired original OpenCode assignment
loads only its assigned output skill before writing locally. Operational
Paperclip is first loaded during automatic disposition recovery; its
early recipe is visible then, but it never saves the missing document.
Explicit candidate loads Paperclip and reads the new reference before
saving successfully. All nine actual runs are counted. Reported LLM
totals are $0.3802537209 baseline and $0.4918990161 candidate; local
runtime is unmetered.
- Initial setup, packaging, cancelled/missing-cell recovery, callback
test and relative-output attempts remain retained. No completed provider
failure was rerun. Frozen measurement branches are unchanged by later
canonical metadata maintenance.
- Run `pnpm test:e2e:runner:stock-harness`, `pnpm test:e2e:runner:unit`,
and `pnpm test:e2e:runner:typecheck`. Select `stock-harness` explicitly
for paid execution; it is excluded from `--all`.

Prior-head integration: `36aa4d81c49a1a8f6f04b1a068fae19aa901955f`
replays this PR on merged hiring #14985
(`862a5758ba0e88a33232c1f1fa645e85c38a3113`), preserving the four
explicit custom-CEO-bundle checks, minimal generic manual boundary, and
both suites. The combined fixture catalog and hiring calibrations pass
67 assertions; exact-head stock prerequisites pass 599 assertions (598
TypeScript + 1 Rust), all six gates and retained-receipt verification,
zero providers/source errors, fingerprint
`a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. E2E
typecheck and 26-cell stock discovery pass. Fresh current-head CI passes
all 52 checks with two intentional skips, and fresh Greptile is 5/5 with
zero unresolved review threads.

The prior source-plan browser failure is retained: a deterministic
process fixture replayed its last `fixture:plan` command on
`chat_task_completed`, writing revision 2 with identical body after the
approval handoff. This was not paid provider execution. Rebased
current-head CI passes the same assertion without an old-head retry or a
change to that browser fixture.

The merged hiring change was measured separately on immutable matched
unions, with this reduced/shared/operational context and native
completion guidance held constant. [Complete original two-profile
report](https://github.com/paperclipai/paperclip/blob/f0512647656be78e48abd8c22a3078db8bf6bcd2/doc/plans/2026-10-02-hiring-template-live-comparison.md):
[candidate](https://github.com/paperclipai/paperclip/actions/runs/37075466208)
/ [historical
baseline](https://github.com/paperclipai/paperclip/actions/runs/37075469463),
705 provider-free prerequisites each. Both pairs are unchanged Fail →
Fail on the exact-five count, with six core delivery checks passing all
four cells; 28 actual successful runs include eight automatic completion
wakes, zero retries, four successful cleanups. Source-read coverage is
uncomparable, actual model charges unknown. Separately versioned
provider-free accounting remains analytical work; original verdicts are
preserved. This does not rerun or qualify the completed default-manual
or native campaigns.

## Risks

- Legacy ACP Claude's additional delivery loss is not closed by any
later matched run and blocks the no-extra-failing-behavior merge
criterion. Legacy document delivery may have relied on the prior
manual/shared prompts. The early skill repair improves Claude in one
trial; the later OpenCode pairs pass in both variants and cannot
establish causality or robust recovery. Both original-case links remain
deficient beneath the storage-only grade. The later issue-derived
reference correction has only provider-free validation. Native
finish/block descriptions must not be supplied to legacy agents.
- The comparison holds merged native Codex fix #14920 constant; it
cannot measure that fix's before/after task performance.
- These bounded skill/context/chat workflows do not measure general
coding quality. Unrepresented providers remain unqualified.
- Saved manuals and old Codex sessions are not automatically migrated.
Codex through ACP still has a separate base-instruction follow-up.

## Model Used

OpenAI Codex, GPT-6 family as identified by this session. The exact
deployment ID and context-window size are not exposed. The assistant
used reasoning, repository tools, code execution, and delegated PR/eval
work.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (relevant suites and all
three unchanged narrow reruns pass; complete-run timing failures
retained in Verification)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green on the new repair head
(prior-head checks retained above)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
on the new repair head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 12:32:42 -05:00
DottaandPaperclip 78e0034498 fix(evals): account for hiring completion notifications (#15007)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Product E2E evals check real hiring and delegated task completion.
> - The hiring fixture requires three requested CEO turns and two coder
executions.
> - The server can also wake the CEO when each delegated task completes.
> - Two exact-five-run guards rejected these valid completion turns in
all four retained cells.
> - This pull request validates bounded completion turns in both guards.
> - The benefit is accurate workflow grading while all actual runs and
coverage failures remain visible.

## Linked Issues or Issue Description

Refs: #14985, #14948, #14961.

**What happened?**

The original hiring comparison reports Codex Fail → Fail and Claude Fail
→ Fail. Each cell has seven successful runs. The five requested work
turns are accompanied by two server task-completion notifications. All
six other delivery checks pass.

**Expected behavior**

Require exactly three distinct user-requested CEO turns and one coder
execution for each of two known tasks. Admit at most two strictly
attributed server completion turns, including one turn that batches both
tasks. Reject unknown, duplicate, failed, retried or extra-work runs.

**Steps to reproduce**

Inspect the retained four-cell report linked below. Each original result
fails `five-successful-turns`. The same exact count was also enforced by
the final chat-flow guard.

## What Changed

- Add one typed lifecycle helper shared by the hiring scorer and the
hiring-only final chat guard.
- Validate public run ledgers, company/user/account identity, request
attribution, task origins, completion deliveries, timing and replies.
- Keep exactly five required work turns; declare seven maximum total
turns for cost and timeout planning.
- Count all actual runs, including notification runs and unexpected
resets. Keep other chat count guards unchanged.
- Version the hiring grader as v3 (turn accounting v2) and include the
helper and chat guard in its definition digest.
- Keep source-read, exact coder-body and all six other delivery checks
unchanged.
- Add 144 focused helper/scorer/settlement calibrations and separately
versioned exact retained-input replay reports.
- Retry complete bracketed observations, await both owed callbacks and
attributed replies, and refresh the final guard consistently.
- Reject unrelated completion writes and failed mutation attempts using
exact canonical/native action IDs. Missing identity mapping is
uncomparable action coverage.

## Verification

- All 977 credential-free E2E support tests pass across 64 files,
including 144 focused lifecycle/action/scorer/settlement calibrations.
- E2E typecheck, ordinary plugin SDK and Runner TypeScript dependency
builds, capability contract/inventory checks and the existing two-cell
hiring discovery pass.
- [Executable replay
report](https://github.com/paperclipai/paperclip/blob/fed1729018cc100f5f4bbfb692777e49009c423b/doc/plans/2026-10-02-hiring-executable-accounting-replay.md)
pins current code revision `e4077ade1818d98b9862ae79ee1d49a007dcf9c1`,
v3 definition digest, exact source/input hashes and each original/new
check.
- The stricter replay verifies both Codex variants through both
executable guards. ACPX Claude action attribution remains
unresolved/uncomparable because provider execution IDs cannot be exactly
joined to native request IDs; guards fail closed. No notification writes
are observed. All six other outcomes and every original source/template
coverage check stay unchanged. Original files and Fail → Fail machine
verdicts remain preserved; zero providers are called.
- Full attempts remain uncomparable in both profiles. Historical Claude
also keeps its six-backtick exact-template mismatch. This grading repair
does not prove model-performance equivalence.
- The limited sidecar-v1 and initial executable-v2 passes checked
notification-created tasks but could miss unrelated document writes.
Those assessments remain preserved and do not prove harmless
notifications. The stricter v3 replay is separate.
- [Original measurement and separate
sidecar](https://github.com/paperclipai/paperclip/blob/8eb517ca1497687237163bdef4dfc4d3332ea916/doc/plans/2026-10-02-hiring-template-live-comparison.md)
retain 28 actual runs, eight automatic notifications, four successful
cleanups and unknown actual model charges. No models are rerun.
- The branch is replayed on master `59c07ede7`. Intervening master
changes are UI-only; eval source bytes and replay verdicts match. The
four-cell provider-free replay was repeated against the reachable code
revision.
- Initial-head normal CI retained browser failures in agent-run denial
feedback and touch-picker scroll position. Those browser paths and
imports were unchanged, but their cause was not established. The
necessary review-fix head passes both browser checks; no blind rerun was
requested.
- Local full repository typecheck/test/build were not repeated.
Exact-head normal CI passes the required repository gates, including
typecheck, tests, build and browser shards. Fresh Greptile review
completed on `fed1729018cc100f5f4bbfb692777e49009c423b` with 5/5 and
zero unresolved threads. An independent rerun of the 144 focused
helper/scorer/settlement tests passes on the unchanged head.


**Merge readiness:** This PR repairs the evaluator. Its positive and
negative calibrations pass, both guards reject missing action
attribution, current-head CI and review pass, and there are no merge
conflicts. The retained ACPX cells remain uncomparable because their
action IDs cannot be joined. That coverage limit remains a separate
follow-up; it does not require relaxing this grader or changing the old
results. No model calls, production instructions, carrier changes, or
historical regrades are part of this readiness update.

## Risks

- Missing or inconsistent public lifecycle evidence fails the bounded
helper. The focused calibrations reject plausible false positives and
malformed observations. Unmatched action IDs fail closed and are
reported as uncomparable rather than a model task regression.
- Source-read evidence remains incomplete. This PR does not change
provider event carriers or relax the coverage oracle.
- The versioned count check differs from original v1 results. Reports
retain both versions and exact input hashes.

## Model Used

OpenAI Codex, GPT-6 family as identified by this session. The exact
deployment ID and context-window size are not exposed. The assistant
used reasoning, repository tools, code execution and delegated
calibration work.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip normal CI gates are green (exact head
`fed1729018cc100f5f4bbfb692777e49009c423b`; fresh review tracked
separately below)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(completed exact-head review; zero unresolved threads)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 06:42:51 -05:00
DottaandPaperclip 862a5758ba fix(agents): reduce hiring templates to role descriptions (#14985)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for work.
> - New agents receive role instructions from onboarding, the hiring skill, or a team package.
> - These sources repeat harness procedures and impose generic work policies.
> - They can crowd out the task and the harness instructions.
> - This pull request reduces those sources to short role descriptions.
> - It preserves configuration, skills, authentication, reporting lines, and approval controls.
> - The benefit is less repeated instruction text with explicit coverage for the default hiring path.

## Linked Issues or Issue Description

Refs #3307. The CEO template can impose a fixed delegation route instead of letting the agent choose how to fulfill the request. This change removes that route. It does not implement autonomous goal selection.

Related work: #14920 preserves stock Codex base instructions. #14948 reduces the generic manual and shared runtime prompts. #14961 improves native completion-tool descriptions. This PR is separate from those changes.

## What Changed

- Select only the short CEO `AGENTS.md` for new default CEO bundles. Keep the three former companion files as compatibility assets.
- Reduce the first-agent chief-of-staff prompt and coder, QA, UX, and security role examples.
- Reduce seven bundled team role bodies. Preserve their role, reporting, and skill metadata. Regenerate the catalog.
- Make hiring examples optional. Replace the long generic role manual with short role drafting guidance. Preserve explicit requester instructions.
- Add configuration and import coverage for native and legacy managed bundles, custom instructions, first-agent rendering, and catalog contents.
- Add an explicit-only hiring eval that starts from the production CEO default and checks one coder hire, independently computed JSON output, saved instructions, and worker reuse.
- Include the full prompt comparison and a separate three-request drafting simulation. Neither is a live provider comparison.

Prompt differences: [before and after](doc/plans/2026-10-02-hiring-template-prompt-diff.md). The CEO default falls from 1,897 to 20 words. The coder example falls from 652 to 18 words. Word counts describe instruction size, not outcome quality or billing.

## Verification

- PASS: 99 focused server tests and eight shipped-catalog tests.
- PASS: catalog generation and validation for four shipped teams.
- PASS: hiring skill validation.
- PASS: `pnpm -r typecheck`.
- PASS: `pnpm build`.
- INCOMPLETE: the full local `pnpm test:run` was stopped before rebase. Its original log is retained. This is not a completed full-suite pass. The full current-head GitHub CI workflow passed: https://github.com/paperclipai/paperclip/actions/runs/37073372419.
- PASS: `pnpm test:e2e:runner:typecheck` and `pnpm test:e2e:runner:unit` (63 files / 843 tests).
- PASS: discovery for the two new hiring cells, 50 existing everyday cells, and the full 438-cell catalog.
- PASS after rebase: 99 server tests, 11 catalog tests, 62 selected E2E support tests, and the E2E typecheck.
- PASS: all current-head PR checks at `57dcee147ed0b2d2e3cc657cd9e50fb16bf9ec25`: 51 successful check runs, two intentional Storybook skips, and successful Snyk status. Fresh Greptile is 5/5 with zero unresolved threads.
- PENDING follow-up: matched live hiring runs on frozen integration refs. No live outcome-quality or non-regression result is claimed from the configuration checks or this merge.

The new suite has two local native cells: Codex and ACPX Claude. It expects five provider turns per cell. It compares source-derived bundles, so the historical long templates remain admissible. Missing successful source-read receipts make a pair uncomparable. They do not establish a behavior regression or equivalence.

## Risks

- New default roles have fewer prescribed procedures. Live checks must determine whether a removed instruction was needed for an outcome.
- Existing custom and saved bundles keep their contents. The retained companion assets avoid a source-file compatibility break.
- Specialized Summarizer, Reflection Coach, and Wiki Maintainer prompts remain unchanged. Their product contracts need separate review.
- The generic non-CEO fallback reduction is in #14948. This PR alone does not provide its eight-word fallback.
- Configuration tests and drafting simulations do not establish live outcome quality. QA, UX, security, and chief-of-staff hiring behavior remain outside the new two-cell comparison.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code editing, shell tools, and delegated verification. The runtime does not expose the exact deployment model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context to this change
- [x] I have specified the model used (with version and capability details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before requesting merge

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-02 17:56:42 -05:00
DottaandPaperclip 018993140f feat: let agents name prompt-only tasks (#14761)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users create tasks with a title and a description.
> - A required title adds work when the prompt already explains the
request.
> - An agent can name the task once it reads that request.
> - This pull request accepts prompt-only tasks and starts them with a
short prompt slice.
> - A scoped title tool lets the assigned agent replace that slice early
without changing execution state.
> - A live browser eval checks the real agent call, saved title, audit
entry, and preservation of user titles.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: task creation, shared contracts, database, server, runner
tools, and board UI.

**Problem or motivation**

Users must currently write a title before they can submit a detailed
task prompt. The agent has enough context to write a useful title
itself.

**Proposed solution**

Make the title optional when a description is present. Save the first
120 characters of the normalized prompt as a provisional title. Ask the
assigned agent to call `set_task_title` early. Use an atomic
provisional-title guard to preserve titles supplied or edited by users.
Keep explicit titles supported.

Related: #14543 and #14556 concern empty-title submission. This change
intentionally enables that submission when a prompt is present, instead
of requiring a title.

## What Changed

- Add the `titleNeedsGeneration` field with an idempotent migration.
Keep existing titles unchanged.
- Add `PUT /api/issues/:id/title` and the native and legacy
`set_task_title` tool. Enforce company access, active-run ownership,
shared, bounded retry receipts across native/HTTP calls, and
transactional audit logging. Refresh external-object links after commit,
with the same feature gate and plugin detectors as ordinary title edits.
- Add early naming guidance in Standard, Ask, and Plan task context.
Preserve the description, status, and assignment.
- Allow prompt-only root and child task creation, plus draft restoration
in the New Task dialog. Keep user titles supported.
- Add an opt-in Product E2E suite for prompt-only Standard and Ask
tasks, plus an explicit-title control. It checks actual provider calls
within the first five tools, persisted state, audit attribution, and the
reloaded UI.
- Preserve a closed vocabulary of API key maintenance phrases in
declared prose while rejecting opaque credential suffixes. Add one
bounded naming retry after wording is rejected, without treating the
rejected call as a saved title.
- Repair the native cleanup receipt check exposed during full
verification: accept matching input digests, retain legacy input checks,
and reject conflicting receipts.

## Verification

- Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3
passed** with native Codex `gpt-5.4-mini`, first attempts only,
automatic retries disabled. Standard and Ask each saved “Rotate expired
API key” on their first tool call, with matching persisted state and a
single same-run audit entry. The explicit-title control retained its
user title with zero title writes. All three verified the reloaded
browser UI.
- Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns
are retained separately; they exposed credential-prose handling and
prompted the naming recovery fix. No failed result was regraded or
deleted.
- Reproduce with `pnpm test:e2e:runner -- --id
task-titles.runner-codex-mini.local.prompt-title-standard --id
task-titles.runner-codex-mini.local.prompt-title-ask --id
task-titles.runner-codex-mini.local.preserve-explicit-title
--max-automatic-retries 0` and an authorized provider key.
- Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit.
The runner build used the configured external eval source tree.
- Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck
and UI token gates passed.
- Title API/native regressions cover prompt-only and explicit child
creation, user edits, ownership/company isolation, external reference
refresh, cross-surface retry replay, and the 64-key limit without
receipt eviction. All passed. Prompt-context coverage: **44 tests
passed**.
- Rust credential regressions: **35 tests passed**, including benign
maintenance qualifiers and opaque credential rejection in every declared
prose field. Catalog/report reconciliation: **28 tests passed**. Native
recovery: **560 tests passed**.
- Broad local `pnpm test:run`: **14,555 tests passed** in the general
server group; two suites failed to initialize embedded PostgreSQL and
the existing 40,000-file Git streaming stress test exceeded its
300-second macOS timeout. All three suites then passed in isolation (**5
tests passed**) without code or timeout changes. The original full local
command exited nonzero and is not being represented as a clean full run.
- Latest-head GitHub checks are green: **53 passed, 4 skipped, zero
failed or pending**, including all test shards and the canary packaging
dry run. Greptile reviewed the same commit at **5/5**, with zero
unresolved review threads.

## Risks

- The additive database field must reach the server and UI together. The
migration uses `IF NOT EXISTS` and defaults existing tasks to a final
title.
- Title generation depends on the assigned agent running. Tasks without
a run keep their provisional title.
- Live qualification covers the native Codex path in Standard and Ask
modes. API/legacy and Plan behavior have deterministic coverage.
- The credential-prose exception validates the entire suffix against a
closed maintenance vocabulary. Unknown suffixes, assignments, quoted
values, credential prefixes, and diagnostics retain strict checks.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, tool use, and code
execution. The exact deployment ID and context window are not exposed in
this session. The live eval uses the native Codex `gpt-5.4-mini`
profile.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 16:48:24 -05:00
DottaandPaperclip 3c561642b4 fix(chat): resolve approvals and preserve unanswered questions (#14613)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents ask for decisions and optional details through cards in chat.
> - A clear approval in a message can leave the matching card pending.
> - An unanswered question can also block an unrelated later reply.
> - Decisions need a saved source message, while optional questions need
to remain answerable in history.
> - This pull request records conversational decisions and lets users
move on from questions and answer them later.

## Linked Issues or Issue Description

**What happened?**

Native Claude and Codex could act on approval in chat while the original
approval card stayed pending. Pending question forms stayed above the
composer, were absent from history, and could suppress later chat
replies. A late native question answer could wait for a finished run to
reconnect.

**Expected behavior**

The active agent records a clear approval or refusal against the exact
card and user message. Ambiguous replies do not grant consent. Users can
send another message without answering a question. The question remains
pending in history and can be reopened and answered later. The saved
answer reaches the agent.

**Steps to reproduce**

1. Ask an agent to propose work with a confirmation card, then approve
it in chat.
2. Check that the original card records that approval before work
starts.
3. Ask an interactive question, send an unrelated message, and reload.
4. Open the unanswered question from history and submit an answer.

Related work: #14408 added completion delivery. #14607 tests completion
reporting turns. Neither records conversational answers on approval
cards.

## What Changed

- Add a confirmation endpoint backed by a user comment, with schema
validation, OpenAPI discovery, and native Plan-mode access. Ask mode
remains read-only.
- Check company, active run, actor, current session, message provenance,
revision, and resolver policy. Save the decision and audit in one
transaction. Retries do not repeat effects. Emit resolution telemetry
after commit.
- Give fresh and resumed chat turns the actual pending confirmation
identities. Teach agents to save clear conversational decisions before
acting and to clarify ambiguity.
- Keep unanswered Agent Chat questions as compact history entries. A
newer user message closes the old form. Question cards never contribute
to composer pending counts or navigation, including after dismissing a
fresh form. The history card is the sole reminder; clicking it restores
that exact form and draft.
- Preserve Agent Chat questions when later messages or questions arrive.
Historical ordinary inputs no longer gate later chat replies.
Current-run requests, task execution, and governed approvals keep their
gates. Remove the special acknowledgement-publication proof helpers that
this rule replaces.
- Route answers to finished native runs through durable fresh-wake
delivery, with existing idempotency and source-question context. Settle
late replies against contiguous completed conversation turns and freeze
their history replay; failed, unhandled, and newly arriving messages
remain actionable.
- Add real-component Storybook scenarios, database and UI regressions,
and a three-turn native Claude/Codex E2E case. Capture distinct,
UI-ready screenshots and report the individual assertions.

## Verification

- Focused decision/publication/UI regressions after merging master: 288
passed; subsequent UI draft, failed-send, and conversation checks: 199
passed.
- Native question and durable delivery regressions: 106 passed,
including all four terminal run states and exactly-once late delivery.
Seven targeted regressions fail against the original implementation and
pass with the fix.
- Latest conversation/decision/native-delivery regressions after the
master merge: 121 passed. Covers completed progress, missing or failed
intervening turns, new messages during a late reply, stale sessions, and
frozen retry/replay boundaries. Four new assertions fail before the
ordering fix.
- E2E support suite after the master merge: 792 passed. Negative
controls reject expired cards, wrong questions/answers, stale or missing
replies, unrelated clarification forms, and unexpected tasks.
- The embedded-browser walkthrough caught one additional defect:
dismissing a fresh question still showed a composer badge. Both Cancel
and close-button regressions failed before the fix. The fix at
`65f2ade12` passes 170 chat-thread tests and 792 E2E support tests.
After merging master, 232 chat-thread/confirmation tests, server/UI
typechecks, and token gates pass. The preview and two-provider live E2E
pass at `e5512a206`; Greptile is 5/5 with zero unresolved threads at
that commit. All 55 checks are now successful at `e5512a206` (four
conditional checks skipped), including the aggregate verification gate
and clean-install canary test. The first attempt was interrupted by
simultaneous CI worker shutdowns; one failed-job rerun passed without
code changes.
- [Published
Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fchat-approval-resolution/?path=/story/chat-comments-agent-chat-unanswered-questions--moved-on):
nine real-component scenarios. Manually exercised move on, reopen,
preserve draft, answer later, answer one of multiple questions, and a
custom mobile answer in the embedded browser. Retested fresh Cancel and
close-button dismissal in the updated build, then reopened and submitted
the preserved Green selection and inspected its answered receipt. Static
preview has no live model/backend; its callbacks are fixture responses.
- [First live
campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36714504406-1/)
reproduced the late-answer completion-state defect on both providers
despite correct saved answers and acknowledgements. It also exposed a
valid imperative clarification rejected by the old oracle. Both issues
are fixed with regression controls; this failing run is retained as
evidence.
- [Four-cell
qualification](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36717804064-1/)
passed 4/4 at `2bf8a1009`: unanswered-question return and ambiguous
confirmation, each on native Claude and Codex. Inspected saved state,
source-message decisions, visible cards, and agent replies. Both
late-answer chats settled to waiting; no unrequested tasks were created.
[Final branch
rerun](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36719666238-1/)
passed 2/2 at `142630720`: the same unanswered-question journey after
merging master, plus an additional screenshot and browser assertion for
the actual late-answer acknowledgement.
- [Composer-reminder
E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36727006818-1/)
passed 2/2 at `5b62c52d9`: native Claude and Codex, three turns each,
with explicit no-badge assertions before and after reload. Inspected
saved pending/answered state, both screenshots with a clear composer,
and actual Blue acknowledgements; all five behavioral matchers passed
per provider and neither created tasks. Cost coverage is partial; this
is bounded workflow qualification.
- [Fresh-dismissal
E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36742773318-1/)
passed 2/2 at `e5512a206`: native Claude and Codex, including fresh
Cancel, clear composer, reopen, unrelated message, reload, late Blue
answer, and actual agent acknowledgement. All five behavioral matchers
pass per provider. Inspected the fresh-dismissal screenshots and saved
pending/answered identity; neither created tasks. Cost coverage is
partial (4/6 runs).
- Prior evidence remains available in [the earlier
campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36642252725-1/).
Its early loading screenshot and overwritten final capture prompted the
UI-ready, distinct screenshot fixes.

## Risks

- The model interprets intent. The server verifies permission and
provenance; it does not infer consent from text. Ambiguous and unrelated
replies are not approvals.
- Historical questions can accumulate. They remain visible, pending, and
answerable; no automatic answer or expiry is invented.
- The change to completion gates is scoped to Agent Chat and ordinary
historical inputs. Current-turn and governed approvals retain their
existing controls.
- Live qualification is limited to the selected stories. Broader native
onboarding finalization remains separate work.
- No database migration. Telemetry adds no fields or values; the
contract and README document the commit boundary. Privacy review was
requested on the PR.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, code
execution, and browser-test orchestration. The exact model ID and
context-window size are not exposed to this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 11:46:46 -05:00
DottaandPaperclip d30b03bd8c test: add persistent E2E coverage for human blocker decisions (#14707)
## Thinking Path

> - Paperclip manages work for AI agents.
> - Agents use the coordination skill when work needs human authority or
a scope decision.
> - PR #14188 replaced automatic manager escalation with direct blocker
handling.
> - This behavior needs real browser, server, database, and provider
tests.
> - The test must verify saved human input, task ownership, and resumed
work.
> - This pull request adds six reusable Product E2E cases and improves
the skill examples that they exercise.

## Linked Issues or Issue Description

Refs #14188.

The merged change needs repeatable behavior coverage. The new suite
tests missing administrator access, missing hiring permission, and
requester scope questions. Searches found no duplicate blocker-guidance
suite. This extends the existing eval system described in ROADMAP.md.

## What Changed

- Add the explicit-only `blocker-guidance` Product E2E suite. It has
three local scenarios on legacy Codex and legacy Claude.
- Use the production UI and public APIs to create work, save a
human-only question or confirmation, answer it after reload, and resume
the same task.
- Check requester identity, ownership history, manager activity, hiring,
saved answers, and completion. Keep direct text input as a separate UX
result.
- Save pending and final screenshots, API checkpoints, skill hashes,
provider evidence, and billing data through the existing report
pipeline.
- Isolate the Claude fixture home. Verify the served skill bytes before
dispatch so an old installed skill cannot silently replace the evaluated
skill.
- Improve the coordination and hiring skill examples. Include the
human-only policy, requester address, wake behavior, and handling of
authorized scope changes.
- Grader v5 requires the exact approved public welcome note as a new
worker comment. Browser input checks reject unwritable scope cards
before clicking, and confirmation direction must be saved in the
resolution before the worker wakes.
- Add grader calibration and browser-input tests. Update the fixture
guide and generated capability inventories.

## Verification

- `pnpm build`: passed after rebasing onto current master.
- `pnpm -r typecheck`: passed.
- `pnpm test:e2e:runner:typecheck`: passed.
- `pnpm test:e2e:runner:unit`: 742 tests passed.
- `pnpm test:e2e:runner:browser-support blocker-input.spec.ts`: 10 tests
passed.
- `pnpm test:e2e:runner -- --list --suite blocker-guidance`: six cells
found.
- Capability inventory and generated-contract checks: passed.
- `pnpm exec vitest run
server/src/__tests__/hiring-operational-examples.test.ts`: four tests
passed after synchronizing the generated API reference and section
anchor.
- Full general and serialized test suites: passed in CI on
`6652cee74517039676bad6a720f213625d265acd`. The redundant local `pnpm
test:run` was interrupted after complete CI coverage passed; it is not
claimed as a completed local full-suite run.
- Final GitHub checks: 54 passed, two optional Storybook checks skipped.
The runtime-exposure startup test hit a 10-second readiness timeout
once, passed a targeted local reproduction, and its CI shard passed the
single retry without code changes.
- Current-head Greptile: 5/5, clean check, zero unresolved threads.
- Historical live measurement on September 29 at
`4edc77ae2b95b10dd61426ce3f042bac00527ad9`: three independent six-cell
runs scored 5/6, 6/6, and 6/6. Claude Sonnet 4.6 passed 9/9. Codex
`gpt-5.6-sol` passed 8/9. These runs predate this rebase.
- Version 5 changes the scope answer to an exact approved publication
draft. The historical runs do not qualify that new requirement; the
two-provider scope pilot at `49a1f4eab369948b9e3b34a6ce436489e875e4ec`
passed Codex and failed Claude. Claude posted the correct salary-free
sentence but omitted its required reference line from that comment,
placing the reference in a separate completion message. The
`public-welcome-note` check correctly failed. An earlier Claude
database-startup failure was retained separately; its fresh-instance
retry reached the model. This pilot is not a six-cell qualification.
- The failed Codex scope case omitted `addresseeUserId`. The strict
routing check remains. All 18 attempts had clean evidence manifests and
passed cleanup.
- To repeat with provider credentials: `pnpm test:e2e:runner -- --suite
blocker-guidance --max-parallel 1`. This is a paid, opt-in suite and is
excluded from `--all`.

## Risks

- The live suite measures variable model behavior. The retained 17/18
historical result and the current 1/2 scope pilot are not all-pass
qualifications. These paid cases are opt-in; their observed model
failures remain visible independently of deterministic CI checks.
- A separate generic task-replacement diagnostic still exposed a Claude
refusal. The ordinary cases use specific business decisions. The
diagnostic is not a standalone catalog case in this change.
- Earlier measurements included an old installed Claude skill and test
defects. Their grades remain retained and are not combined with the
three final repetitions.
- Skill examples can affect when agents ask for human input. Downstream
permission checks still apply.
- Native runners, Daytona, agent-requester routing, and real external
connection authorization are outside this suite.
- Raw provider traces and credentials remain private. No screenshots,
raw reports, secrets, workflow changes, or lockfile changes are
committed.

## Model Used

OpenAI GPT-6 through Codex assisted with this change. The exact deployed
variant and context window size are not exposed in this session. The
assistant used reasoning, repository edits, tool use, and shell
execution. The evaluated models were `gpt-5.6-sol` and
`claude-sonnet-4-6`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 07:56:54 -05:00
DottaandPaperclip 83076d7e7c feat: return completed handoffs to Agent Chat (#14408)
Return completed Agent Chat handoffs through a durable outbox and scope each generated update to its supplied tasks. Add recovery, browser delivery, result access, and calibrated quality coverage.

Validated with two consecutive ten-case Claude/Codex campaigns, all CI checks, and a 5/5 review.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 10:25:52 -05:00
DottaandPaperclip 3f4f8b37ba fix: grade Codex clarification and refusal outcomes from evidence (#14570)
Grade clarification lists, obsolete unstarted wakes, and refusal cancellation from persisted evidence. Preserve execution and ownership assertions, add boundary regressions, and version the affected eval definitions.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 09:40:55 -05:00
DottaandPaperclip 24beb00575 feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification.

Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 08:56:21 -05:00
DottaandFry 3ca196b0a6 feat(agents): persist agent files across tasks without revision history (#14420)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - An agent needs personal files across tasks and sessions.
> - AGENTS.md is one file in that directory. Supporting files need the
same persistence.
> - The Instructions Editor and agent runs must share one current
directory.
> - Concurrent runs should apply only the files they change. The last
sync of the same file wins.
> - This pull request uses existing file transport and removes temporary
copies after sync.
> - Old instruction-only sessions keep their restore contract. New saves
do not create revision history.

## Linked Issues or Issue Description

Refs #14325. This replaces its revision-oriented design with persistent
agent files. Keep #14325 unmerged.

Transport prerequisite #14416 merged first at
`d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master
and remains below 100 changed files.

Related work: #4513 and #8798 cover instruction tooling. This change
handles run synchronization, cross-task personal files, browser editing,
and old-session restoration.

## What Changed

- Keep one current directory per company and agent. Point AGENT_HOME at
a temporary working copy for each active run. Keep task files and
provider HOME separate.
- Restore text, binary files, and nested folders through workspace
transport. Exclude remote agent files from task Git snapshots with a
self-ignoring file inside the reserved runtime directory; never write
through repository-controlled Git metadata.
- Collect after the provider and child processes have stopped. Keep
resumable conversation state.
- Apply changed and deleted files under the agent lock. The last sync
wins for the same file. Unrelated concurrent changes survive.
- Remove temporary copies after successful sync, rejected sync, and
staging failure. Register ownership before copying so restart recovery
can remove interrupted preparation. Retry transient synchronization up
to three times. Preserve the original remote lease reference until
deletion succeeds; restart cleanup never acquires a replacement sandbox.
Do not create captured directories or a conflict-review queue for new
runs.
- Keep browser editing, stale-draft protection, and streaming binary
downloads. Keep the instruction entry and text editor limited to 1 MiB.
- Keep historical agent-folder sync failures on their affected runs
instead of repeating them above current saved instructions. Preserve
legacy candidate review and current browser-save errors. Avoid duplicate
quota warnings while retaining separate sync failures when they describe
a different problem.
- Require target-scoped caller grants for peer instruction access, while
preserving self edits, responsible-user checks, and protected-change
consent.
- Treat full storage as a nonblocking run warning, never an agent pause
or run-admission failure. Restore already-over-quota saved folders so
ordinary agent cleanup can recover; warn on each run until cleanup. The
run detail view shows the warning.
- Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash
large files as streams. Check editor-save quotas with metadata instead
of hashing unrelated files.
- Preserve old native inputs, instruction-only copies, paths, digests,
and pending legacy candidates. Adopt old revision heads once. New writes
do not append history rows.
- Add idempotent migration 0287 and verify upgrades from the preview
tables and receipts.
- Add nine interactive stories under **Agents / Persistent files**,
including automatic incoming edits, stale browser drafts, and
storage-limit diagnostics.

## Verification

- Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after
merging current master and the landed transport prerequisite.
Integration required no manual conflict resolution; the feature remains
99 changed files. Full workspace typecheck, production build, token
gates, and 715 focused tests passed on this merge candidate. Fresh
Greptile review is 5/5 with no unresolved findings. All 55 checks
passed, with four conditional skips, including the build, typecheck,
browser E2E, and canary dry run. A single retry recovered four jobs
interrupted by runner shutdowns; no source changes were required.
- Historical-warning UI fix: all 6,834 UI tests across 640 files passed,
including regression coverage for three old failures, legacy preserved
edits, and warnings scoped to the affected run. Full workspace
typecheck, production build, Storybook build, and token gates passed.
Browser-verified Storybook playtests passed for Historical Failures
After Successful Save, Storage Limit, and Full Storage Run Warning.
- Review follow-ups at `4e20c9fb2`: all 18 focused tests passed,
including external Git directories, linked worktrees, symlinks,
hardlinks, and distinct I/O failures alongside storage warnings. Server
and UI typechecks, token gates, and the production build passed.
- Storage warning regressions at `0724f3012`: all 33 directory tests and
all five heartbeat-list tests passed, with no skips in their successful
runs. They cover repeated runs while full, an already-over-quota saved
folder, cleanup, warnings retained after unrelated save failures, and
bounded warnings in large result JSON. Server typecheck passed after the
final warning fixes.
- Full workspace typecheck, production build, and token gates passed
during this follow-up. Product E2E harness: 631 tests passed across 52
files; harness typecheck passed. Earlier native session/context and
directory/legacy collection suites passed 537 tests; Runner
unit/transport suites passed 329 tests.
- **Real E2E at `0724f3012` (before this follow-up):** legacy local
Codex and native Daytona Codex each passed six tasks, one server
restart, seven independent assertions, and cleanup verification. Both
prove browser-to-agent edits, agent-to-browser edits, nested/binary
restoration, per-file last-sync-wins, a successful run after an
oversized save rejection, and cleanup clearing the warning.
- Native local Codex also passed the six-task quota flow before the
final warning-retention fixes. That pass began at `918d1ed02` while the
bounded-result warning fix was being edited, so it is not claimed as
exact-final-head evidence. Its final-head rerun failed during embedded
PostgreSQL bootstrap before any provider run: the macOS host had 87,365
of 87,381 SysV semaphores occupied. No unrelated services or kernel
limits were changed.
- The final-source report intentionally records **2/3 cells passed**,
preserving the blocked native-local attempt:
`tests/runner-e2e/results/agent-files-quota-final-20260928-report/`.
Earlier failed attempts and provenance notes remain under
`tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and
the original campaign directories.
- Daytona used immutable image
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf`
and its exact Linux runner binary. Controller source is `0724f3012`;
image source is recorded separately.
- Legacy-session compatibility and all three ACP Stop/resume browser
regressions passed on the prior validated feature head
`169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same
provider session is retained and interrupted writes are not replayed.
Migration upgrade tests also passed earlier.
- Nine interactive stories are under **Agents / Persistent files**,
including **Full Storage Run Warning**. Its playtest and visual browser
inspection passed; the warning states that runs continue and the editor
remains available.
- Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs
skipped, no failures or pending checks. All eight browser E2E shards and
their aggregate passed. Fresh Greptile review is 5/5 with no findings;
all review threads are resolved, the security scan passed, and GitHub
reports no merge conflicts.
- The broad local follow-up test run was interrupted after host
semaphore exhaustion affected isolated PostgreSQL instances. It also
encountered the existing macOS long-path fixture failure and two timeout
failures. This is not a claim that the full local suite passed. Logs are
retained; focused storage/warning tests passed.

## Risks

- A later sync can overwrite an earlier edit to the same file, including
a saved browser edit. There is no text merge or retained version. This
is the intended last-sync-wins policy.
- A save that exceeds a storage limit is rejected and its temporary copy
is discarded. The run itself continues normally, and later runs restore
the last saved files with a warning until cleanup. Transient sync
failures get bounded retries. An I/O failure partway through a sync can
leave some files updated; a failed receipt does not claim whole-folder
success.
- Larger folders increase copy time, network traffic, and temporary disk
usage. Active runs still need working copies. Terminal runs do not
accumulate archives. Operators must provision disk for agents and
configured concurrency; these limits are not company-wide quotas.
- A restored old native session remains instruction-only until a fresh
session starts. Its original conflict fence and existing pending
candidates remain compatible.
- Provider processes close at the collection boundary. Conversation
resume remains available, but warm process reuse is lost.
- Backups must include the instance filesystem and database. External
bundles keep their existing behavior until explicitly moved to managed
storage.

## Model Used

OpenAI Codex, GPT-6 family. The session does not expose a more specific
model ID or context-window size. Reasoning, code execution, and browser
tools assisted this change. Real provider E2E uses `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing>
2026-09-29 08:25:56 -05:00
DottaandPaperclip 992f720262 fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English
(ASD-STE100). -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task descriptions, comments, continuation data, skills, and
execution rules enter several agent adapters.
> - The same source can be rendered by more than one automatic input
carrier.
> - Failed resumes can also rebuild input from stale or compact context.
> - This pull request gives each Paperclip-owned source one delivery
owner and preserves the required transport boundaries.
> - It adds deterministic adapter, interaction, runner, and browser
tests for these boundaries.
> - The benefit is more predictable context delivery with explicit
evidence for later live qualification.

## Linked Issues or Issue Description

Related: #13144 removes a duplicate environment payload and bounds wake
lists. Related: #11360 addresses Hermes resume behavior. This pull
request preserves compatible active-session formats while repairing
context ownership and stale question creation.

**What happened?**

Task descriptions and comments could enter more than one automatic
context block. Native transports could wrap a complete model input in a
second task envelope. Some legacy and gateway adapters could omit the
owned assignment on ordinary tasks or rebuild a failed resume with stale
compact context. A continuation could also request a question after
newer human comments had arrived.

**Expected behavior**

Each task or comment source has one automatic model-facing owner.
Distinct comment IDs and repeated wording remain distinct. Fresh
fallback attempts rebuild the required full context. A question request
is rejected when newer queued human direction makes it stale. Harness
access policy remains owned by execution configuration.

**Steps to reproduce**

1. Build a task with a description and current comments.
2. Capture the actual adapter or runner input.
3. Compare source ownership and task-envelope nesting.
4. Queue a human comment before a continuation requests a question.
5. Trigger a failed resume and inspect the fresh retry input.
6. Run the focused adapter, interaction, runner, and browser checks.

## What Changed

- Add shared prompt-section selection at the provider-attempt boundary.
- Deliver owned assignment context through native, legacy CLI, ACP,
gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and
Hermes paths.
- Rebuild full or compact context after resume recovery changes the
attempt. Add native and Claude ACP tests of actual recovery requests.
- Preserve custom templates, loaded instruction files, execution
policies, and older active-session formats.
- Record continuation source metadata and reject stale question creation
under the issue-row lock.
- Add explicit Product E2E context-integrity profiles, prerequisite
gates, credential-isolation checks, and report fixtures.
- Bypass service-worker forwarding for same-origin Vite development
modules. A real Chromium test fails with resource exhaustion before the
repair and passes after it. Production asset caching keeps its existing
policy.
- Add browser diagnostics and service-worker module-loading regressions.
- Add an explicit zero-retry eval option. The default retry behavior
remains unchanged. Each campaign records its effective policy.
- Remove the model-facing working-directory sentence from four prompt
builders. Existing workspace, sandbox, permission, and custom-template
configuration remains unchanged.
- Align the everyday workflow assertion with the current 47-entry
catalog.

Compared with current upstream master, the branch carries the
context-ownership implementation and its tests, the explicit
context-integrity catalog and evidence harness, and the focused browser
regression checks.

## Verification

**Merge assessment:** focused regression evidence supports merge. This
is not full completion of the original broad qualification matrix. The
maintainer has authorized merge after fresh verification of the master
integration.

- Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This
integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`.
All 14 conflicts are resolved. Cancellation checks, workspace
finalization, native Grok support, and both sets of tests are retained.
- Current-head Greptile: **5/5**, with no blocking findings. The review
names this exact commit. All **59 reported checks are terminal: 55
successful, 4 skipped, zero pending or failing**. This includes the full
root general and serialized suites, separate runner checks, typecheck,
build, canary, browser E2E, Docker, and security checks. The successful
legacy security status is included in that total.
- After integration: workspace typecheck and full build passed. Separate
runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust
tests, and 39 preparation checks**. Other passing checks include 621
Product E2E harness units, 376 focused shared/adapter tests, 160
real-database/API tests, 86 Hermes tests, 18 browser-support checks, and
Product E2E typechecking. The complete root suite passed in CI. The
duplicate local monolithic root run was stopped after that CI result; it
is not counted as a completed local pass.
- New native recovery coverage retains full assignment, completion
contract, and explicit skill selection after safe replacement, for old
and prepared input formats. Full native session test file: **136/136
passed**.
- New Claude ACP coverage captures actual fresh, resumed, and
missing-session fallback requests. It verifies one assignment copy,
comment order, identical text under distinct comment IDs, and full
fallback context. Full file: **33/33 passed**. Both affected TypeScript
checks passed.
- Existing deterministic tests cover source revisions, approval and
trust boundaries, completion validation, custom templates, compatible
sessions, standalone driver wrapping, and maintained adapter transport
requests.
- Provider-free browser support: **17/17 passed** after the master
merge. Service-worker unit tests: **33/33 passed**. The module-overload
regression failed before the repair and passed after it in real
Chromium.

### Fresh live comparisons

The new batch ran exactly four Product E2E attempts. **All four passed
on the first attempt; no retries.** Each has six terminal matchers plus
the existing browser lifecycle and invariant checks.

| Exact case ID | Control | Candidate |
|---|---|---|
| `core-compatibility.runner-codex.local.plan-revise-accept` | Passed |
Passed |
|
`local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume`
| Passed | Passed |

The plan case checks a revised canonical plan and revision-bound
approval before completion. The question case restarts the server before
submitting the answer, then verifies the continuation completes.

Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate
source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical
frozen definitions and provider versions: Codex `0.156.0` with
`gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with
`claude-sonnet-5`. The September 24 head added master browser recovery
and test-only changes. The September 28 head also integrates newer
master changes, including cancellation, workspace finalization, and
native Grok. These are frozen-source live results, not exact-head live
runs.

The candidate received one description copy where the control initially
received three. The submitted initial plan envelopes were 7,969 versus
19,097 characters. Question envelopes were 7,592 versus 18,919. These
are structural measurements, not whole-provider token or dollar savings.

### Earlier evidence and failed attempts

- The preceding fresh batch has four effective passing pairs: OpenCode
comment continuation and assigned skill, native Codex comment
continuation, and native Claude comment continuation. It retains **11
attempts: eight passed and three failed**.
- Original failures remain recorded: missing local PostgreSQL library
links before task creation; host-sleep cleanup after task/page checks
passed; and a Claude **control** session-open rejection before a model
turn. Setup was repaired identically on both worktrees. The permitted
unchanged infrastructure retries passed. The underlying Claude provider
startup error was not retained and remains unknown.
- Older R2 retains **17 passes and one failure** across 18 attempts,
including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode
blank-page failure led to the service-worker repair. R2 is historical
evidence: master changed the native fixed prompt and removed duplicate
wake environment data afterward.
- The September 24 CI run initially failed one unrelated preview
readiness test (`ECONNREFUSED` on its local fixture). Its test and
production code match master. Isolated local verification passed **28
tests, 3 skipped**. One unchanged CI retry passed the full shard: **831
passed, 1 skipped**, including all **31 preview-exposure tests**. The
aggregate CI gate passed afterward. The precise startup cause remains
unknown; a port race is a hypothesis, not a proved cause.

### Limits

The original wider profile/workflow matrix, repeated trials, and remote
Daytona qualification are incomplete. These results support a focused
merge recommendation, not statistical equivalence or universal harness
qualification. Some usage receipts are missing in both variants, so no
token or dollar savings are claimed. The $500 ceiling was preserved
using conservative allowances; failed attempts and unknown charges
remain in the ledger.

Reproduce the focused additions with `pnpm exec vitest run
packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm
--filter @paperclipai/paperclip-runner exec vitest run
src/native-session-runtime.test.ts`. Full checks use `pnpm -r
typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner
checks. Paid evals require the frozen definitions, profiles, and
credentials; do not use `--all` as a substitute for the selected cases.

## Risks

- Context placement changes can affect model behavior. Deterministic
checks cover the selected paths, but live qualification remains
incomplete.
- The stale-question guard can reject a request when queued human
comments arrived during the run. This is intended.
- New stored inputs and model envelopes retain compatibility readers for
older active sessions.
- Custom templates may intentionally repeat content.
- Removing a model-facing working-directory sentence does not change
filesystem, command, sandbox, or permission configuration.
- The worker bypass applies only to same-origin development module
paths. Cache-policy tests preserve private-response handling and
production asset caching. Mounted HTTP fixture changes remain test-only.
- This PR does not claim measured token savings or statistical
equivalence across every harness.

## Model Used

OpenAI Codex, exact model gpt-6-astra, with repository tools and code
execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The
serving context-window size is not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR using the required issue fields
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run the focused local checks and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect these changes
- [x] I have considered and documented risks above
- [x] All current-head Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
for the current head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:49:14 -05:00
DottaandPaperclip f1a394bd30 feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path

> - Paperclip manages AI agents and governs their work.
> - Its native runner uses structured provider protocols for sessions
and tools.
> - Grok Build supports ACP over stdio, but the runner did not expose
it.
> - Native execution requires company-scoped credentials, verified
identities, and permission gates.
> - This change adds Grok through ACPX for local and Daytona execution.
> - Subscription login and explicit API-key execution have separate
credential paths.
> - Qualification grades real tool outcomes, durable state, and browser
workflows.

## Linked Issues or Issue Description

Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979.

Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`,
`acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents
keep their adapter. Merge the three companion fixes (#13973, #13977,
#13979) before treating the integrated Product qualification as deployed
behavior.

## What Changed

- Synchronize shared, TypeScript, Rust, server, validation, and UI
provider contracts.
- Run Grok native ACP stdio through ACPX and the authenticated Paperclip
MCP bridge. Verify the pinned executable and exact ACP model identity.
- Prefer company subscription login. Support an explicit company-secret
API key without automatic paid fallback. Fence refresh and copyback to
the same account and remove private runtime credentials after
containment.
- Preserve selected permissions, cancellation, durable session identity,
resume, and restart recovery. Keep unsupported steering and goals
unavailable. Preserve missing usage and cost as unknown.
- Package checksum-verified Grok Build 1.0.13 for Daytona with an
immutable, signed image built on EC2.
- Add deterministic admission, protocol, permissions, identity,
credential, failure, and cleanup checks. Add the maintained 39-case
protocol roster and separate subscription/API Product profiles.
- Fix live-test findings in reasoning events, reloads, idle-owner
retirement, credential-home cleanup, expired-login model discovery,
launcher pinning, and rerun evidence selection.
- Align control-plane state readers with the transport's 64 MiB bound
while retaining identity, ownership, lifecycle, and size rejection
checks.
- Stabilize two asynchronous CI assertions while retaining actual
outcome and filesystem-evidence checks.

## Verification

Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)).
Trunk code-owner requirements remain enforced. The review summary’s
non-blocking saved-asset offset classification note concerns code
already merged in #14301; those runtime files are identical to master
and outside this stack’s diff. Historical live evidence below retains
its original source revisions.


Earlier integration checkpoint:
`24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master
`0f14d2612`, preserving Grok qualification alongside the new accounting
and lifecycle suites. All 124 focused catalog, evidence, and
service-worker checks pass. The current base workflow includes the
explicitly selected public-install verification lane; follow-up #14024
supplies its verifier script. CI at that earlier checkpoint was green
(56 successful checks/statuses, four intentional skips), and the review
is 5/5 with no unresolved findings. Prior feature CI at
`fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run
36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902));
that is historical evidence, not a current-head result.

Paid Product measurements use frozen integrated source
`2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature
with #13973, #13977, and #13979. That source passed all 52 CI checks and
clean 5/5 review. Later master syncs incorporate upstream changes. Their
checks remain separate from these pinned live measurements.

| Check | Result and source-pinned report |
| --- | --- |
| Subscription protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html),
runtime `bc6833f7`, evals `92bb4b8c` |
| API protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html),
runtime `4a1061c8`, evals `3213dbec` |
| Subscription full Product matrix | [16/16 first attempts; 144
assertions; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html),
source `2d939a92` |
| Subscription core repetitions | 18/18: tool use, planning approval,
and Stop/resume each passed three times in local and Daytona profiles.
The full matrix contains repetition one; [repeat
two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html)
and [repeat
three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html)
each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. |
| API smoke and question continuation | [4/4 first attempts; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html),
both environments at `2d939a92` |
| Historical API Product coverage | [16/16 full
matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html)
and 18/18 core repetitions at `4a1061c8`; retained as measurements of
that revision |
| Native Daytona proof | Three subscription and three API
MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login
admission and fenced refresh checks passed without inference. All test
sandboxes were removed. |
| Inspectable artifacts and UI | Current-source screenshots verify
planning approval, direct Ask completion, question continuation after
controller restart, and two downloadable project revisions. The project
downloads pass 12 and 18 tests; all 40 independent artifact oracle
checks pass. |
| Provider-free checks | 116 eval-validator tests, 39 Grok definitions,
and 359 enabled/external campaign cells pass. Continuation regressions
above 2 MiB and 16 MiB failed before their fixes; 32 focused
recovery/ownership/size checks pass. |

The 32 unique current-source Product attempts have no failures, retries,
or skipped cells, and all cleanup checks pass. Whole-workflow timing,
model identity, image and provider-pack provenance, attempts, and
accounting coverage are retained in the canonical reports. The report
publisher's conservative `complete=false` flag is preserved; independent
audits verify the exact selected source catalog and immutable result
rows.

Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model
`grok-4.7`. Linux binary SHA-256:
`edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`.
Launcher SHA-256:
`f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`.
Image:
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`.
Image build source is `4196a4cd`, recorded separately from application
source `2d939a92`; each campaign verifies the image signature and
provider pack.

Original failed campaigns remain available: [continuation
bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059),
[scheduler/event
capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537),
and [startup cleanup plus EC2
interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743).
They retain their original grades. No Docker or Rust builds ran on the
developer laptop for these follow-ups.

## Risks

Merge packaging follow-up #14024 with this base before public release.
The follow-up replaces the private Grok bridge package with a built-in
launcher and makes the native binary an explicit sandbox prerequisite.

Three separate, reviewed fixes are part of the tested integrated
behavior: #13973 serializes task-run admission; #13977 captures complete
event evidence; #13979 durably reconciles failed Daytona creation. Each
has green CI and clean 5/5 review. Failed-create recovery has 277 plugin
tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona
lost-deletion-receipt proof. The live proof uses a private file for
journal persistence; database durability is covered by host tests.
Worker death before delivery of a failure envelope remains outside that
recovery mechanism.

Subscription fixtures stage an authorized company login; interactive
browser sign-in is not qualified. Local Product profiles ran on EC2
Linux. The temporary subscription credential was removed from the
protected GitHub environment after all subscription audits, with absence
verified. Runtime homes and refresh copyback remain ownership-fenced.

Protocol results remain pinned to their original revisions; they are not
relabeled as tests of the latest feature commit. New binary/model
versions require qualification. Missing token usage and model cost
remain unknown; runtime estimates do not establish a full bill.
Automatic paid Grok scheduling remains disabled pending separate
reviewed enablement. The 64 MiB bound can increase memory use for
verbose sessions, and larger files still fail closed. No automatic
legacy-agent migration occurs.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:54:36 -05:00
DottaandPaperclip 3447609d22 fix(runner): stream and page large API responses within capture budgets (#14301)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.

## Linked Issues or Issue Description

**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.

**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.

**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.

Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.

## What Changed

- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.

## Verification

- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.

## Risks

- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 10:36:32 -05:00
DottaandPaperclip cea8dda472 test: evaluate completion updates after native task handoffs (#13969)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users can delegate work through onboarding and Agent Chat.
> - A completed task does not prove that its result reached the original
conversation.
> - Existing tests do not isolate completion after the source chat
becomes idle.
> - This pull request adds four explicit native-runner probes across
Claude and Codex.
> - The probes preserve the result and reply so we can separate delivery
failures from inaccurate answers.

## Linked Issues or Issue Description

Refs #13775. Refs #13813.

These evals extend native-runner qualification. They measure completion
updates before we choose a product change.

## What Changed

- Add the opt-in `completion-updates` suite with two stories for each
native provider.
- Test completion in the existing onboarding task flow and after an
Agent Chat handoff becomes idle.
- Gate the chat worker on a brief inside its managed project workspace.
Prove the source is idle before releasing the worker.
- Check durable task completion, saved output, a subsequent source
reply, and rendered access to the result.
- Preserve replies, task state, screenshots, run events, and a separate
semantic review rubric.
- Add grader regression tests and update the documented eval contract.
- Preserve the suites added on master and include four completion cases
in the 306-cell catalog.

Production behavior and prompts are unchanged.

## Verification

- Passed all 565 eval support tests across 45 files after merging
current master: `node node_modules/vitest/vitest.mjs run --config
tests/runner-e2e/vitest.config.ts`.
- Passed eval TypeScript: `node node_modules/typescript/bin/tsc -p
tests/runner-e2e/tsconfig.json`.
- Confirmed four selected cells: `node cli/node_modules/tsx/dist/cli.mjs
tests/runner-e2e/launch.ts --list --suite completion-updates`.
- Four-cell behavior campaign on source
`ad47cf1da2b1e36f19f4227cfeb53998720b0b5b`:
https://github.com/paperclipai/paperclip/actions/runs/36072337485.
- A screenshot-only follow-up waits for the restored source reply to
render after result-link navigation. Its one-cell Claude onboarding
verification passed on final head:
https://github.com/paperclipai/paperclip/actions/runs/36075716141.
Corrected report:
https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36075716141-1/.
The original four-cell onboarding screenshots caught navigation loading;
its saved reply evidence remains valid. The follow-up again found stale
wording: "That work will run next" was posted 38 seconds after the child
was Done. The four-cell campaign keeps its original source and
measurements.
- Published evidence:
https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36072337485-1/.
- Suite definition:
`afba4d85d6c53d9f64c08b37a2e9cc20481b78f5bd7e2fa012045e2c69444d9d`,
version 6. Models: native `gpt-5.6-sol` and `claude-sonnet-5`, local
execution, one attempt per cell. All four cleanup checks passed.
Onboarding billing coverage is partial; reported zero cost must not be
read as a free run.

| Story | Automated delivery/access | Separate semantic review |
| --- | --- | --- |
| Codex onboarding | Pass | Pass: accurate completion reply with an
accessible result |
| Claude onboarding | Pass | Fail: reply says it will save the note once
the task runs, after the note is already saved and the task is Done |
| Codex idle chat handoff | Fail | Worker completed and saved the note;
no completion reply during the full observation window |
| Claude idle chat handoff | Fail | Worker completed and saved the note;
no completion reply during the full observation window |

Both chat cases positively recorded the source waiting and the worker at
the brief gate before release. Both saved outputs include the brief-only
start time. The opt-in campaign is red because it exposes current
behavior. It is not a required merge gate. The PR does not fix that
product behavior. Semantic review is a recorded human/agent assessment
of retained evidence; it is not an automated prose-quality judge.
- Second campaign:
https://github.com/paperclipai/paperclip/actions/runs/36071065098. Codex
chat reached the idle boundary and completed its task, then received no
completion reply during the full window. Claude onboarding again
returned a stale handoff answer. Claude chat exceeded the prior
110-second handoff setup budget; this revision raises that bounded setup
window to 180 seconds.
- Retained baseline:
https://github.com/paperclipai/paperclip/actions/runs/36069427676.
Onboarding passed delivery/access for both providers, but Claude gave a
stale handoff answer. Chat cases stopped at fixture problems; they do
not establish a completion-delivery failure. This revision fixes the
workspace path and competing reference requirements.
- On the previous head `4023a2a3c28d45c9eb2c42d452ce99ffba5c7b73`, 54 PR
checks passed and two were skipped, including typecheck, tests, and
build. Broad checks ran in CI, not locally. That head received Greptile
5/5 with no unresolved findings. The unchanged mobile
repository-settings browser test passed on one targeted retry after a
detached/disabled Save-button timeout.

- Merged current master in `9b4491e1f` and resolved the catalog-count
conflict. Eval support tests and eval TypeScript pass locally. All
individual CI jobs passed on this merge commit, including build,
typecheck, server tests, runner checks, and browser shards. The final
aggregate check also passed: 54 checks passed and two were skipped.
Greptile reviewed this exact commit at 5/5 with no unresolved findings.

## Risks

- These explicit probes can expose current product failures. They do not
change the default paid test selection.
- Mechanical delivery and result access do not establish answer
accuracy. The preserved reply still requires semantic review.
- A fixture failure before the idle boundary or worker completion cannot
establish a completion-update failure.
- The handoff setup window lasts three minutes. The worker brief wait is
bounded at four minutes. The observation window lasts two minutes after
worker completion. It retains later replies without erasing earlier
accessible delivery.

## Model Used

OpenAI Codex, GPT-6 (`gpt-6-astra`), with reasoning, repository
inspection, code execution, and GitHub tool use. The runtime does not
expose the context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 09:58:14 -05:00
DottaandPaperclip 96bf004a79 fix: use persisted state for lifecycle continuation and retry budgets (#13888)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its control plane decides when a task can continue, wait, stop, or
complete.
> - Legacy continuation could change when an agent changed its wording
without changing task state.
> - Shared attempt counts also let repair and infrastructure retries
affect each other's limits.
> - This pull request uses persisted state and separate, bounded
allowances for these decisions.
> - If automatic repair stops, the task explains what happened and
offers a guarded retry.
> - Paired tests and real-provider evaluations verify that Stop,
approvals, ownership, and spending limits remain authoritative.

## Linked Issues or Issue Description

Related work: Refs #13761, Refs #11126, Refs #13610. These cover
obsolete continuation dispatch and retry storms. Open and closed issues
and PRs were searched for related lifecycle, continuation, and retry
work.

**What happened?**
Legacy continuation depended on English wording and progress heuristics.
Repair, failure retry, and productive continuation could consume shared
counts. When bounded repair stopped, the task showed a technical
recovery message without a clear next action.

**Expected behavior**
Persisted disposition and owned execution paths determine the next
action. Missing disposition prompts bounded agent repair. Explicit work
mode determines planning mode. Narrative changes and raw activity counts
cannot replenish allowances. An exhausted repair shows a readable
notice. An explicit retry checks current controls and preserves the
assigned agent.

**Steps to reproduce**
Run `pnpm test:lifecycle-baseline`. The paired probes keep structured
state constant while varying completion, planning, blocker, and progress
prose. Run the explicit `lifecycle-baseline` and
`continuation-accounting` Product E2E suites for real-provider coverage.
In Storybook, open **Design previews / Recovery notice** to inspect the
production component's normal, pending, acknowledged, unavailable,
failure, and mobile states.

## What Changed

- Hide the image attachment button, icon, and drop/paste hint in answer
composers. Image paste and drop support remains available.
- Merge current master and retain both browser regression sets. Use a
production-stamped service worker in the offline recovery browser
fixture.
- Share one state-based legacy continuation decision across immediate,
delayed, and recovered dispatch. Bind bounded repairs to their source
run and episode.
- Remove title and description wording from work-mode authority. Agents
can still write requested plans in execution mode.
- Persist separate failure-retry and productive-continuation counters.
Disposition repair and resource waits cannot consume or reset those
allowances.
- Validate delayed repair identity, then recheck current gates before
provider dispatch. Fence native startup cancellation.
- Show **Agent needs attention**, a plain-language explanation, **Retry
agent**, and expandable details in both task interfaces. Report request
progress, acknowledgement, and errors inline.
- Store typed recovery notice metadata. Recognize older active notices
only through exact stored action and run IDs. Notice text never grants
retry authority.
- Use the existing recovery-action endpoint for retry. Recheck current
action, status, owner, agent availability, dependencies, active runs,
pending questions and confirmations, approvals, pause controls, and
budget. Duplicate requests do not wake twice.
- Add component, page, route, database, contract, and Storybook
coverage. Keep the scenario inventory and executable evals here.
Historical reports and snapshots live in the [commit-pinned
paperclip-evals
archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md).
- Preserve unsaved project fields while the same project URL changes to
its canonical alias. Do not reuse data across projects or companies.
This separate fix addresses the repeated repository-editor browser
failure without changing the browser test.
- Keep the development service worker from intercepting Vite module
reloads. Update the connection-intent browser fixture to record progress
and completion through the agent API.

## Verification

Merge preparation on September 25, commit
`c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`:

- Merged master `bd2030932` and resolved the browser test-list conflict
by keeping both sets of regressions.
- Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no
failures, skips, or missing selected evidence. Unit 423, runner 184,
database integration 397, grading 86.
- Browser support: 17/17 passed. The offline recovery test first failed
with an unstamped development worker, then passed with the production
stamp. Its assertions are unchanged.
- Focused interaction UI and offline fallback tests: 19/19 passed.
Verified the custom-answer composer in Storybook: no attachment controls
or hint; entering an answer enables Next.
- Recursive typecheck, production build, token gates, and diff checks
passed. The worktree is clean. No new real-provider campaign was run.
- Current CI and review: [Current PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243):
55 successful checks and two optional Storybook skips. Greptile scored
this exact commit 5/5. Hiding the question attachment controls is an
intentional UI change; paste/drop remains available.

Earlier recovery UI verification, commit
`21be0fec0e90e86b6d662b8ee4831847cd041cdb`:

- Recursive typecheck, production build, token gates, and diff checks
passed.
- Focused UI coverage: 338 tests passed across six suites (336 before
the interaction guard, with the two affected suites rerun at 149 passed
after it). Covers both task interfaces, the real page mutation,
pending/error acknowledgement, stale state, and unavailable controls.
- Recovery database integration: 352 tests passed before the interaction
guard. The complete recovery-action and mutation-route suites passed 181
tests after it. The two new pending question/confirmation regressions
failed before the fix and passed afterward, including
resolved-interaction controls. Shared validator suite: 31 passed. E2E
catalog suites: 34 passed.
- Browser inspection passed for light/dark themes, mobile layout,
expandable details, pending retry, acknowledgement, failure, and
disabled retry. Storybook renders the production component; its request
is simulated.
- The broad local run hit two chat callback-order wait failures and was
stopped after all CI unit/database/runner shards passed. Both local
failures passed when rerun without the competing full-suite process.
- CI exposed a repeated project-repository draft-loss race during
canonical redirects. A new unit regression failed before the fix; all
nine project-page tests now pass, including controls for other projects
and companies. Both unchanged repository browser tests passed against a
fresh local server. UI typecheck, production UI build, and token gates
passed after this fix.
- [Earlier PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798)
on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two
optional Storybook skips, and no failed or pending checks. The
repository browser shard passed with the production fix. Greptile is 5/5
on this exact commit with no unresolved review threads. The PR is
mergeable.

Historical, source-qualified lifecycle evidence:

- Lifecycle baseline: 1,074 assertions. Native session coverage: 447
tests. Product E2E support: 515 tests. Browser support: 11 tests. Full
earlier verification is retained in the archive.
- [Real-provider campaign: 8/8 passed, zero
retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html),
source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup
checks passed. This includes deliberately exhausted repair cases that
correctly remain blocked; it does not mean every task finished Done.
This campaign predates the recovery UI change.
- Archive migration verified all 16 original JSON files byte-for-byte
and all 24 checksum entries. App tests do not need private archive
access. [Archive PR
#27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged.

## Risks

- Agents that omit durable disposition receive at most two repair
attempts by default. Prose-only completion exposes missing state rather
than silently changing scheduling.
- A retry is an explicit board action. The server rechecks current
controls. A successful response confirms the task returned to To do; it
does not claim that the provider has already started.
- Existing notice metadata remains valid. Only older active notices with
matching structured evidence receive the new UI. Historical notices
without that evidence keep their existing rendering. No schema migration
is required.
- Old run records require conservative retry accounting. Tests cover old
counters, alternating retry lanes, restarts, and exhausted repairs.
- Historical snapshots require private `paperclip-evals` access. The app
index retains public campaign links. Live campaigns qualify specific
sources and scenarios; no new real-provider campaign has run for the
recovery UI commit.

> This fixes existing lifecycle and recovery behavior and does not
duplicate planned core work.

## Model Used

OpenAI GPT-6 through Codex assisted implementation, reasoning, code
execution, and review. The exact serving model ID and context window are
not exposed in this task. Historical real-provider evaluations used
Codex model `gpt-5.6-sol`, separately from the implementation assistant.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 15:28:11 -07:00
DottaandPaperclip 5842185e4f fix: surface native cleanup quarantine and add chat qualification evals (#13775)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agent Chat needs reliable native execution before native runners
become the onboarding default.
> - Existing stories covered idle reassignment and controller restart,
but not an executing worker handoff or worker process loss.
> - Status answer tests also need to reject stale claims and invented
facts.
> - This pull request adds six opt-in full-stack cells with independent
state assertions and retained evidence.
> - The probes exposed a misleading Retry across server projection and
recovery-banner paths; the fix reports the blocked recovery honestly.
> - The tests preserve failures without changing recovery policy,
production prompts, or onboarding defaults.

## Linked Issues or Issue Description

Refs: #13762. Related: #13765 (Retry targets the latest failed attempt),
#13753 (task context ownership), #13746 (native recovery work).

## What Changed

- Add active reassignment with saved draft and plan preservation,
old-worker cancellation, and successor completion checks.
- Preserve recovery-needed projection when native cleanup fails before
its coordinator exists, refuse a generic retry that would immediately
fail again, and replace the recovery banner's misleading Retry with
Inspect run.
- Add verified local worker process loss with a required successful
continuation; retain a failing qualification result when recovery is
unavailable, while independently verifying the UI/API refuse doomed
retries.
- Add two-turn factual answer checks for current blockers, stale claims,
inactive backlog work, and unknown facts. Retain prose for separate
semantic review.
- Add positive and negative oracle calibration and document fault
isolation, cleanup, billing, and qualification limits.

## Verification

- Eval TypeScript check passes.
- All 442 eval support tests pass locally. The 89 focused server tests
and server typecheck pass. Six recovery-banner UI tests and token gates
pass.
- Initial new-cell campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657128077. All
six results are retained; four failed on fixture-contract issues and two
exposed real worker cleanup quarantine.
- All 26 existing native onboarding cells:
https://github.com/paperclipai/paperclip/actions/runs/35656761484 (26/26
passed on master 846336e5a, all cleanup passed).
- Intermediate handoff/fault campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657945095 (four
retained failures: two overly strict draft oracles, two real crash
quarantines).
- Final active handoff:
https://github.com/paperclipai/paperclip/actions/runs/35659014397 (2/2
passed on cf6d4ae3a; both cleanup passed).
- Clarified answer-quality fixtures:
https://github.com/paperclipai/paperclip/actions/runs/35658262695 (2/2
passed on 4a26f10be; both cleanup passed; all four answers semantically
reviewed).
- Quarantine guard regression campaign:
https://github.com/paperclipai/paperclip/actions/runs/35659580100 (both
API requests correctly refused with 409/no second run, but exposed a
separate misleading Retry in the recovery banner and a fixture wait on a
non-admitted run).
- Final quarantine guard verification:
https://github.com/paperclipai/paperclip/actions/runs/35661067305
(147e42f7e: both providers verify Inspect run/no Retry, HTTP 409, one
retained run, unchanged saved plan, and successful disposable cleanup.
Both evals intentionally remain red with
`worker_crash_recovery_unqualified`; no successful continuation exists).
The preceding campaign 35658772755 never ran provider cases because
GitHub artifact finalization returned HTTP 403.
- Full repository CI passes on 147e42f7e: typecheck, tests, build, and
browser gates. One unchanged local-service-supervisor readiness test
failed initially; its six-test file passed in isolation and the failed
shard passed on its single rerun. Latest-head rollup: 54 successful, 2
intentionally skipped, no failed or pending checks. Greptile is 5/5 with
zero unresolved findings.
- See tests/runner-e2e/QUALIFICATION-2026-09-21.md for retained attempts
and semantic review. Published reports:
[onboarding](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35656761484-1/),
[handoff](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35659014397-1/),
[grounded
answers](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35658262695-1/),
[crash guards and unqualified
recovery](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35661067305-1/).

## Risks

- Paid cells are explicit-only and local-only. The fault fixture signals
only the exact native run PID after checking its identity.
- Live worker-loss probes currently fail on cleanup quarantine for both
providers. The eval must remain red until there is a usable recovery,
even when preservation and refusal checks pass. Verified cleanup with a
fresh attempt versus exact-session resume remains a product decision.
- Structured facts alone do not qualify prose quality; semantic review
remains separate.
- Onboarding uses the existing runtime switch after the real wizard and
before provider execution. Native UI selection and public defaults
remain unchanged.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact served snapshot and context-window size are not
exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 18:13:01 -05:00
DottaandPaperclip 846336e5a0 test: harden agent chat setup, interruptions and restart evals (#13762)
## Thinking Path

> - Paperclip lets people manage agents through ongoing conversations.
> - Chat users can change instructions while a provider is already
working.
> - Existing chat evals wait for each turn to settle before the next
message.
> - They cannot prove delivery during active work or the saved effect of
a correction.
> - Existing fixtures also enable Agent Chat through the API rather than
the settings UI.
> - This PR adds bounded browser workflows and checks their persisted
outcomes.

## Linked Issues or Issue Description

Refs #13741, #13752, #13750.

**What happened?**

The chat suites cover planning, delegation, status, and recovery. They
lack active-turn follow-ups and the experimental settings lifecycle. A
sequential conversation can pass even if messages sent during work are
lost.

**Expected behavior**

A follow-up submitted during a provider turn survives and affects the
final reply. A changed launch day appears in the saved plan. Disabling
Agent Chat rejects new messages while preserving history; re-enabling
resumes the same conversation.

**Steps to reproduce**

Run the explicit `agent-chat-stories` suite. It selects three local
cases for each native Claude and Codex profile. An ordinary provider
command waits for a fixture brief file so the browser can send the
follow-up at an observed active-run boundary.

## What Changed

- Add six opt-in Product E2E cells for settings, active follow-ups, and
plan corrections.
- Drive experimental settings through the UI and verify disabled sends
are rejected by the public API.
- Use a bounded file wait in the actual isolated agent workspace, with
provider-written readiness and an undisclosed brief reference.
- Grade persisted user messages, final replies, native run outcomes, and
exact saved plan fields.
- Accept active-turn steering or one queued successor; reject lost
input, duplicate input, and stale outputs.
- Allow one steered run or two sequential runs throughout the shared
harness, while preserving exact counts for other cases.
- Require a single marker-bearing response attributed to the final
provider run.
- Unload the development browser client before restarting the server,
avoiding reconnect/navigation races without weakening the post-restart
memory check.
- Add browser regressions for restart isolation and asynchronously saved
settings switches.
- Document prepared-agent setup, native onboarding limits, and the
separate API-tool rollout gate.

## Verification

- Eval TypeScript check passed.
- Eval support suite: 436 tests passed in 39 files.
- New oracle calibration: six tests passed, including plausible invalid
outcomes.
- Browser support regressions: seven tests passed; the restart
regression was observed failing before the fix.
- Catalog discovery selects exactly six local native cases and leaves
default paid selection unchanged.
- [Consolidated existing native chat
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35643286055-1/):
master `b82661b56`, 33/34 passed, all cleanup passed. The failure was a
browser navigation timeout across restart; the page request returned 200
and the chat rendered.
- [Nine targeted restart/replay
cells](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35645850088-1/)
passed on `1fe2fe275`, including the original failure, across native
Claude/Codex and local/Daytona; all cleanup passed.
- [Initial six-story
campaign](https://github.com/paperclipai/paperclip/actions/runs/35644832817)
retained all six failures: asynchronous switch assertions, unavailable
fixture paths, and rich-text escaping in raw command comparisons. The
corrected fixtures preserve the same behavioral assertions.
- [Six-story campaign
v2](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35646270035-1/)
on `8232773a0`: 4/6 passed (both settings cases and both Claude
interruptions). Codex could not see the host-temp fixture outside its
workspace; this failed before follow-up delivery was exercised.
- [Four affected interruption
cases](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35647760635-1/)
all passed, including cleanup, on definition v3 / `ad6ac0545`. Files
live inside the actual agent workspace and the observed run workspace is
verified. Both providers saved Friday in the real plan with the
undisclosed brief reference; follow-ups persisted while the original run
was active. Together with both unchanged settings cases from v2, all six
new scenario variants have passing live evidence.
- Final head `ad6ac05456646c09d3452e320279457625353948`: 54 successful
checks, two intentional skips, zero pending/failing checks; mergeable
and clean. Fresh Greptile 5/5, zero unresolved findings.
- Full typecheck, tests, build, and browser CI passed remotely. One
earlier head encountered a signoff-policy browser timing failure; the
final head passed that shard.
- Local pnpm wrapper could not fetch its version/signature metadata in
the restricted environment; local eval checks used the installed Node
executables. Repo-wide validation was completed by GitHub Actions.

## Risks

These are eval-only changes. The file wait is a timing fixture in the
isolated agent workspace, not a production runner hook. Native Codex
host-filesystem isolation stays unchanged. It has a two-minute limit and
is released in `finally`. The prepared-agent settings case is not full
native onboarding: the wizard currently offers legacy adapters. The
disabled-entry assertion uses full document navigation, which clears the
prior React Query cache; preserved history is checked through the public
API and re-enabled chat. No production prompt, rollout default, adapter
behavior, or credential policy changes. Active-task reassignment and
worker-crash recovery remain outside these new cases.

## Model Used

OpenAI Codex, GPT-6, with repository tools and code execution. The exact
deployment model ID and context window are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 16:18:06 -05:00
DottaandPaperclip 9d19f98b50 fix: harden native chat recovery and add coordination evals (#13741)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - Agent chat uses native runner sessions to plan, delegate, and track
that work.
> - A user can press Stop while the native session is still starting.
> - The server can acknowledge that Stop without dispatching it, then
let the session submit a turn.
> - This leaves chat recovery waiting for an execution that the user
expected to stop.
> - This PR waits for the startup handle, dispatches cancellation, and
prevents a late startup from submitting a turn.
> - New full-stack evals check the resulting records and outputs across
Claude and Codex.
> - Those evals also exposed missing ACPX readiness fields, unbounded
polling, and an old-run identity check that rejected valid warm
handoffs.

## Linked Issues or Issue Description

**What happened?**

Stop during native startup could record an acknowledged cancellation
with `dispatched: false`. The provider could then begin work. A
subsequent `/new` stayed queued. A remote Claude follow-up also
exhausted the command journal while probing warm-session readiness: ACPX
never returned the readiness fields required by the shared transport.
Once readiness worked, attachment incorrectly compared the next run
descriptor against the old run ID. The 25 ms polling loop could issue
4,800 commands during its two-minute wait, beyond the 500-command bound.
The existing chat eval treated lifecycle logs as proof of an active
provider turn, so it did not distinguish startup cancellation from
active-turn cancellation.

**Expected behavior**

A Stop during startup must reach the pending session. A late session
must not submit a prompt after Stop. Recovery must retain control when
startup exceeds the bounded wait. Chat evals must check saved task
state, document contents, worker identity, account binding, and
duplicate effects.

**Steps to reproduce**

1. Start a native Claude or Codex chat turn.
2. Press Stop after process startup is requested but before the provider
turn starts.
3. Send `/new`, then send a fresh message.
4. On the affected base, cancellation can be acknowledged without
dispatch and the reset stays queued.

**Paperclip version or commit**

The live Claude baseline reproduced this on `29d6b3509`. The branch also
includes master commit `0f5fafe16`.

Related work: #13678, #13686, #13693, #13291, #13738. A separate runner
reliability branch also contains a startup-wait fix. Its overlap must be
reconciled before merging; this branch additionally prevents prompt
submission after a late startup.

## What Changed

- Wait for a pending native startup before acknowledging a run-scoped
Stop. Preserve the existing recovery error when that wait expires.
- Keep a Stop guard on startup. Cancel a late handle before it can
submit a provider turn.
- Add regression tests for normal handle publication and publication
after the Stop deadline.
- Back off blocked warm-attachment probes. Keep the fast two-snapshot
barrier, fail closed, and record changed blockers.
- Add red/green tests for delayed readiness, persistent blockers,
alternating readiness, and readiness near the deadline.
- Publish ACPX readiness and blockers. Preserve the old authority’s
event acknowledgement barrier; only settled sessions can proceed to
attachment.
- Bind warm ACPX descriptors to the validated next authority while
retaining old-run event correlation until activation. Preserve session
identity and provider profile checks.
- Exercise two consecutive run rotations through a qualified fake
sidecar, verifying checkpointing, provider identity, pre-activation
rejection, and new-run work admission.
- Separate startup and active-turn cancellation checkpoints in the
browser eval.
- Add 18 explicit native chat eval cells: 12 local and 6 Daytona cells
across Claude and Codex.
- Cover hiring and reuse through managed AI accounts, source-based
review, current blocked-task status, request replay after a lost HTTP
acknowledgement, server restart continuity, and Stop/reset continuity.
- Use ordinary production agent instructions. Enable API tools only for
the two coordination cases that need them.
- Calibrate the matchers with invalid records and outputs. Require
remembered context after restart and a structured status snapshot that
distinguishes the current blocker from history and task status from
active execution. Compare the public issue mutation contract and
relationships during read-only reporting. Preserve before/after source
records in failed eval evidence.
- Fix the lost-ack browser harness and verify it against a real HTTP
server. Check the chat composer after restart instead of waiting for an
unrelated document lifecycle event.
- Document the scope and limits of each case.

## Verification

- The startup regression failed on the unfixed executor and passed after
the fix.
- `pnpm test:e2e:runner:typecheck` passed.
- `pnpm test:e2e:runner:unit` passed: 424 tests in 37 files.
- `pnpm exec vitest run
server/src/services/native-runtime/native-session-executor.test.ts`
passed: 385 tests.
- [Baseline live
campaign](https://github.com/paperclipai/paperclip/actions/runs/35608208868):
Claude Stop reproduced the bug. Codex Stop and Claude hire/reuse passed.
Codex delegation was blocked by provider capacity.
- [Eval-only startup
campaign](https://github.com/paperclipai/paperclip/actions/runs/35609479786):
both providers failed as expected. Both persisted `dispatched: false`
and left `/new` queued.
- [First fixed
campaign](https://github.com/paperclipai/paperclip/actions/runs/35610533706)
on `c9e95797d`: 10/18 cells passed. Startup Stop passed for both
providers. Failed cases exposed eval harness defects and remote
continuity failures. All attempts remain available.
- [Original workflows and stronger memory
checks](https://github.com/paperclipai/paperclip/actions/runs/35611896649)
on `c04324fab`: 9/12 passed. Reassignment, local restart memory, and
startup Stop passed for both providers; Codex remote restart passed.
Claude remote restart exposed the missing readiness contract. Two Codex
planning cells hit provider capacity.
- [Unchanged-model
retry](https://github.com/paperclipai/paperclip/actions/runs/35613854548):
Codex planning and backlog creation both passed.
- [18-cell campaign with ACPX
readiness](https://github.com/paperclipai/paperclip/actions/runs/35614586963)
on `6a98ef743`: 16/18 passed, including all local/remote Stop and
committed-send cases. Claude remote continuity exposed the
next-authority check, now fixed. Codex hiring produced its checklist,
but the runner redacted the requested marker after it appeared as
“Tracking token: …”. That content-redaction policy is unchanged and
remains an explicit limitation.
- [Structured status
grading](https://github.com/paperclipai/paperclip/actions/runs/35614954725)
on `50448c228`: both providers passed on their first attempt, including
cleanup.
- [Complete read-only state
grading](https://github.com/paperclipai/paperclip/actions/runs/35616089011)
on `551e13892`: both providers passed.
- [Final ACPX handoff and hiring
retry](https://github.com/paperclipai/paperclip/actions/runs/35617045456)
on `cbd637587`: all three Claude Daytona cases passed (restart
continuity, active Stop/reset, and lost-ack replay). Codex hiring
reproduced the content-redaction failure: the saved checklist contained
`Tracking token: [REDACTED]` instead of the required business marker.
All four cases completed cleanup successfully. [Published
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35617045456-1/).
The only subsequent commit adds the qualified-sidecar integration test;
production code is identical to this live proof.
- `pnpm test:e2e:runner:browser-support` passed: 5 browser tests without
paid models.
- Runner TypeScript typecheck passed. All 5 warm-readiness tests pass;
two failed with the prior fixed-rate loop, and the late-readiness test
failed before the pacing correction.
- ACPX readiness and warm-identity regressions each failed before their
fixes. All 292 runner-core Rust library tests passed. The
qualified-sidecar integration test passes. Rust formatting is checked.
- Status-grader regressions for misleading historical mentions and
previously unchecked mutations each failed before tightening the oracle
and pass now.
- [Latest-head
CI](https://github.com/paperclipai/paperclip/actions/runs/35617522307)
passed on `a4093c8f1`: full build, type checks, test partitions, browser
E2E, and native runner checks. Two unrelated tests initially failed
(Sentry fixture release attribution and local-service fixture
readiness); both passed locally together (35 passed, 5 optional SDK
tests skipped) and on the failed-job retry. No changes were made to
those tests.
- Greptile reviewed `a4093c8f1` at 5/5; both earlier findings are fixed
and all review threads are resolved.
- The paid live suite is not fully green: the reproducible
content-redaction case remains red. This is separate from the passing PR
merge checks. No production content-redaction, prompt, model, or
completion-policy change is included.
- Managed-account hiring and review cases explicitly enable API tools;
these do not qualify default new-user onboarding.

## Risks

- Stop can wait up to 30 seconds for startup, then use the existing
pending-recovery path. This does not prove that remote cleanup has
finished.
- Blocked warm readiness adds up to 750 ms between later probes with the
two-minute remote budget, or about 32 ms with the default five-second
budget. Ready sessions retain the short second barrier.
- Paid evals can fail because of provider capacity or agent decisions.
Each failure needs evidence-based classification.
- The HTTP request replay case checks comment idempotency and duplicate
effects. It does not prove replay safety for an ambiguous provider tool
call.
- The new suite is opt-in. It does not increase the default paid
campaign.
- No production prompts or model selection change. Review-handoff
behavior and content-redaction policy remain separate product decisions.
The latter can remove harmless business content that looks like
credential syntax; the failing attempt is retained.

## Model Used

OpenAI Codex, GPT-6, with repository tools and code execution. The exact
deployment model ID and context window are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 10:35:38 -05:00
DottaandPaperclip 18989a9e73 docs: add eval guide, authoring skills, and public history hub (#13535)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its evaluations test both the Runner and complete product workflows.
> - The guides and run histories are in separate places.
> - The shared Evalbook viewer can make the test boundary unclear.
> - This pull request names the two families and adds a guide, authoring
skills, and a public hub.
> - Contributors can choose the correct test and inspect its history.

## Linked Issues or Issue Description

**Issue type**

Missing documentation.

**Where is the issue?**

Runner and Product E2E evaluation guides, case-authoring procedures, and
public result navigation.

**What's wrong?**

There is no single entry point. A report format can be mistaken for an
execution boundary. There are no dedicated case-authoring skills for
these two families.

**Suggested fix**

Add a guide and three skills. Link both existing histories from a public
hub. Keep existing campaign URLs and grading unchanged.

## What Changed

- Add `doc/evals.md` and links from existing guides.
- Add the `paperclip-evals`, `add-runner-eval`, and
`add-product-e2e-eval` skills. Install copies in
`~/paperclipai/.agents/skills`.
- Add a static hub builder that reads the existing public history feeds.
- Show a dated snapshot for each family. Label partial campaigns and
preserve measurement dates across report refreshes.
- Document publication and refresh commands for
https://pages.paperclip.ing/evals/.

## Verification

- Seven Python summary tests pass: `python3 -m unittest discover -s
scripts/evals-hub -p 'test_*.py'`. Run these checks directly; this PR
does not modify package scripts.
- All three skills pass the skill-creator `quick_validate.py` check with
`/usr/bin/python3`.
- Build tested with saved history fixtures and the live public feeds.
- Desktop and mobile browser checks pass. The mobile page has no
horizontal overflow.
- Published https://pages.paperclip.ing/evals/. Browser check: HTTP 200,
no page errors, all eight links return HTTP 200, no mobile overflow.
- Independent skill exercises found the existing Notion-decline case and
a direct Runner permission-denial case. Roster validation with an
explicit run ID passes.
- Missing refresh measurement date: regression fails before the fix and
passes after it.
- `git diff --check` passes.
- No paid evals were run for this documentation and reporting change.
The preceding head passed typecheck, build, server/workspace tests,
runner verification, browser E2E, and the canary dry run. Checks for the
latest commit are pending. Local repository-wide typecheck, test, and
build were not repeated because no product code changed.

## Risks

The hub is a dated static snapshot. It can lag behind the linked
histories until an operator refreshes it. A changed history schema stops
the build. Existing archives and grades are not modified. The published
guide link is pinned to the reviewed commit so branch deletion cannot
break it. Later builds can use master.

## Model Used

OpenAI gpt-6-astra for implementation and review. OpenAI gpt-5.6-luna
for documentation and independent skill checks. Both used repository
tools and code execution. Context window sizes are not exposed by this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (latest commit pending)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(preceding head was 5/5; latest commit pending)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 08:41:06 -05:00