Commit Graph
2018 Commits
Author SHA1 Message Date
DottaandPaperclip 7f452cbd75 Align the second registry fixture with the package-local probe boundary
Carry the existing downstream probe import and partial mock, preserving all original exports and readiness assertions. All 112 affected server tests pass. The other prerequisites and readiness branch already have this exact test blob, so no shipping or downstream source changes are needed.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-05 14:01:46 -05:00
DottaandPaperclip 89f4448509 Complete existing package-local Runner probe export boundary
Carry the companion Runner and server shim exports with the already-forwarded import repair. Full Runner TypeScript typecheck passes, and the real shim resolves all three source function identities without mocks or probe calls. Final shipping source remains unchanged.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-05 13:47:26 -05:00
DottaandPaperclip 32b5e275d5 Carry existing Pi admission assertions and package-local probe import
Keep Pi source assertions consistent with the held qualification candidate. Use the existing vendored runtime boundary for installed probes. All four changes already exist downstream; no final shipping input or profile pin changes.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-05 13:37:35 -05:00
DottaandPaperclip 89284e2905 feat(runner): prepare held Pi 1 production admission
Replay the inactive Pi-only admission patch on the frozen Pi 1.0.0 source, preserving profile 11, its exact digest and all native closure inputs. Cursor and Copilot remain gated. This local preparation requires complete qualification and rebuilt normal-mode acceptance before activation.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 15:45:57 -05:00
DottaandPaperclip 5eba61ece9 Merge recorded master baseline into Pi 1 production candidate
Integrate 8ec4b84e1c before runtime qualification, preserving the new workspace restore lock, continuation and Docker packaging fixes. Pi profile 11 source and distribution inputs are unchanged by this merge.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 15:17:17 -05:00
DottaandPaperclip 3ab022d7df test: align merged directory and continuation fixtures
Remove a duplicate service import, register the warm remote fixture leases required by the cleanup ownership guard, and assert the server-owned bounded continuation for an unauthorized unfinished response wait. Keep runtime implementation and qualified inputs unchanged.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 14:48:54 -05:00
DottaandPaperclip 8ec4b84e1c fix(chat): resume messages after failed runs without duplicate delivery (#14857)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A user can send a new message after a native run fails.
> - The server checks that the old execution has stopped before it
starts a fresh turn.
> - A failed run can retain a result accepted before checkpoint or
cleanup failed.
> - The continuation gate treated that saved result as active recovery
and held the new message forever.
> - This pull request removes that false liveness signal while retaining
controller, process, environment, and authorization checks.
> - Live staging then exposed a second defect: chat admission created a
successor without consuming the original deferred receipt, so completion
delivered the message again.
> - Consume that exact receipt atomically with admission, while
preserving separate turns for later chat messages.

## Linked Issues or Issue Description

**What happened?**

A new user message stayed in the queue with `controller_settling` after
the previous run had reached `terminal_failure`. The old coordinator had
no lease owner but still had a `resultId`. Its remote environment had a
verified stop receipt.

**Expected behavior**

Start one fresh turn after execution has stopped and normal admission
checks pass. Preserve the failed run and its accepted result as history.

**Steps to reproduce**

1. Accept a native result, then fail checkpoint or cleanup and exhaust
recovery.
2. Retain the result ID on the terminal failure record and stop the
execution environment.
3. Send a new user message. Before this fix, it waits forever for the
finished controller.

**Paperclip version or commit**

Reproduced in a database-backed regression test on `26900655b`.

**Deployment mode**

Server with a native runner and remote sandbox. Local process stop
checks also apply.

Related: https://github.com/paperclipai/paperclip/pull/14775. Searched
existing PRs for retained-result continuation fixes; no duplicate found.

## What Changed

- Remove the retained-result veto for terminal failures.
- Keep controller ownership, successor, process, environment cleanup,
pending decision, and ordinary admission checks.
- Add regressions for retained results, active execution, missing stop
evidence, and delayed remote cleanup.
- Atomically consume the resumed receipt in agent chat, even though chat
does not coalesce other queued messages.
- Reproduce completion-time duplicate promotion, race cleanup against
periodic recovery, and prove a subsequent chat message keeps its own
turn.
- Document that a saved result does not make a terminal failure active.
- Keep exhausted workspace export on its separate repair path, tested
through the production finalizer.

## Verification

- Red: retained-result admission failed with `controller_settling`
before the original fix. The new chat-specific regression then
reproduced duplicate promotion when the first reply finished.
- Green: 406 tests across native continuation, workspace-export
recovery, and the wake-queue module passed on `cbc531cc0`.
- The chat regressions exercise real Postgres transactions, simultaneous
recovery callbacks, successful completion, the production queue-drain
use case, and repeated drain attempts. A distinct follow-up remains a
separate turn.
- `pnpm -r typecheck` and `pnpm build` passed on `cbc531cc0`.
- The earlier full local test run encountered a timeout and follow-on
failure in unchanged AI connection-adoption tests; all 50 tests passed
on isolated rerun. That local run was stopped after the full CI test
matrix passed on the earlier head.
- All 54 CI checks passed on `cbc531cc0` (2 skipped), including the full
test matrix and browser shards. One unchanged interaction-route test
returned HTTP 500 on its first CI attempt; its full 84-test file passed
locally, and the failed shard passed on one targeted rerun.
- Greptile reviewed `cbc531cc0`: 5/5, no unresolved findings.
- Live staging first verified that the original saved message resumes
and receives a successful response; that test exposed the duplicate now
covered above.
- Deployed exact commit `cbc531cc0410e1ef6e8811c6c5c014c3528351ed` to
the affected staging workspace; deployment verification, health,
authentication, and startup recovery passed.
- Submitted a fresh message through the browser. The agent replied in 39
seconds; server records show exactly one successful run, native phase
`committed`, no error, and an empty queue. A later check more than a
minute after completion found no duplicate run.

## Risks

The change affects admission after native execution failure and
consumption of a resumed deferred receipt. A fresh turn must never
overlap the prior execution, and consuming one chat receipt must not
absorb later messages. Tests retain the controller, process, and
remote-stop guards. This change does not migrate data, apply an old
result, or reset the old retry budget.

## Model Used

OpenAI Codex (GPT-6). The exact runtime model identifier and context
window are not exposed in this session. Used reasoning, repository
inspection, code execution, database-backed tests, and browser
inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 14:43:02 -05:00
dependabot[bot] 8f7baf2f72 chore(deps): bump @aws-sdk/client-s3 from 3.1122.0 to 3.1141.0 (#13475)
Bumps
[@aws-sdk/client-s3](https://github.com/aws/aws-sdk-js-v3/tree/HEAD/clients/client-s3)
from 3.1122.0 to 3.1141.0.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/aws/aws-sdk-js-v3/releases">@​aws-sdk/client-s3's
releases</a>.</em></p>
<blockquote>
<h2>v3.1141.0</h2>
<h4>3.1141.0(2026-09-25)</h4>
<h5>Chores</h5>
<ul>
<li><strong>codegen:</strong> smithy-aws-typescript-codegen 0.54.0 (<a
href="https://redirect.github.com/aws/aws-sdk-js-v3/pull/8314">#8314</a>)
(<a
href="https://github.com/aws/aws-sdk-js-v3/commit/ad80ce3ebaf394679aabc6e26b2dcd023ce8e010">ad80ce3e</a>)</li>
</ul>
<h5>New Features</h5>
<ul>
<li><strong>client-connect:</strong> Agent Privacy During Hold is a new
privacy capability for Amazon Connect Voice that prevents agent audio
from being captured in call recordings or Contact Lens conversational
analytics during hold. When enabled, agents are automatically muted on
entering hold and unmuted on resuming the contact (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/03527f9ea153365c1e3654ac6d3f3e064d06b5d0">03527f9e</a>)</li>
<li><strong>client-qconnect:</strong> Release shapes for the proactive
agentic recommendations and the multi-knowledge base search features.
Increases the maximum length of QuickResponseContent. (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/e008332b0b70c55d22dfca8a9c2317b67d06a5f3">e008332b</a>)</li>
<li><strong>client-bedrock-agent:</strong> Adds support for calling VPC
configuration API's in Bedrock. These configurations allow the use of On
Prem connectors in Bedrock Managed Knowledge bases (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/18524dc69cf36ebbb8bc7bc33e0bce6311dcdb23">18524dc6</a>)</li>
<li><strong>client-mediaconnect:</strong> This release adds support for
RTMP push router outputs in AWS Elemental MediaConnect. (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/5bd8d80bbca2c1c2545521b03d741b46fccd09b4">5bd8d80b</a>)</li>
<li><strong>client-securityagent:</strong> This release adds the
ListActorMessages operation, which returns the multi-factor
authentication messages received at an actor's server-generated email
address (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/10e53d506db74c48b09b94d9b9387b84dd40ed58">10e53d50</a>)</li>
<li><strong>client-arc-region-switch:</strong> Adds a service quota
checker to Region switch to verify quota parity between your primary and
standby Region, and automatically submit quota limit increases. Adds an
optional EC2 Auto Scaling and ECS setting that waits for instances or
tasks in the scaled-up Region to be healthy in target groups. (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/cca33e380a8985e23a0e3fbb420577aab4b60ac7">cca33e38</a>)</li>
<li><strong>client-bedrock-agentcore-control:</strong> Amazon Bedrock
AgentCore Payments now supports credential rotation for payment
connectors, letting you rotate API and wallet secrets for Quick Create
payment auths from the console. This release also adds Type and Creation
type columns to the payment managers views. (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/adca591f07c448603de2876faaf7914cad441436">adca591f</a>)</li>
<li><strong>client-neptune-graph:</strong> Add GraphIdentifier filter
for ListImportTasks (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/9b9aea9b553ba84f006f95da5d8ffd5381cc8a17">9b9aea9b</a>)</li>
<li><strong>client-rekognition:</strong> This release adds support for
Feedback and Metadata in the GetFaceLivenessSessionResults response.
Feedback returns codes explaining why a Face Liveness check produced its
result. Metadata includes the client SDK type. (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/0831c361bb69d44357ff57360db39afdbb149337">0831c361</a>)</li>
<li><strong>client-glue:</strong> add support for table level federation
(<a
href="https://github.com/aws/aws-sdk-js-v3/commit/a44458b77853cbb25a9fcb362b0a275d7dc1c69b">a44458b7</a>)</li>
<li><strong>client-wellarchitected:</strong> This change releases the
Well-Architected Agent, a generative AI service that analyzes a
customer's AWS environment and delivers personalized, prioritized
recommendations across cost, security, performance, and resilience. (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/d0656586067f70b8d50000300f04512dd089df96">d0656586</a>)</li>
</ul>
<hr />
<p>For list of updated packages, view
<strong>updated-packages.md</strong> in
<strong>assets-3.1141.0.zip</strong></p>
<h2>v3.1140.0</h2>
<h4>3.1140.0(2026-09-24)</h4>
<h5>Documentation Changes</h5>
<ul>
<li><strong>client-route53resolver:</strong> Documentation updates for
Route 53 Resolver. Clarifies which Outpost Resolver operations apply to
first-generation AWS Outposts and that Resolver is managed automatically
on second-generation Outposts. Adds Local Network Interface subnet
compatibility notes for Resolver endpoints. (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/b4432abaaaf19bf4ac5d4d2b09bcc83e4d5440a0">b4432aba</a>)</li>
<li><strong>client-iot:</strong> Fixed ListV2LoggingLevels and
DeleteV2LoggingLevel documentation to include all supported target-types
(<a
href="https://github.com/aws/aws-sdk-js-v3/commit/4dcf76d527e0c1fe76d64cb8ea444ce62d64135b">4dcf76d5</a>)</li>
</ul>
<h5>New Features</h5>
<ul>
<li><strong>clients:</strong> update client endpoints as of 2026-09-24
(<a
href="https://github.com/aws/aws-sdk-js-v3/commit/29a8566cb4c6eeb0cc554f4ae9bf985160556523">29a8566c</a>)</li>
<li><strong>client-eventbridgev2:</strong> Introducing Amazon
EventBridge enhanced Custom event bus, a new shareable event bus for
organizational-scale event-driven applications feature ordered delivery,
deduplication, open event formats, and cross-account bus sharing. (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/69fbe6a22fd7810332b0b0356803a0eee3f563cf">69fbe6a2</a>)</li>
<li><strong>client-datazone:</strong> Amazon DataZone now supports the
TOOLING blueprint category on CreateEnvironmentBlueprint,
UpdateEnvironmentBlueprint, GetEnvironmentBlueprint, and
ListEnvironmentBlueprints, for custom tooling blueprints.
CreateConnection now accepts roleArn in iamProperties. (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/591cd6f5a7b714427cbe87b36b13357d560d48ea">591cd6f5</a>)</li>
<li><strong>client-elasticache:</strong> Added tagging support for
ElastiCache Global DataStore. (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/0fa9da5946312f09942dad6f711885855f0d0b8e">0fa9da59</a>)</li>
<li><strong>client-marketplace-discovery:</strong> AWS Marketplace
Discovery API now supports localized responses and SigV4a request
signing. It returns new fulfillment details, including AMI architecture,
EBS volume and security group information, SaaS quick-launch status, and
SageMaker input and output MIME types. (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/510673e376bbc578d318308136ecb5a841416bc3">510673e3</a>)</li>
<li><strong>client-redshift-data:</strong> Updates to the ListDatabases
and WorkgroupName validation (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/218c24e106efe1ad64658984123af6634fc7278d">218c24e1</a>)</li>
<li><strong>client-securityagent:</strong> Added support for Confluence
export, enabling customers to publish security findings to Confluence
pages. (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/81408527778af52f0c604a4eaca440f445082f78">81408527</a>)</li>
<li><strong>client-cloudwatch:</strong> This release adds Create, Get,
Update, and DeleteResourceMetricsConfiguration to enable detailed metric
collection for an AWS resource, and adds UpdateOTelEnrichment plus
include and exclude filters on StartOTelEnrichment so you can choose
which metric namespaces CloudWatch enriches. (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/765cc1ce8f95a4f62d83dc07b81a927d74e09b52">765cc1ce</a>)</li>
<li><strong>client-eventbridge:</strong> Adds a ManagedBy field to the
DescribeEventBus and ListEventBuses responses, identifying the AWS
service that created an event bus on your behalf. (<a
href="https://github.com/aws/aws-sdk-js-v3/commit/28a639b27585c85376a4b5db528c31efb80e874d">28a639b2</a>)</li>
</ul>
<h5>Tests</h5>
<ul>
<li><strong>undici-http-handler:</strong> update bidi stream e2e test to
nova-2-sonic model (<a
href="https://redirect.github.com/aws/aws-sdk-js-v3/pull/8313">#8313</a>)
(<a
href="https://github.com/aws/aws-sdk-js-v3/commit/d9a37d9d318f2ef7f5bcf6286bf3c7b475e4175b">d9a37d9d</a>)</li>
</ul>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/aws/aws-sdk-js-v3/blob/main/clients/client-s3/CHANGELOG.md">@​aws-sdk/client-s3's
changelog</a>.</em></p>
<blockquote>
<h1><a
href="https://github.com/aws/aws-sdk-js-v3/compare/v3.1140.0...v3.1141.0">3.1141.0</a>
(2026-09-25)</h1>
<p><strong>Note:</strong> Version bump only for package
<code>@​aws-sdk/client-s3</code></p>
<h1><a
href="https://github.com/aws/aws-sdk-js-v3/compare/v3.1139.0...v3.1140.0">3.1140.0</a>
(2026-09-24)</h1>
<p><strong>Note:</strong> Version bump only for package
<code>@​aws-sdk/client-s3</code></p>
<h1><a
href="https://github.com/aws/aws-sdk-js-v3/compare/v3.1138.0...v3.1139.0">3.1139.0</a>
(2026-09-23)</h1>
<p><strong>Note:</strong> Version bump only for package
<code>@​aws-sdk/client-s3</code></p>
<h1><a
href="https://github.com/aws/aws-sdk-js-v3/compare/v3.1137.0...v3.1138.0">3.1138.0</a>
(2026-09-22)</h1>
<p><strong>Note:</strong> Version bump only for package
<code>@​aws-sdk/client-s3</code></p>
<h1><a
href="https://github.com/aws/aws-sdk-js-v3/compare/v3.1136.0...v3.1137.0">3.1137.0</a>
(2026-09-21)</h1>
<p><strong>Note:</strong> Version bump only for package
<code>@​aws-sdk/client-s3</code></p>
<h1><a
href="https://github.com/aws/aws-sdk-js-v3/compare/v3.1135.0...v3.1136.0">3.1136.0</a>
(2026-09-18)</h1>
<p><strong>Note:</strong> Version bump only for package
<code>@​aws-sdk/client-s3</code></p>
<h1><a
href="https://github.com/aws/aws-sdk-js-v3/compare/v3.1134.0...v3.1135.0">3.1135.0</a>
(2026-09-17)</h1>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/aws/aws-sdk-js-v3/commit/5bc8d9a96936723ac90d721e0c5a2bff7ee8520d"><code>5bc8d9a</code></a>
Publish v3.1141.0</li>
<li><a
href="https://github.com/aws/aws-sdk-js-v3/commit/6050a3813c26795562b5ada8b0d9ea498eb9f8a1"><code>6050a38</code></a>
Publish v3.1140.0</li>
<li><a
href="https://github.com/aws/aws-sdk-js-v3/commit/03d54a858f80012bbc60046a77242223e8dfd9d9"><code>03d54a8</code></a>
Publish v3.1139.0</li>
<li><a
href="https://github.com/aws/aws-sdk-js-v3/commit/c68e50e4a6e0469a20c2894fe8a29c140553ebb8"><code>c68e50e</code></a>
Publish v3.1138.0</li>
<li><a
href="https://github.com/aws/aws-sdk-js-v3/commit/9a104768684e8f22d4373fcc5d910711e62676d6"><code>9a10476</code></a>
chore(codegen): sync for MetricsRecorder support and core error/retry
fixes (...</li>
<li><a
href="https://github.com/aws/aws-sdk-js-v3/commit/6b432472f9bdf5437319b9186f706e3af5c9a748"><code>6b43247</code></a>
Publish v3.1137.0</li>
<li><a
href="https://github.com/aws/aws-sdk-js-v3/commit/d6b94db8f4a00cc452dbe0aacb247e8ece3897ea"><code>d6b94db</code></a>
Publish v3.1136.0</li>
<li><a
href="https://github.com/aws/aws-sdk-js-v3/commit/2d5f18d08aa373d95692d83cb3d60a6a79248fae"><code>2d5f18d</code></a>
Publish v3.1135.0</li>
<li><a
href="https://github.com/aws/aws-sdk-js-v3/commit/0d6310bf6979ddbf737a7e15cfd8d0e7cec07063"><code>0d6310b</code></a>
Publish v3.1134.0</li>
<li><a
href="https://github.com/aws/aws-sdk-js-v3/commit/615a1ca4661ec0e4cb34b8da89fe60c2419b94d0"><code>615a1ca</code></a>
Publish v3.1133.0</li>
<li>Additional commits viewable in <a
href="https://github.com/aws/aws-sdk-js-v3/commits/v3.1141.0/clients/client-s3">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-10-01 11:13:22 -07:00
DottaandPaperclip efc2e6810e fix: show each task once in dashboard agent cards (#14847)
## Thinking Path

> - Paperclip helps people manage AI agents and their tasks.
> - The dashboard shows recent agent activity in compact cards.
> - Those cards use run records, so two runs for one task can create
duplicate task cards.
> - An operator needs to see each task once when scanning the dashboard.
> - This pull request selects one run per linked task before it applies
the card limit.
> - The live runs page still shows each run for run inspection.

## Linked Issues or Issue Description

**What happened?**

The dashboard showed the same task in two agent cards when that task had
both an active run and a completed run.

**Expected behavior**

The dashboard should show a linked task at most once. It should keep the
active run card when one is present.

**Steps to reproduce**

1. Start an agent run for a task that already has a completed run.
2. Open the company dashboard.
3. Observe two cards linked to the same task.

**Paperclip version or commit**

Reproduced on the pre-change master at `8b4aa0692`.

**Deployment mode**

Local dev, built from source. The bug is in the core dashboard UI and
does not depend on an agent adapter or database mode.

## What Changed

- Select distinct linked tasks from capped active and recent run samples
before applying the dashboard card limit.
- Keep separate cards for runs without a linked task.
- Preserve the dashboard's count of additional distinct cards behind the
live-runs link.
- Add UI and embedded Postgres regression tests for duplicate runs and
document the dashboard rule.
- Give the existing multi-request cross-tenant authorization test enough
time on loaded CI runners.

## Verification

- `pnpm --filter @paperclipai/ui exec vitest run
src/components/ActiveAgentsPanel.test.tsx`
- `pnpm --filter @paperclipai/ui exec vitest run
src/api/heartbeats.test.ts`
- `pnpm exec vitest run server/src/__tests__/dashboard-service.test.ts
server/src/__tests__/agent-live-run-routes.test.ts`
- `pnpm exec vitest run
server/src/__tests__/agent-cross-tenant-authz-routes.test.ts`
- `pnpm --filter @paperclipai/ui typecheck`
- `pnpm --filter @paperclipai/server typecheck`
- `pnpm --filter @paperclipai/ui build`
- `pnpm -r typecheck`
- `pnpm build`
- `pnpm check:token-gates`
- Review the dashboard with an active and a completed run on the same
task. Confirm that it shows one card. Open Live agent runs to inspect
both run records.

## Risks

- A very high volume of recent runs for one task can fill the capped
sample and leave older tasks off the dashboard. The Live runs page
remains available for full run inspection.
- The dashboard may fetch up to 50 distinct run representatives to
preserve its overflow count. The default run API response and persisted
data are unchanged.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex, GPT-6. The runtime does not expose the exact model ID or
context window size to this task. The model used reasoning, tool calls,
and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 12:18:35 -05:00
Devin FoleyandPaperclip 6f2ce27ca7 fix(workspaces): prepare checkouts without a local seed config (#14810)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task preparation can create an isolated Git worktree and run its
setup script.
> - The Paperclip repository setup script also prepares a seeded
development instance.
> - A server configured through environment variables can have no local
seed config.
> - This stops ordinary task preparation before the agent starts.
> - This pull request prepares checkout dependencies when no seed source
exists, while preserving errors for invalid sources and existing
development instances.
> - Tasks can start without creating or claiming a seeded development
runtime.

## Linked Issues or Issue Description

**What happened?**

A task with the Paperclip repository fails during setup when the host
has no repository-local or default instance config. The automatic
worktree provisioner requires a seed source even when the task only
needs the checkout.

**Expected behavior**

A plain checkout should prepare its dependencies without a local
development database. A missing custom source, invalid source path, or
existing development instance with a missing source should still fail.
Starting a seeded runtime must still require a valid source.

**Steps to reproduce**

1. Run an environment-configured Paperclip server without a local
instance config.
2. Add the Paperclip repository to a project.
3. Start a task that uses an isolated Git worktree without a custom
provision command.
4. Observe the setup error before agent execution.

**Paperclip version or commit**

Reproduced against `0d3e7bf6ac` with a real script subprocess and
workspace realization regression.

**Deployment mode**

Environment-configured server with external PostgreSQL.

**Additional context**

Searched open and closed GitHub PRs and issues. Related work: Refs
#14795 (seed-source diagnostics) and Refs #11733 (source validation).
This change keeps source validation and seed-readiness checks in place.

## What Changed

- Permit dependency setup when the default seed config is absent
(including the Docker image config path) and the worktree has no
development-instance state.
- Keep missing custom configs, invalid paths, and lost sources for
existing instances as errors.
- Create no config, environment file, or seed manifest for a plain
checkout.
- Keep dependency install failures visible and allow normal instance
setup once a source becomes available.
- Cover the setup script, seed-runtime refusal, and automatic server
worktree realization.
- Document the difference between checkout preparation and
seeded-runtime readiness.

## Verification

- Regression tests failed before the fix for absent-source checkout
preparation and dependency setup.
- `bash -n scripts/provision-worktree.sh`
- `node --test scripts/__tests__/provision-worktree-self-heal.test.mjs`
— 34 passed; 1 existing flock-dependent test skipped on macOS.
- Server regression — 2 passed, covering an unset config and the Docker
image default path.
- `pnpm build` — passed.
- `pnpm -r typecheck` — passed.
- All CI checks passed, including the full test shards, build,
typecheck, browser tests, and canary dry run.
- The first local `pnpm test:run` encountered two chat-test failures
because skill discovery selected an unrelated parent directory. Both
tests pass at the PR commit in a clean temporary checkout. The full
local run was not completed; the redundant clean run was stopped after
the complete CI suite passed.
- `git diff --check` and added-line secrets/PII scan passed.
- Greptile: 5/5, no comments. The branch has no merge conflicts.
- No live tenant deployment or task retry was performed.

## Risks

- A new checkout with no implicit seed config now completes dependency
setup. It has no seeded development instance. A runtime request still
fails until a valid source exists.
- Existing instances and custom source paths retain their failure
behavior. The script does not synthesize a source from environment
credentials or copy a live database.
- No schema, API, or task-setting changes. Revert the commit to restore
the previous setup behavior.

## Model Used

OpenAI Codex (GPT-6), with tool-assisted analysis, code edits, and local
tests. The runtime did not expose a verified model variant or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 10:00:25 -07:00
DottaandPaperclip 6395cae072 fix(runner): ship provider pack in the standard Docker image (#14854)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Remote OpenCode and ACPX runs need a provider pack from the
application build.
> - Cloud now builds its application image from the standard production
image.
> - The provider pack was added only to the legacy cloud image target.
> - The standard image therefore cannot supply the pack to downstream
Cloud images.
> - This pull request adds the pack to the production image and lets the
cloud target inherit it.
> - Remote runs can then use the pack that matches the application
source commit.

## Linked Issues or Issue Description

Refs #13827. Refs #14024.

The standard production image does not include the remote provider pack.
Downstream Cloud images inherit that omission. Remote OpenCode and ACPX
runs fail with `runner_remote_provider_artifact_incompatible` and ask
for `PAPERCLIP_RUNNER_REMOTE_PROVIDER_PACK_PATH`.

## What Changed

- Build and copy the provider pack into the standard production image.
- Set the pack path and check that an unprivileged user can read its
artifacts and execute Node.
- Let the legacy cloud target inherit the pack from production.
- Add regression checks for production packaging and cloud inheritance.
- Document stamped image behavior and the default pack path.

## Verification

- The 12 focused Docker stamp and provider-pack reuse tests pass on
commit `4a11f8d52aead55f85527c8e82c6d7f2644ce0da`.
- The new packaging regression failed against the old Dockerfile and
passed with the fix.
- On the current commit, `pnpm build` and `pnpm -r typecheck` pass. All
seven standard-image contract tests also pass.
- The current-head CI build, typecheck, test, browser, and native Runner
checks passed. The local full suite hit one chat-channel assertion
failure; that exact test passed in isolation. The remaining local run
was stopped after CI completed to avoid duplicating its full suite. An
earlier run on the pre-rebase base had a heartbeat comment batching
timeout; the external chat wait integration suite passed all 142 tests
in isolation.
- [The stamped preview image build
passed](https://github.com/paperclipai/paperclip/actions/runs/36885002850/job/110446106393),
including the production-stage provider pack build, copy, and
unprivileged artifact readability/executable check. Publication,
compatibility validation, and deployment of this exact commit to a
staging QA instance passed.
- Reproduced the exact missing-pack error on an existing staging image
with Paperclip Runner, ACPX, and Claude in a remote Daytona computer.
The legacy Claude adapter succeeds with the same account and computer.
After deploying this commit, the same native task succeeded: it computed
`5050` with a real remote shell command, wrote a proof file, read it
back in a separate call, uploaded the file as a deliverable, and
completed the task. The uploaded file contents and Done state persisted
after a page reload. The run trace confirms Paperclip Runner, ACPX, and
Claude. The first run took 2m 59s, including approximately 97s of remote
artifact preparation. A second native run read the unchanged file from
the prior run and completed successfully. Its startup took about 120s;
this verifies repeated execution and file persistence, not fast
provider-pack reuse.

## Risks

- Stamped standard images now include the provider pack and its build
cost. A pack build failure now fails the production image build.
- Unstamped local builds still skip pack generation. Setting the path
alone does not create a pack.
- No database, provider authentication, or runner verification rules
change.

## Model Used

OpenAI Codex, GPT-6. The exact serving model identifier and context
window are not exposed in this session. Capabilities used: repository
inspection, code editing, shell verification, and browser testing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 11:07:50 -05:00
Dotta f6406e7e55 Merge frozen master into the rich ACP production candidate
Preserve incremental Codex checkpoints and ACP unchanged-directory ownership as separate warm-session paths. Combine cancellation commit fencing, terminal outcome recovery, and both native fixture catalogs. Keep all candidate qualification states and provider profile identities unchanged.

Validation: 629 controller unit checks, 5 heartbeat cancellation checks, 210 runner/profile/sidecar checks, 44 catalog checks; token gates, Rust source formatting, and generated protocol manifest pass. Database tests and builds intentionally deferred. Source-aliased no-emit checking is blocked only by the borrowed ACPX SessionRecord declaration lacking the already-patched cursor_prompt_usage field.
2026-10-01 10:41:38 -05:00
DottaandPaperclip 4ac374103f fix(connections): repair Asana MCP and add shared-app sign-in (#14756)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections let agents use provider tools through the permission
gateway.
> - Asana provides an official remote MCP server, but its v2 server
requires a registered MCP OAuth app.
> - Setup can discover retired v1 endpoints and send a callback that
differs from the displayed URL.
> - This pull request repairs custom app setup and adds sign-in through
Paperclip's shared app.
> - Users can choose their own app without enrolling with Paperclip
Cloud.
> - Agents can use Asana tools after the user connects their account and
sets action permissions.

## Linked Issues or Issue Description

Related: #14739 supplies the personal credential repair used by resumed
Asana setup. No duplicate Asana authentication PR was found.

**What happened?**

Asana setup failed even with a user-created app. Root discovery metadata
still points at v1. MCP v2 uses the Asana OAuth issuer and requires an
MCP app with a client secret. Local setup also displayed a localhost
callback while an Origin header could make authorization use a numeric
loopback callback.

**Expected behavior**

Sign in with Paperclip's app when its broker profile is available. Keep
custom MCP app setup available without Cloud enrollment. Use the correct
issuer, callback, client credentials, and resource throughout setup.

**Steps to reproduce**

1. Open Asana in the connection catalog.
2. Supply an Asana MCP app's client ID and secret.
3. Start OAuth on a local instance opened with a numeric loopback
address, or resume a draft that cached v1 metadata.
4. Observe the wrong discovery endpoint or callback mismatch.

**Paperclip version or commit**

Reproduced from b54b2dc35c. Rebased onto
master at `0829d94af` after the single-screen setup change in #14811.

**Deployment mode**

Local development from source. The managed path also supports enrolled
self-hosted instances.

## What Changed

- Add the `asana.mcp` managed profile and the default Sign in with Asana
method.
- Use Asana's reviewed v2 protected-resource metadata before cached
endpoints.
- Require a custom MCP app's client secret and retain saved credentials
during setup or reconnect. Repair only the known Asana v1
issuer/resource binding, retaining company and callback checks.
- Expose a boolean for the acting user's saved client secret. The form
offers secret reuse only when that user has an active grant with the
required reference.
- Let users select their own app from the enrollment and
shared-app-unavailable screens, or from Advanced on the single-screen
setup page.
- Canonicalize HTTP loopback callbacks even when the request includes an
Origin header.
- Extend signed broker claims and provider URL validation for Asana.
Require refresh credentials on managed authorization.
- Document setup, distribution, and shared-app rollout requirements.
- Resolve permission-profile name collisions when finishing another
account. The live staging test found this after renaming the first Asana
connection; OAuth succeeded but profile finalization failed.

## Verification

- Final live staging proof used app commit
`c6053157c4e42ac017727117b754ae77fa5c45fa` and the real Cloud broker at
`767b63835170f664542afd0df99a76615e204b62`. In the embedded browser,
default shared sign-in required no client credentials, returned through
the central Cloud callback to the tenant, and discovered 39 actions. Get
me succeeded through the gateway as the selected QA agent (2.1 seconds).
The custom-app connection also returned a real result on this final
build (0.9 seconds).
- Retried the shared draft that failed during the first staging test. It
completed after the profile-name fix, retained the selected agent, and
kept the existing custom connection intact. Two database regressions
reproduced the collision before the fix and passed afterward. The
updated transaction rollback test also passes.
- Shared reconnect returned to the same staging connection with 39
actions. Earlier staging checks verified the custom-app fallback when
the shared profile was unavailable, saved-secret reuse on reconnect, and
Off blocking the action test. Allowed was restored after that check.
- Local live-provider checks also repaired a saved Asana v1
issuer/resource binding without reentering the secret and verified that
numeric-loopback setup uses the displayed localhost callback. Expiring
the local managed access-token timestamp triggered a real Asana refresh
and a successful Get me call. These early local broker tests used
enrollment/authentication and storage fixtures; the final staging proof
used deployed Cloud identity and persistent storage.
- The new production app is registered and configured, but production
sign-in has not been deployed or verified. Live provider revocation was
not run because the existing staging test app is shared with other
connections.

- After rebasing onto the single-screen setup flow, full `pnpm -r
typecheck`, `pnpm build`, and `pnpm check:token-gates` pass. Focused
verification passes 365 service and 36 broker-client tests. Broader
checks pass all 833 shared-package tests and all 530 connector-page
tests. The shared suite uses `TMPDIR=/private/tmp` to avoid macOS
temporary-directory symlinks in its canonical-path tests. The UI tests
verify the shared-app default and switching to a custom app with its
required client secret.
- Embedded-browser smoke on the current rebased build verified the
shared sign-in default, Advanced → custom app (client ID and secret
required), and switching back to Paperclip. Both existing Asana
connections remained connected after restart. No new provider
authorization was performed during this smoke.
- The full local `pnpm test:run` was interrupted when the execution
session restarted. Before interruption, it reported one runtime-slot
restart test failure. That test passed on an isolated retry after
clearing two unused PostgreSQL shared-memory segments. The full local
suite did not complete; CI must pass on the current head before merge.
- All 52 CI and security checks pass on
`9318fd4e9b616cdc3de12f40cdb9bd32d865af4c` (CI run `36882064080`),
including all eight browser shards, nine serialized-server shards,
build, typecheck, and canary dry run. Two optional Storybook checks were
skipped. Greptile review 4 reports 5/5 on this exact commit, with all
review threads resolved. Its updated summary identifies the current SHA;
this comment-triggered review did not publish a separate GitHub check
run.
- Provider revocation is unit-tested in the companion broker. Live
provider revocation was not run because the existing test app is shared
with other connections.

## Risks

- The shared Paperclip Asana MCP app has been registered with its
production callback and Any workspace distribution. Its secret is
provisioned in the production secret store, and the runtime client ID
and secret reference are configured. The production profile is enabled
in the saved deployment configuration. The companion broker has merged
and passed staging deployment; production sign-in still requires a
production deployment and live verification. Custom setup remains
available.
- Asana MCP uses the provider's fixed `default` grant. Paperclip action
policies limit agent tool use; they do not narrow provider consent.
- The callback correction affects HTTP loopback OAuth flows. Public
HTTPS callbacks retain their existing behavior.
- Reviewed discovery URLs now override stale cached endpoints. Tests
cover the Asana v1-to-v2 repair.
- No schema migration. Connection removal retains the existing
local-only revocation behavior.

## Model Used

OpenAI GPT-6 through Codex, with code execution, browser testing, and
GitHub tooling. The exact model variant and context window are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 10:24:44 -05:00
dependabot[bot]andPriya Raman 8b4aa06920 build(deps): bump multer from 2.2.0 to 2.4.0 (#14493)
Bumps [multer](https://github.com/expressjs/multer) from 2.2.0 to 2.4.0.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/expressjs/multer/releases">multer's
releases</a>.</em></p>
<blockquote>
<h2>v2.4.0</h2>
<h2>Highlights</h2>
<p><strong>multer finally supports Google Cloud Functions and Firebase
🎉</strong></p>
<p>These platforms read the request body before your code runs, so
multer's classic <code>req.pipe(busboy)</code> received nothing: empty
<code>req.body</code>, empty <code>req.files</code>, and nearly a decade
of duplicated issues.</p>
<p>The new <code>streamHandler</code> option closes that gap: you decide
how the body reaches the parser, so the pre-read <code>rawBody</code>
just works (see image).</p>
<pre lang="js"><code>const multer = require('multer')
<p>const upload = multer({<br />
storage: multer.memoryStorage(),<br />
streamHandler: (req, busboy) =&gt; {<br />
// Cloud Functions / Firebase expose the pre-read body here<br />
if (req.rawBody) busboy.end(req.rawBody)<br />
else req.pipe(busboy)<br />
}<br />
})</p>
<p>app.post('/upload', upload.single('file'), (req, res) =&gt; {<br />
res.json({ name: req.file.originalname, size: req.file.size })<br />
})<br />
</code></pre></p>
<p>This landed thanks to community PRs going back to 2017; their authors
are credited as co-authors in the release.</p>
<h2>Important: Security</h2>
<ul>
<li>Fix <a
href="https://www.cve.org/CVERecord?id=CVE-2026-88932">CVE-2026-88932</a>
(<a
href="https://github.com/expressjs/multer/security/advisories/GHSA-3pph-fpjx-jg34">GHSA-3pph-fpjx-jg34</a>)</li>
</ul>
<h2>What's Changed</h2>
<ul>
<li>docs: remove README translations by <a
href="https://github.com/UlisesGascon"><code>@​UlisesGascon</code></a>
in <a
href="https://redirect.github.com/expressjs/multer/pull/1463">expressjs/multer#1463</a></li>
<li>ci: add macOS to the test matrix by <a
href="https://github.com/kilisamemarisaaa"><code>@​kilisamemarisaaa</code></a>
in <a
href="https://redirect.github.com/expressjs/multer/pull/1464">expressjs/multer#1464</a></li>
<li>feat. improve wording for LIMIT_UNEXPECTED_FILE error code by <a
href="https://github.com/flashbag"><code>@​flashbag</code></a> in <a
href="https://redirect.github.com/expressjs/multer/pull/426">expressjs/multer#426</a></li>
<li>feat: add filename to file errors by <a
href="https://github.com/UjjwalKumar239"><code>@​UjjwalKumar239</code></a>
in <a
href="https://redirect.github.com/expressjs/multer/pull/1416">expressjs/multer#1416</a></li>
<li>refactor: remove concat-stream dependency by <a
href="https://github.com/Phillip9587"><code>@​Phillip9587</code></a> in
<a
href="https://redirect.github.com/expressjs/multer/pull/1356">expressjs/multer#1356</a></li>
<li>fix: reject non-integer fileSize limits by <a
href="https://github.com/abhu85"><code>@​abhu85</code></a> in <a
href="https://redirect.github.com/expressjs/multer/pull/1395">expressjs/multer#1395</a></li>
<li>fix: allow exactly limits.parts parts by <a
href="https://github.com/deepakganesh78"><code>@​deepakganesh78</code></a>
in <a
href="https://redirect.github.com/expressjs/multer/pull/1446">expressjs/multer#1446</a></li>
<li>fix: do not consume maxCount for files skipped by fileFilter by <a
href="https://github.com/Sagargupta16"><code>@​Sagargupta16</code></a>
in <a
href="https://redirect.github.com/expressjs/multer/pull/1426">expressjs/multer#1426</a></li>
<li>fix: validate all limits at construction time by <a
href="https://github.com/ShubhamOulkar"><code>@​ShubhamOulkar</code></a>
in <a
href="https://redirect.github.com/expressjs/multer/pull/1335">expressjs/multer#1335</a></li>
<li>test: cover storage engine _removeFile invocation semantics by <a
href="https://github.com/kilisamemarisaaa"><code>@​kilisamemarisaaa</code></a>
in <a
href="https://redirect.github.com/expressjs/multer/pull/1460">expressjs/multer#1460</a></li>
<li>feat: accept a function for limits by <a
href="https://github.com/hossein-zare"><code>@​hossein-zare</code></a>
in <a
href="https://redirect.github.com/expressjs/multer/pull/1133">expressjs/multer#1133</a></li>
<li>fix: add flush option to disk storage to fsync files before
completion by <a
href="https://github.com/kilisamemarisaaa"><code>@​kilisamemarisaaa</code></a>
in <a
href="https://redirect.github.com/expressjs/multer/pull/1458">expressjs/multer#1458</a></li>
<li>feat: expose busboy defCharset, highWaterMark and fileHwm options by
<a
href="https://github.com/UlisesGascon"><code>@​UlisesGascon</code></a>
in <a
href="https://redirect.github.com/expressjs/multer/pull/1465">expressjs/multer#1465</a></li>
<li>docs: describe the file stream contract for storage engines by <a
href="https://github.com/UlisesGascon"><code>@​UlisesGascon</code></a>
in <a
href="https://redirect.github.com/expressjs/multer/pull/1468">expressjs/multer#1468</a></li>
</ul>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/expressjs/multer/blob/main/CHANGELOG.md">multer's
changelog</a>.</em></p>
<blockquote>
<h2>2.4.0</h2>
<ul>
<li>Fix <a
href="https://www.cve.org/CVERecord?id=CVE-2026-88932">CVE-2026-88932</a>
(<a
href="https://github.com/expressjs/multer/security/advisories/GHSA-3pph-fpjx-jg34">GHSA-3pph-fpjx-jg34</a>)</li>
<li>Add <code>filename</code> to <code>LIMIT_FILE_SIZE</code> and
<code>LIMIT_UNEXPECTED_FILE</code> errors (<a
href="https://redirect.github.com/expressjs/multer/pull/1416">#1416</a>)</li>
<li>Accept a function for <code>limits</code>, called with the request,
to set limits per request (<a
href="https://redirect.github.com/expressjs/multer/pull/1133">#1133</a>)</li>
<li>Add opt-in <code>flush</code> option to <code>DiskStorage</code> to
fsync files before the callback runs (<a
href="https://redirect.github.com/expressjs/multer/pull/1458">#1458</a>)</li>
<li>Expose busboy's <code>defCharset</code>, <code>highWaterMark</code>
and <code>fileHwm</code> options (<a
href="https://redirect.github.com/expressjs/multer/pull/1465">#1465</a>)</li>
<li>Add <code>streamHandler</code> option to feed busboy from
pre-consumed bodies (Google Cloud Functions, Firebase) (<a
href="https://redirect.github.com/expressjs/multer/pull/1466">#1466</a>)</li>
<li>Allow <code>multer.diskStorage()</code> to be called without options
(<a
href="https://redirect.github.com/expressjs/multer/pull/1471">#1471</a>)</li>
<li>Decode WHATWG-escaped characters (<code>%0A</code>,
<code>%0D</code>, <code>%22</code>) in field names, matching
<code>file.originalname</code> since 2.3.0: <code>req.body</code> keys,
<code>file.fieldname</code> and <code>err.field</code> now carry the
real name. If you matched the escaped spelling as a workaround, use the
real name now (<a
href="https://redirect.github.com/expressjs/multer/pull/1473">#1473</a>)</li>
<li>Report the decoded filename in <code>err.filename</code> on
<code>LIMIT_FILE_SIZE</code> errors, matching
<code>file.originalname</code> (<a
href="https://redirect.github.com/expressjs/multer/pull/1478">#1478</a>)</li>
<li>Reject non-integer or negative <code>limits</code> values at
construction time; a float limit silently disabled the check (<a
href="https://redirect.github.com/expressjs/multer/pull/1395">#1395</a>,
<a
href="https://redirect.github.com/expressjs/multer/pull/1335">#1335</a>)</li>
<li>Accept requests with exactly <code>limits.parts</code> parts;
<code>LIMIT_PART_COUNT</code> now fires only when the limit is exceeded.
If you set <code>parts</code> one higher to work around this, you can
drop the extra one (<a
href="https://redirect.github.com/expressjs/multer/pull/1446">#1446</a>)</li>
<li>Files skipped by <code>fileFilter</code> no longer count towards
<code>maxCount</code> (<a
href="https://redirect.github.com/expressjs/multer/pull/1426">#1426</a>)</li>
<li>Change the <code>LIMIT_UNEXPECTED_FILE</code> message to
&quot;Unexpected file field&quot; (<a
href="https://redirect.github.com/expressjs/multer/pull/426">#426</a>)</li>
<li>Remove the <code>concat-stream</code> dependency (<a
href="https://redirect.github.com/expressjs/multer/pull/1356">#1356</a>)</li>
<li>Docs: add JSDoc to the public API and document the storage engine
stream contract (<a
href="https://redirect.github.com/expressjs/multer/pull/1467">#1467</a>,
<a
href="https://redirect.github.com/expressjs/multer/pull/1468">#1468</a>)</li>
<li>Docs: add FormData upload examples (<a
href="https://redirect.github.com/expressjs/multer/pull/896">#896</a>)</li>
<li>Docs: remove the translated READMEs (<a
href="https://redirect.github.com/expressjs/multer/pull/1463">#1463</a>)</li>
<li>Internal: run the test suite on macOS (<a
href="https://redirect.github.com/expressjs/multer/pull/1464">#1464</a>)</li>
</ul>
<h2>2.3.0</h2>
<ul>
<li>Fix <a
href="https://www.cve.org/CVERecord?id=CVE-2026-77078">CVE-2026-77078</a>
(<a
href="https://github.com/expressjs/multer/security/advisories/GHSA-wc9g-mqfw-jrwm">GHSA-wc9g-mqfw-jrwm</a>)</li>
<li>Fix <a
href="https://www.cve.org/CVERecord?id=CVE-2026-77037">CVE-2026-77037</a>
(<a
href="https://github.com/expressjs/multer/security/advisories/GHSA-qfvm-cv95-jqjf">GHSA-qfvm-cv95-jqjf</a>)</li>
<li>Fix <a
href="https://www.cve.org/CVERecord?id=CVE-2026-77063">CVE-2026-77063</a>
(<a
href="https://github.com/expressjs/multer/security/advisories/GHSA-qvfw-j98x-7q72">GHSA-qvfw-j98x-7q72</a>)</li>
<li>Fix <a
href="https://www.cve.org/CVERecord?id=CVE-2026-82333">CVE-2026-82333</a>
(<a
href="https://github.com/expressjs/multer/security/advisories/GHSA-535w-7cp7-47q4">GHSA-535w-7cp7-47q4</a>)</li>
<li>Add <code>MulterError</code> codes <code>INVALID_FIELD_NAME</code>
and <code>STREAM_DESTROYED</code></li>
<li>Add opt-in <code>limits.fieldArrayIndexLimit</code> to bound numeric
array indexes in field names (<a
href="https://redirect.github.com/expressjs/multer/pull/1438">#1438</a>)</li>
<li>Accept files whose size is exactly <code>limits.fileSize</code> (<a
href="https://redirect.github.com/expressjs/multer/pull/1407">#1407</a>)</li>
<li>Preserve the caller's async context (<code>AsyncLocalStorage</code>)
when calling <code>next()</code> (<a
href="https://redirect.github.com/expressjs/multer/pull/1124">#1124</a>)</li>
<li>Decode WHATWG-escaped characters (<code>%0A</code>,
<code>%0D</code>, <code>%22</code>) in <code>file.originalname</code>
(<a
href="https://redirect.github.com/expressjs/multer/pull/1421">#1421</a>)</li>
<li>Do not crash when <code>fileFilter</code> invokes its callback more
than once (<a
href="https://redirect.github.com/expressjs/multer/pull/1427">#1427</a>)</li>
<li>Use a fallback message for <code>MulterError</code> codes without a
mapping (<a
href="https://redirect.github.com/expressjs/multer/pull/1448">#1448</a>)</li>
<li>Docs: clarify <code>preservePath</code> and <code>parts</code>, use
<code>crypto.randomBytes</code> in the <code>DiskStorage</code> example
(<a
href="https://redirect.github.com/expressjs/multer/pull/1414">#1414</a>,
<a
href="https://redirect.github.com/expressjs/multer/pull/1430">#1430</a>,
<a
href="https://redirect.github.com/expressjs/multer/pull/1436">#1436</a>)</li>
<li>Docs: add Indonesian, Japanese and Tamil translations and refresh
all translations from the current README (<a
href="https://redirect.github.com/expressjs/multer/pull/1431">#1431</a>,
<a
href="https://redirect.github.com/expressjs/multer/pull/1354">#1354</a>,
<a
href="https://redirect.github.com/expressjs/multer/pull/1462">#1462</a>)</li>
<li>Internal: run the test suite on Windows (<a
href="https://redirect.github.com/expressjs/multer/pull/1334">#1334</a>)</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/expressjs/multer/commit/35979e5afbb814bdb4b750ce028b125eb84c53af"><code>35979e5</code></a>
2.4.0 (<a
href="https://redirect.github.com/expressjs/multer/issues/1469">#1469</a>)</li>
<li><a
href="https://github.com/expressjs/multer/commit/b888532fe2e10ceb13da44286448cb2bb4ce9720"><code>b888532</code></a>
chore(deps): bump github/codeql-action/upload-sarif to 4.37.9 (<a
href="https://redirect.github.com/expressjs/multer/issues/1474">#1474</a>)</li>
<li><a
href="https://github.com/expressjs/multer/commit/e6bcd7db69315714fb9a38b1cd1e0f582cbce65a"><code>e6bcd7d</code></a>
chore(deps): bump github/codeql-action/analyze from 4.37.4 to 4.37.9 (<a
href="https://redirect.github.com/expressjs/multer/issues/1475">#1475</a>)</li>
<li><a
href="https://github.com/expressjs/multer/commit/00dec43b394c5bfaf4efb6d0fe8467bad05525ed"><code>00dec43</code></a>
chore(deps): bump github/codeql-action/init from 4.37.4 to 4.37.9 (<a
href="https://redirect.github.com/expressjs/multer/issues/1476">#1476</a>)</li>
<li><a
href="https://github.com/expressjs/multer/commit/8d5c3b72e430e7edaa2749e96dcb75bf18d84733"><code>8d5c3b7</code></a>
feat: allow diskStorage without options (<a
href="https://redirect.github.com/expressjs/multer/issues/1471">#1471</a>)</li>
<li><a
href="https://github.com/expressjs/multer/commit/02f6e8265b6bf6b6921a819db3b84276efa03ed2"><code>02f6e82</code></a>
fix: report the decoded filename on LIMIT_FILE_SIZE (<a
href="https://redirect.github.com/expressjs/multer/issues/1478">#1478</a>)</li>
<li><a
href="https://github.com/expressjs/multer/commit/bc3f72d5edaa19b993771348e3fb47b366316a88"><code>bc3f72d</code></a>
fix: decode escaped field names, not just filenames (<a
href="https://redirect.github.com/expressjs/multer/issues/1473">#1473</a>)</li>
<li><a
href="https://github.com/expressjs/multer/commit/2661325ba8ca2a72b7fb554b63d2df51da9290c4"><code>2661325</code></a>
docs: add JSDoc to the public API (<a
href="https://redirect.github.com/expressjs/multer/issues/1467">#1467</a>)</li>
<li><a
href="https://github.com/expressjs/multer/commit/53337f9713619ef3381ee6b4e541f926dbaac305"><code>53337f9</code></a>
fix: remove late-completing uploads aborted before the engine names
them</li>
<li><a
href="https://github.com/expressjs/multer/commit/7f2c9ab5f35c0478a7b3bb0f5e1af9b536780132"><code>7f2c9ab</code></a>
feat: add streamHandler option to feed busboy from pre-consumed bodies
(<a
href="https://redirect.github.com/expressjs/multer/issues/1466">#1466</a>)</li>
<li>Additional commits viewable in <a
href="https://github.com/expressjs/multer/compare/v2.2.0...v2.4.0">compare
view</a></li>
</ul>
</details>
<details>
<summary>Maintainer changes</summary>
<p>This version was pushed to npm by <a
href="https://www.npmjs.com/~GitHub%20Actions">GitHub Actions</a>, a new
releaser for multer since your current version.</p>
</details>
<br />

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Priya Raman <noreply@paperclip.ing>
2026-10-01 15:00:48 +00:00
DottaandPaperclip 0829d94af2 fix(auth): derive low-trust human direction from existing execution records (#14775)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Low-trust review contains work that may include hostile input.
> - Its default intake boundary currently blocks direct human chat and
tasks outside that boundary.
> - Human direction should authorize the assigned work while preserving
containment.
> - Existing conversations and execution requests already identify
direct human instructions.
> - This pull request derives exact-task authority from those records
and the current assignee.
> - The agent can perform that work without gaining access to unrelated
tasks or privileged tools.

## Linked Issues or Issue Description

**What happened?**

A low-trust agent with a project boundary rejects its owner's direct
Agent Chat before provider execution. Human-assigned tasks outside that
project fail the same check.

**Expected behavior**

An authorized human can talk to the agent or assign it a task. The exact
task runs with its existing sandbox, credential, and tool restrictions.

**Steps to reproduce**

1. Enable Agent Chat and isolated workspaces. Configure a sandbox agent
with low-trust review scoped to an intake project.
2. Send the agent a direct board chat message, or assign it a
projectless task.
3. Observe `low_trust_boundary_mismatch` before execution.

Related: #14766 adds private task directories for repo-free low-trust
execution. It is now merged into master and included in the branch base,
so CI and staging verify the combined behavior.

## What Changed

- Derive owner-chat access from existing conversation identity.
- Derive exact-task access from the existing human requester and
server-owned request origin, including coalesced requests. Plugin and
external sender attribution do not authorize work.
- Follow existing `retryOfRunId` database links for automatic
continuations, checking company, agent, and task throughout; cancelled
ancestors cannot grant authority.
- Require a live run and current assignment. Preserve sandbox,
credential, privileged-tool, responsible-user, and quarantined-output
checks.
- Retain board backlog assignments in existing request records without
starting execution. Reassignment cancels old human requests in the
common service transaction, including plugin writes; late settlement
cannot revive them.
- Add real database and HTTP coverage for request provenance, retry
ancestry, cancelled runs, concurrent reassignment, spoofing, and
containment. Document the rule.
- Preserve legacy board assignment requests through their existing
source, reason, and human requester.
- Use the existing wrapped-error helper for concurrent chat-question
idempotency; a deterministic race test reproduces the CI failure before
the fix and passes after it.
- No new schema, migrations, or user-identity fields. Existing requester
columns hold attribution.

## Verification

- Passed the focused database, policy-retention, HTTP, and reassignment
tests locally. The HTTP test creates a task through the real board route
and checks the resulting persisted wakeup before exercising agent reads,
comments, mutations, and review handoff.
- Database tests hold a reassignment transaction open to verify coherent
authorization before and after commit, with a two-connection pool. They
cover retries, coalesced requests, cancelled ancestry, invalid
cross-company/agent/task links, cycles, and forged attribution.
- Full local `pnpm -r typecheck` and `pnpm build` passed on the final
commit (`d107c26df`). [Latest-head
CI](https://github.com/paperclipai/paperclip/actions/runs/36815589542)
passed: 54 successful checks, two expected skips, including all eight
browser-test shards. Greptile is 5/5 on this exact commit with no
unresolved threads. Local tests were targeted; the full test suite ran
through CI’s test matrix.
- The revised HTTP suite passed all 13 tests; database authorization
tests passed all 11, including legacy compatibility and late watchdog
settlement; the backlog route contract passed all 3 tests. Another 102
tests covering durable chat admission, wake queues, and Cursor execution
passed.
- All 90 interaction-service tests passed with both create calls
deliberately held until their optimistic reads complete, forcing
duplicate-key recovery. That forced race failed before switching to the
shared wrapped-error helper.
- Previous staging proof covered owner chat and projectless task
persistence. The simplified revision has not been redeployed; that
earlier proof is not claimed for the new implementation.

## Risks

- This is an authorization change: only the live run's exact task
qualifies, and normal responsible-user restrictions still apply.
- Existing request and retry records are authoritative. Merely naming a
responsible/originating user or an external connector sender does not
qualify.
- Reassignment invalidates existing human request records
transactionally. A cancelled run or request cannot regain authority when
the task is assigned back.
- Ordinary task exceptions require server-owned origin or the legacy
board assignment source/reason/actor combination. Existing owner chats
use conversation identity.

## Model Used

OpenAI GPT-6 in Codex, with reasoning, repository tools, code execution,
and browser testing. The runtime does not expose a more specific model
ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 09:45:18 -05:00
467125fafb feat(connections): one-screen connector setup with stated defaults (#14811)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents use Connections (the Apps catalog) to act in services like
Notion, GitHub, Google Workspace and Railway
> - Each connector asked the user to answer setup questions before it
went to the provider. Most of the questions already had the correct
answer selected
> - ROADMAP.md lists "simpler setup" for Apps and Connections as ongoing
work. This change continues that work
> - This pull request removes the questions that Paperclip can answer
itself. It states the defaults in one line and moves the choices behind
"Change" and onto the Permissions tab
> - The benefit is that most connectors take one click in Paperclip and
then the provider's own consent screen

## Linked Issues or Issue Description

No public issue exists. This is the description, from the enhancement
template.

**What existing behavior does this improve?**
The setup flow for tool connectors in the Apps catalog.

**Subsystem affected**
Apps and Connections: `ui/src/features/connections`,
`ui/src/pages/apps`, the `packages/shared` app definitions, and the
OAuth routes in `server/src/routes/tool-access.ts`.

**Current behavior**
Every connector opened with an Access step. The step asked who can use
the connection and which agents get it, and both answers were already
selected. 18 connectors also asked "How do you want to connect?" when
Paperclip could rank the methods. The Google apps and Postman also asked
"What should Paperclip be able to do?" before sign-in. The four gateway
connectors (Zapier, Arcade, Composio, Executor) used a separate two-step
wizard. Asana was pinned to a customer-owned OAuth app, so the user had
to register an app in Asana's developer console. The "Set all" control
on the Permissions tab changed only one action. After the user approved
access, Railway's consent page showed "you can close this window" and
did not return to Paperclip.

**Proposed behavior**
One screen per connector, with one primary button. The screen states the
defaults in one sentence, for example "Connects for everyone in your
organization, available to all agents". A "Change" link opens one
Advanced panel. When the provider's metadata allows dynamic client
registration, Paperclip registers a client itself. Connecting lands on
the Permissions tab. On that tab, "Set all" changes every action in the
group.

**Reason and benefit**
The user makes fewer decisions before the connection exists. Most
choices are easier to make after the connection, on the Permissions tab,
where a change has an immediate effect.

**Breaking changes**
None. No schema or API change. Existing connections keep their settings.

## What Changed

- **No Access step.** `ConnectionSetupFlow` no longer has the Access
step. The flow shows the resolved default above the primary button and
on the completion screen. The access controls moved into one Advanced
panel. The panel opens automatically only when a setting in it is
required.
- **A default method for every app.** The flow always picks the ranked
default method. Alternate methods are in the Advanced panel. The Google
and Postman capability choice is not asked before sign-in. The
write-capable method is the default.
- **Gateway connectors.** `RemoteMcpProductionSetup` (Zapier, Arcade,
Composio, Executor) no longer has its own Access step. Its commit path
and the main commit path use one helper, `askFirstCatalogEntryIdsFor`,
for server-suggested defaults.
- **Dynamic registration from live metadata.**
`canRegisterOAuthClientDynamically` now allows registration when the
provider advertises a registration endpoint, even if the catalog entry
lists only customer-owned clients. The Asana and Linear definitions and
catalog text match live probes. Asana issues clients for loopback
callbacks only, so a hosted deployment still needs an Asana app.
- **Connection setup states.** New
`packages/shared/src/connection-setup-state.ts` sorts each method into
`instant`, `authorize`, `paste` or `register`. The gallery card verb
("Connect" or "Add key") comes from this resolver and the instance's
ownership availability.
- **Generic MCP.** The generic path no longer asks "Does it need a key?"
first. A credential challenge from the server shows the key field.
- **Permissions tab.** Each action row shows its risk level. Each group
has a "Set all" control. The control sends one change for the whole
group. Before, each row's save started from the same render, so the
saves overwrote each other. The Zapier/Arcade/Composio/Executor setup
screen had the same defect.
- **OAuth callback interstitial.** A cross-site browser navigation to
`/api/tools/oauth/callback` gets a small same-origin "Finishing your
connection…" page. That page repeats the request, and the repeat does
the code exchange. Railway's consent page replaces itself after about
two seconds, and the code exchange plus tool discovery takes longer than
that. The interstitial uses only a meta refresh, because the OAuth code
is single-use. Requests without `Sec-Fetch-Site: cross-site` take the
old path.
- **Linear registers through its MCP server.** Linear pins the console
endpoints at `linear.app`. Pinned endpoints now replace discovery only
when the method cannot register, or when the connection has an
operator-entered client. So a Linear connection now finds the
registration endpoint at `mcp.linear.app`.
- **Own-OAuth-app recovery stays on the one-click screen.** When the
method also accepts a customer-owned client, the client fields are in
the Advanced panel. The panel opens after a failed sign-in. "Try again"
resumes the draft with the operator's client.
- **E2E specs** follow the one-screen flow. The Access-step clicks are
removed, the specs open **Change** before they pick agents, and they
expect GitHub's **Add key** verb.
- **Default permissions do not change.** New connections still allow
every action. The user can set actions to Ask first or Off on the
Permissions tab.

## Verification

- `cd ui && npx vitest run src/pages/apps src/features/connections
--no-file-parallelism`
- `cd packages/shared && npx vitest run src/app-definitions.test.ts
src/connection-setup-state.test.ts`
- `cd server && npx vitest run src/__tests__/tool-access-service.test.ts
src/__tests__/remote-mcp-connectors.test.ts`
- `pnpm check:token-gates`
- New tests:
- `PermissionsPanel.group.test.tsx` checks that "Set all" sends one
change for the whole group. It fails on the old code.
  - `action-permissions.test.ts` checks the group update.
  - `connection-setup-state.test.ts` checks the four setup states.
- A server test checks that a cross-site callback gets the interstitial
and does not use the OAuth state, and that the same-origin repeat
completes the connection.
- Manual check on a hosted staging deployment. GitHub, Google Drive,
Composio, Notion, PostHog and Railway each connected from one screen and
returned to the Permissions tab. On Railway, "Set all" changed all 65
write actions, and the change remained after a reload.
- Visual changes: snapshot baselines are intentionally not updated. See
the `doc/design/DECISION-SHEET.md` entry "Per-change snapshot
verification demoted to dormant (Jul 13 2026)".

## Risks

- **Fewer confirmation clicks.** Organization-wide access is the
default, and the user does not confirm it on a separate step. This was
already the preselected answer. The flow shows the default before the
user clicks and again after the connection.
- **Google write scope.** Google apps now request the write-capable
scope by default. A narrower scope needs a new sign-in.
- **Dynamic registration from live metadata.** A provider can advertise
registration and then reject a redirect URI. Asana rejects hosted
callbacks, for example. In that case registration fails, and the
customer-owned client path remains available for recovery.
- **Callback interstitial.** The OAuth callback adds one same-origin
step for cross-site browser navigations. Browsers without `Sec-Fetch-*`
headers use the old direct path.
- Chat and bot connectors (Discord, Telegram, Microsoft Teams, iMessage)
do not change.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Claude Opus 5.5 (Anthropic), model ID `claude-opus-5-5`, used through
Claude Code with tool use (shell, file editing, browser automation) and
extended thinking. It wrote the code, the tests and this description. A
human product owner directed the work and tested it by hand.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: scotttong <squadbot000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 23:13:47 -07:00
Dotta 3665daf9dc fix(runner): bind normalized semantic receipt inputs 2026-09-30 23:28:02 -05:00
DottaandPaperclip a06fa49483 chore(runner): integrate pending v10 correlation for qualification only
Preserve both reviewed source and historical qualification evidence. CI-only branch; do not merge.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 22:14:15 -05:00
DottaandPaperclip 041d84760f fix(runner): align Cursor control IDs and retain native receipt evidence
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 22:11:14 -05:00
DottaandPaperclip 4eca3e8a02 fix(runner): correlate Copilot semantic tool results with bounded receipts
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 21:51:20 -05:00
Devin FoleyandPaperclip 4b9a6000f7 Add bounded evidence for directory lock timeouts (#14787)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent files use directory locks during collection and cleanup.
> - A lock timeout can fail finalization after the model turn completes.
> - The timeout currently identifies no owner state or waiting
operation.
> - This pull request adds bounded evidence to the existing run failure
report.
> - Operators can distinguish a known local holder from a possible old
lock without changing lock safety.

## Linked Issues or Issue Description

**What happened?**

A directory lock timeout does not distinguish active local work from an
owner record left by an earlier process. The stored execution stage can
also precede the cleanup operation that failed.

**Expected behavior**

The failure report should identify the waiting operation and expose
bounded ownership clues. It must preserve the timeout and keep unknown
ownership protected.

**Steps to reproduce**

Hold a directory merge lock while a second caller reaches its
acquisition deadline. The regression tests exercise a live holder and an
older owner record with a live PID.

Related: #9667 proposes stale-lock recovery under a single-server
assumption. This change only adds evidence and does not adopt that
assumption. #14575 and #14665 add other run failure diagnostics.

## What Changed

- Record lock owner state, capped age and wait duration, same-process
and process-age comparisons, and whether this module holds the lock.
- Label agent-directory release, collection, checkpoint, and warm
handoff timeouts with a fixed operation code.
- Validate each field before the existing event-local Sentry report
accepts it. Exclude owner records, PIDs, paths, and absolute timestamps.
- Limit the extra diagnostic owner read to 100 ms with best-effort
abort; malformed JSON is `invalid` and unreadable owner records remain
`unknown`.
- Document the diagnostic limits and verify that contenders never
reclaim protected locks.

## Verification

- Focused lock, diagnostic, real Sentry SDK, and database-backed
agent-directory tests: 126 passed, including stalled-read and
malformed/missing/unreadable-owner regression coverage.
- Final revision `0691613dcc`: all 54 reported checks successful, with
two intentionally skipped Storybook checks. Greptile: 5/5, zero
unresolved review threads; no merge conflicts.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm test:run`: complete suite coverage ran with the existing
repository shard flags: four general-server shards, four serialized
shards, two general-workspaces-a shards, and general-workspaces-b. The
full run is not green because of the base failures below.
- The broad run found 13 failures in the unchanged macOS skill-cache
tests. All 13 reproduce on the clean base revision. Open PR #14290
covers that existing failure.
- Two unchanged CLI archive tests hit their five-second limits during
the broad run; all 17 tests in that file pass on recheck. A CLI auth
socket error also cleared on recheck (19 tests), and its full serialized
shard passed on rerun.

## Risks

This is a diagnostic change, not a stale-lock fix. Owner observations
can race with release. Wall-clock shifts can affect the age comparison.
A local-holder flag covers only this module instance. None of these
fields authorizes reclamation or proves a file save. Lock acquisition,
release, retries, task status, and recovery guards retain their current
behavior. No schema change or deployment action is required.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The exact model build and context window were not exposed to this agent.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass (focused checks pass;
existing base failures are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 18:35:12 -07:00
DottaandPaperclip f7e36ba3e2 fix: isolate repository-free low-trust tasks in private directories (#14766)
## Thinking Path

> - Paperclip manages work by agents within company boundaries.
> - Email tasks can run under the low-trust review preset.
> - These tasks must use an isolated workspace and a sandbox.
> - The default workspace strategy assumed that the project had a Git
repository.
> - A project without a configured workspace failed before the agent
could start.
> - This change gives each such task a private directory and keeps the
sandbox requirement.

## Linked Issues or Issue Description

**What happened?**
An inbound email assigned to a low-trust agent failed with
`git_worktree_base_not_git_checkout` when its boundary project had no
configured workspace. Setup had accepted the project and sandbox.

**Expected behavior**
The agent can process email without a repository. Its workspace stays
isolated from other tasks and the shared agent home.

**Steps to reproduce**
1. Select a low-trust agent with an active sandbox and a project
boundary.
2. Leave the project without a configured workspace.
3. Receive an email through AgentMail.
4. Observe that startup fails before provider work starts.

**Paperclip version or commit**
Reproduced against `5edf55d73`.

**Deployment mode**
Hosted staging with sandbox execution.

Related: #13256 added email tasks. #13636 fixed default isolation for
projects without workspaces; the explicit isolation used by low-trust
tasks still needed this path.

## What Changed

- Select private task directories for low-trust sandbox tasks with no
configured workspace or explicit workspace strategy.
- Keep each directory scoped to its company and task. Retain files
across turns and reassignment and reject symlink paths and mismatched
workspace reuse.
- Preserve Git validation for configured workspaces and explicit
strategies, plus the existing authorization and remote gates for
referenced projects.
- Add a startup regression and directory isolation tests. Document the
supported repository-free path.

## Verification

- The startup regression failed before the fix with the same Git
validation error.
- 260 targeted email, workspace policy, heartbeat, referenced-project
and directory tests pass.
- Full `pnpm -r typecheck` and `pnpm build` pass on the latest commit.
- All CI checks, including the complete sharded test suite and canary
dry run, pass on `b4ccd9802b09b2e95499df72d48b4a3906b8c328`.
- The final commit also passes the same server shard locally: 60 files,
1,024 passed / 6 skipped tests. The earlier all-groups local run was
interrupted during follow-up edits; complete-suite verification comes
from CI on the final commit.
- Deployed the reviewed commit to staging and independently verified the
full serving SHA. Two real Codex runs in Daytona succeeded and finalized
the same private company/task workspace. The first wrote a 35-byte
marker; the second read the existing file without modifying it and
returned the independently verified SHA-256
`ce3bbeb44d07ca6822826d3a5945752a38d30b356d10829f3159a191e5aa92a6`.
- Live runtime caveat: Codex reported a nested `bwrap` loopback
permission error and used its configured escalated execution inside
Daytona. The outer Daytona sandbox remained active for both runs.
- The startup regression uses a real database, production trust checks,
workspace persistence, sandbox lease acquisition and realization, and a
fake provider. It checks reassignment and allows only an authorized
referenced project.
- The transfer regression runs production archive/sync-back/merge code
against distinct filesystem roots: create output in one sandbox, restore
it, then read and update it in a fresh sandbox. The provider I/O is
emulated; live staging verification is separate.

## Risks

- The new default applies only to low-trust sandbox tasks without
workspace configuration. Standard agents and explicit Git strategies
keep their existing behavior.
- Task directories retain work across turns and consume instance
storage. The change does not migrate or copy existing shared files.
- No database migration or credential changes are required.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
execution, and browser tools. The exact runtime model revision and
context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 18:06:07 -05:00
Devin FoleyandPaperclip 3bbb8d0f69 Add bounded AgentMail failure diagnostics (#14768)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - AgentMail connections create inboxes and handle email tasks.
> - A failed provider request currently records only its HTTP status.
> - The same 403 can mean a permission denial, a resource limit, or
another provider restriction.
> - This pull request records a fixed operation name and a documented,
allowlisted error code.
> - Operators can distinguish these failures without exposing provider
payloads or changing retry behavior.

## Linked Issues or Issue Description

**What happened?**

An AgentMail inbox creation failure reports only `AgentMail request
failed (403)`. The response body is deliberately excluded because it can
contain private mail or credentials. That also discards the provider
code needed to identify the cause.

**Expected behavior**

Keep the HTTP failure visible with a fixed operation name and a safe
provider code. Never copy arbitrary error text, resource identifiers,
suggested fixes, or URLs into diagnostics.

**Steps to reproduce**

1. Make an inbox creation request through `agentmailApi` with a fake
provider returning HTTP 403 and `code: "missing_permission"`.
2. Observe that the old error lacks the operation and provider code.
3. With this change, verify the error includes `operation=create_inbox,
code=missing_permission`, preserves status 403, and excludes all other
response fields.

Related work: #13256 introduced the AgentMail connection. The provider
documents stable codes in its [error
reference](https://docs.agentmail.to/errors).

## What Changed

- Add a fixed method/route-to-operation map and an allowlist of
documented provider codes.
- Read at most 8 KiB for diagnostics, with a one-second deadline. Cancel
unread bodies and preserve the HTTP error if reading or parsing fails.
- Keep the existing error prefix, status, retry delay, and failure
handling.
- Add regression coverage and document the diagnostic limits.

## Verification

- `pnpm exec vitest run server/src/__tests__/agentmail-api.test.ts` — 44
tests passed.
- `pnpm build` — passed.
- `pnpm -r typecheck` — passed before the review correction. Final `pnpm
--filter @paperclipai/server exec tsc --noEmit` also passed.
- Full local `pnpm test:run` did not finish successfully; three
company-skills-service failures were observed outside the changed
module. The final-head CI server suites passed. The remaining
workspaces-b CI retry covers an unrelated HTTP/2 port collision.
- The diff passed a scan for configured secrets, private deployment
references, and non-fixture email addresses.

## Risks

- A failed request can now wait up to one extra second while reading its
diagnostic code.
- New, missing, malformed, or oversized provider codes report `unknown`.
A future provider code needs an explicit allowlist update.
- This is a diagnostics change. It does not establish or repair the
cause of an existing provider denial.
- No schema, credential policy, or retry behavior changes.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, and code execution. The
exact served model ID and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 at 20853e31abacf53a32f1a63467ee208a719bbf00; the
locked-body finding is fixed and its thread resolved
- [x] I will address all Greptile and reviewer comments before
requesting merge


Final verification (September 30): all final-head GitHub checks pass at
`20853e31abacf53a32f1a63467ee208a719bbf00`, including the targeted
workspaces-b rerun after the unrelated EADDRINUSE failure. Greptile
scored 5/5 on this head and no review threads remain unresolved. The
branch is mergeable. The local full-suite run did not yield a passing
completion; CI completed successfully across all suites. This public PR
remains open for maintainer merge.

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 15:29:15 -07:00
DottaandPaperclip e2908fff5c fix: enforce terminal outcomes during recovered sandbox cleanup (#14767)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native runners can keep a reusable sandbox warm after successful
turns.
> - Failed turns must stop their sandbox before a later retry resumes
it.
> - Workspace recovery can release a lease outside the executor's normal
teardown.
> - Successful file copy-back can then retain a sandbox even when its
run failed.
> - This pull request checks the durable run outcome at the shared
release boundary.
> - Ordinary teardown and recovery now apply the same retention rule.

## Linked Issues or Issue Description

**What happened?**

A real Daytona verification on `5edf55d73` produced a failed native
turn. The runner process exited, but workspace recovery retained the
sandbox without stopping it. The recovery callback used successful
workspace copy-back to select warm retention. It bypassed the
terminal-state check in normal heartbeat teardown.

**Expected behavior**

Keep a sandbox running only after a successful run. Failed, cancelled,
timed-out, and interrupted runs must convert requested warm retention to
stop-and-retain. Preserve the existing ownership hold before release.

**Steps to reproduce**

1. Persist a terminal failed native run whose workspace copy-back
succeeds.
2. Release its environment lease from the recovery path with a stored
`keep_running` disposition.
3. Observe that the provider receives `keep_running` on the base commit.
4. With this fix, the provider receives `stop_and_retain` and the lease
status follows the durable run outcome.

**Paperclip version or commit**

Base: `5edf55d7350c7f08c9dd132c7e0f1421fa0bf2fb`.

**Deployment mode**

Native runner with a reusable Daytona sandbox.

Related: #14747 retires unsuccessful warm runner sessions. This change
closes the separate recovery lease-release path. Related search found no
duplicate fix.

## What Changed

- Read the durable run status at the shared lease-release boundary.
- Apply the existing terminal-outcome retention rule before calling the
environment runtime.
- Use that same durable status for the lease-state mapping.
- Add five database-backed regressions for four unsuccessful outcomes
and successful warm retention.
- Document the recovery rule.

## Verification

- Before the fix: the four unsuccessful-outcome regressions fail; the
successful case passes.
- After the fix: 163 tests pass across the lease-release, native
lifecycle, and explicit continuation suites.
- Repository typecheck and build pass.
- `pnpm test:run` encountered the existing local
`native-session-resume.test.ts:1125` assertion failure; a focused rerun
reproduced the same failure. This was also recorded before this
follow-up with the unmodified base executor. The full command was
stopped after that confirmation, so later local groups were not
completed.
- All 56 latest-head checks are successful or intentionally skipped (54
passed, 2 skipped), including the complete CI test groups and browser
E2E suite. Greptile is 5/5 on `e7da3b3ad`, with no unresolved review
comments.
- Deployed exact PR head `e7da3b3adbf7a13642c0e56f5b0f0c4666adf358` to
the staging workspace and independently verified the serving commit. A
real failed native turn persisted `keep_running` and completed workspace
finalization, reproducing the recovery-path conditions; Paperclip
automatically issued stop-and-retain, and an independent Daytona read
confirmed `stopped`. No manual stop was used.
- Retried that failed run through the public API after correcting its
temporary API-key credential. The same stopped sandbox resumed
successfully. Three successful turns retained one live runner PID/start
time and the same native, runner, and provider sessions.
- Verified exact canonical note contents, deletion persistence, and an
unchanged 8 MiB binary after every turn. Subsequent checkpoints
copied/hashed only 58 and 87 bytes. After sandbox deletion, canonical
files still matched. Temporary secrets were removed and
agent/configuration policy restored.
- Live acceptance used temporary API-key authentication. The separate
managed-subscription authentication issue and browser Retry control were
not tested by this campaign.

## Risks

- Recovery callers can no longer use stale success status to retain a
failed run's sandbox.
- The existing native ownership hold still blocks release while
ownership is unresolved.
- Successful warm turns and explicit destroy dispositions keep their
existing behavior.
- No database migration, API contract, dependency, or UI change is
included.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and test
analysis. The exact served model ID and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 17:20:25 -05:00
DottaandPaperclip 33f2b3a159 fix: separate GitHub tools and code review bot connections (#14750)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Connectors catalog lets people give agents tools or connect
agents to conversations.
> - GitHub put these two uses behind one card and an extra choice.
> - People should choose the connection they need from the catalog.
> - This pull request keeps GitHub for tools and adds GitHub Code Review
Bot as a separate card.
> - Each card opens its setup directly. Both use the existing connection
code.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

GitHub connector discovery and setup.

**Current behavior**

With chat connectors enabled, GitHub opens a menu that asks whether to
use tools or create a bot. Saved tools and bots share the same catalog
entry.

**Proposed behavior**

GitHub opens tool account access. GitHub Code Review Bot opens agent
selection. Saved bots and drafts appear under the bot card.
Chat-disabled instances show only GitHub tools.

**Reason and benefit**

The catalog names the two uses and removes an extra setup choice. The
bot keeps the existing GitHub provider, credentials, endpoint IDs, setup
steps, and runtime.

**Additional context**

Related work: https://github.com/paperclipai/paperclip/pull/12843 and
https://github.com/paperclipai/paperclip/pull/14594 established GitHub
account identity. This change preserves that tool flow. No duplicate
catalog split was found.

## What Changed

- Split the generated app definitions into GitHub tools and GitHub Code
Review Bot. Reuse the existing GitHub logo and channel method.
- Open bot setup directly, including old resume and reconnect links.
- Put existing bot endpoints and drafts under the bot card. Hide
duplicate internal chat applications.
- Keep pasted GitHub URLs mapped to the tool connection.
- Add seven Storybook states for the catalog, saved connections,
disabled chat, both setup paths, mobile, and light mode.
- Fix narrow-screen bot rows so the label cannot overlap status and
setup actions.
- Update catalog, route, browser, and API tests, plus the GitHub
connector guide.

## Verification

- [Hosted
Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fgithub-review-connection/?path=/story/connections-github-and-code-review-bot--catalog):
seven states built from this branch. The deployment passed its
public-file verification.
- All GitHub checks pass on `d13a2cd53561645bb2a15c6f8e75a61a936d6459`.
Two optional Storybook jobs skip under their normal trigger rules; the
manual Storybook deployment passes. The branch has no merge conflicts.
- Greptile: 5/5 on the current head, with no review comments or
unresolved threads.
- `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and `pnpm
build-storybook` passed. The final Storybook fixture also passed UI
typecheck and the hosted build.
- Targeted catalog, URL matching, routing, grouping, brand, and chat UI
contract tests passed.
- GitHub provider browser tests: 2 passed. These cover direct tool setup
and the bot setup and management lifecycle with provider responses
mocked.
- Embedded-browser test on an isolated local instance: opened both
cards, selected an agent, saved a bot draft, and resumed the same
endpoint under the bot card after a reload.
- Storybook Tool Setup and Bot Setup assertions pass in the published
preview. Chat Disabled assertions pass locally. Inspected mobile and
light mode, including the draft-row layout and official GitHub marks.
- Local full-suite limitation: `pnpm test:run` was not clean. A
cross-company route assertion failed in the aggregate run and passed in
isolation; a workspace-runtime test reached its 30-second hook timeout.
Some isolated database reruns skipped when the embedded-PostgreSQL
availability probe failed. The local aggregate was stopped after CI
completed. The corresponding full CI suites pass all 360 tool-access
tests and all 162 workspace-runtime tests.
- No live GitHub authorization or installation was performed. The
isolated instance correctly stopped at the cloud enrollment or public
HTTPS prerequisites.

## Risks

- Low scope: catalog presentation and routing change. There is no
database migration or provider credential change.
- Existing GitHub bot URLs now open bot setup directly. The tool route
remains `/apps/connect?source=github`.
- The bot remains behind the existing chat-connectors feature flag.
Existing endpoints retain `provider: github`.
- Channel applications are represented by endpoint rows. Regression
tests cover legacy bot applications, tools, active bots, and drafts
together.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, and
embedded-browser tools. The exact deployed model ID, context window
size, and reasoning setting are not exposed to this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 17:15:02 -05:00
DottaandPaperclip 6ce62cac75 test(server): report routine telemetry mock failure state
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 17:07:28 -05:00
DottaandPaperclip 018993140f feat: let agents name prompt-only tasks (#14761)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users create tasks with a title and a description.
> - A required title adds work when the prompt already explains the
request.
> - An agent can name the task once it reads that request.
> - This pull request accepts prompt-only tasks and starts them with a
short prompt slice.
> - A scoped title tool lets the assigned agent replace that slice early
without changing execution state.
> - A live browser eval checks the real agent call, saved title, audit
entry, and preservation of user titles.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: task creation, shared contracts, database, server, runner
tools, and board UI.

**Problem or motivation**

Users must currently write a title before they can submit a detailed
task prompt. The agent has enough context to write a useful title
itself.

**Proposed solution**

Make the title optional when a description is present. Save the first
120 characters of the normalized prompt as a provisional title. Ask the
assigned agent to call `set_task_title` early. Use an atomic
provisional-title guard to preserve titles supplied or edited by users.
Keep explicit titles supported.

Related: #14543 and #14556 concern empty-title submission. This change
intentionally enables that submission when a prompt is present, instead
of requiring a title.

## What Changed

- Add the `titleNeedsGeneration` field with an idempotent migration.
Keep existing titles unchanged.
- Add `PUT /api/issues/:id/title` and the native and legacy
`set_task_title` tool. Enforce company access, active-run ownership,
shared, bounded retry receipts across native/HTTP calls, and
transactional audit logging. Refresh external-object links after commit,
with the same feature gate and plugin detectors as ordinary title edits.
- Add early naming guidance in Standard, Ask, and Plan task context.
Preserve the description, status, and assignment.
- Allow prompt-only root and child task creation, plus draft restoration
in the New Task dialog. Keep user titles supported.
- Add an opt-in Product E2E suite for prompt-only Standard and Ask
tasks, plus an explicit-title control. It checks actual provider calls
within the first five tools, persisted state, audit attribution, and the
reloaded UI.
- Preserve a closed vocabulary of API key maintenance phrases in
declared prose while rejecting opaque credential suffixes. Add one
bounded naming retry after wording is rejected, without treating the
rejected call as a saved title.
- Repair the native cleanup receipt check exposed during full
verification: accept matching input digests, retain legacy input checks,
and reject conflicting receipts.

## Verification

- Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3
passed** with native Codex `gpt-5.4-mini`, first attempts only,
automatic retries disabled. Standard and Ask each saved “Rotate expired
API key” on their first tool call, with matching persisted state and a
single same-run audit entry. The explicit-title control retained its
user title with zero title writes. All three verified the reloaded
browser UI.
- Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns
are retained separately; they exposed credential-prose handling and
prompted the naming recovery fix. No failed result was regraded or
deleted.
- Reproduce with `pnpm test:e2e:runner -- --id
task-titles.runner-codex-mini.local.prompt-title-standard --id
task-titles.runner-codex-mini.local.prompt-title-ask --id
task-titles.runner-codex-mini.local.preserve-explicit-title
--max-automatic-retries 0` and an authorized provider key.
- Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit.
The runner build used the configured external eval source tree.
- Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck
and UI token gates passed.
- Title API/native regressions cover prompt-only and explicit child
creation, user edits, ownership/company isolation, external reference
refresh, cross-surface retry replay, and the 64-key limit without
receipt eviction. All passed. Prompt-context coverage: **44 tests
passed**.
- Rust credential regressions: **35 tests passed**, including benign
maintenance qualifiers and opaque credential rejection in every declared
prose field. Catalog/report reconciliation: **28 tests passed**. Native
recovery: **560 tests passed**.
- Broad local `pnpm test:run`: **14,555 tests passed** in the general
server group; two suites failed to initialize embedded PostgreSQL and
the existing 40,000-file Git streaming stress test exceeded its
300-second macOS timeout. All three suites then passed in isolation (**5
tests passed**) without code or timeout changes. The original full local
command exited nonzero and is not being represented as a clean full run.
- Latest-head GitHub checks are green: **53 passed, 4 skipped, zero
failed or pending**, including all test shards and the canary packaging
dry run. Greptile reviewed the same commit at **5/5**, with zero
unresolved review threads.

## Risks

- The additive database field must reach the server and UI together. The
migration uses `IF NOT EXISTS` and defaults existing tasks to a final
title.
- Title generation depends on the assigned agent running. Tasks without
a run keep their provisional title.
- Live qualification covers the native Codex path in Standard and Ask
modes. API/legacy and Plan behavior have deterministic coverage.
- The credential-prose exception validates the entire suffix against a
closed maintenance vocabulary. Unknown suffixes, assignments, quoted
values, credential prefixes, and diagnostics retain strict checks.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, tool use, and code
execution. The exact deployment ID and context window are not exposed in
this session. The live eval uses the native Codex `gpt-5.4-mini`
profile.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 16:48:24 -05:00
Devin FoleyandPaperclip 0ea6b10967 Record ACP activity and workspace restore failure evidence (#14665)
## Thinking Path

Paperclip records terminal run failures for operators. A timeout's own
log and cleanup output update the run's last-output timestamp, so that
timestamp can make a long-silent provider look active. Snapshot runtime
activity before finalization and include the saved workspace restore
classification to make the next failure actionable without copying tool
payloads.

## Linked Issues or Issue Description

**What existing behavior does this improve?**
Terminal run diagnostics in the existing opt-in Sentry integration.

**Current behavior**
Reports cannot distinguish runtime events from finalization logging and
omit the already-persisted workspace restore code. A later successful
run also does not establish that earlier workspace files were restored.

**Proposed behavior**
Record runtime-event age/count and pending-tool inventory at
finalization, before status reads and cleanup. Forward only finite
counts, a completeness boolean, and known restore codes through the
existing reporter.

**Reason and benefit**
Operators can distinguish a silent turn with unfinished tools from
recent runtime activity and see restore failures without retrieving
private run output. Neither signal certifies productive work or
successful recovery.

**Breaking changes**
None. Error grouping, execution deadlines, cancellation, recovery
policy, and the Sentry opt-in remain unchanged.

Related diagnostic work: #14573, #14575, #14639.

## What Changed

- Snapshot ACP activity before success/failure finalization, including
thrown relay failures.
- Forward bounded numeric/boolean evidence and shared workspace restore
codes; exclude commands, tool IDs, paths, and arbitrary result data.
- Document limitations and test silence, empty streams, timeout, cleanup
delay, incomplete tool inventory, and privacy.

## Verification

- `pnpm -r typecheck` passed after the final implementation.
- Changed suites: 252 tests passed; all 29 database reporter tests
subsequently passed after restoring the embedded-Postgres package
library symlinks. The migration test also passed (30 database cases
total).
- `pnpm build` passed during implementation. Final head
`fe78dba6f592b1abccac7cdbf341bd2e0b0d30cb` passed all 53 CI checks,
including complete test coverage, typecheck, build, and browser/runner
gates; two unrelated checks intentionally skipped.
- Full local test attempts initially hit missing embedded-Postgres
library symlinks; the dependency setup was repaired and database tests
passed. Duplicate local full-suite runs were stopped after full CI
completed. This PR does not claim a completed full local suite.
- Review regression: completed, failed, and cancelled tools are excluded
from the pending count; focused activity/timeout tests and adapter-utils
typecheck passed.
- Reviewed the diff for secrets, customer data, and internal references.

## Risks

Low risk, diagnostic-only. The existing tool inventory is incomplete for
some runtime events, so the report carries its completeness flag. Event
age is measured at finalization and does not prove useful work or
identify the underlying provider failure. No schema changes or new
capture gate.

## Model Used

OpenAI GPT-6 via Codex, with repository inspection, code execution, and
tests. Exact model build identifier is not exposed by this session.

## Checklist

- [x] Thinking path and model are specified
- [x] Checked ROADMAP.md; this is a maintenance correction, not planned
feature work
- [x] Searched for duplicate and related PRs
- [x] Described the issue using the enhancement template
- [x] No internal issue references, customer data, or private instance
links
- [x] Descriptive branch name
- [x] Focused regression tests pass
- [x] Added tests and updated documentation
- [x] Risks documented
- [x] Required validation and CI gates are green (full suite validated
in CI; local scope documented above)
- [x] Greptile is 5/5 with no unresolved findings
- [x] I will address review comments before requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 14:27:53 -07:00
DottaandPaperclip 5edf55d735 fix: retire failed warm sessions before sandbox stop (#14747)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native runner sessions can stay warm between turns in a reusable
sandbox.
> - Heartbeat stops that sandbox when a turn fails or is cancelled.
> - The native executor treated every returned terminal result as a
successful warm release.
> - A failed session could therefore retain a transport for a stopped
sandbox.
> - This pull request retires unsuccessful sessions before heartbeat
stops their sandbox.
> - A retry can start without inheriting that stale transport.
Successful turns stay warm.

## Linked Issues or Issue Description

**What happened?**

A structured failed or cancelled terminal result did not throw. The host
kept its native session in the warm cache even though heartbeat stopped
the reusable sandbox. Later session retirement could use the stopped
provider transport and reject the retry.

**Expected behavior**

Retire the unsuccessful session and collect its managed files before
returning to heartbeat. Keep successful sessions warm.

**Steps to reproduce**

1. Run a native session with a warm lifecycle and a reusable sandbox.
2. Return a structured failed or cancelled result from the provider.
3. Stop the sandbox after the executor returns.
4. Retry with a changed native session identity.
5. Verify that the previous session was closed before step 3 and is not
closed again during retry.

**Paperclip version or commit**

Developed from `b54b2dc35`, the current `origin/master` at
implementation time.

**Deployment mode**

Native runner with a reusable Daytona sandbox. The same warm-session
release path also serves local providers.

Related work: #14735 added incremental managed-file checkpoints for warm
turns. #14734 preserves tool outcomes during shutdown. This change fixes
the host's handling of unsuccessful terminal results.

## What Changed

- Retain a warm session only when its terminal run state is `succeeded`.
- Use the existing failed-session retirement path for structured
failures and cancellations.
- Preserve the existing checkpoint-before-retirement path, including
when provider shutdown fails. Collect stopped files after successful
shutdown, including edits made during shutdown.
- Retire the warm owner if the checkpoint or its receipt callback
rejects, then propagate the initiating error.
- Add ten regression cases for failure and cancellation, with and
without managed files. They check cleanup order, cache removal, retry
with a new session identity, and edits preserved when close rejects.
- Document the lifecycle rule.

## Verification

- The failed and cancelled managed-file regressions fail on unpatched
master because the provider is not closed.
- All 516 native executor tests pass, including ten new regressions.
- `pnpm -r typecheck` and `pnpm build` pass. `pnpm test:run` finished
its general-server group with 14,537 passed, 75 skipped, and one
existing failure in `native-session-resume.test.ts`
(`retainedNativeCleanupJournalMatches`, line 1125). A separate run using
the unmodified `origin/master` executor fails the same assertion. The
command stops before later local groups; all equivalent CI groups pass
on this head.
- All 56 PR checks are successful or intentionally skipped. Greptile
reviewed this head at 5/5 with no remaining findings.
- A real Daytona public-API probe forced a Codex authentication failure,
waited for a confirmed sandbox stop, corrected the credential, and
retried successfully in the same resumed sandbox. It verified the agent
note in canonical storage and deleted the test sandbox.
- Pumpkin staging on this exact commit: forced a structured
authentication failure, confirmed sandbox stop, corrected the
credential, and retried successfully in that same sandbox. Three
successful turns kept the same live runner PID/start ticks, native
session, runner instance, and provider session. Canonical downloads
verified the note, all 8 MiB binary bytes, and deletion persistence.
Subsequent checkpoints hashed/copied only 52 and 78 bytes. After sandbox
deletion, canonical files still matched. Temporary secrets, agent
configuration, and test policy were cleaned up or restored.
- Unpatched live probes with both 5-second and 90-second retry delays
also succeeded. The unit regressions prove incorrect retention; the
original stopped-lease exception was not reproduced in those probes. The
staging recovery campaign used the public retry API after an explicit
reset of the disposable task’s already-deleted old sandbox/session. The
UI did not expose a Retry control on the inspected task or run detail
views.
- Temporary API-key authentication was used for the campaign. This
change does not repair the original managed-subscription authentication
401.

## Risks

- Failed and cancelled turns now close their provider sooner. Successful
warm-turn behavior is unchanged.
- Cleanup uses the existing failed-session path. Its existing handling
of cleanup errors is unchanged.
- No database migration, API contract change, or dependency change is
included.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and test
analysis. The exact served model ID and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (516 relevant executor
tests; the full local suite has the independently reproduced baseline
failure documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 15:36:06 -05:00
DottaandPaperclip ad55d0a281 fix(connections): repair personal credentials and request write access (#14739)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use Apps through a gateway that checks identity, company
access, and action policies.
> - Personal pasted credentials can point to company secrets. Setup can
show success while the gateway rejects every call.
> - Several OAuth methods also omit the scopes needed for their
supported write actions.
> - This pull request gives setup, health checks, and invocation the
same credential rules. Owners repair existing connections by
reconnecting.
> - New connections request reviewed permissions for their supported
actions. Read-only choices remain available under Advanced.
> - Agents can use the connections people give them, while existing
consent, identity boundaries, and action restrictions remain enforced.

## Linked Issues or Issue Description

Refs #14009 and #14008. This addresses the personal-credential defect.
The separate GitHub organization-identity selection defect is outside
this change.

Related work: #13942 fixed part of new personal-key setup. #14200
independently fixes legacy personal reconnect and protects managed-agent
profile credentials during removal. This PR covers that ownership
invariant across key and secret-URL setup, reconnect, health, discovery,
and invocation, and keeps owner reconnect as the repair path. #14059
tracks requested versus provider-asserted OAuth scopes; it remains
separate work. I searched open PRs and issues for Zapier, Airtable
scopes, connector writes, and personal credential failures.

**What happened?**

A Zapier secret URL saved through personal setup can become a company
secret referenced by a user grant. Health checks bypass the gateway's
ownership check, so the connection appears healthy but calls fail with
`grant_credential_invalid`. Custom-header paths can also receive a
duplicate `credentials.` prefix. Omitted OAuth scopes make write access
depend on provider defaults.

**Expected behavior**

Personal invocation credentials belong to the selected user. Setup,
health, and actual calls enforce the same rule. New connections request
documented permissions for supported read and write actions. Existing
tokens gain no permissions without provider consent.

**Steps to reproduce**

1. Connect Zapier or a generic secret URL with the personal identity.
2. Allow an agent to use the connection and complete setup.
3. Invoke a tool through a run-scoped gateway. The legacy layout fails
ownership validation despite successful setup.

**Paperclip version or commit**

The implementation started from
`44736c9c7c67b7b646ead9d51721db10f5b83835` and was rebased onto master
at `94e8dec56`.

**Deployment mode**

Built from source. Regression tests use isolated PostgreSQL fixtures and
controlled MCP transports.

## What Changed

- Share credential writing, ownership validation, and canonical paths
across initial setup, resume, reconnect, rotation, health, discovery,
and gateway calls. Keep OAuth client-registration secrets separate from
invocation credentials.
- Existing personal connections with company-scoped credentials require
owner reconnect with a fresh key or secret URL. Reconnect creates a
correctly owned value and updates the existing grant and declarations.
There is no automatic ownership backfill or new startup hook.
- Preserve PostgreSQL timestamp precision when reconnect checks whether
a grant changed. Previously, converting the timestamp to a JavaScript
Date could reject reconnect with a false concurrent-change error.
- Protect credentials used by other grants, connections, bindings,
managed-agent profiles, routine triggers, or secret proposals from
connection removal.
- Review all 117 tool methods, including 84 OAuth methods. Record
explicit scopes or documented provider-default exceptions with official
evidence. Add Airtable's seven scopes, Hugging Face repository/job
scopes, and other documented MCP permissions.
- Prefer available write-capable methods. Put explicit read-only choices
under Advanced. Explain pasted-key permissions and offer reconnect for
missing OAuth consent. Preserve existing grants, policies, Google
availability gates, and curated scope allowlists.
- Reconnect generic secret URLs and custom headers using their stored
credential fields. Refresh the catalog after setup, correct reconnect
feedback and error guidance, and let Cancel exit invalid setup while
Save & exit retains draft-saving behavior.
- Apply ownership checks to the new GitHub repository/skill connection
picker. Align the permission audit with the Google scope reductions
merged on master.
- Add run-scoped gateway, ownership, owner-reconnect, OAuth URL,
insufficient-scope, UI, and catalog-wide regression coverage. Update the
connector playbook and permission audit.

## Verification

Latest commit `97bc0b86e0eae0ec892e4ac44beff1a66164b20e` passes all
CI/status gates (55 completed check runs, no failures or pending checks)
and has a completed Greptile review at **5/5 with no outstanding
findings**. GitHub reports the PR as mergeable/CLEAN.

- **Embedded browser:** used the actual server and built UI from this
worktree, a fresh isolated database, and local HTTP MCP fixtures.
Completed personal bearer-key, secret-URL, and custom-header setup;
reproduced the legacy ownership failure; reconnected through the owner’s
form; and completed writes afterward. Read-back was verified for
bearer-key and secret-URL connections. Public organization-wide setup
appeared immediately in Browse without reload. Zapier URL
validation/Cancel and Google’s enrollment gate were also exercised.
- **Persistence and invocation:** verified user ownership, canonical
`credentials.authorization` / `remote.url` / `headers.X-Api-Key`
declarations, and unchanged connection/grant identity. The old company
secrets retain their ownership. Separate HTTP calls through an actual
run-scoped gateway session completed a write and read-back.
- **Backend coverage:** the final gateway suite passes all 82 cases,
including catalog Zapier and generic inline reconnect. It checks
company/user isolation, canonical declarations, same-endpoint URL
validation, fresh credentials, retained restrictions, and real gateway
read/write execution using fixture transport. A timestamp with
PostgreSQL microseconds covers the former false reconnect conflict.
- **Local checks:** 368 catalog, gateway, repository, and UI tests
passed before the final extra Zapier case; 49 GitHub skill access tests
also passed. All three Apps browser regressions pass, including
reconnect through the actual form and catalog visibility without reload.
Full `pnpm -r typecheck`, `pnpm build`, server typecheck after the final
patch, and token gates passed. Full tool-access service runs hit varying
15-second Google fixture timeouts; both affected cases and the updated
reconnect assertion pass in isolation (3 tests). The complete test
matrix passes in CI on this head.
- **Verification limits:** no live provider account was available for
Zapier/Airtable/OAuth consent or account-bound write proof. Public
metadata and local fixtures do not establish provider consent. The
original development database clone failed on a pre-existing missing
`tool_connections_transport_check` constraint; browser acceptance used a
fresh isolated database created by the normal CLI onboarding flow.

## Risks

- Existing broken personal connections stay unusable until their owner
reconnects. Health, discovery, and invocation return an actionable
ownership error; startup does not rewrite credential ownership.
- Scope changes affect new authorization requests. Providers may still
require resource selection, account roles, paid plans, or app
verification. Existing consent and action restrictions remain unchanged.
- Shared credentials are retained rather than reassigned or revoked.
Provider-default exceptions and unavailable live checks are documented
in `doc/connections/CONNECTOR-PERMISSION-AUDIT.md`.
- No new endpoint, database table, lockfile change, or CI workflow
change is included.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code editing, shell
execution, web research, and browser tools. The exact deployment model
ID and context window were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 14:32:34 -05:00
DottaandPaperclip cbd278dc03 fix(interactions): derive chat recipients and validate explicit users (#14742)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - Agents use saved questions to get human input and continue the same
task.
> - The standard question example recently told models to copy a user
ID.
> - A model can omit an identity prefix and create a question its
intended recipient cannot answer.
> - Agent Chat already knows the conversation owner, so the server can
supply that identity.
> - This pull request removes the blanket instruction and validates
explicit recipients before saving.
> - Ordinary questions stay simple, and explicit addressing remains
available for decisions that need a particular person.

## Linked Issues or Issue Description

Refs #14707, #14188. Related: #14238 handles legacy email recipients;
this change prevents invalid recipients in new cards and retains exact
ID matching.

**What happened?**

A model copied a Cloud user ID without its prefix into
`addresseeUserId`. Creation succeeded. The intended user's answer then
failed the exact recipient check.

**Expected behavior**

Ordinary chat questions use the saved conversation owner. A task may
optionally name a specific recipient. The API rejects an unknown or
unauthorized recipient before it creates a card.

**Steps to reproduce**

Create a chat question for a user whose ID is `paperclip-id:example`.
Supply `example` as the addressee. Before this change, creation accepts
the invalid recipient and the owner cannot answer. With this change,
creation returns 422. Omitting the field saves the full owner ID and
allows that owner to answer.

## What Changed

- Remove `addresseeUserId` from standard question examples and remove
the blanket requester-ID instruction.
- Derive the recipient of ordinary chat questions from the persisted
conversation owner. Reject conflicting explicit user IDs.
- Keep explicit task recipients optional. Validate supplied user IDs
with the existing board mutation policy, including company, viewer, and
Cloud restrictions.
- Preserve explicit agent routing, connector intents, confirmations,
exact recipient checks, idempotent retries, and no-login local-board
authority in local-trusted mode.
- Update the blocker grader to accept an omitted recipient and verify
the actual requester answered.
- Add database and HTTP tests for prefixed identities, denied
recipients, concurrent retries, saved answers, and response delivery.

## Verification

- Database interaction service suite: 90 tests passed, including
implicit local-board creation/answering and authenticated/Cloud denial.
- Interaction HTTP route suite: 84 tests passed.
- Affected interaction/native/connector/documentation suites: 231 tests
passed across six files after valid-user fixtures were updated.
- Resolver and interaction unit suites: 29 tests passed.
- Product E2E unit/calibration suite: 793 tests passed; Product E2E
typecheck and blocker catalog discovery passed.
- Generated API-reference and capability contract checks passed.
- `pnpm -r typecheck` and `pnpm build` passed.
- Full local `pnpm test:run` did not finish green: its initial
general-server pass had 14,416 passing assertions, one unrelated
native-resume assertion failure on macOS, and three teardowns from an
intermediate fixture cleanup fixed above. Separate broad local groups
also encountered timeout/live-port failures under host load. Local UI
(7,026), CLI (502), shared (817), and skills-catalog (20) tests passed;
the complete final-head CI matrix is the broad verification gate.
- After two CI cold-start readiness timeouts, a separate test-only
commit gives the first exposure lifecycle fixture the existing normal
30-second readiness budget. Its real HTTP, ordering, and cleanup
assertions remain intact; the targeted case and final Linux CI shard
passed. Production deadlines are unchanged.
- A separate OpenCode fixture failed twice on GitHub-hosted Ubuntu
because its cached Node executable was group-writable; the same case
passed on AWS runners. The fixture now qualifies its own Linux copy with
mode `0500` and the actual copy digest. Host files and production
security checks are unchanged. The focused macOS case passed; the new
Linux-copy branch also passed on the final AWS-hosted Linux runner
(1,125 passing Runner tests, 3 skipped). The final run was not on a
GitHub-hosted runner.
- Final-head [CI run
36762078176](https://github.com/paperclipai/paperclip/actions/runs/36762078176)
passed for `116b968b24fa0a8c5724a7bf96e73a8dda5f0425`: 54 successful
checks and two conditional Storybook skips, with no pending or failed
checks. The 27 general/serialized test jobs reported 28,635 passing
tests. Typecheck, build, Runner, browser E2E, and Canary gates passed.
Greptile reviewed that exact head at 5/5; both review threads are
resolved, with no open follow-ups.
- No live provider replay is claimed by this PR.

## Risks

- New explicitly addressed cards reject users who cannot mutate the
issue, including viewers, inactive members, and invalid IDs. Callers
that supplied invalid recipients must correct their request.
- Existing addressed cards are not rewritten. Existing authorization
checks remain strict.
- Chat inference applies only to questions without an agent addressee.
Connector intents and governed confirmations retain their own recipient
paths.
- No schema change or migration is required.

## Model Used

OpenAI Codex, GPT-6 (exact serving variant and context window are not
exposed in this environment). Used reasoning, tool use, code editing,
and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 14:16:15 -05:00
Dotta e4bf459e31 fix(server): fence native retry cancellation through final commit
Recheck retry eligibility under the coordinator lock before creating a cancellation intent. Preserve later run and coordinator outcomes with a failed-only acknowledged-result CAS, while retaining same-intent recovery and NOWAIT conflict handling.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
(cherry picked from commit 29230e2000)
2026-09-30 13:54:30 -05:00
Dotta 45b977bca0 fix(server): admit correlated Stop for durable native retries
Check failed-run retry eligibility under the run and coordinator locks, fail closed on concurrent coordinator claims, and preserve same-caller recovery after the audited intent disables a retry. Keep terminal failures and foreign actors or intents rejected.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
(cherry picked from commit 79db1c2a6f)
2026-09-30 13:36:15 -05:00
DottaandPaperclip b54b2dc35c fix: preserve warm Codex turns with incremental managed file checkpoints (#14735)
## Thinking Path

> - Paperclip manages AI agents and keeps their instructions and files
durable.
> - Native Codex runners can keep a process alive between compatible
turns.
> - Managed file collection stopped that process after each turn, which
defeated warm reuse.
> - Agent folders can contain large images and other files, so full
copies on every turn are expensive.
> - This change keeps one managed directory for the live session and
saves only file changes after each turn.
> - Ownership, authorization, instruction changes, and process
retirement still control when reuse is safe.

## Linked Issues or Issue Description

Related: #13710 introduced native warm session reuse. This fixes managed
file collection that still forced those sessions to stop. No duplicate
open PR or issue was found.

**What happened?**

With managed instructions and warm native Codex enabled, consecutive
turns reused a Daytona sandbox but started a new runner process each
time. The managed directory collector required process termination
before saving files.

**Expected behavior**

Compatible turns keep the same process and managed `AGENT_HOME`. Each
completed turn saves added, changed, and deleted files before the next
turn starts. Unchanged large files do not transfer again.

**Steps to reproduce**

1. Use a native Codex agent with managed instructions and a reusable
Daytona environment.
2. Enable warm session reuse and run three turns on the same task.
3. Write a large binary on the first turn, edit a small note on each
turn, and delete a file on the second turn.
4. Compare process identity across turns and read the canonical files
through the public agent-files API.

**Paperclip version or commit**

Reproduced on `d30b03bd8c17604cdab1533eeeeb087aba30e8b1`.

**Deployment mode**

Local server with remote Daytona execution; cloud native runner uses the
same path.

## What Changed

- Retain the managed directory only for the verified owner of a live
native Codex session.
- Checkpoint each completed turn before releasing the session for reuse.
Retry unstable captures, then stop and collect when a warm checkpoint
cannot be validated.
- Compare metadata and cached hashes, stream only changed file payloads,
record deletions, and validate path, content, quota, and authorization
before saving.
- Rotate sessions when canonical files, loaded instructions,
credentials, or launch policy change. Fence stale collection and cleanup
callbacks from later owners.
- Keep cleanup and recovery aware of the current session owner. Recheck
canonical files under the writer lock at handoff, attach the successor
collector before fallible bookkeeping, and emit one final save receipt
on checkpoint fallback. Preserve storage warnings across unchanged
checkpoints.
- Add regression coverage and a three-turn Daytona test with independent
public API file checks, an unchanged 8 MiB binary, deletion checks, and
strict process identity checks.
- Document checkpoint consistency, lifecycle behavior, and local run-log
counters.
- Replace a timing assumption in the Daytona teardown test with explicit
transfer-arrival gates after CI exposed an unset release callback.

## Verification

- Full local `pnpm -r typecheck` and `pnpm build` passed. Server checks
were repeated after the final storage-warning fix.
- Runner E2E typecheck and 749 runner E2E unit tests passed.
- Focused file checkpoint, directory ownership, instruction collection,
native session, and merge tests passed. After review fixes, the
managed-directory and native-session suites passed 550 tests, including
intervening canonical edits, same-run fresh restore, failed handoff
collection, and one-call fallback collection. Server typecheck passed
again. The Daytona plugin suite passed 218 tests. The quota-warning
regression failed before the fix and passed afterward.
- Three real Daytona campaigns passed before the final handoff review
fixes. The latest kept PID 547 across all three turns. The first
checkpoint copied 8,388,635 bytes; the next two copied 36 and 54 bytes.
Public API reads verified the binary, note contents, and deletion after
every turn. Test cleanup deleted the sandbox.
- The final head was also deployed to an isolated cloud staging instance
and passed three UI-triggered native Codex turns with managed
instructions. All three retained the same process ID/start time, native
session, provider session, runner instance, and Daytona sandbox.
Checkpoints copied 8,388,643 bytes on turn 1, then only 52 and 78 bytes
on turns 2 and 3; those warm captures also hashed only 52 and 78 bytes.
Independent canonical API reads verified every byte of the unchanged 8
MiB binary and the exact note contents after every turn; the deleted
file returned 404 after turns 2 and 3. After restoring the original
lifecycle and agent-auth configuration, removing the temporary secret,
pausing the test agent, and deleting both test sandboxes, independent
canonical API reads still verified the entire binary, the final 78-byte
three-line note, and the deletion. The native runner flag remained
enabled and the final serving revision remained the PR head.
- Two earlier staging attempts are preserved as failures and are
excluded from the acceptance result: a saved ChatGPT login failed with a
provider routing 401, and its subsequent stopped-sandbox retry failed
before provider startup with a closed-lease admission error. The
successful campaign used a fresh sandbox and a temporary encrypted
API-key binding. The stopped-lease retry remains unexplained; this
campaign does not establish recovery of that failed sandbox.
- All [Paperclip CI
gates](https://github.com/paperclipai/paperclip/actions/runs/36750397355)
pass on `26ef2ef56a389259246809805c0b34a4747eb86b`, including full test
partitions, build, typecheck, runner verification, E2E shards, and the
Canary clean public-npm install. Greptile reviewed that exact head at
5/5 with no unresolved review threads or outstanding findings.
- Full local repository coverage used the existing CI partitions, but
the 40,000-file Git streaming stress test timed out and its local retry
was interrupted by macOS thermal emergency sleep; this is not a green
full local suite claim. The exact stress test passed on the final head
in [CI server shard
2/12](https://github.com/paperclipai/paperclip/actions/runs/36750397355/job/110008294290),
in 111.9 seconds.
- Repeat the live test with configured credentials and a Linux runner
artifact: `pnpm test:e2e:runner -- --id
daytona-warm-continuity.runner-codex.daytona.warm-three-turn`.

## Risks

- This is a file-level checkpoint, not an atomic snapshot of the whole
folder. Background writes after a capture are saved by the next
checkpoint or final stopped collection.
- Metadata scans still visit all paths. Modified files transfer in full;
unchanged files do not rehash or transfer.
- Incorrect ownership or reuse could collect the wrong directory. Run
ownership fences, current authorization, stable capture validation, and
stopped collection fallbacks are covered by tests.
- Warm reuse remains opt-in. No database migration or fleet default
changes.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code editing, tool use, and
test execution. The exact serving model ID and context-window size are
not exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 13:33:37 -05:00
DottaandPaperclip d432dc7fa3 Add GitHub-synced skill sources (#14713)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Company skills supply instructions and files to those agents.
> - GitHub imports already exist, but users cannot manage repositories
as skill sources.
> - Repository refresh also needs caller-authorized access and complete
local packages.
> - This pull request adds Sources inside Skills and reuses GitHub
connections from Apps.
> - Installed snapshots let agents use skills without fetching GitHub
during a run.
> - Manual refresh preserves skill identity and leaves failed imports on
their last good version.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: skills UI, server, database, shared contracts, and
runtime materialization.

**Problem or motivation**

Users keep skills in GitHub repositories. They need a clear way to
select, import, and refresh those skills. Existing imports do not expose
repository management or consistently preserve supporting files.

**Proposed solution**

Add company-scoped skill sources. Browse repositories from all
accessible GitHub connections, or paste a public repository or branch
URL. Select whole skill packages, inspect included files and reference
warnings, and install complete, immutable snapshots. Refresh each source
manually.

**Alternatives considered**

Project repository settings hide the workflow from Skills. A second
GitHub connector would duplicate credentials and grants. Upstream
editing and PR creation are separate work.

**Roadmap alignment**

This implements the Skills Manager direction in ROADMAP.md. The
maintainer requested this scope and reviewed the component and full-app
journey stories before implementation.

Related reports: Refs #10285, Refs #10949, Refs #13464. Related work:
#14356, #13656, #9268.

## What Changed

- Add source and entry records, an idempotent migration, company-scoped
APIs, and legacy GitHub import adoption.
- Reuse current caller grants and credential refresh. Combine and
deduplicate repository inventories across accessible connections. Pasted
public URLs also prefer the active user’s authorized connections. Tokens
stay in the Git child environment, never argv or disk.
- Fetch a shallow Git snapshot at one immutable commit. Scan the full
local tree, including hidden and nested folders. Read Git objects
without checkout or archive transformations and enforce nested package
boundaries.
- Bound Git downloads to 128 MiB and three minutes. Cancel active
process groups and remove incomplete downloads. Preserve cancellation
and deadlines while progress drains; close stalled HTTP progress streams
after 30 seconds. Reuse caller-scoped temporary snapshots for
preview/import after reauthorization.
- Index repository package boundaries once and cap expanded work at
1,000 packages, 10,000 files, and 100 MiB, including repeated copies of
shared blobs. Bound path depth and the shared path index. Discovery
keeps audited manifests without retaining all package bodies.
- Resolve moving branches before fetching so unchanged discovery reuses
caller-scoped snapshots. Limit active scans, scan frequency, and new
downloads per caller and company; quotas apply before metadata reads and
across connections, and cached scans do not consume the download quota.
- Store complete versions with script content, binary bytes, and
executable modes. Preserve these through copies, runtime caches, and
runner packaging.
- Stage downloads before publication. Use source leases, revision
checks, and transactional activity records. Keep installed versions
after failures, upstream deletion, deselection, and disconnect.
- Add the approved import flow, Sources page, selection tree,
provenance, read-only Studio behavior, and saved return from GitHub
setup.
- Add package manifests, commit-pinned file previews, and separate
runtime requirements and reference warnings. Supporting files are
included together; nested skills remain independently selectable.
Preview requests reauthorize the caller and re-audit package content.
- Show installed skills as compact links beneath each source. Repository
titles open GitHub. Keep Refresh, Select skills, and Disconnect source
in a three-dot menu. Source rows omit the branch, imported count, and
refresh timestamp; action alignment and repository titles work at narrow
widths.
- Stream discovery metadata over an opt-in NDJSON response. Show
measured Git download progress and real package/file counts, animate
newly checked skills, support cancellation, and require a complete scan
before selection. Keep the existing JSON API.
- Retain component stories and add a separate full-app journey story
group. Include fixed progress states and interactive scan,
large-repository, interruption, and saving stories.
- Update Skills documentation and product contracts. Suppress private
GitHub skill references in telemetry. Privacy review requested for the
telemetry changes.

## Verification

- Local repository typecheck, full build, token gates, and Storybook
build passed during this work. Focused transport, authorization,
scanner, persistence, route, and UI tests pass. The final UI refinement
passes all eight focused UI tests, UI typecheck/build, and token gates.
The scanner resource and repeated-discovery fixes pass 132 focused
scanner, transport, authorization, source-service, route, and rate-limit
tests, plus server typecheck/build. Full-suite verification comes from
CI; the older full local Vitest run was stopped after unrelated chat
failures and a font-test failure, all of which passed in fresh focused
runs. At commit `1098d5996`, all 54 active checks pass; two optional
Storybook jobs are skipped. CI covers repository typecheck, build, the
full test suites, browser shards, and the canary dry run. Greptile is
5/5 with no open findings; the security scan also passes.
- Adversarial scanner tests verify repeated-blob byte accounting with
and without declared sizes, package/file/path caps, one-time repository
indexing, metadata-only discovery audits, and nested package boundaries.
Additional tests cover branch movement, snapshot reuse, caller/company
quotas, isolation across connections, active-lease cleanup, quota
recovery, and rejection before any metadata API call.
- Real Git tests verify hidden paths, exact binary bytes, executable
modes, export-ignore preservation, symlink/submodule reporting, pinned
commits, caller-scoped cache reuse, cancellation, cleanup, and
credential isolation. Regression tests hold both download slots with
permanently blocked progress callbacks, verify timeout/cancellation
cleanup and retry, and exercise HTTP backpressure cancellation. Access
tests cover automatic public-URL connection selection and revoked
grants. Database tests verify company and grant audiences.
- Live isolated browser test: the public `anthropics/skills` scan now
completes and discovers all 20 skills without connecting an account.
Imported canvas-design with all 83 files, opened it from Sources, and
verified the installed binary-font preview/download control. Package
previews also expose the complete file inventory before import.
Cancelled an active Git download and retried successfully to all 20
discovered skills; the browser displayed measured download progress. The
current audits reject four other packages; eligible selections remain
importable.
- Browser checks verify the simplified source rows at desktop and narrow
widths, keyboard navigation into the actions menu, Refresh from the
menu, selection, and fixture disconnect with installed skills retained.
Storybook includes a menu-open checkpoint and a 320px layout.
- Storybook includes receiving/preparing download checkpoints and a
timed full-app import journey, plus cancellation, retry,
large-repository, and saving states. Streaming tests cover split UTF-8
frames, incomplete streams, late responses, cross-company requests, HTTP
errors, and JSON compatibility.
- Earlier live acceptance on this PR imported `stitch-skill` with
`DESIGN.md`, assigned it to an agent, disconnected its source, and ran a
successful Studio test that read both installed files. An editable copy
changed independently. Both Skills variants, mobile selection, and
return from GitHub setup were exercised.
- Private access, revoked credentials, OAuth success return,
binary/script preservation, concurrent refresh, transaction rollback,
version pins, and legacy adoption have automated coverage. A real
private-repository OAuth grant was not created during this test.

## Risks

- The migration groups recognizable legacy imports without provider
calls. Their first successful refresh completes the local package
snapshot.
- Reference checks are advisory. They cover Markdown links and explicit
relative resource paths, not arbitrary runtime dependency graphs.
Preview text is capped at 64 KiB; imported bytes remain complete.
- Git must be installed on the server. Shallow fetches still download
the branch snapshot, including files outside selected packages.
Downloads have size/time/concurrency limits. Temporary caches are
bounded and caller-scoped. GitHub API quota still applies to repository
metadata and the connection picker; content no longer uses per-file API
requests. Failed scans retain installed content.
- Sources depend on the current caller's GitHub access. A saved
connection does not grant access to another person's token.
- GitHub script support and immediate manual refresh are explicit
maintainer-approved requirements. The operator trusts the selected
repository and accepts upstream script and executable-mode changes on
refresh. Static audits are not a sandbox or a guarantee of safe code;
agents may later invoke installed helpers under their runtime
permissions. Import and refresh do not execute scripts, hooks, package
installation, or builds. Raw URL and skills.sh imports keep their prior
script restrictions.
- Source originals remain read-only. Refresh affects subsequent unpinned
runs; explicit pins and active runs retain their versions.
- The telemetry change removes source-managed GitHub identifiers from
skill-reference events. It introduces no event or field. Please review
the privacy boundary.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code execution, and
browser tools. The exact serving model ID and context-window size are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 13:32:58 -05:00
scotttongandClaude Opus 5.5 6aa908ef15 fix(heartbeat): cancel queued runs whose claim is permanently rejected (#14738)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The heartbeat scheduler claims queued runs for each agent in
`startNextQueuedRunForAgent`, and startup recovery calls it through
`resumeQueuedRuns`
> - `claimQueuedRun` can throw a 4xx `HttpError` that the run's own rows
decide, for example the 403 from #14043 ("Queued-message interrupt
authority is unavailable")
> - The claim loop does not catch that error. The run stays `queued`, so
every recovery pass and every restart claims it again and gets the same
error
> - During periodic recovery the error is logged, but it stops the claim
loop, so all queued runs of every agent after it wait. During startup
recovery the error is rethrown, so the server cannot boot
> - We saw this on a self-hosted instance: one card-answer run from
#14043 put the server in a crash loop for about 11 hours (216 failed
boots) until the run was cancelled by hand in the database
> - This pull request cancels a queued run when its claim fails with a
403 that cannot clear, keeps a run queued when its 4xx rejection can
clear later, and in both cases continues with the rest of the queue
> - The benefit is that one bad queued run can no longer stop the server
from starting or stall other agents. #14510 fixes the specific identity
mismatch; this change is the safety net for this whole class of error

## Linked Issues or Issue Description

Refs #14043

Refs #14510, #13539, #13315

## What Changed

- `server/src/services/heartbeat.ts`: The claim loop in
`startNextQueuedRunForAgent` catches errors from `claimQueuedRun` for
each run.
- A 403 is a permanent rejection. In the claim path it only comes from
the run's persisted identity (an unverifiable interrupt receipt, a
manual wake without a user), and those rows do not change. The loop
cancels that run with error code `queued_run_claim_rejected`.
`cancelRunInternal` also cancels the wakeup request and releases the
issue execution lock.
- The cancel runs after the agent start lock is released (`.finally()`),
because `cancelRunInternal` promotes the next queued run under the same
lock. A cancel inside the lock waits 30 seconds for the stale-lock
timeout.
- Any other 4xx (for example a 422 `responsible_user_unresolved`, or a
409) can clear later. The run stays queued, the loop logs a warning, and
it continues with the next run. A later recovery pass tries the run
again.
- Non-HTTP errors and 5xx errors (for example database errors) keep
propagating, as before. Startup still fails when the database or another
dependency is broken.
- If a cancel itself fails, the loop logs the error. The run stays
queued for the next recovery pass.
- `server/src/__tests__/heartbeat-queued-run-claim-isolation.test.ts`:
New embedded-Postgres tests with a mocked adapter:
- A queued-comment interrupt run with an unverifiable receipt is
cancelled, `resumeQueuedRuns()` resolves, and a second pass stays clean.
- A claimable run behind a rejected run of the same agent is claimed and
executed.
- A run with an unresolvable responsible user (422) stays queued, and
the claimable run behind it is executed.
- A rejected run for one agent does not stop the queue of another agent.

## Verification

- `cd server && npx vitest run
src/__tests__/heartbeat-queued-run-claim-isolation.test.ts` passes.
- Without the change in `heartbeat.ts`, all four tests fail (the 403
`Queued-message interrupt authority is unavailable` and the 422
`responsible_user_unresolved` escape `resumeQueuedRuns()`).
- All `heartbeat*`, `*queued*` and `run-identity` server suites pass
locally (1252/1253). The one local failure, `heartbeat-process-recovery`
"redacts opaque environment-bound credentials from Sentry diagnostics",
fails the same way on unchanged `master` on my machine and passes in CI.
- `pnpm -r typecheck`, `pnpm test:run` and `pnpm build` pass locally.

## Risks

- Behavior shift: a queued run whose claim fails with a 403 is now
cancelled, not left queued. Such a run could never start before this
change, so no runnable work is lost. The cancel reason and the
`queued_run_claim_rejected` error code make the cancel visible on the
run.
- A run with another 4xx rejection stays queued, as before, but it no
longer stops the rest of the queue. It logs a warning on each recovery
pass until the cause clears.
- Low risk otherwise: no schema, API or UI change.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Claude Opus 5.5 (`claude-opus-5-5`) in Claude Code, with tool use and
code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 10:41:17 -07:00
DottaandPaperclip 647e4adf27 fix(native): bind operator Stop to caller cancellation intent
Reserve an optional board request UUID under the run lock and retain it
through default Stop joins and native dispatch. Reject prior or competing
intents and require the same actor for idempotent retries.

Bind the Copilot denial fixture to its exact request and audited intent,
with versioned causal receipts and negative controls for earlier Stop.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 12:03:34 -05:00
DottaandPaperclip 94e8dec56b fix(runner): preserve tool outcomes through shutdown and restart (#14734)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The runner sends authorized tool calls to the server and saves their
results.
> - A provider turn can stop while a server write is still running.
> - The old shutdown path invented a failed result that could conflict
with the real result.
> - Truncated execution input and incomplete recovery records made the
failure harder to diagnose.
> - This pull request preserves exact inputs and actual outcomes through
shutdown and restart.
> - Tests force the race and crash boundaries so safe retries do not
repeat writes.

## Linked Issues or Issue Description

**What happened?**

Stopping a turn during a server tool call could record a false failure,
then reject the actual result as a conflict. The diagnostic input
formatter could truncate instruction content before execution. A crash
during saved-result delivery could leave that delivery permanently
indeterminate. Cleanup could hide the first failure, and a retry could
overwrite earlier run logs.

**Expected behavior**

Keep dispatched tools pending until their actual result is known.
Preserve accepted input bytes. Accept identical result delivery without
failing the task. Reject conflicting results with enough evidence to
diagnose them. Recover saved-result delivery without repeating the
business operation.

**Steps to reproduce**

1. Hold an instruction update at the filesystem commit barrier.
2. Stop its provider turn before the server returns the result.
3. Release the write, deliver its result, and replay the same result.
4. Repeat with a restart before and after the delivery receipt is saved.
5. Check that there is one write and one audit row, and that the exact
result survives.

**Paperclip version or commit**

The change was developed from `44736c9c7` and rebased onto `0e5830887`.

**Deployment mode**

Self-hosted server with the native runner. Tests use local runner
processes, scripted providers, and PostgreSQL.

Related work: #12353 added durable semantic tool receipts; #12384 added
durable Codex tool recovery; #12404 bound semantic tools to ACPX
sessions. #14633 covers separate native-provider cancellation and
qualification work. This PR addresses server semantic-tool outcomes and
their durable delivery. No duplicate fix was found. AgentMail discovery
is outside this PR.

## What Changed

- Close turn admission without inventing results for dispatched tools.
Keep pending calls and accept late actual results.
- Accept identical result replay with a diagnostic warning. Include call
identity and both result hashes in real conflict errors.
- Preserve exact execution arguments. Reject prohibited or oversized
input before dispatch. Keep diagnostic previews redacted and bounded.
- Commit instruction-attempt evidence before the filesystem write. Save
completed mutation receipts so concurrent and restarted duplicates
return the first result. Recheck authorization before replay. An attempt
without a completed result stays unknown and cannot execute again.
Definite pre-write failures save and replay their original error without
another write.
- Recover an interrupted saved-result delivery only for backends with
durable result receipts. Never replay an ordinary business operation
with an unknown outcome.
- Preserve the initiating error when cleanup also fails. Record
incomplete settlement evidence. Propagate typed unknown-outcome errors
through the native tool wrapper without creating a false completed tool
result.
- Append run-log attempts and restore the durable log before appending
after local file loss. Reject incomplete restores. Publish a restored
prefix only if the destination is absent so concurrent attempts cannot
overwrite new lines.
- Add deterministic race, crash, replay, authorization, exact-content,
and log-restoration tests. Document their assertions in
`packages/paperclip-runner/docs/durable-recovery.md`.

## Verification

- Current head: `7e088f4c7fba8ebabf98ae95485a5753b013d489`. All 55
applicable checks pass; four conditional/manual checks are skipped. This
includes build, typecheck, Rust, both runner TypeScript shards, server
and workspace tests, all eight browser shards, isolated runner
compilation, and the clean-install release dry run. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36746101110).
- Greptile reviewed this exact head at 5/5 with zero new findings. All
three earlier review threads are resolved.
- Focused local verification includes 11 instruction integration tests,
23 surrounding authority/tool tests, 26 run-log tests, and 169
controller/driver tests. The post-rebase controller/transport/runtime
selection passed 415 tests. The full Rust release suite passed 617 tests
with two ignored. The real-process SIGKILL recovery test passed three
consecutive runs.
- The fault matrix in
`packages/paperclip-runner/docs/durable-recovery.md` uses explicit
barriers, real PostgreSQL rollback, durable journal reloads, and killed
runner processes. It covers late results, identical and conflicting
replay, exact long content, concurrent log restoration, lost commit
acknowledgements, and definite failure replay after the original CAS
base becomes valid again. No paid model calls are needed.
- Full local recursive typecheck and build passed during implementation.
Server typecheck and the runner TypeScript build passed after the review
fixes. The broad local repository test run was stopped after repeated
database startup timeouts. Four timing/launch failures in an earlier
broad runner run passed focused reruns without changed assertions or
timeouts. These are local verification limitations; the complete
current-head CI suite is green. An earlier CI workspace job received an
infrastructure shutdown signal; its current-head replacement passed.

## Risks

- A stopped turn can remain blocked when a dispatched operation has no
proven result. The system does not guess its outcome or rerun its
effect.
- Conflicting results still fail settlement. Existing failed or
conflicting journals are not repaired automatically.
- Accepted semantic input is limited to 480 KiB of encoded JSON to fit
the encrypted transport. Larger input fails before execution.
- Instruction filesystem writes and database receipts are not one atomic
storage operation. A separately committed attempt and audit record
survive rollback. An attempt without a completed success or definite
pre-write failure receipt remains blocked as an unknown outcome. It is
not replayed or reported as success.
- Run-log restoration now reads the durable object before appending when
the local log is missing. Failed or incomplete reads reject the append.
- No schema migration, dependency change, workflow change, or AgentMail
change is included.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and test
analysis. The exact served model ID and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; the broad
local run limitation is recorded above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 11:54:56 -05:00
DottaandPaperclip 3c561642b4 fix(chat): resolve approvals and preserve unanswered questions (#14613)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents ask for decisions and optional details through cards in chat.
> - A clear approval in a message can leave the matching card pending.
> - An unanswered question can also block an unrelated later reply.
> - Decisions need a saved source message, while optional questions need
to remain answerable in history.
> - This pull request records conversational decisions and lets users
move on from questions and answer them later.

## Linked Issues or Issue Description

**What happened?**

Native Claude and Codex could act on approval in chat while the original
approval card stayed pending. Pending question forms stayed above the
composer, were absent from history, and could suppress later chat
replies. A late native question answer could wait for a finished run to
reconnect.

**Expected behavior**

The active agent records a clear approval or refusal against the exact
card and user message. Ambiguous replies do not grant consent. Users can
send another message without answering a question. The question remains
pending in history and can be reopened and answered later. The saved
answer reaches the agent.

**Steps to reproduce**

1. Ask an agent to propose work with a confirmation card, then approve
it in chat.
2. Check that the original card records that approval before work
starts.
3. Ask an interactive question, send an unrelated message, and reload.
4. Open the unanswered question from history and submit an answer.

Related work: #14408 added completion delivery. #14607 tests completion
reporting turns. Neither records conversational answers on approval
cards.

## What Changed

- Add a confirmation endpoint backed by a user comment, with schema
validation, OpenAPI discovery, and native Plan-mode access. Ask mode
remains read-only.
- Check company, active run, actor, current session, message provenance,
revision, and resolver policy. Save the decision and audit in one
transaction. Retries do not repeat effects. Emit resolution telemetry
after commit.
- Give fresh and resumed chat turns the actual pending confirmation
identities. Teach agents to save clear conversational decisions before
acting and to clarify ambiguity.
- Keep unanswered Agent Chat questions as compact history entries. A
newer user message closes the old form. Question cards never contribute
to composer pending counts or navigation, including after dismissing a
fresh form. The history card is the sole reminder; clicking it restores
that exact form and draft.
- Preserve Agent Chat questions when later messages or questions arrive.
Historical ordinary inputs no longer gate later chat replies.
Current-run requests, task execution, and governed approvals keep their
gates. Remove the special acknowledgement-publication proof helpers that
this rule replaces.
- Route answers to finished native runs through durable fresh-wake
delivery, with existing idempotency and source-question context. Settle
late replies against contiguous completed conversation turns and freeze
their history replay; failed, unhandled, and newly arriving messages
remain actionable.
- Add real-component Storybook scenarios, database and UI regressions,
and a three-turn native Claude/Codex E2E case. Capture distinct,
UI-ready screenshots and report the individual assertions.

## Verification

- Focused decision/publication/UI regressions after merging master: 288
passed; subsequent UI draft, failed-send, and conversation checks: 199
passed.
- Native question and durable delivery regressions: 106 passed,
including all four terminal run states and exactly-once late delivery.
Seven targeted regressions fail against the original implementation and
pass with the fix.
- Latest conversation/decision/native-delivery regressions after the
master merge: 121 passed. Covers completed progress, missing or failed
intervening turns, new messages during a late reply, stale sessions, and
frozen retry/replay boundaries. Four new assertions fail before the
ordering fix.
- E2E support suite after the master merge: 792 passed. Negative
controls reject expired cards, wrong questions/answers, stale or missing
replies, unrelated clarification forms, and unexpected tasks.
- The embedded-browser walkthrough caught one additional defect:
dismissing a fresh question still showed a composer badge. Both Cancel
and close-button regressions failed before the fix. The fix at
`65f2ade12` passes 170 chat-thread tests and 792 E2E support tests.
After merging master, 232 chat-thread/confirmation tests, server/UI
typechecks, and token gates pass. The preview and two-provider live E2E
pass at `e5512a206`; Greptile is 5/5 with zero unresolved threads at
that commit. All 55 checks are now successful at `e5512a206` (four
conditional checks skipped), including the aggregate verification gate
and clean-install canary test. The first attempt was interrupted by
simultaneous CI worker shutdowns; one failed-job rerun passed without
code changes.
- [Published
Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fchat-approval-resolution/?path=/story/chat-comments-agent-chat-unanswered-questions--moved-on):
nine real-component scenarios. Manually exercised move on, reopen,
preserve draft, answer later, answer one of multiple questions, and a
custom mobile answer in the embedded browser. Retested fresh Cancel and
close-button dismissal in the updated build, then reopened and submitted
the preserved Green selection and inspected its answered receipt. Static
preview has no live model/backend; its callbacks are fixture responses.
- [First live
campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36714504406-1/)
reproduced the late-answer completion-state defect on both providers
despite correct saved answers and acknowledgements. It also exposed a
valid imperative clarification rejected by the old oracle. Both issues
are fixed with regression controls; this failing run is retained as
evidence.
- [Four-cell
qualification](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36717804064-1/)
passed 4/4 at `2bf8a1009`: unanswered-question return and ambiguous
confirmation, each on native Claude and Codex. Inspected saved state,
source-message decisions, visible cards, and agent replies. Both
late-answer chats settled to waiting; no unrequested tasks were created.
[Final branch
rerun](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36719666238-1/)
passed 2/2 at `142630720`: the same unanswered-question journey after
merging master, plus an additional screenshot and browser assertion for
the actual late-answer acknowledgement.
- [Composer-reminder
E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36727006818-1/)
passed 2/2 at `5b62c52d9`: native Claude and Codex, three turns each,
with explicit no-badge assertions before and after reload. Inspected
saved pending/answered state, both screenshots with a clear composer,
and actual Blue acknowledgements; all five behavioral matchers passed
per provider and neither created tasks. Cost coverage is partial; this
is bounded workflow qualification.
- [Fresh-dismissal
E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36742773318-1/)
passed 2/2 at `e5512a206`: native Claude and Codex, including fresh
Cancel, clear composer, reopen, unrelated message, reload, late Blue
answer, and actual agent acknowledgement. All five behavioral matchers
pass per provider. Inspected the fresh-dismissal screenshots and saved
pending/answered identity; neither created tasks. Cost coverage is
partial (4/6 runs).
- Prior evidence remains available in [the earlier
campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36642252725-1/).
Its early loading screenshot and overwritten final capture prompted the
UI-ready, distinct screenshot fixes.

## Risks

- The model interprets intent. The server verifies permission and
provenance; it does not infer consent from text. Ambiguous and unrelated
replies are not approvals.
- Historical questions can accumulate. They remain visible, pending, and
answerable; no automatic answer or expiry is invented.
- The change to completion gates is scoped to Agent Chat and ordinary
historical inputs. Current-turn and governed approvals retain their
existing controls.
- Live qualification is limited to the selected stories. Broader native
onboarding finalization remains separate work.
- No database migration. Telemetry adds no fields or values; the
contract and README document the commit boundary. Privacy review was
requested on the PR.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, code
execution, and browser-test orchestration. The exact model ID and
context-window size are not exposed to this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 11:46:46 -05:00
DottaandPaperclip 0e5830887b perf: reveal task content sooner and parallelize issue reads (#14727)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - A task page must show saved replies quickly so a person can read the
work.
> - The title could appear while the conversation waited for unrelated
metadata and transcripts.
> - Thread requests also waited for enriched task details, while the
server read several independent fields in sequence.
> - This change shows saved content as soon as it is ready, starts
thread reads earlier, and runs independent server reads together.
> - Native event history takes priority over legacy log fallback, and
mentioned tasks load on intent.
> - Content-free timing spans make the remaining server delays visible
without recording task content.

## Linked Issues or Issue Description

**What happened?**

Task titles and properties appeared quickly, but saved conversation
content stayed hidden for several more seconds while metadata and run
transcripts loaded.

**Expected behavior**

Saved replies and the task description should be readable without
waiting for supporting history. Returning to a cached task should show
content within a frame or two.

**Steps to reproduce**

1. Open a task with saved comments and completed runs.
2. Delay the task activity and runs responses by five seconds in the
browser.
3. Observe whether saved content remains hidden until those responses
finish.
4. Navigate away and return to the task to check cached navigation.

**Paperclip version or commit**

The change was developed from `1b48e73e0` and rebased onto `44736c9c7`.

**Deployment mode**

Built from source, tested in an authenticated staging deployment and
with local response replay.

Related: #14667 overlaps the transcript reveal behavior and adds
separate retry UX. This PR also changes navigation prefetch, parent
metadata gates, native log fallback, server read scheduling, and timing
spans. #12647 proposes a separate SQL predicate optimization in the runs
service. #13597 and #13095 are earlier loading fixes.

## What Changed

- Start activity and runs when the task page mounts, alongside task
details and comments, using the route reference for shared query keys.
Hover/focus prefetch does not start full history reads.
- Reveal saved comments and descriptions while metadata and transcripts
load. Preserve strict waits for linked-comment navigation and tasks with
only runtime content.
- For settled native runs, fetch legacy logs only when event history is
empty or fails. Preserve live-log subscriptions for queued and running
native runs. Fetch mentioned-task details on hover or focus.
- Run independent issue-detail enrichment and run metadata reads in
parallel while preserving recovery dependencies.
- Add `Server-Timing` phases and opt-in OpenTelemetry spans, plus
regression tests and observability documentation.

## Verification

- **313 tests passed** across the initial seven focused component,
cache, timing, and scroll suites. After review fixes, **313 tests
passed** across five task-page, cache, prefetch, and live-transcript
suites (`pnpm exec vitest run` with `--maxWorkers=1`; these sets
overlap).
- `pnpm -r typecheck` and `pnpm build` passed locally before the final
UI-only review fixes. UI typecheck/build and `pnpm check:token-gates`
passed after those fixes. The final commit also passes full typecheck
and build in CI.
- The full local Vitest run was attempted. Several unrelated embedded
PostgreSQL fixtures failed to start, and parallel test workers hit
timeouts. Focused reruns passed. A later local full-suite rerun was
stopped after the complete CI suite passed; it is not claimed as a local
full-suite pass.
- Final commit `d1e147145`: **54 checks passed, 2 skipped**, including
all server/workspace test shards, all eight browser E2E shards, full
build/typecheck, Runner checks, release registry, and canary
clean-install verification. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36734642034).
Greptile **5/5** after two reviews; both findings fixed and all review
threads resolved. No merge conflicts.
- Live browser tests on an existing task with two saved replies: median
full reload to visible content fell from **1.92 s** (3 samples) to
**1.40 s** (5 samples). Cached return fell from **421 ms** (1 sample) to
**29 ms** (3 samples). These are observed samples, not a performance
guarantee.
- With activity and runs delayed by five seconds, saved content appeared
in **1.38 s** on desktop and **1.33 s** on mobile. The inspected comment
did not move when metadata arrived. Verified history expansion, task
properties, pending-input navigation, dashboard return, and mobile
layout.
- A separate local replay with fixed responses reduced visible-content
time from **5.12 s** to **2.15 s**. This isolates frontend behavior and
is not a live-server benchmark.

## Risks

Progressive history can change the thread after first paint. Existing
anchor behavior is retained and covered by tests and delayed-response
browser checks. Query aliases must stay aligned for invalidation.
Parallel reads can increase short bursts of database work; dependent
recovery operations remain ordered. Full reloads still depend on network
and task-detail latency.

No schema or authorization change. OpenTelemetry remains disabled
without an operator endpoint. The added spans use a closed set of phase
names and carry no task IDs, task content, or exception text.

I checked `ROADMAP.md`; this is a performance fix within the existing
task page.

## Model Used

OpenAI GPT-6 in Codex. The runtime does not expose a more specific model
ID or context-window size. The agent used reasoning, code editing,
terminal tools, and Chrome performance profiling.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 10:26:28 -05:00
DottaandPaperclip a36cbffa9e fix(connections): broaden natural-language and aggregator search (#14725)
## Thinking Path

> - Paperclip manages agents and the services they need for work.
> - Agents use connection search to discover a setup path before they
request access.
> - Tool-only filtering hid channel and AI methods from this search.
> - Requiring every query word to match rejected normal task
descriptions.
> - A small aggregator index also omitted supported apps such as
Circleback.
> - This pull request broadens retrieval and returns purpose-specific
setup guidance.
> - Agents can choose a relevant result while existing access and
provider-choice checks still apply.

## Linked Issues or Issue Description

**What happened?**

A search such as “AgentMail create an email address and manage an agent
mailbox” returned no usable result. “Help me find tools for circle back”
also missed Composio's supported Circleback toolkit. Queries longer than
200 characters failed validation.

**Expected behavior**

Return useful native and verified aggregator matches from
natural-language queries. Include channel/email methods when Chat
connectors is enabled. Identify each method's purpose and the correct
setup path.

**Steps to reproduce**

Use the queries above with `connections_search` from an active task.
Enable Chat connectors for the AgentMail case. The regression suite
reproduces these misses before the change.

Related routing work: #13941. This change does not change the runner
failure path or add channel setup to tool-only connection cards.

## What Changed

- Rank name and capability matches. Accept extra words, split names,
small spelling errors, and queries up to 4,000 characters.
- Include tool, channel/email, and AI methods. Return company-prefix
setup links for channel and AI flows.
- Add a dated snapshot of 1,583 official Composio toolkit names and a
refresh script. Merge duplicate MCP variants for search and link each
support claim to official evidence.
- Find authorized indexed aggregator namespaces within longer queries.
Return multiple app matches when the agent needs to choose.
- Prefer exact app names over fuzzy matches for other apps; retain
existing AI readiness.
- Preserve native preference, company and identity boundaries,
administrative denials, and saved provider consent.
- Add relevance and database regressions, extend native tool-authority
coverage, and document search behavior.

## Verification

- Red: 15 new assertions failed against the previous implementation; the
existing baseline passed. Added red-green regressions for Motion versus
fuzzy Notion and existing AI access during review. A further regression
covers mixed ready/unconfigured AI results and their per-result setup
guidance.
- Green: all 103 tests in the eight focused shared, database,
runtime-tool, fixture, and route suites pass on the latest commit.
- `pnpm -r typecheck` and `pnpm build` passed.
- Latest-commit CI passed: 54 successful checks and two skipped checks,
including the full test matrix, browser E2E, typecheck, build, and
canary dry run.
- The long local `pnpm test:run` attempt began before the review fixes
and retained transformed pre-fix search code; it also hit an unrelated
timing failure. Fresh serial reruns of the affected search suites and
three timeout cases passed all 124 tests. Parallel local route shards
hit two additional database setup timeouts; both suites passed all 17
tests on a fresh serial rerun. The complete corresponding CI suites also
passed. Duplicate broad local runs were stopped after CI completed. The
local UI suite independently passed all 7,007 tests.
- Greptile: 5/5 on `cc6a0180d`; all review findings resolved.
- Browser verification passed in a disposable local instance through
real process-agent search requests: AgentMail opened its setup flow with
the requester selected; the saved Circleback choice produced the
Composio setup card; a paragraph-length Notion query produced its setup
card. No provider credentials or external accounts were created.
- The browser test caught an invalid UUID-based setup URL. The fix uses
the company prefix and has a regression assertion.
- A 3,971-character catalog query found Circleback first in a local 10
ms spot check after sharing query preparation across the catalog scan.
This is a single measurement, not a performance guarantee.

## Risks

- Broader retrieval can return extra candidates. Named services rank
first; agents must select the relevant method.
- The public support snapshot can age. It proves catalog support, not
account authorization or the availability of every requested action.
- Channel and AI methods use existing setup links. The tool connection
card still accepts tool methods only.
- No schema, migration, credential, or runner lifecycle changes.

## Model Used

OpenAI Codex (GPT-6). The session does not expose a more specific model
identifier or context-window size. Used reasoning, repository search,
code execution, tests, and browser tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 10:20:50 -05:00
DottaandPaperclip dd7fc1f90a fix: raise the native journal read limit to 256 MiB (#14711)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native sessions persist control-plane state so they can resume
safely.
> - The state includes committed provider history needed for recovery.
> - The server, runnerd recovery, and durable control plane validate
this state before trusting its identity.
> - Their differing 64 MiB and 192 MiB limits can reject a valid journal
before recovery.
> - This pull request aligns all three local state limits at 256 MiB.
> - Larger files remain bounded, while recovery can read larger valid
histories.

## Linked Issues or Issue Description

Refs #13882
Refs #14312

## What Changed

- Raise the server and runnerd recovery limits from 64 MiB to 256 MiB,
and align the durable control-plane limit from 192 MiB to 256 MiB.
- Add coverage for a valid history above 64 MiB and rejection above 256
MiB.

## Verification

- Matching server recovery passed with more than 64 MiB of actual
committed event payloads (128 events with 512 KiB deltas).
- The real runnerd exact-authority resume regression with the test Codex
provider passed with 193 MiB of valid JSON whitespace appended. It
crosses the former 192 MiB core limit and confirms the same provider
identity. This exercises the runner process and durable control plane
with a simulated provider, not a live OpenAI API call. This test used
approximately 1.15 GiB peak RSS.
- The actual runnerd reader accepted valid 256 MiB JSON and rejected
valid 256 MiB + 1 byte. The reader call took 231 ms; the fresh process
peaked at 1,244 MiB RSS.
- Server tests reject mismatched identity above 64 MiB and files above
256 MiB.
- `pnpm -r typecheck`, `pnpm build`, and `git diff --check` passed.
- Full local `pnpm test:run`: 13,730 passed, 575 skipped, 7 failed
across 6 files. All failures were embedded PostgreSQL startup errors
after five attempts. They affected agent hiring, instruction revisions,
environment images, reviewed chat bindings, issue monitoring, and legacy
continuation authority. The focused journal tests passed; the latest
pushed head passed all ordinary CI checks. Superagent is the only
blocking check.

## Risks

- **Open review concern:** Greptile is 5/5, but Superagent is
`ACTION_REQUIRED` with two P2 findings on the server and runnerd
readers. Both flag the increased synchronous parsing and memory cost.
This PR keeps the requested fixed-limit change small. It does not add a
worker parser or a process-wide memory budget. This resource tradeoff
needs review before merge.

- Large state parsing is synchronous and can consume several times the
file size in memory.
- Remote checkpoint archive and expanded-size limits remain 64 MiB, so
this change alone does not make larger remote checkpoint transfers
portable.

## Model Used

- OpenAI Codex, GPT-6, with delegated assistance from `gpt-6-luna` at
high reasoning effort; tool use and code execution. The GPT-6 context
window is not exposed in this task runtime.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 10:08:25 -05:00
DottaandPaperclip d30b03bd8c test: add persistent E2E coverage for human blocker decisions (#14707)
## Thinking Path

> - Paperclip manages work for AI agents.
> - Agents use the coordination skill when work needs human authority or
a scope decision.
> - PR #14188 replaced automatic manager escalation with direct blocker
handling.
> - This behavior needs real browser, server, database, and provider
tests.
> - The test must verify saved human input, task ownership, and resumed
work.
> - This pull request adds six reusable Product E2E cases and improves
the skill examples that they exercise.

## Linked Issues or Issue Description

Refs #14188.

The merged change needs repeatable behavior coverage. The new suite
tests missing administrator access, missing hiring permission, and
requester scope questions. Searches found no duplicate blocker-guidance
suite. This extends the existing eval system described in ROADMAP.md.

## What Changed

- Add the explicit-only `blocker-guidance` Product E2E suite. It has
three local scenarios on legacy Codex and legacy Claude.
- Use the production UI and public APIs to create work, save a
human-only question or confirmation, answer it after reload, and resume
the same task.
- Check requester identity, ownership history, manager activity, hiring,
saved answers, and completion. Keep direct text input as a separate UX
result.
- Save pending and final screenshots, API checkpoints, skill hashes,
provider evidence, and billing data through the existing report
pipeline.
- Isolate the Claude fixture home. Verify the served skill bytes before
dispatch so an old installed skill cannot silently replace the evaluated
skill.
- Improve the coordination and hiring skill examples. Include the
human-only policy, requester address, wake behavior, and handling of
authorized scope changes.
- Grader v5 requires the exact approved public welcome note as a new
worker comment. Browser input checks reject unwritable scope cards
before clicking, and confirmation direction must be saved in the
resolution before the worker wakes.
- Add grader calibration and browser-input tests. Update the fixture
guide and generated capability inventories.

## Verification

- `pnpm build`: passed after rebasing onto current master.
- `pnpm -r typecheck`: passed.
- `pnpm test:e2e:runner:typecheck`: passed.
- `pnpm test:e2e:runner:unit`: 742 tests passed.
- `pnpm test:e2e:runner:browser-support blocker-input.spec.ts`: 10 tests
passed.
- `pnpm test:e2e:runner -- --list --suite blocker-guidance`: six cells
found.
- Capability inventory and generated-contract checks: passed.
- `pnpm exec vitest run
server/src/__tests__/hiring-operational-examples.test.ts`: four tests
passed after synchronizing the generated API reference and section
anchor.
- Full general and serialized test suites: passed in CI on
`6652cee74517039676bad6a720f213625d265acd`. The redundant local `pnpm
test:run` was interrupted after complete CI coverage passed; it is not
claimed as a completed local full-suite run.
- Final GitHub checks: 54 passed, two optional Storybook checks skipped.
The runtime-exposure startup test hit a 10-second readiness timeout
once, passed a targeted local reproduction, and its CI shard passed the
single retry without code changes.
- Current-head Greptile: 5/5, clean check, zero unresolved threads.
- Historical live measurement on September 29 at
`4edc77ae2b95b10dd61426ce3f042bac00527ad9`: three independent six-cell
runs scored 5/6, 6/6, and 6/6. Claude Sonnet 4.6 passed 9/9. Codex
`gpt-5.6-sol` passed 8/9. These runs predate this rebase.
- Version 5 changes the scope answer to an exact approved publication
draft. The historical runs do not qualify that new requirement; the
two-provider scope pilot at `49a1f4eab369948b9e3b34a6ce436489e875e4ec`
passed Codex and failed Claude. Claude posted the correct salary-free
sentence but omitted its required reference line from that comment,
placing the reference in a separate completion message. The
`public-welcome-note` check correctly failed. An earlier Claude
database-startup failure was retained separately; its fresh-instance
retry reached the model. This pilot is not a six-cell qualification.
- The failed Codex scope case omitted `addresseeUserId`. The strict
routing check remains. All 18 attempts had clean evidence manifests and
passed cleanup.
- To repeat with provider credentials: `pnpm test:e2e:runner -- --suite
blocker-guidance --max-parallel 1`. This is a paid, opt-in suite and is
excluded from `--all`.

## Risks

- The live suite measures variable model behavior. The retained 17/18
historical result and the current 1/2 scope pilot are not all-pass
qualifications. These paid cases are opt-in; their observed model
failures remain visible independently of deterministic CI checks.
- A separate generic task-replacement diagnostic still exposed a Claude
refusal. The ordinary cases use specific business decisions. The
diagnostic is not a standalone catalog case in this change.
- Earlier measurements included an old installed Claude skill and test
defects. Their grades remain retained and are not combined with the
three final repetitions.
- Skill examples can affect when agents ask for human input. Downstream
permission checks still apply.
- Native runners, Daytona, agent-requester routing, and real external
connection authorization are outside this suite.
- Raw provider traces and credentials remain private. No screenshots,
raw reports, secrets, workflow changes, or lockfile changes are
committed.

## Model Used

OpenAI GPT-6 through Codex assisted with this change. The exact deployed
variant and context window size are not exposed in this session. The
assistant used reasoning, repository edits, tool use, and shell
execution. The evaluated models were `gpt-5.6-sol` and
`claude-sonnet-4-6`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 07:56:54 -05:00
DottaandPaperclip 22721c1ab0 fix(server): serialize standalone agent directory probes
Include the production tsx name helper in generated Node programs so unchanged instruction copies remain warm. Exercise both generated programs through the actual tsx loader while retaining mutation and containment checks.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 07:55:15 -05:00
DottaandPaperclip 1b48e73e0b feat(ui): add secondary navigation for agent chat (#14706)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent Chat already provides a persistent conversation with each
agent.
> - Its shortcuts share the primary navigation and do not give chats a
dedicated place.
> - People need to find agents, start a chat, and switch conversations
without moving the page layout.
> - This pull request adds a secondary chat sidebar and a landing page
around the existing chat surface.
> - The same conversation, composer, history, and context panel remain
in use.

## Linked Issues or Issue Description

Refs #13283 and #13420. This extends the existing experimental Agent
Chat navigation after review of the component and page stories. It
supports the CEO Chat roadmap item through the existing task-backed
conversation model.

**Subsystem affected**

The board UI and the company-scoped conversation list API.

**Current behavior**

Chat shortcuts sit inside the primary navigation. There is no dedicated
landing page with a searchable conversation list. A separate landing
header also moves the sidebar when an agent is selected.

**Proposed behavior**

Show a Chat entry in primary navigation. Keep a searchable agent sidebar
beside the chat content. The plus button starts or reopens the current
user's single conversation with that agent. Keep the header and sidebar
in the same positions before and after selection.

**Reason and benefit**

People can find agents and return to persistent conversations without
leaving the chat area or creating duplicate chats.

**Breaking changes**

The experimental chat navigation changes. Explicitly adding a chat now
resolves its conversation immediately. Direct visits to unused agent
chat URLs remain read-only. The existing per-agent routes and message
contracts remain compatible. No database migration is required.

## What Changed

- Add an account- and company-scoped conversation list endpoint with the
existing access checks, feature gate, and OpenAPI entry.
- Add the live secondary sidebar, landing page, avatars, search, loading
states, errors, and retry controls.
- Make the agent picker wait for chat creation and display failures.
Existing agents reopen the same conversation. A dismissed selection
cannot close a reopened picker or navigate over a newer choice.
- Preserve recent-activity ordering and terminated agents’ chat history.
Scope live list refreshes to the current user’s conversation events. A
failed historical-agent lookup leaves healthy chats usable and offers a
focused retry.
- Keep the sidebar and header stable across chat routes. Keep mobile
selection in the navigation drawer.
- Use the production components in Storybook. Prepare the theme and
mobile viewport before mounting the page to avoid the startup flash.
- Update product documentation, the design guide, and navigation tests.
Replace old browser expectations for stars and recent shortcuts with
persistent conversation and layout coverage.

## Verification

- `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` pass
after rebasing onto master.
- Focused UI tests pass, including 105 sidebar, picker, and live-update
checks after review fixes. The 33 conversation service and route tests
and 10 OpenAPI checks pass, including ownership, feature gating, and
concurrent creation.
- The full local test run passed 14,163 server tests before three
environment or timeout failures. The embedded Postgres startup,
connector socket, and native runner failures all passed direct reruns.
- Browser test-drive verification covers a real provider reply, add and
reopen, persisted history after reload, no-match search recovery, mobile
drawer dismissal, and top-aligned context panels.
- Browser measurements confirm that the sidebar has the same position
and dimensions on the landing page and an agent conversation.
- Storybook builds and its add-and-reopen interaction passes.
- The revised browser regression passes locally against a freshly built
throwaway instance. It covers stable sidebar geometry, add/reopen
uniqueness, drafts, search, history, and terminated-agent history after
reload. The full CI browser suite also passes.
- Latest commit `b323577d9523180104df4000eaceedea2772608c`: all 54
completed checks pass, including the complete server/workspace/browser
suites, aggregate verification, build/typecheck, security scans, and
canary packaging. The two Storybook jobs are skipped by their workflow
conditions. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36714052050).
- Greptile reviewed this same commit at 5/5 with no remaining actionable
findings; all review threads are resolved.
- Reviewer path: enable Agent Chat, click Chat, use plus to choose an
agent, send a message, switch away, and reopen that agent. One
conversation must remain, with its history intact.

## Risks

- The new sidebar lists persistent conversations instead of starred and
recent shortcuts.
- Chat creation is asynchronous. Errors stay visible in the picker, and
delayed responses cannot navigate into a previous company or account.
- The shell adjustment is limited to chat routes and preserves the
existing conversation implementation.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, and browser
tools. The session does not expose the exact API model ID or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#123` / `Refs #123` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 07:35:32 -05:00
DottaandPaperclip d72389bee2 feat: add Browser Use Cloud connector and live task browsers (#14627)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Apps gateway gives agents governed access to external tools.
> - Browser Use Cloud can run browser work, but a tool result alone does
not let a person watch or take over.
> - A task needs a durable browser session, a visible viewer, and
recorded costs.
> - This pull request adds a Browser Use Cloud v4 connection and
interactive browser tabs on tasks.
> - People can follow the work, interact with the page, and retain the
browser after the agent finishes.

## Linked Issues or Issue Description

**Problem or motivation**

Agents need governed access to Browser Use Cloud. People need to see and
interact with the same browser from the task. A browser must remain
available after a run finishes and appear at the correct point in the
task feed.

**Proposed solution**

Add a native REST connection for the v4 API. Bind each session to its
company, task, agent, and credential grant. Open its interactive viewer
in the task side panel. Record provider costs as financial events. Use
`browser-use-cloud` as the app and connector key. Keep its skill with
the connector and deliver it only with authorized connection tools.

**Alternatives considered**

A v3 MCP connection would expose tools without the v4 lifecycle
integration. An external viewer link would leave the task. A fixed
viewer size would prevent pages from responding to changes in the task
pane.

**Roadmap alignment**

This extends the governed Apps gateway and Connected Apps roadmap. It
uses the existing task, grant, secret, approval, and financial records.
The work was requested by the maintainer. A search found no duplicate
Browser Use connector PR or issue.

## What Changed

- Add the Browser Use Cloud app, brand asset, API-key connection, and
profile settings under the `browser-use-cloud` key.
- Bundle the `browser-use-cloud` skill with the connector. Keep it out
of global `skills/` discovery. Deliver it only with authorized task/run
connection tools. Remove retired connector skill keys from runtime
overlays and preserve unrelated browser skills.
- Expose seven v4 tools through the governed gateway and deliver them to
native and CLI agents.
- Persist sessions, browsers, runs, event and recovery cursors, shutdown
leases, and cumulative cost accounting. Recover uncertain paid starts
without replaying them.
- Enforce task ownership, credential grants, approvals, revoked access,
and budget limits.
- Add interactive task browser tabs and compact chronological feed
entries. Retain the viewer across tab switches and keep visible idle
browsers open.
- Add debounced automatic viewport fitting, standard size presets, and a
viewer ownership lease.
- Add lifecycle, authorization, accounting, viewport, UI, and Storybook
coverage.
- Add an idempotent database migration after the current master
migration. Preserve deployed migration hashes. Migrate pre-release Cloud
connection and financial keys without replacing grants, credentials, or
browser history.
- Document provider behavior, live acceptance results, and the lack of
documented passkey forwarding.

## Verification

- Full workspace typecheck and production build pass on the updated
branch.
- Token gates, brand asset validation, module boundaries, and migration
ordering pass.
- Cloud tests verify global skill exclusion, authorized task/run
delivery, unassigned agents, disabled connections, revocation, adapter
isolation, and secret exclusion. The existing AgentMail connector
assignment test also passes.
- Migration replay runs twice against existing browser work and
financial records. It preserves the records and avoids duplicate costs.
- The focused provider, app catalog, OpenAPI, connection gateway, and
migration regression suites pass. Recovery coverage includes lost
replies, process crashes, provider rejection, and browser arrival
acknowledgement.
- All 54 checks pass on `2974b5f03641ad0cea3c941d8c02579316fa8c92`,
including the full test matrix, browser E2E shards, build, typecheck,
security, and release canary. Two optional Storybook jobs are skipped.
- Greptile is 5/5 on the same commit, with zero unresolved review
threads. The corrected review uses the actual master-to-head diff.
- The local `pnpm test:run` started and was stopped after the full CI
matrix passed. It did not complete locally; the full-suite result above
comes from CI.
- Earlier live acceptance used an isolated company with a capped
provider credential. The agent opened paperclip.ing, the embedded viewer
accepted navigation, and the same browser stayed available after
completion and tab switches.
- The local Storybook build passes. Stories cover the panel, footer,
feed entries, settings, lifecycle failures, and viewport modes with an
offline viewer fixture.

## Risks

- Browser Use charges for hosted work. Provider caps and local budget
checks reduce exposure; reported costs can arrive after work completes.
- Viewer and CDP URLs grant access to the browser. The server validates
and restricts them. They are excluded from agent results and durable
event data.
- Runtime resizing of v4 agent browsers uses a provider option confirmed
by live testing but absent from its published agent schema. Resizing
during a click may invalidate coordinates. Fixed presets remain
available.
- Viewport ownership is process-local and resets on restart. The
lifecycle and accounting records remain in the database.
- The original intermittent embedded-viewer stall has not been fully
diagnosed. A bounded reconnect and active-session recovery cover the
observed failure paths.
- Live tests did not cover every revocation, approval, rate-limit, or
restart case. Deterministic integration tests cover those paths. Passkey
forwarding is not claimed.
- Unknown create outcomes keep the credential available for cleanup.
Run-list absence cannot prove a paid POST was rejected, so recovery
stays pending until it can identify provider work.

## Model Used

OpenAI Codex, GPT-6. Used reasoning, repository search, code execution,
browser interaction, and test tools. The exact serving model ID and
context-window size were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 07:13:16 -05:00
DottaandPaperclip 2f6fa3b6dc fix: recover provider authentication inside tasks (#14629)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents need a working model provider connection to run a task.
> - A provider can reject a stored credential after the task starts.
> - The failed run must ask the responsible user to repair that
connection.
> - This pull request adds that request directly to the task and reuses
Connections sign-in.
> - The user can choose an API key or subscription, then continue the
task with a fresh session.

## Linked Issues or Issue Description

**What happened?**

A run that ended with `acpx_auth_required` or another known provider
authentication error did not immediately offer an inline way to connect
the provider. A repair form could also lock the user to the failed
account's sign-in method.

**Expected behavior**

Show a provider connection card in the task as soon as the
authentication failure is saved. Allow the responsible user to connect
or repair the provider with any supported sign-in method. Keep the
connection name automatic and resume the task after successful setup.

**Steps to reproduce**

1. Run a task with a supported provider and an expired or invalid
credential.
2. Let the run fail with a provider authentication error.
3. Open the task and attempt to repair the connection.

Related work: Refs #13724 and #13726. This change adds the inline task
repair flow and method choice.

## What Changed

- Classify provider authentication failures and create one connection
request for the current task. A persisted blocked classification
suppresses automatic retries only after the repair card is created;
unsupported providers retain their existing recovery path.
- Mark only the attributed, unchanged credential as needing sign-in.
Preserve credentials that were refreshed after the failed run started.
- Reuse the provider sign-in controls inside the task. Allow API key and
subscription choices for Claude, Codex, and Grok. Keep names hidden and
generate a default from the user, provider, and method.
- Keep the existing account when reconnecting with the same method.
Create and select another account when the method changes. Validate
updates to explicit agent bindings through the normal agent save path.
- Require explicit adoption for legacy agent authentication. Validate in
the agent environment, then commit the binding, connection install,
audit, and card completion in one transaction. Keep failed setup and
account selection visible and retryable.
- Add regression tests and update the specification and Connections
documentation.

## Verification

- Fresh local verification: 199 tests passed across the inline form,
provider method selector, default naming, authentication and recovery
classifiers, run liveness, OpenAPI routes, database adoption/rollback,
and Cursor execution suites. The adoption database suite also passed
against disposable Docker PostgreSQL.
- Full repository `pnpm build` and `pnpm -r typecheck` passed on the
latest commit. Token gates are clean.
- Embedded browser: opened real task cards from seeded authentication
failures; switched Claude from API key to subscription and back;
switched Codex from subscription to API key; confirmed the name field
stays hidden. Provider sign-in was not completed with real credentials.
- The broad local `pnpm test:run` started before review fixes and was
interrupted after the working tree changed; it is not counted as a
passing full run. Fresh focused tests passed. CI supplies the full test
and browser suite results for the current commit.
- CI is green on commit `4b97a4e447045ff3d7516525a187a5d1d21e0d4c`: 54
checks passed and two Storybook checks were skipped by their path rules.
The workspace preview job passed on one rerun after a local-server
startup timeout; its rerun passed 835 tests.
- Greptile is 5/5 on the same commit with no actionable findings and no
unresolved review threads.

## Risks

- Incorrect authentication classification could prompt for a connection
unnecessarily. Tests exclude tool authorization, quota, and unrelated
runtime failures.
- A method change selects the new personal provider default, which also
applies to other agents that use that user's default. Explicit account
bindings use the existing permission and runtime validation path.
- Credential invalidation must not race with refresh or reconnect. The
code compares the saved credential generation and grant update time
under locks.
- No database migration or new credential storage format is required.

## Model Used

OpenAI GPT-6 through Codex. The exact model ID and context window size
were not exposed in this session. Capabilities used: reasoning,
repository editing, shell commands, database tests, and embedded-browser
interaction.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused tests listed
above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 06:39:37 -05:00
Dotta 5c69b69ef2 Integrate Cursor child counter provenance
Co-Authored-By: Paperclip <noreply@paperclip.ing>

* commit '986f8cc5344107cc199c69e0581cba5401c54282':
  Verify Cursor child usage lifecycle and fence profile v9

# Conflicts:
#	doc/architecture/runner-cursor-capabilities.md
2026-09-30 05:34:15 -05:00
DottaandPaperclip 986f8cc534 Verify Cursor child usage lifecycle and fence profile v9
Mark missing and invalid child run ordinals as partial observations. Exercise the pinned vendor child construction and reuse methods on all three platforms and regenerate the Cursor-only execution identity.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 05:33:44 -05:00
DottaandPaperclip 3518902733 Integrate bounded warm retirement recovery
Co-Authored-By: Paperclip <noreply@paperclip.ing>

* commit '84a50124a997f4256932eab15fe79dd9ce875940':
  fix(runner): bound instruction collection and preserve pending retirement
2026-09-30 05:29:03 -05:00