Commit Graph
220 Commits
Author SHA1 Message Date
DottaandPaperclip 95dfde7c5c Merge current Cursor ACP runtime into Copilot v4
Preserve both profile-specific guards and complete terminal diagnostics; bind Copilot v4 to the regenerated ACPX patch.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 11:28:38 -05:00
DottaandPaperclip 3a08122558 Merge current master into Cursor ACP provider branch
Preserve upstream ACPX terminal diagnostics and rich extension delivery hooks.

Co-Authored-By: Paperclip <hello@paperclip.ing>
2026-09-29 10:36:20 -05:00
DottaandPaperclip 83076d7e7c feat: return completed handoffs to Agent Chat (#14408)
Return completed Agent Chat handoffs through a durable outbox and scope each generated update to its supplied tasks. Add recovery, browser delivery, result access, and calibrated quality coverage.

Validated with two consecutive ten-case Claude/Codex campaigns, all CI checks, and a 5/5 review.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 10:25:52 -05:00
Dotta 304472946d Merge commit 'abad2bd0d' into codex/runner-copilot-acp
* commit 'abad2bd0d':
  Bind Daytona image identity to Cursor isolation patch
2026-09-29 09:58:05 -05:00
DottaandPaperclip abad2bd0da Bind Daytona image identity to Cursor isolation patch
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 09:53:20 -05:00
DottaandPaperclip 3f4f8b37ba fix: grade Codex clarification and refusal outcomes from evidence (#14570)
Grade clarification lists, obsolete unstarted wakes, and refusal cancellation from persisted evidence. Preserve execution and ownership assertions, add boundary regressions, and version the affected eval definitions.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 09:40:55 -05:00
DottaandPaperclip 55fa510d5e Integrate final rich ACP foundation for production qualification
Co-Authored-By: Paperclip <noreply@paperclip.ing>

* codex/runner-cursor-acp:
  fix(runner): stop settlement publication after persistence failure
  feat(agents): persist agent files across tasks without revision history (#14420)
  test(runner): synchronize recovered ACP fixture cleanup
  fix(runner): persist ACP response settlement across crashes
  feat(storage): add plain directory sync with conflict preflight (#14416)
  fix: retry sandbox ACP input delivery after gateway failures (#14485)
  fix: retain project defaults in partial workspace overrides (#14502)
  test: control Telegram subscription retry timing (#14501)
  fix: retain diagnostic reasons for native runner failures (#14481)
  fix: prevent run identity locks from blocking audit checks (#14478)
  fix: preserve company context across browser hot reload (#14482)
  feat(ui): improve task composer controls and pending input (#14322)
  Improve task artifacts with rich cards and editable stories (#14469)
2026-09-29 09:33:14 -05:00
DottaandPaperclip 3d79b2b1f3 Integrate final rich ACP foundation for production qualification
Co-Authored-By: Paperclip <noreply@paperclip.ing>

* codex/runner-rich-acp:
  fix(runner): stop settlement publication after persistence failure
  feat(agents): persist agent files across tasks without revision history (#14420)
  test(runner): synchronize recovered ACP fixture cleanup
  fix(runner): persist ACP response settlement across crashes
  feat(storage): add plain directory sync with conflict preflight (#14416)
  fix: retry sandbox ACP input delivery after gateway failures (#14485)
  fix: retain project defaults in partial workspace overrides (#14502)
  test: control Telegram subscription retry timing (#14501)
  fix: retain diagnostic reasons for native runner failures (#14481)
  fix: prevent run identity locks from blocking audit checks (#14478)
  fix: preserve company context across browser hot reload (#14482)
  feat(ui): improve task composer controls and pending input (#14322)
  Improve task artifacts with rich cards and editable stories (#14469)
2026-09-29 09:33:13 -05:00
DottaandPaperclip 24beb00575 feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification.

Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 08:56:21 -05:00
DottaandPaperclip 8747917668 Merge master agent-file persistence into rich ACP foundation
Preserve the extended harness and instruction-persistence suites, generated
contracts, and cleanup/collection ordering. Record the unqualified ACPX
agent-directory capability without broadening environment or path policy.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 08:36:37 -05:00
DottaandFry 3ca196b0a6 feat(agents): persist agent files across tasks without revision history (#14420)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - An agent needs personal files across tasks and sessions.
> - AGENTS.md is one file in that directory. Supporting files need the
same persistence.
> - The Instructions Editor and agent runs must share one current
directory.
> - Concurrent runs should apply only the files they change. The last
sync of the same file wins.
> - This pull request uses existing file transport and removes temporary
copies after sync.
> - Old instruction-only sessions keep their restore contract. New saves
do not create revision history.

## Linked Issues or Issue Description

Refs #14325. This replaces its revision-oriented design with persistent
agent files. Keep #14325 unmerged.

Transport prerequisite #14416 merged first at
`d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master
and remains below 100 changed files.

Related work: #4513 and #8798 cover instruction tooling. This change
handles run synchronization, cross-task personal files, browser editing,
and old-session restoration.

## What Changed

- Keep one current directory per company and agent. Point AGENT_HOME at
a temporary working copy for each active run. Keep task files and
provider HOME separate.
- Restore text, binary files, and nested folders through workspace
transport. Exclude remote agent files from task Git snapshots with a
self-ignoring file inside the reserved runtime directory; never write
through repository-controlled Git metadata.
- Collect after the provider and child processes have stopped. Keep
resumable conversation state.
- Apply changed and deleted files under the agent lock. The last sync
wins for the same file. Unrelated concurrent changes survive.
- Remove temporary copies after successful sync, rejected sync, and
staging failure. Register ownership before copying so restart recovery
can remove interrupted preparation. Retry transient synchronization up
to three times. Preserve the original remote lease reference until
deletion succeeds; restart cleanup never acquires a replacement sandbox.
Do not create captured directories or a conflict-review queue for new
runs.
- Keep browser editing, stale-draft protection, and streaming binary
downloads. Keep the instruction entry and text editor limited to 1 MiB.
- Keep historical agent-folder sync failures on their affected runs
instead of repeating them above current saved instructions. Preserve
legacy candidate review and current browser-save errors. Avoid duplicate
quota warnings while retaining separate sync failures when they describe
a different problem.
- Require target-scoped caller grants for peer instruction access, while
preserving self edits, responsible-user checks, and protected-change
consent.
- Treat full storage as a nonblocking run warning, never an agent pause
or run-admission failure. Restore already-over-quota saved folders so
ordinary agent cleanup can recover; warn on each run until cleanup. The
run detail view shows the warning.
- Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash
large files as streams. Check editor-save quotas with metadata instead
of hashing unrelated files.
- Preserve old native inputs, instruction-only copies, paths, digests,
and pending legacy candidates. Adopt old revision heads once. New writes
do not append history rows.
- Add idempotent migration 0287 and verify upgrades from the preview
tables and receipts.
- Add nine interactive stories under **Agents / Persistent files**,
including automatic incoming edits, stale browser drafts, and
storage-limit diagnostics.

## Verification

- Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after
merging current master and the landed transport prerequisite.
Integration required no manual conflict resolution; the feature remains
99 changed files. Full workspace typecheck, production build, token
gates, and 715 focused tests passed on this merge candidate. Fresh
Greptile review is 5/5 with no unresolved findings. All 55 checks
passed, with four conditional skips, including the build, typecheck,
browser E2E, and canary dry run. A single retry recovered four jobs
interrupted by runner shutdowns; no source changes were required.
- Historical-warning UI fix: all 6,834 UI tests across 640 files passed,
including regression coverage for three old failures, legacy preserved
edits, and warnings scoped to the affected run. Full workspace
typecheck, production build, Storybook build, and token gates passed.
Browser-verified Storybook playtests passed for Historical Failures
After Successful Save, Storage Limit, and Full Storage Run Warning.
- Review follow-ups at `4e20c9fb2`: all 18 focused tests passed,
including external Git directories, linked worktrees, symlinks,
hardlinks, and distinct I/O failures alongside storage warnings. Server
and UI typechecks, token gates, and the production build passed.
- Storage warning regressions at `0724f3012`: all 33 directory tests and
all five heartbeat-list tests passed, with no skips in their successful
runs. They cover repeated runs while full, an already-over-quota saved
folder, cleanup, warnings retained after unrelated save failures, and
bounded warnings in large result JSON. Server typecheck passed after the
final warning fixes.
- Full workspace typecheck, production build, and token gates passed
during this follow-up. Product E2E harness: 631 tests passed across 52
files; harness typecheck passed. Earlier native session/context and
directory/legacy collection suites passed 537 tests; Runner
unit/transport suites passed 329 tests.
- **Real E2E at `0724f3012` (before this follow-up):** legacy local
Codex and native Daytona Codex each passed six tasks, one server
restart, seven independent assertions, and cleanup verification. Both
prove browser-to-agent edits, agent-to-browser edits, nested/binary
restoration, per-file last-sync-wins, a successful run after an
oversized save rejection, and cleanup clearing the warning.
- Native local Codex also passed the six-task quota flow before the
final warning-retention fixes. That pass began at `918d1ed02` while the
bounded-result warning fix was being edited, so it is not claimed as
exact-final-head evidence. Its final-head rerun failed during embedded
PostgreSQL bootstrap before any provider run: the macOS host had 87,365
of 87,381 SysV semaphores occupied. No unrelated services or kernel
limits were changed.
- The final-source report intentionally records **2/3 cells passed**,
preserving the blocked native-local attempt:
`tests/runner-e2e/results/agent-files-quota-final-20260928-report/`.
Earlier failed attempts and provenance notes remain under
`tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and
the original campaign directories.
- Daytona used immutable image
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf`
and its exact Linux runner binary. Controller source is `0724f3012`;
image source is recorded separately.
- Legacy-session compatibility and all three ACP Stop/resume browser
regressions passed on the prior validated feature head
`169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same
provider session is retained and interrupted writes are not replayed.
Migration upgrade tests also passed earlier.
- Nine interactive stories are under **Agents / Persistent files**,
including **Full Storage Run Warning**. Its playtest and visual browser
inspection passed; the warning states that runs continue and the editor
remains available.
- Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs
skipped, no failures or pending checks. All eight browser E2E shards and
their aggregate passed. Fresh Greptile review is 5/5 with no findings;
all review threads are resolved, the security scan passed, and GitHub
reports no merge conflicts.
- The broad local follow-up test run was interrupted after host
semaphore exhaustion affected isolated PostgreSQL instances. It also
encountered the existing macOS long-path fixture failure and two timeout
failures. This is not a claim that the full local suite passed. Logs are
retained; focused storage/warning tests passed.

## Risks

- A later sync can overwrite an earlier edit to the same file, including
a saved browser edit. There is no text merge or retained version. This
is the intended last-sync-wins policy.
- A save that exceeds a storage limit is rejected and its temporary copy
is discarded. The run itself continues normally, and later runs restore
the last saved files with a warning until cleanup. Transient sync
failures get bounded retries. An I/O failure partway through a sync can
leave some files updated; a failed receipt does not claim whole-folder
success.
- Larger folders increase copy time, network traffic, and temporary disk
usage. Active runs still need working copies. Terminal runs do not
accumulate archives. Operators must provision disk for agents and
configured concurrency; these limits are not company-wide quotas.
- A restored old native session remains instruction-only until a fresh
session starts. Its original conflict fence and existing pending
candidates remain compatible.
- Provider processes close at the collection boundary. Conversation
resume remains available, but warm process reuse is lost.
- Backups must include the instance filesystem and database. External
bundles keep their existing behavior until explicitly moved to managed
storage.

## Model Used

OpenAI Codex, GPT-6 family. The session does not expose a more specific
model ID or context-window size. Reasoning, code execution, and browser
tools assisted this change. Real provider E2E uses `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing>
2026-09-29 08:25:56 -05:00
DottaandPaperclip 9774086a73 Integrate current master before rich ACP foundation review
Co-Authored-By: Paperclip <noreply@paperclip.ing>

* commit '53aad90b9':
  fix: retry sandbox ACP input delivery after gateway failures (#14485)
  fix: retain project defaults in partial workspace overrides (#14502)
  test: control Telegram subscription retry timing (#14501)
  fix: retain diagnostic reasons for native runner failures (#14481)
  fix: prevent run identity locks from blocking audit checks (#14478)
  fix: preserve company context across browser hot reload (#14482)
  feat(ui): improve task composer controls and pending input (#14322)
  Improve task artifacts with rich cards and editable stories (#14469)
2026-09-29 07:50:49 -05:00
Devin FoleyandPaperclip faf72cb1d5 fix: preserve company context across browser hot reload (#14482)
## Thinking Path

> - Paperclip uses a shared company context for browser providers and
consumers.
> - Vite can load a new consumer module while an older provider is
mounted.
> - Recreating the context disconnects that consumer from the mounted
provider.
> - The consumer then reports a missing provider even though one is in
its React ancestry.
> - This change preserves the context object across development module
refreshes.
> - Bounded global error diagnostics distinguish development bundles and
otherwise context-free promise rejections.

## Linked Issues or Issue Description

**What happened?**
A refreshed company consumer can throw `useCompany must be used within a
CompanyProvider`. A real Vite and Chromium reproduction confirms that a
retained provider and a refreshed consumer can hold different context
objects. Global promise rejections also lack the bounded document state
already attached to React boundary errors.

**Expected behavior**
A refreshed consumer should read the mounted provider. Error reports
should identify the loaded bundle mode and bounded browser state while
preserving monitoring opt-in, sign-out, and privacy behavior.

**Steps to reproduce**
Run `pnpm test:e2e:browser-context`. The isolated Vite fixture renders
the real CompanyProvider, imports a new timestamped consumer module, and
renders that consumer below the retained provider. The test fails before
the context change and passes after it. The SDK regression invokes its
real unhandled-rejection handler with an undefined reason.

## What Changed

- Keep the React context object in Vite's per-module `hot.data`. Account
values stay in React.
- Add an isolated browser regression with mocked API responses and no
live instance, discovered by the existing Chrome CI shards.
- Add document-state diagnostics to global errors while preserving
earlier boundary snapshots.
- Tag events with development or production bundle mode and the type of
an unhandled rejected value.
- Document the new test command and diagnostic fields.

## Verification

- Company context, browser context, and Sentry suites: 59 passed.
- `pnpm test:e2e:browser-context`: passed in Chromium. The original
context code fails the reproduction.
- Real SDK tests preserve DSN/sign-out behavior and omit request
context, breadcrumbs, and private DOM data.
- `pnpm check:token-gates`: passed.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- Local `pnpm test:run` exited in the general-server phase: 13,819
passed, 86 skipped, 21 failures in unchanged filesystem and host-port
suites. Cache permission and long-path failures also reproduce on the
unmodified base; seven other failures involve local runtime port
ownership. This is not a full local pass.
- Full Linux CI and review passed at a0b5270662 (53 successful checks;
two conditional checks skipped). Three jobs interrupted by a runner
shutdown passed on rerun. Greptile is 5/5 with no unresolved comments.
The real browser regression also passed in the normal Chrome CI shard
(3.2 seconds).
- Code-owner approval remains required because the dedicated test
command changes `package.json`.
- Scanned the diff and PR text for credentials and private data before
pushing.

## Risks

The context cache applies only to development hot reload. It retains the
context object, not account state. Production continues to create an
ordinary React context. The diagnostic hook adds only bounded state and
type values, preserves boundary snapshots, and returns the original
event if enrichment fails. It does not suppress errors or restore raw
breadcrumbs. The global rejection diagnostics do not identify the
promise's originating operation by themselves.

## Model Used

OpenAI GPT-6 through Codex, with repository inspection, code editing,
and command execution. The runtime does not expose a more specific model
revision or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites and
Chromium regression; full local limitations documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 16:15:14 -07:00
0be2afcca6 feat(ui): improve task composer controls and pending input (#14322)
## Thinking Path

> - Paperclip lets operators assign tasks to AI agents and review their
work.
> - The task composer controls the next message and its assigned agent.
> - Operators needed a way to choose that agent's model and effort
without leaving the composer.
> - The old mode selector, upload button, and input cards made the
mobile composer crowded and hid normal messaging during a pending
decision.
> - Harnesses publish different model and effort capabilities, so the
picker must follow the selected agent.
> - This pull request adds one responsive composer flow, keeps pending
cards visible above it, and protects Codex ACP authentication in the
local test path.
> - Operators can choose run settings, send a message, and answer a
pending card as separate actions.

## Linked Issues or Issue Description

**Subsystem affected**

Task composer UI, issue thread interactions, Codex ACP credential
handling, and Storybook.

**Problem or motivation**

The composer did not expose model or effort for the selected agent.
Mobile actions wrapped poorly. Pending questions and confirmations
replaced the composer. A local Codex ACP test could also reuse host
authentication after the managed key was removed.

**Proposed solution**

Put assignee search, model search, exact model IDs, effort, and fast
mode in one picker. Use a mobile dialog. Replace the direct-upload plus
action and separate mode selector with an Add menu and removable Plan or
Ask chips. Place pending interaction cards above the usable composer.
Keep these cards pending after an ordinary message unless their creator
asks for comment superseding. Replace managed ACP auth files atomically
and isolate the test key from host credentials.

**Roadmap alignment**

ROADMAP.md does not list an overlapping composer milestone. This change
improves the existing task and review flows.

## What Changed

- Added the combined assignee, model, and effort picker to both task
composers. Search matches agent name, role, and harness. The server uses
a curated Codex list by default and honors instance-declared models.
Manual IDs remain available.
- Added an effort slider for known model capabilities, a conditional
Codex fast control, and reset. The picker opens in a modal on mobile.
- Added the Add menu for files, supported goals, Plan mode, and Ask
mode. Plan and Ask are exclusive removable chips. Keyboard mode cycling
remains available.
- Adjusted mobile spacing, avatars, wrapping, and Send placement.
Removed the composer divider.
- Moved pending question, confirmation, review, and related cards above
the composer. Ordinary comments now leave question and confirmation
cards pending by default. The onboarding prompt retains explicit comment
superseding.
- Updated the Storybook composer group with responsive states and the
production picker. Added UI, service, route, and browser regression
coverage.
- Isolated Codex ACP API-key authentication, skipped subscription auth
merge and shared-home copy-back for remote API-key runs, and replaced
the managed auth file atomically.

## Verification

- `pnpm -r typecheck` — passed on the final local head.
- `pnpm check:token-gates` — passed on the final local head.
- `pnpm exec vitest run server/src/__tests__/adapter-models.test.ts
ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` — 31
tests passed, including role and harness search, declared Codex models,
and filtering general OpenAI models.
- `pnpm exec vitest run
server/src/__tests__/issue-thread-interactions-service.test.ts` — 74
tests passed.
- `pnpm exec vitest run
packages/adapters/codex-local/src/server/acp.test.ts` — 42 tests passed,
including remote API-key copy-back isolation.
- `pnpm test:run` — attempted locally; the embedded PostgreSQL test
database could not initialize on macOS. The isolated
`heartbeat-run-event-sequencing` suite reproduced that environment
failure. GitHub CI runs the full test matrix for this head.
- `pnpm build` — passed on the final head. `pnpm build-storybook` passed
after the last UI change; only server code, tests, and docs changed
afterward.
- Live local test drive — Codex ACP ran a task with a managed API key.
The test agent was restored to its default ACP configuration afterward.
- Review the interactive stories under the top-level Composer group with
`pnpm storybook`. Check a narrow desktop width and mobile Plan, Ask,
picker, and pending-question states.

## Risks

- A pending card stays open when an ordinary comment changes the
discussion. Its creator can set `supersedeOnUserComment: true` when a
new comment should replace it.
- Model and effort overrides persist on the task until reset or changed.
An unlisted manual model ID may fail when the provider runs it.
- Some harness catalogs do not report effort support. The picker hides
effort for those models.
- No database migration is required.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI GPT-6 via Codex. This runtime does not expose the exact model ID
or context window to the task. The model used code execution and browser
tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI Codex <codex@openai.com>
2026-09-28 23:02:39 +00:00
Dotta eecd3e9041 Merge branch 'codex/runner-cursor-acp' into codex/runner-copilot-acp
* codex/runner-cursor-acp:
  fix(runner): pin patched Node for remote provider packs
2026-09-28 17:36:04 -05:00
Dotta 1edb2173a9 Merge branch 'codex/runner-rich-acp' into codex/runner-cursor-acp
* codex/runner-rich-acp:
  fix(runner): pin patched Node for remote provider packs
2026-09-28 17:36:01 -05:00
Dotta 3a5637732b fix(runner): pin patched Node for remote provider packs 2026-09-28 17:35:56 -05:00
DottaandPaperclip a55422a48d Integrate Cursor prerequisite without losing Copilot registrations
Combine both pinned native builders, isolated profile installers, capabilities, extension adapters and complete Daytona content inputs. Verify dispatch remains isolated between the two native installation factories.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 16:26:55 -05:00
DottaandPaperclip 24c58e479a Improve task artifacts with rich cards and editable stories (#14469)
Render eight artifact card types from real task records and share them with editable Storybook stories. Preserve document review and media/file actions, and load bounded CSV previews on request.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 15:38:06 -05:00
DottaandPaperclip 6f7e2d20d9 Audit Copilot provider packs through verified offline launch
Hash the exact Copilot builder and materializer into Daytona image identity, and add a credential-free initialize smoke that exercises the packaged installation registry and native snapshot launch.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 15:17:09 -05:00
DottaandPaperclip 747ca61ad0 Bind Cursor candidate assets to Daytona image content
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 15:16:00 -05:00
DottaandPaperclip 4b20957878 Merge public Grok packaging while preserving ACP candidate boundaries
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 15:03:59 -05:00
DottaandPaperclip c3b7e9ecd0 Merge explicit task context ownership while preserving candidate eval lanes
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 14:55:20 -05:00
DottaandPaperclip 18e8c121d9 fix(runner): include Grok support in public installs with sandbox prerequisites (#14024)
## Thinking Path

> - Paperclip manages agents through a shared native runner.
> - Built-in harness support should ship with Paperclip's public
distribution.
> - Grok already speaks ACP; it does not require a new public bridge
package.
> - Sandbox provisioning owns the native executable and its pinned
version.
> - The runner must verify that prerequisite without downloading it
during npm installation.
> - This change separates built-in launcher identity from external
runtime identity.
> - Clean npm installation and live staging checks verify the
distribution boundary.

## Linked Issues or Issue Description

Refs #13882, #13973, #13977, #13979.

This follow-up now targets master after #13882 was squash-merged. It
replaces the private `@paperclipai/grok-acp` workspace package with
runner-owned assets. Current master is included so the branch also
contains the merged scheduler, complete-event capture, and durable
cleanup fixes.

## What Changed

- Ship Grok launcher and qualification metadata inside the runner's
compiled output and the public server's vendored runner tree.
- Remove the separate Grok npm package and all package-manager install
hooks for this runtime.
- Require the checksum-verified Grok Build 1.0.13 binary at
`/opt/paperclip/providers/grok/1.0.13/grok` in the selected execution
environment. Provision it explicitly in the Daytona image and CI setup.
- Keep native binaries outside the provider pack. Bind the built-in
launcher into the pack manifest.
- Preserve executable leases, descriptor-backed startup, credential
fences, permissions, and exact ACP model admission.
- Use `builtin:grok-acp` and `native:grok` as profile identities.
Historical package-profile sessions fail closed on resume rather than
being silently reinterpreted.
- Resolve built-in assets from the authenticated sidecar location,
including public server npm layouts. Keep the controller path out of
provider environments.
- Add clean npm tarball installation verification to the existing
trusted canary CI job and the admitted manual EC2 verification path. It
stages a unified release version and runs npm lifecycle scripts, then
verifies missing-prerequisite rejection and admission after separate
provisioning without credentials or inference.
- Include the controller-owned provider pack in stamped Cloud images.
Unstamped local images omit the pack and remain usable; remote ACPX
requires full source provenance.
- Correct CLI approval-page metadata for an already authenticated Cloud
board user; approval authorization remains unchanged.
- Honor explicit native-runner enablement in the Cloud agent picker and
direct setup page, keeping the flag disabled by default.
- Allow selecting the execution environment before connecting
credentials. Include Grok in the existing authenticated hello-probe
flow, targeting its pinned native prerequisite for runner setup.
- Recover an existing subscription sign-in conflict through an explicit
cancel-and-retry action, serialized after cancellation succeeds.
- Preserve the selected ACPX harness before normalizing config fields,
so new Grok agents use the Grok default model.
- Keep the credential-free Cloud provider pack root-owned and readable
after runtime UID remapping; verify manifest and referenced asset access
under an unrelated unprivileged UID during image builds.
- Archive prior failover backups alongside explicitly replaced harness
state, preserving evidence while preventing stale backups from blocking
a fresh replacement.
- Update Daytona image content inputs and contract tests for the
built-in assets and explicit provisioner.
- Document and regression-test the shared `approve-all` default for Grok
setup, saved configuration, and native execution. Explicitly saved
restrictions remain unchanged.

## Verification

Current merge-repair head `df09eb3e1a619430ad8419a0ee9aedd486689b05`
incorporates master `f1a394bd30cb56fb9e479f98b9f50176fe921858` after the
base PR was squash-merged. All 12 conflicts came from incoming files
identical to the tested pre-squash base. The final tree exactly matches
a three-way merge using that original base, preserving built-in Grok
distribution and removal of the obsolete private package. All 252
focused runner/UI tests, six npm-isolation tests, and token gates pass.
Fresh exact-head Greptile review is 5/5 with no outstanding findings;
security scans and EC2 native compilation pass. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([run
36468768035](https://github.com/paperclipai/paperclip/actions/runs/36468768035)).
The repository owner explicitly authorized bypassing code-owner approval
after all checks passed; no CI checks or repository protection settings
are bypassed or changed. The only remaining PR was removed from the
completed stack metadata to permit native auto-merge.

Earlier integration head `78cb306ecc41b5c96577c26c1d89153b0ef865a1`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
attempts](https://github.com/paperclipai/paperclip/actions/runs/36447124691)).
The initial attempt lost two EC2 runners to shutdown signals and stalled
a third shard during dependency preparation; all three passed the
same-commit failed-job-only retry. Trunk code-owner requirements remain
enforced. The review summary’s non-blocking saved-asset offset
classification note concerns code already merged in #14301; those
runtime files are identical to master and outside this stack’s diff.
Historical live evidence below retains its original source revisions.
[Final public npm
verification](https://github.com/paperclipai/paperclip/actions/runs/36445542764)
passed on `76ea70cd4d13786a042af9df82f0fd7a8c85ae30`: 17 public
packages, an executed offline lifecycle sentinel, unchanged consumer
lock, built-in launcher, missing-prerequisite rejection, and verified
separately provisioned binary/command lease. Provisioning and cleanup
require no host privilege elevation; only the positive probe mounts the
temporary native binary read-only. The verifier is unchanged by the
final master merge. All six isolation tests and an offline npm smoke
test pass. The prior head had 56 green CI checks and a 5/5 review after
two unchanged tests timed out and passed a failed-job-only retry ([CI
attempts](https://github.com/paperclipai/paperclip/actions/runs/36444597313)).
All 56 recovery-display/lineage tests pass; re-review cleared the
already-covered missed-retry concern. Earlier EC2 failures remain
retained: [npm lockfile
rejection](https://github.com/paperclipai/paperclip/actions/runs/36436311203),
[missing compiler in the slim
image](https://github.com/paperclipai/paperclip/actions/runs/36440210984),
and the aggregate 15-minute test timeouts in those broad runs. Both
broad attempts passed typecheck, token gates, Product E2E type/unit
checks and build. The focused EC2 lane preserves the existing
trusted-actor and immutable-source gates.


Earlier documentation/test checkpoint
`ff244c4fd78a7ede5a3e00efe09f475f133ef33e` leaves runtime behavior
unchanged. 154 focused tests pass across configuration building, native
provider resolution, permission policy, credentials, UI configuration,
and new-agent setup (including both Grok auth modes); token gates pass.
All fresh CI is green for this head: 56 successful checks/statuses and
two intentional skips ([run
36367065119](https://github.com/paperclipai/paperclip/actions/runs/36367065119)).
Greptile is 5/5 with no new findings. Grok already inherits the shared
`approve-all` default, so unattended setup requires no manual permission
change.

Runtime head `bb5a9307991f1ac567b781970ef11b39d518e19b` fixes a final
staging continuation failure before provider startup: explicit
replacement archived the old harness but left its failover backups
active, which caused `runner_harness_state_mismatch`. The regression
fails before the fix and passes after it; all eight adjacent
recovery-safety cases also pass. Old backups remain inspectable inside
the continuity archive. All fresh CI is green at this head ([run
36360839248](https://github.com/paperclipai/paperclip/actions/runs/36360839248)),
with a 5/5 review. One unrelated Cursor test timed out in the initial
server shard; the same-commit failed-job rerun passed, and both attempts
are retained. Staging deployment is confirmed healthy on this revision.
The controller image is
`ghcr.io/paperclipai/paperclip@sha256:6ad91c487910ccd2596ff7aed0a3a3ea5233d12b51b83cd6e1402237749b9673`.
The final browser-created staging task passed on this exact revision
with API authentication: context read → structured human question →
controller restart → answer submission → same native provider session
resumed → document saved → task Done. The two turns took approximately
119s and 77s. The actual write receipt was applied, and the saved
document has exactly one revision containing the selected answer and
requested marker. Usage and cost were not reported. [Controller image
build](https://github.com/paperclipai/paperclip/actions/runs/36360889243).

- Previous integration head `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`:
all CI green (53 successful checks/statuses, two intentional skips),
including repository typecheck/build/tests, native Runner tests, browser
shards, and canary installation checks. [CI run
36358672529](https://github.com/paperclipai/paperclip/actions/runs/36358672529).
Greptile is 5/5 with no unresolved findings.
- Focused checks cover Grok credentials, executable admission, launcher
assets, provider-pack paths/permissions, workflow contracts, setup
defaults, CLI authorization, and subscription conflict recovery. All 39
protocol definitions validate. Final integration checks pass 124
catalog/evidence/cache tests and nine project-form tests; token gates
pass. Some local dependency checks could not load the stale installed
dependency tree; the corresponding fresh EC2 checks pass.
- Clean public npm installation passed on EC2 at
`8b172ebcf8e02e30662d830c00f3961e3bd459ec` ([run
36164964900](https://github.com/paperclipai/paperclip/actions/runs/36164964900)):
17 unified-version packages, lifecycle scripts enabled, built-in
launcher present, no separate Grok package or npm-downloaded binary,
missing prerequisite rejected, separately provisioned native executable
and command lease verified. No credentials or inference were used.
Subsequent changes preserve this npm asset layout.
- The immutable Daytona prerequisite image is
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:98957d5be0ac774d086b6402b5849e8e6356fec70fb8c09fca6eb4ed6de918e0`,
built from `5a2db471f3ddabe77f9f80e76ed27f996cb97fba`. The previous
Cloud controller image was
`ghcr.io/paperclipai/paperclip@sha256:fd914e1ab1e45f741e8e078ff452d16f082d7ac05f9b4b3506d3a3c64150d204`,
built from `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`; it is superseded
by the latest image above. Its EC2 build verified provider-pack access
under an unrelated unprivileged UID.
- Browser staging at `40f898bc4cba73c1dff4e6344a3983ba0fb247ef` passed
full Grok onboarding with the correct `grok-4.7` model, saved credential
delivery, and pinned Daytona execution. A browser-created task read
context and asked the structured human question. After a controller
restart, answering the persisted question resumed the same native
provider session, saved the requested document, and completed the task.
Actual tool outcomes and durable state agree: one question and one
document revision. The two successful turns took 42.7s and 63.1s; usage
and cost were not reported.
- Restricted policy returned the expected `approval_required` outcome.
Functional staging tests explicitly selected `approve-all`; controller
authorization and governed approvals remain enforced. Temporary board
CLI access was revoked and verified rejected (HTTP 401), and the
disposable onboarding agent was paused.

Failures remain retained: the pre-fix continuation failure (its task
remains blocked; the passing final task is fresh), the original Cloud
provider-pack permission failure, the expected restricted-policy denial,
the superseded npm staging failure, and an earlier monolithic CI
infrastructure timeout. Browser CI exposed a project alias/form race;
the final stack uses master's stronger draft-preservation fix and all
browser shards pass. Historical full subscription/API protocol and
Product rosters retain their original source revisions and do not
qualify this packaging revision. No local Docker or Rust build was used.

## Risks

The branch includes master’s draft-preservation fix for project URL
aliases. It keeps the same project’s edit form mounted and clears prior
data when the project or company changes.

Custom sandboxes and local execution hosts must provision the pinned
binary before Grok starts. Missing, changed, unsupported-platform, and
symlinked executables fail admission. The new builtin profile cannot
resume sessions created with the former private-package profile.
Existing Claude/Codex npm bridge profiles retain their package pins.
Grok restricted modes preserve the selected policy but cannot
automatically admit Paperclip calls: ACP permission metadata does not
independently bind tool authority, so those calls stop with
`approval_required`. New Grok configurations default to `approve-all`,
including API configurations that omit the mode. Existing explicitly
restricted configurations remain restricted; controller authorization
and governed approvals remain enforced.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:54:44 -05:00
DottaandPaperclip 768c70f69d Merge mainline Grok runner support without losing rich ACP qualification
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 14:54:04 -05:00
DottaandPaperclip 992f720262 fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English
(ASD-STE100). -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task descriptions, comments, continuation data, skills, and
execution rules enter several agent adapters.
> - The same source can be rendered by more than one automatic input
carrier.
> - Failed resumes can also rebuild input from stale or compact context.
> - This pull request gives each Paperclip-owned source one delivery
owner and preserves the required transport boundaries.
> - It adds deterministic adapter, interaction, runner, and browser
tests for these boundaries.
> - The benefit is more predictable context delivery with explicit
evidence for later live qualification.

## Linked Issues or Issue Description

Related: #13144 removes a duplicate environment payload and bounds wake
lists. Related: #11360 addresses Hermes resume behavior. This pull
request preserves compatible active-session formats while repairing
context ownership and stale question creation.

**What happened?**

Task descriptions and comments could enter more than one automatic
context block. Native transports could wrap a complete model input in a
second task envelope. Some legacy and gateway adapters could omit the
owned assignment on ordinary tasks or rebuild a failed resume with stale
compact context. A continuation could also request a question after
newer human comments had arrived.

**Expected behavior**

Each task or comment source has one automatic model-facing owner.
Distinct comment IDs and repeated wording remain distinct. Fresh
fallback attempts rebuild the required full context. A question request
is rejected when newer queued human direction makes it stale. Harness
access policy remains owned by execution configuration.

**Steps to reproduce**

1. Build a task with a description and current comments.
2. Capture the actual adapter or runner input.
3. Compare source ownership and task-envelope nesting.
4. Queue a human comment before a continuation requests a question.
5. Trigger a failed resume and inspect the fresh retry input.
6. Run the focused adapter, interaction, runner, and browser checks.

## What Changed

- Add shared prompt-section selection at the provider-attempt boundary.
- Deliver owned assignment context through native, legacy CLI, ACP,
gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and
Hermes paths.
- Rebuild full or compact context after resume recovery changes the
attempt. Add native and Claude ACP tests of actual recovery requests.
- Preserve custom templates, loaded instruction files, execution
policies, and older active-session formats.
- Record continuation source metadata and reject stale question creation
under the issue-row lock.
- Add explicit Product E2E context-integrity profiles, prerequisite
gates, credential-isolation checks, and report fixtures.
- Bypass service-worker forwarding for same-origin Vite development
modules. A real Chromium test fails with resource exhaustion before the
repair and passes after it. Production asset caching keeps its existing
policy.
- Add browser diagnostics and service-worker module-loading regressions.
- Add an explicit zero-retry eval option. The default retry behavior
remains unchanged. Each campaign records its effective policy.
- Remove the model-facing working-directory sentence from four prompt
builders. Existing workspace, sandbox, permission, and custom-template
configuration remains unchanged.
- Align the everyday workflow assertion with the current 47-entry
catalog.

Compared with current upstream master, the branch carries the
context-ownership implementation and its tests, the explicit
context-integrity catalog and evidence harness, and the focused browser
regression checks.

## Verification

**Merge assessment:** focused regression evidence supports merge. This
is not full completion of the original broad qualification matrix. The
maintainer has authorized merge after fresh verification of the master
integration.

- Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This
integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`.
All 14 conflicts are resolved. Cancellation checks, workspace
finalization, native Grok support, and both sets of tests are retained.
- Current-head Greptile: **5/5**, with no blocking findings. The review
names this exact commit. All **59 reported checks are terminal: 55
successful, 4 skipped, zero pending or failing**. This includes the full
root general and serialized suites, separate runner checks, typecheck,
build, canary, browser E2E, Docker, and security checks. The successful
legacy security status is included in that total.
- After integration: workspace typecheck and full build passed. Separate
runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust
tests, and 39 preparation checks**. Other passing checks include 621
Product E2E harness units, 376 focused shared/adapter tests, 160
real-database/API tests, 86 Hermes tests, 18 browser-support checks, and
Product E2E typechecking. The complete root suite passed in CI. The
duplicate local monolithic root run was stopped after that CI result; it
is not counted as a completed local pass.
- New native recovery coverage retains full assignment, completion
contract, and explicit skill selection after safe replacement, for old
and prepared input formats. Full native session test file: **136/136
passed**.
- New Claude ACP coverage captures actual fresh, resumed, and
missing-session fallback requests. It verifies one assignment copy,
comment order, identical text under distinct comment IDs, and full
fallback context. Full file: **33/33 passed**. Both affected TypeScript
checks passed.
- Existing deterministic tests cover source revisions, approval and
trust boundaries, completion validation, custom templates, compatible
sessions, standalone driver wrapping, and maintained adapter transport
requests.
- Provider-free browser support: **17/17 passed** after the master
merge. Service-worker unit tests: **33/33 passed**. The module-overload
regression failed before the repair and passed after it in real
Chromium.

### Fresh live comparisons

The new batch ran exactly four Product E2E attempts. **All four passed
on the first attempt; no retries.** Each has six terminal matchers plus
the existing browser lifecycle and invariant checks.

| Exact case ID | Control | Candidate |
|---|---|---|
| `core-compatibility.runner-codex.local.plan-revise-accept` | Passed |
Passed |
|
`local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume`
| Passed | Passed |

The plan case checks a revised canonical plan and revision-bound
approval before completion. The question case restarts the server before
submitting the answer, then verifies the continuation completes.

Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate
source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical
frozen definitions and provider versions: Codex `0.156.0` with
`gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with
`claude-sonnet-5`. The September 24 head added master browser recovery
and test-only changes. The September 28 head also integrates newer
master changes, including cancellation, workspace finalization, and
native Grok. These are frozen-source live results, not exact-head live
runs.

The candidate received one description copy where the control initially
received three. The submitted initial plan envelopes were 7,969 versus
19,097 characters. Question envelopes were 7,592 versus 18,919. These
are structural measurements, not whole-provider token or dollar savings.

### Earlier evidence and failed attempts

- The preceding fresh batch has four effective passing pairs: OpenCode
comment continuation and assigned skill, native Codex comment
continuation, and native Claude comment continuation. It retains **11
attempts: eight passed and three failed**.
- Original failures remain recorded: missing local PostgreSQL library
links before task creation; host-sleep cleanup after task/page checks
passed; and a Claude **control** session-open rejection before a model
turn. Setup was repaired identically on both worktrees. The permitted
unchanged infrastructure retries passed. The underlying Claude provider
startup error was not retained and remains unknown.
- Older R2 retains **17 passes and one failure** across 18 attempts,
including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode
blank-page failure led to the service-worker repair. R2 is historical
evidence: master changed the native fixed prompt and removed duplicate
wake environment data afterward.
- The September 24 CI run initially failed one unrelated preview
readiness test (`ECONNREFUSED` on its local fixture). Its test and
production code match master. Isolated local verification passed **28
tests, 3 skipped**. One unchanged CI retry passed the full shard: **831
passed, 1 skipped**, including all **31 preview-exposure tests**. The
aggregate CI gate passed afterward. The precise startup cause remains
unknown; a port race is a hypothesis, not a proved cause.

### Limits

The original wider profile/workflow matrix, repeated trials, and remote
Daytona qualification are incomplete. These results support a focused
merge recommendation, not statistical equivalence or universal harness
qualification. Some usage receipts are missing in both variants, so no
token or dollar savings are claimed. The $500 ceiling was preserved
using conservative allowances; failed attempts and unknown charges
remain in the ledger.

Reproduce the focused additions with `pnpm exec vitest run
packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm
--filter @paperclipai/paperclip-runner exec vitest run
src/native-session-runtime.test.ts`. Full checks use `pnpm -r
typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner
checks. Paid evals require the frozen definitions, profiles, and
credentials; do not use `--all` as a substitute for the selected cases.

## Risks

- Context placement changes can affect model behavior. Deterministic
checks cover the selected paths, but live qualification remains
incomplete.
- The stale-question guard can reject a request when queued human
comments arrived during the run. This is intended.
- New stored inputs and model envelopes retain compatibility readers for
older active sessions.
- Custom templates may intentionally repeat content.
- Removing a model-facing working-directory sentence does not change
filesystem, command, sandbox, or permission configuration.
- The worker bypass applies only to same-origin development module
paths. Cache-policy tests preserve private-response handling and
production asset caching. Mounted HTTP fixture changes remain test-only.
- This PR does not claim measured token savings or statistical
equivalence across every harness.

## Model Used

OpenAI Codex, exact model gpt-6-astra, with repository tools and code
execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The
serving context-window size is not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR using the required issue fields
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run the focused local checks and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect these changes
- [x] I have considered and documented risks above
- [x] All current-head Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
for the current head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:49:14 -05:00
DottaandPaperclip 97bd3112de Bound candidate Product E2E provider execution to two minutes
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 14:31:10 -05:00
DottaandPaperclip af06ada99f Add explicit extended ACP Product E2E qualification suite
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 14:21:44 -05:00
DottaandPaperclip f1a394bd30 feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path

> - Paperclip manages AI agents and governs their work.
> - Its native runner uses structured provider protocols for sessions
and tools.
> - Grok Build supports ACP over stdio, but the runner did not expose
it.
> - Native execution requires company-scoped credentials, verified
identities, and permission gates.
> - This change adds Grok through ACPX for local and Daytona execution.
> - Subscription login and explicit API-key execution have separate
credential paths.
> - Qualification grades real tool outcomes, durable state, and browser
workflows.

## Linked Issues or Issue Description

Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979.

Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`,
`acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents
keep their adapter. Merge the three companion fixes (#13973, #13977,
#13979) before treating the integrated Product qualification as deployed
behavior.

## What Changed

- Synchronize shared, TypeScript, Rust, server, validation, and UI
provider contracts.
- Run Grok native ACP stdio through ACPX and the authenticated Paperclip
MCP bridge. Verify the pinned executable and exact ACP model identity.
- Prefer company subscription login. Support an explicit company-secret
API key without automatic paid fallback. Fence refresh and copyback to
the same account and remove private runtime credentials after
containment.
- Preserve selected permissions, cancellation, durable session identity,
resume, and restart recovery. Keep unsupported steering and goals
unavailable. Preserve missing usage and cost as unknown.
- Package checksum-verified Grok Build 1.0.13 for Daytona with an
immutable, signed image built on EC2.
- Add deterministic admission, protocol, permissions, identity,
credential, failure, and cleanup checks. Add the maintained 39-case
protocol roster and separate subscription/API Product profiles.
- Fix live-test findings in reasoning events, reloads, idle-owner
retirement, credential-home cleanup, expired-login model discovery,
launcher pinning, and rerun evidence selection.
- Align control-plane state readers with the transport's 64 MiB bound
while retaining identity, ownership, lifecycle, and size rejection
checks.
- Stabilize two asynchronous CI assertions while retaining actual
outcome and filesystem-evidence checks.

## Verification

Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)).
Trunk code-owner requirements remain enforced. The review summary’s
non-blocking saved-asset offset classification note concerns code
already merged in #14301; those runtime files are identical to master
and outside this stack’s diff. Historical live evidence below retains
its original source revisions.


Earlier integration checkpoint:
`24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master
`0f14d2612`, preserving Grok qualification alongside the new accounting
and lifecycle suites. All 124 focused catalog, evidence, and
service-worker checks pass. The current base workflow includes the
explicitly selected public-install verification lane; follow-up #14024
supplies its verifier script. CI at that earlier checkpoint was green
(56 successful checks/statuses, four intentional skips), and the review
is 5/5 with no unresolved findings. Prior feature CI at
`fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run
36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902));
that is historical evidence, not a current-head result.

Paid Product measurements use frozen integrated source
`2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature
with #13973, #13977, and #13979. That source passed all 52 CI checks and
clean 5/5 review. Later master syncs incorporate upstream changes. Their
checks remain separate from these pinned live measurements.

| Check | Result and source-pinned report |
| --- | --- |
| Subscription protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html),
runtime `bc6833f7`, evals `92bb4b8c` |
| API protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html),
runtime `4a1061c8`, evals `3213dbec` |
| Subscription full Product matrix | [16/16 first attempts; 144
assertions; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html),
source `2d939a92` |
| Subscription core repetitions | 18/18: tool use, planning approval,
and Stop/resume each passed three times in local and Daytona profiles.
The full matrix contains repetition one; [repeat
two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html)
and [repeat
three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html)
each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. |
| API smoke and question continuation | [4/4 first attempts; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html),
both environments at `2d939a92` |
| Historical API Product coverage | [16/16 full
matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html)
and 18/18 core repetitions at `4a1061c8`; retained as measurements of
that revision |
| Native Daytona proof | Three subscription and three API
MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login
admission and fenced refresh checks passed without inference. All test
sandboxes were removed. |
| Inspectable artifacts and UI | Current-source screenshots verify
planning approval, direct Ask completion, question continuation after
controller restart, and two downloadable project revisions. The project
downloads pass 12 and 18 tests; all 40 independent artifact oracle
checks pass. |
| Provider-free checks | 116 eval-validator tests, 39 Grok definitions,
and 359 enabled/external campaign cells pass. Continuation regressions
above 2 MiB and 16 MiB failed before their fixes; 32 focused
recovery/ownership/size checks pass. |

The 32 unique current-source Product attempts have no failures, retries,
or skipped cells, and all cleanup checks pass. Whole-workflow timing,
model identity, image and provider-pack provenance, attempts, and
accounting coverage are retained in the canonical reports. The report
publisher's conservative `complete=false` flag is preserved; independent
audits verify the exact selected source catalog and immutable result
rows.

Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model
`grok-4.7`. Linux binary SHA-256:
`edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`.
Launcher SHA-256:
`f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`.
Image:
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`.
Image build source is `4196a4cd`, recorded separately from application
source `2d939a92`; each campaign verifies the image signature and
provider pack.

Original failed campaigns remain available: [continuation
bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059),
[scheduler/event
capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537),
and [startup cleanup plus EC2
interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743).
They retain their original grades. No Docker or Rust builds ran on the
developer laptop for these follow-ups.

## Risks

Merge packaging follow-up #14024 with this base before public release.
The follow-up replaces the private Grok bridge package with a built-in
launcher and makes the native binary an explicit sandbox prerequisite.

Three separate, reviewed fixes are part of the tested integrated
behavior: #13973 serializes task-run admission; #13977 captures complete
event evidence; #13979 durably reconciles failed Daytona creation. Each
has green CI and clean 5/5 review. Failed-create recovery has 277 plugin
tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona
lost-deletion-receipt proof. The live proof uses a private file for
journal persistence; database durability is covered by host tests.
Worker death before delivery of a failure envelope remains outside that
recovery mechanism.

Subscription fixtures stage an authorized company login; interactive
browser sign-in is not qualified. Local Product profiles ran on EC2
Linux. The temporary subscription credential was removed from the
protected GitHub environment after all subscription audits, with absence
verified. Runtime homes and refresh copyback remain ownership-fenced.

Protocol results remain pinned to their original revisions; they are not
relabeled as tests of the latest feature commit. New binary/model
versions require qualification. Missing token usage and model cost
remain unknown; runtime estimates do not establish a full bill.
Automatic paid Grok scheduling remains disabled pending separate
reviewed enablement. The 64 MiB bound can increase memory use for
verbose sessions, and larger files still fail closed. No automatic
legacy-agent migration occurs.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:54:36 -05:00
DottaandPaperclip d9d2147171 fix(auth): keep Cloud tenants on the Cloud sign-in flow (#14407)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud owns human identity and passes a verified identity to each
tenant.
> - The tenant can report no session while the Cloud session is still
valid.
> - The access gate and direct `/auth` route then show the instance
password form.
> - This pull request sends those users through the configured Cloud
entry endpoint.
> - Cloud can renew the tenant session or show its login page, then
return to the original task.

## Linked Issues or Issue Description

**What happened?**

A Cloud tenant can display the self-hosted email/password form after an
instance session check returns no session. This gives Cloud users the
wrong login method.

**Expected behavior**

An active Cloud session renews tenant access automatically. A signed-out
user signs in through Cloud. Staging and production use their own
configured Cloud origins. Self-hosted instances keep their instance
login form.

**Steps to reproduce**

1. Open a Cloud tenant task or an `/auth?next=...` link.
2. Keep the Cloud session active but make the instance session check
return 401.
3. Observe the instance password form instead of Cloud session recovery.

**Deployment mode**

Cloud-managed authenticated instances. No database or server API
changes.

Searched related authentication PRs. Native self-hosted OIDC support in
#10411 is a separate feature; this change uses the existing Cloud entry
contract.

## What Changed

- Wait for deployment metadata before showing an instance login form.
- Use the health response's Cloud origin and stack slug for session
recovery.
- Preserve the tenant path, query, and fragment. Reject external and
recursive login return targets.
- Limit automatic recovery per tab. Show a manual Cloud retry after
failed recovery. Show service failures as errors.
- Keep self-hosted login and local trusted access. Add focused tests,
browser regressions, deployment documentation, and an unavailable-state
design example.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed.
- UI typecheck and `pnpm check:token-gates` passed after the final UI
edits.
- All UI tests passed: 639 files, 6,777 tests.
- 63 focused Vitest tests passed across Auth, CloudAccessGate, Cloud
links, and recovery coordination.
- `pnpm exec playwright test --config tests/e2e/playwright.config.ts
tests/e2e/cloud-auth.spec.ts`: 5 passed. The tests use the real tenant
UI and database with a simulated Cloud HTTP endpoint. They cover both
Cloud origins, direct auth/task links, no password-form flash, preserved
URLs, reload, and self-hosted login.
- Hands-on browser test used the real Cloud gateway and a fresh tenant
build with disposable local data. Active Cloud session plus a forced
missing instance session returned to the task. An expired tenant cookie
also renewed automatically and returned to the task. Removing both
sessions reached the real Cloud email/social login UI. Persistent
failure stopped at the retry screen; retry succeeded after removing the
injected fault. The fixture used a loopback transport adapter and a
simulated signed-out OIDC issuer. No production session or deployment
was changed.
- The default local browser startup hit the host's embedded PostgreSQL
resource limit. The passing run used a separate disposable database on
the test PostgreSQL process.
- The full local `pnpm test:run` sweep was stopped after about 31
minutes once CI completed the full suite. It had reported 33 failures in
the unchanged runner API unit/integration files; both files pass in
isolation (1,749 + 28 tests). The local sweep did not reach the later
workspace/serialized groups. CI completed all of those groups
successfully.
- Greptile reviewed commit `d603fd4e39455de44da9dae81b72197096c0e1e8` at
5/5 with its only thread resolved. All CI gates are green on this
commit, including all general/serialized server groups, workspace tests,
Runner checks, typecheck, build, and all eight browser shards
([run](https://github.com/paperclipai/paperclip/actions/runs/36446230697)).

## Risks

- Recovery depends on valid Cloud origin and stack metadata. Incomplete
metadata shows an unavailable message instead of a password form.
- Browsers with session storage disabled use the manual Cloud link,
since automatic retries cannot be bounded across documents.
- The external identity provider's email/social login was not completed
in this local test. Existing Cloud authentication owns that flow.
- No migration, credential format, membership rule, or production
deployment changes.

## Model Used

OpenAI GPT-6 through Codex. The exact served variant and context-window
limit are not exposed in this session. Used reasoning, repository tools,
code execution, and browser testing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:45:27 -05:00
DottaandPaperclip 2ed1540570 Merge current master and validate negotiated ACP capability envelopes
Synchronize with master 14795136f5 while retaining the recorded implementation base. Regenerate protocol contracts and preserve both input-presence and permission-scope checks.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 12:07:57 -05:00
DottaandPaperclip a3a71ebdb1 Bind Daytona candidate image identity to selected provider assets
Refresh the reviewed resolved dependency digest and fix ACP selector typing.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:59:18 -05:00
DottaandPaperclip 3447609d22 fix(runner): stream and page large API responses within capture budgets (#14301)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.

## Linked Issues or Issue Description

**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.

**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.

**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.

Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.

## What Changed

- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.

## Verification

- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.

## Risks

- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 10:36:32 -05:00
DottaandPaperclip cea8dda472 test: evaluate completion updates after native task handoffs (#13969)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users can delegate work through onboarding and Agent Chat.
> - A completed task does not prove that its result reached the original
conversation.
> - Existing tests do not isolate completion after the source chat
becomes idle.
> - This pull request adds four explicit native-runner probes across
Claude and Codex.
> - The probes preserve the result and reply so we can separate delivery
failures from inaccurate answers.

## Linked Issues or Issue Description

Refs #13775. Refs #13813.

These evals extend native-runner qualification. They measure completion
updates before we choose a product change.

## What Changed

- Add the opt-in `completion-updates` suite with two stories for each
native provider.
- Test completion in the existing onboarding task flow and after an
Agent Chat handoff becomes idle.
- Gate the chat worker on a brief inside its managed project workspace.
Prove the source is idle before releasing the worker.
- Check durable task completion, saved output, a subsequent source
reply, and rendered access to the result.
- Preserve replies, task state, screenshots, run events, and a separate
semantic review rubric.
- Add grader regression tests and update the documented eval contract.
- Preserve the suites added on master and include four completion cases
in the 306-cell catalog.

Production behavior and prompts are unchanged.

## Verification

- Passed all 565 eval support tests across 45 files after merging
current master: `node node_modules/vitest/vitest.mjs run --config
tests/runner-e2e/vitest.config.ts`.
- Passed eval TypeScript: `node node_modules/typescript/bin/tsc -p
tests/runner-e2e/tsconfig.json`.
- Confirmed four selected cells: `node cli/node_modules/tsx/dist/cli.mjs
tests/runner-e2e/launch.ts --list --suite completion-updates`.
- Four-cell behavior campaign on source
`ad47cf1da2b1e36f19f4227cfeb53998720b0b5b`:
https://github.com/paperclipai/paperclip/actions/runs/36072337485.
- A screenshot-only follow-up waits for the restored source reply to
render after result-link navigation. Its one-cell Claude onboarding
verification passed on final head:
https://github.com/paperclipai/paperclip/actions/runs/36075716141.
Corrected report:
https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36075716141-1/.
The original four-cell onboarding screenshots caught navigation loading;
its saved reply evidence remains valid. The follow-up again found stale
wording: "That work will run next" was posted 38 seconds after the child
was Done. The four-cell campaign keeps its original source and
measurements.
- Published evidence:
https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36072337485-1/.
- Suite definition:
`afba4d85d6c53d9f64c08b37a2e9cc20481b78f5bd7e2fa012045e2c69444d9d`,
version 6. Models: native `gpt-5.6-sol` and `claude-sonnet-5`, local
execution, one attempt per cell. All four cleanup checks passed.
Onboarding billing coverage is partial; reported zero cost must not be
read as a free run.

| Story | Automated delivery/access | Separate semantic review |
| --- | --- | --- |
| Codex onboarding | Pass | Pass: accurate completion reply with an
accessible result |
| Claude onboarding | Pass | Fail: reply says it will save the note once
the task runs, after the note is already saved and the task is Done |
| Codex idle chat handoff | Fail | Worker completed and saved the note;
no completion reply during the full observation window |
| Claude idle chat handoff | Fail | Worker completed and saved the note;
no completion reply during the full observation window |

Both chat cases positively recorded the source waiting and the worker at
the brief gate before release. Both saved outputs include the brief-only
start time. The opt-in campaign is red because it exposes current
behavior. It is not a required merge gate. The PR does not fix that
product behavior. Semantic review is a recorded human/agent assessment
of retained evidence; it is not an automated prose-quality judge.
- Second campaign:
https://github.com/paperclipai/paperclip/actions/runs/36071065098. Codex
chat reached the idle boundary and completed its task, then received no
completion reply during the full window. Claude onboarding again
returned a stale handoff answer. Claude chat exceeded the prior
110-second handoff setup budget; this revision raises that bounded setup
window to 180 seconds.
- Retained baseline:
https://github.com/paperclipai/paperclip/actions/runs/36069427676.
Onboarding passed delivery/access for both providers, but Claude gave a
stale handoff answer. Chat cases stopped at fixture problems; they do
not establish a completion-delivery failure. This revision fixes the
workspace path and competing reference requirements.
- On the previous head `4023a2a3c28d45c9eb2c42d452ce99ffba5c7b73`, 54 PR
checks passed and two were skipped, including typecheck, tests, and
build. Broad checks ran in CI, not locally. That head received Greptile
5/5 with no unresolved findings. The unchanged mobile
repository-settings browser test passed on one targeted retry after a
detached/disabled Save-button timeout.

- Merged current master in `9b4491e1f` and resolved the catalog-count
conflict. Eval support tests and eval TypeScript pass locally. All
individual CI jobs passed on this merge commit, including build,
typecheck, server tests, runner checks, and browser shards. The final
aggregate check also passed: 54 checks passed and two were skipped.
Greptile reviewed this exact commit at 5/5 with no unresolved findings.

## Risks

- These explicit probes can expose current product failures. They do not
change the default paid test selection.
- Mechanical delivery and result access do not establish answer
accuracy. The preserved reply still requires semantic review.
- A fixture failure before the idle boundary or worker completion cannot
establish a completion-update failure.
- The handoff setup window lasts three minutes. The worker brief wait is
bounded at four minutes. The observation window lasts two minutes after
worker completion. It retains later replies without erasing earlier
accessible delivery.

## Model Used

OpenAI Codex, GPT-6 (`gpt-6-astra`), with reasoning, repository
inspection, code execution, and GitHub tool use. The runtime does not
expose the context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 09:58:14 -05:00
DottaandPaperclip 890d11137f fix(ui): register artifact tabs without opening the panel (#14193)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Tasks keep agent outputs in the Artifacts tab.
> - An output can arrive while the user writes a message or reads a
document.
> - Opening the side panel on arrival interrupts that work, especially
on mobile.
> - This pull request adds the Artifacts tab without opening the panel
or changing the selected tab.
> - Users can open their outputs when they choose.

## Linked Issues or Issue Description

**What happened?**
New agent outputs opened the task side panel or mobile drawer. An
arrival could also replace the selected document or workspace file.
Existing outputs did not always register an Artifacts tab.

**Expected behavior**
Register one Artifacts tab for existing and new outputs. Keep a closed
panel closed. Preserve composer focus, the selected tab, and document or
file links.

**Steps to reproduce**
Open a task from the inbox. Close its side panel. Enter a message draft.
Create an agent output in that task. The panel must stay closed and the
draft must keep focus. Open the panel to see the Artifacts tab. Repeat
on a mobile viewport.

**Paperclip version or commit**
Base commit: 0f14d2612.

Related work: #11226 and #11551.

## What Changed

- Register existing outputs and later arrivals without opening the panel
or selecting Artifacts.
- Keep open documents, workspace-file links, and the tab launcher
unchanged.
- Handle each document deep-link request once so query refreshes
preserve later manual selection.
- Deduplicate attachment and work-product arrivals. Preserve dismissed
tabs across repeated refreshes.
- Add desktop and mobile browser regression tests. Update artifact
presentation documentation.

## Verification

- The closed-panel regression failed before the fix in unit and
real-browser tests.
- All 158 focused UI tests pass, including the original deep-link cases.
- UI typecheck and token gates pass on this branch. Full typecheck and
production build passed on the passive-arrival candidate before the
existing PR integration.
- The two local desktop/mobile browser cases pass against real
API-created artifacts.
- Desktop and mobile staging checks pass on the combined staging
candidate. Artifact arrival preserved a closed pane, draft text, and
composer focus. Explicitly opening the pane showed the Artifacts tab.
- All 56 current-head check contexts are successful or intentionally
skipped at `17ad904455b9378552f07a6f6e51402c6d164688`, including full
typecheck, test shards, build, and browser suites. Greptile is 5/5 on
that commit with no unresolved review threads.
- The interrupted local broad validation was resumed; the remaining
serialized 64 files and 1,035 tests pass. Local database startup
failures passed after stale test resources were released.

## Risks

- Outputs no longer reveal the panel automatically. Users open the panel
to view them.
- The document request guard must still allow a new explicit deep link.
The regression tests cover this case.
- There are no API or database changes.

## Model Used

- OpenAI Codex, GPT-6, with code editing, shell tools, GitHub tools, and
browser verification. The exact serving model ID and context-window size
are not exposed in this session. Earlier implementation model metadata
is not available.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 08:38:00 -05:00
DottaandPaperclip f2ed0b65c4 fix(runner): enable API tools by default (#14186)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner exposes tools for company tasks.
> - API search and call tools cover operations without a dedicated tool.
> - The current default hides these tools unless an operator sets an
environment variable.
> - This pull request enables the tools when that variable is absent.
> - Operators can still disable the tools or restrict them to selected
companies.

## Linked Issues or Issue Description

Refs #13003, which added the guarded API tools.

**What happened?**
The native runner does not advertise `search_api` or `call_api` with the
default server configuration.

**Expected behavior**
The tools are available without a special environment variable. Existing
authorization checks still apply.

**Steps to reproduce**
Remove `PAPERCLIP_RUNNER_API_TOOLS_ENABLED` and
`PAPERCLIP_RUNNER_API_TOOLS_COMPANY_IDS`. Create a normal runner
authority. Inspect its tool definitions.

**Paperclip version or commit**
a6c4e7a.

## What Changed

- Enable API tools when the server flag is absent.
- Keep explicit disable, invalid-value rejection, company restrictions,
and binding restrictions.
- Test default tool definitions and run HTTP integration tests without
the enabling flag.
- Update operator and hiring documentation. The existing shared gate
also controls `hire_agent`.
- Use the same policy in the E2E evidence summary so an unset flag is
not reported as disabled.
- Isolate default-availability tests from operator environment
variables.

## Verification

- Red test: two new rollout policy assertions failed before the fix.
- Focused policy, authority, and HTTP tests: 3 files passed; 41 tests
passed, 2 skipped. The two runnerd transport cases require a Rust-built
binary absent from this workspace.
- Command: `pnpm exec vitest run
server/src/services/native-runtime/runner-api-rollout.test.ts
server/src/services/native-runtime/paperclip-runner-tool-authority.test.ts
server/src/services/native-runtime/runner-api.integration.test.ts`.
- Attempted `pnpm -r typecheck`: blocked by missing `cargo` in this
workspace.
- Attempted `pnpm build`: terminated at the 4 GiB memory limit.
- Attempted `pnpm test:run`: stopped after memory pressure to run
focused tests alone.
- Repeated the 41 passing focused tests with an inherited disabled flag
and a foreign-company restriction; test isolation passed.
- `pnpm test:e2e:runner:unit`: 44 files and 543 tests passed.
- `pnpm test:e2e:runner:typecheck` exceeded the workspace memory limit,
including a retry with bounded Go memory settings.
- CI results will be recorded before handoff.

## Risks

- More native runs can discover API tools and the existing `hire_agent`
tool by default.
- The change does not remove company, run, mode, credential, lifecycle,
or approval checks.
- Explicit operator restrictions still take precedence. No database
migration is required.

## Model Used

OpenAI Codex agent. The runtime does not expose the exact model ID or
context-window size. Used reasoning, repository tools, shell execution,
and tests.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-26 21:16:33 -05:00
DottaandPaperclip 96bf004a79 fix: use persisted state for lifecycle continuation and retry budgets (#13888)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its control plane decides when a task can continue, wait, stop, or
complete.
> - Legacy continuation could change when an agent changed its wording
without changing task state.
> - Shared attempt counts also let repair and infrastructure retries
affect each other's limits.
> - This pull request uses persisted state and separate, bounded
allowances for these decisions.
> - If automatic repair stops, the task explains what happened and
offers a guarded retry.
> - Paired tests and real-provider evaluations verify that Stop,
approvals, ownership, and spending limits remain authoritative.

## Linked Issues or Issue Description

Related work: Refs #13761, Refs #11126, Refs #13610. These cover
obsolete continuation dispatch and retry storms. Open and closed issues
and PRs were searched for related lifecycle, continuation, and retry
work.

**What happened?**
Legacy continuation depended on English wording and progress heuristics.
Repair, failure retry, and productive continuation could consume shared
counts. When bounded repair stopped, the task showed a technical
recovery message without a clear next action.

**Expected behavior**
Persisted disposition and owned execution paths determine the next
action. Missing disposition prompts bounded agent repair. Explicit work
mode determines planning mode. Narrative changes and raw activity counts
cannot replenish allowances. An exhausted repair shows a readable
notice. An explicit retry checks current controls and preserves the
assigned agent.

**Steps to reproduce**
Run `pnpm test:lifecycle-baseline`. The paired probes keep structured
state constant while varying completion, planning, blocker, and progress
prose. Run the explicit `lifecycle-baseline` and
`continuation-accounting` Product E2E suites for real-provider coverage.
In Storybook, open **Design previews / Recovery notice** to inspect the
production component's normal, pending, acknowledged, unavailable,
failure, and mobile states.

## What Changed

- Hide the image attachment button, icon, and drop/paste hint in answer
composers. Image paste and drop support remains available.
- Merge current master and retain both browser regression sets. Use a
production-stamped service worker in the offline recovery browser
fixture.
- Share one state-based legacy continuation decision across immediate,
delayed, and recovered dispatch. Bind bounded repairs to their source
run and episode.
- Remove title and description wording from work-mode authority. Agents
can still write requested plans in execution mode.
- Persist separate failure-retry and productive-continuation counters.
Disposition repair and resource waits cannot consume or reset those
allowances.
- Validate delayed repair identity, then recheck current gates before
provider dispatch. Fence native startup cancellation.
- Show **Agent needs attention**, a plain-language explanation, **Retry
agent**, and expandable details in both task interfaces. Report request
progress, acknowledgement, and errors inline.
- Store typed recovery notice metadata. Recognize older active notices
only through exact stored action and run IDs. Notice text never grants
retry authority.
- Use the existing recovery-action endpoint for retry. Recheck current
action, status, owner, agent availability, dependencies, active runs,
pending questions and confirmations, approvals, pause controls, and
budget. Duplicate requests do not wake twice.
- Add component, page, route, database, contract, and Storybook
coverage. Keep the scenario inventory and executable evals here.
Historical reports and snapshots live in the [commit-pinned
paperclip-evals
archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md).
- Preserve unsaved project fields while the same project URL changes to
its canonical alias. Do not reuse data across projects or companies.
This separate fix addresses the repeated repository-editor browser
failure without changing the browser test.
- Keep the development service worker from intercepting Vite module
reloads. Update the connection-intent browser fixture to record progress
and completion through the agent API.

## Verification

Merge preparation on September 25, commit
`c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`:

- Merged master `bd2030932` and resolved the browser test-list conflict
by keeping both sets of regressions.
- Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no
failures, skips, or missing selected evidence. Unit 423, runner 184,
database integration 397, grading 86.
- Browser support: 17/17 passed. The offline recovery test first failed
with an unstamped development worker, then passed with the production
stamp. Its assertions are unchanged.
- Focused interaction UI and offline fallback tests: 19/19 passed.
Verified the custom-answer composer in Storybook: no attachment controls
or hint; entering an answer enables Next.
- Recursive typecheck, production build, token gates, and diff checks
passed. The worktree is clean. No new real-provider campaign was run.
- Current CI and review: [Current PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243):
55 successful checks and two optional Storybook skips. Greptile scored
this exact commit 5/5. Hiding the question attachment controls is an
intentional UI change; paste/drop remains available.

Earlier recovery UI verification, commit
`21be0fec0e90e86b6d662b8ee4831847cd041cdb`:

- Recursive typecheck, production build, token gates, and diff checks
passed.
- Focused UI coverage: 338 tests passed across six suites (336 before
the interaction guard, with the two affected suites rerun at 149 passed
after it). Covers both task interfaces, the real page mutation,
pending/error acknowledgement, stale state, and unavailable controls.
- Recovery database integration: 352 tests passed before the interaction
guard. The complete recovery-action and mutation-route suites passed 181
tests after it. The two new pending question/confirmation regressions
failed before the fix and passed afterward, including
resolved-interaction controls. Shared validator suite: 31 passed. E2E
catalog suites: 34 passed.
- Browser inspection passed for light/dark themes, mobile layout,
expandable details, pending retry, acknowledgement, failure, and
disabled retry. Storybook renders the production component; its request
is simulated.
- The broad local run hit two chat callback-order wait failures and was
stopped after all CI unit/database/runner shards passed. Both local
failures passed when rerun without the competing full-suite process.
- CI exposed a repeated project-repository draft-loss race during
canonical redirects. A new unit regression failed before the fix; all
nine project-page tests now pass, including controls for other projects
and companies. Both unchanged repository browser tests passed against a
fresh local server. UI typecheck, production UI build, and token gates
passed after this fix.
- [Earlier PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798)
on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two
optional Storybook skips, and no failed or pending checks. The
repository browser shard passed with the production fix. Greptile is 5/5
on this exact commit with no unresolved review threads. The PR is
mergeable.

Historical, source-qualified lifecycle evidence:

- Lifecycle baseline: 1,074 assertions. Native session coverage: 447
tests. Product E2E support: 515 tests. Browser support: 11 tests. Full
earlier verification is retained in the archive.
- [Real-provider campaign: 8/8 passed, zero
retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html),
source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup
checks passed. This includes deliberately exhausted repair cases that
correctly remain blocked; it does not mean every task finished Done.
This campaign predates the recovery UI change.
- Archive migration verified all 16 original JSON files byte-for-byte
and all 24 checksum entries. App tests do not need private archive
access. [Archive PR
#27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged.

## Risks

- Agents that omit durable disposition receive at most two repair
attempts by default. Prose-only completion exposes missing state rather
than silently changing scheduling.
- A retry is an explicit board action. The server rechecks current
controls. A successful response confirms the task returned to To do; it
does not claim that the provider has already started.
- Existing notice metadata remains valid. Only older active notices with
matching structured evidence receive the new UI. Historical notices
without that evidence keep their existing rendering. No schema migration
is required.
- Old run records require conservative retry accounting. Tests cover old
counters, alternating retry lanes, restarts, and exhausted repairs.
- Historical snapshots require private `paperclip-evals` access. The app
index retains public campaign links. Live campaigns qualify specific
sources and scenarios; no new real-provider campaign has run for the
recovery UI commit.

> This fixes existing lifecycle and recovery behavior and does not
duplicate planned core work.

## Model Used

OpenAI GPT-6 through Codex assisted implementation, reasoning, code
execution, and review. The exact serving model ID and context window are
not exposed in this task. Historical real-provider evaluations used
Codex model `gpt-5.6-sol`, separately from the implementation assistant.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 15:28:11 -07:00
DottaandPaperclip bd20309323 fix(ui): keep mobile task pickers above the keyboard (#14022)
## Thinking Path

> - Paperclip helps people manage AI agents and their tasks.
> - The New Task dialog lets a person select an assignee and a project.
> - A mobile browser reduces the visible viewport when the software
keyboard opens.
> - The dialog accepted invalid viewport data, and the pickers stayed
inside a transformed container.
> - This behavior could collapse the dialog or move a picker search
field above the visible area.
> - This pull request validates viewport data and puts mobile pickers in
the visible viewport.
> - The benefit is that a mobile user can see and use each picker while
the keyboard is open.

## Linked Issues or Issue Description

**What happened?**

On a mobile device, the Assignee and Project pickers in the New Task
dialog could move above the visible viewport. A short invalid viewport
value could also collapse the dialog to a line.

**Expected behavior**

The dialog and each open picker must stay in the visible viewport while
the software keyboard is open.

**Steps to reproduce**

1. Open the New Task dialog in a mobile browser.
2. Open the Assignee picker or the Project picker.
3. Focus the picker search field so that the software keyboard opens.
4. Observe that the picker can move above the visible viewport.

**Paperclip version or commit**

`efce9356b5`

**Deployment mode**

Local dev and hosted browser UI.

## What Changed

- Ignore zero, negative, and non-finite Visual Viewport measurements.
- Keep the last safe dialog geometry until the browser gives a valid
measurement.
- Put entity pickers outside the transformed dialog container.
- Size and position the mobile picker from the dialog Visual Viewport
values.
- Add unit tests for invalid viewport recovery.
- Add Chromium tests for the Assignee and Project picker states.

## Verification

- `pnpm exec vitest run ui/src/components/NewIssueDialog.test.tsx`
passed with 34 tests.
- The Chromium viewport test passed 35 of 35 runs with five repeats and
no retries.
- `pnpm --filter @paperclipai/ui typecheck` passed.
- `pnpm check:token-gates` passed.
- `pnpm build-storybook` passed.
- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- A full local Vitest run reached restricted workspace-runtime tests
that require sibling worktree and runtime writes. Remote CI will run the
supported test environment.

## Risks

- Risk is low because the new picker layout applies only to mobile
widths.
- The layout depends on Visual Viewport data when the browser supplies
valid values.
- Unit and browser tests cover invalid data, mobile pickers, tablet
layout, and desktop layout.

> This change fixes a focused UI bug. It does not duplicate a core
feature in `ROADMAP.md`.

## Model Used

- OpenAI Codex with GPT-5. The agent used reasoning, repository tools,
code execution, and browser automation. The runtime did not expose the
context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 08:37:26 -07:00
DottaandPaperclip f56aee5423 fix(evals): capture complete durable run event streams (#13977)
## Thinking Path

> - Paperclip manages AI agents and records their durable work outcomes.
> - Product E2E checks those outcomes through the browser and public
API.
> - A successful long run can emit more than 1,000 durable events.
> - The harness read one page and missed the later completion evidence.
> - This pull request reads every page before it checks runtime
invariants.
> - Invalid or incomplete capture still fails. A missing page cannot
produce a pass.

## Linked Issues or Issue Description

Related: #13882 (Grok qualification). A duplicate search found no
existing event-pagination fix.

**What happened?**

The local structured-question case in [campaign
36071063537](https://github.com/paperclipai/paperclip/actions/runs/36071063537)
completed the task and passed its six outcome matchers. It failed native
runtime invariants because the capture contained exactly 1,000 events.
The last captured event preceded the run's completion by more than a
minute. The API caps each response at 1,000 rows. The harness did not
request the next page.

**Expected behavior**

Read the complete durable event stream through the public API before
checking semantic-result and terminal-event counts. Reject incomplete or
malformed evidence.

**Steps to reproduce**

1. Complete a native task that emits more than 1,000 durable events.
2. Place the semantic-result and terminal events after row 1,000.
3. Capture the run with the Product E2E harness.
4. Before this fix, the invariant checker sees only the first page.

**Paperclip version or commit**

Observed at `4196a4cd76db434854b679035e4146c7f69689ce`. The same
single-page capture exists on master. The original failed result remains
unchanged; missing historical tail evidence is not reconstructed or
graded as a pass.

## What Changed

- Add a bounded event collector that advances through the public
`afterSeq` cursor.
- Use it for task success/failure evidence and shared chat run evidence.
- Reject invalid pages, missing or non-increasing sequence numbers,
repeated cursors, failed later requests, and an exhausted page limit.
- Test completion events beyond the first page, exact page boundaries,
and malformed evidence.
- Correct the existing Everyday catalog test from 38 to the maintained
47 cells. The suite stays explicit-only.
- Document the complete-capture requirement and its bound.

## Verification

- Product harness typecheck passes.
- The full credential-free harness suite passed 477 tests. After adding
the chat integration regression, all 34 chat evidence tests pass.
- Seventeen pagination tests cover the valid tail and malformed-evidence
cases.
- `git diff --check` passes.
- [Full repository
CI](https://github.com/paperclipai/paperclip/actions/runs/36080683422)
passes at `aed6f79c089226be79e75dcf390969561ec4f787`: 52 successful
checks and two intentional skips, including Rust, typecheck, build,
server tests, and browser shards. Greptile gives this exact head 5/5
with no findings. No local Docker, Rust build, browser suite, or model
invocation was used for this change.
- [Live Grok
requalification](https://github.com/paperclipai/paperclip/actions/runs/36080870743)
is running on combined source
`1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`, with the runner, scheduler,
and evidence fixes. Its full credential-free harness passes all 493
tests and typecheck. Live results are pending; the original campaign
remains a failed measurement.

## Risks

Long runs need more read-only API requests and larger private evidence
files. Capture stops with an explicit error after 100 full pages. This
changes neither production APIs nor provider behavior. It does not relax
an invariant or change a historical grade.

## Model Used

OpenAI GPT-6 through Codex, with repository tools and code execution.
The exact serving identifier and context-window size are not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 10:06:33 -05:00
DottaandPaperclip efce9356b5 fix(ui): offer recovery when the app fails before React starts (#13970)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The browser must load its JavaScript before React can render a task.
> - A failed import can stop that process before the React error
boundary exists.
> - The HTML entry then leaves an empty page with no recovery action.
> - This PR adds a small recovery screen that works without React.
> - The user can retry the same page and return to saved task content.

## Linked Issues or Issue Description

Refs #13824 and #13895. This is a follow-up to their browser startup
investigation.

**What happened?**

Interrupting the app bundle or a required import leaves an empty React
root. A startup exception has the same effect. A React error boundary
cannot handle these failures because React has not started.

**Expected behavior**

The page must explain the startup failure and offer a manual retry. A
late successful load must dismiss the recovery message without a reload.

**Steps to reproduce**

1. Open a saved task in the browser.
2. Abort the application bundle request, or make a required module
return HTTP 503.
3. Observe the empty page before this change. With this change, use
Reload page after the fault clears and verify the saved task and
comment.

**Paperclip version or commit**

The failing regression baseline used master at `8781f06a8`.

**Deployment mode**

Local source build and compiled UI. Tests cover both initial navigation
and a page controlled by the production service worker.

The exact cause of the older intermittent Vite stall remains
unconfirmed. Forty app loads and thirty replays of retained responses
did not reproduce it. This PR fixes the missing recovery path; it does
not claim to remove that historical cause. A normal HTTP 304 response is
not a failure.

## What Changed

- Add an inline startup guard and recovery screen in the HTML entry. It
does not depend on the app module graph.
- Show a manual reload action after a startup error or after 30 seconds
without rendered root content.
- Remove the notice, timer, observer, and error listeners when the app
starts. Never reload automatically.
- Keep the recovery screen outside the React root so it cannot satisfy
app-readiness checks.
- Add browser tests for interrupted imports, a stalled import, an
evaluation error, service-worker-controlled retry, repeated offline
retry, and cleanup after successful startup.
- Return a static, uncached HTML retry screen when a
service-worker-controlled navigation fails offline. It contains no task
content.
- Add a full-app test that retries an interrupted compiled bundle and
checks the saved task, comment, composer, route, and absence of agent
runs.
- Document the coverage and the limits of the historical diagnosis.

## Verification

- Red baseline: four recovery cases failed; the normal-startup case
passed. After the change, all five recovery cases passed. The review
found an offline retry gap; that additional case failed before the
worker fix and passed afterward.
- Full provider-free browser-support suite: 16 passed.
- Compiled-app browser tests: four passed, including saved-task reload,
interrupted-bundle recovery, slow-CPU service-worker reload, and sidebar
navigation.
- Expanded service-worker, offline response, PWA, and worker build-ID
unit tests: 37 passed. The two old plain-text offline expectations were
reproduced as failures and updated for the HTML retry contract.
- UI production build, full local repository typecheck (`pnpm -r
typecheck`), runner-E2E typecheck, and design token checks passed.
- Manual browser check: a temporary server failed the compiled bundle
once. The recovery screen appeared. Clicking Reload page restored the
same saved task, comment, and composer.
- Full local `pnpm build` passed.
- Full local `pnpm test:run` was attempted with a bounded deadline and
stopped after it timed out. Workspace runtime/cleanup tests reported
timeouts on this host. The monolithic local run is not a pass. The
focused tests above and the complete Linux CI run provide the successful
verification.
- Final-head [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36072201966)
passed. All 53 check runs succeeded; the two Storybook jobs were
intentionally skipped. The legacy security status also passed.
- Greptile reviewed `f83e0f51fb760541d83353f2c1df4e182f3948f9`: 5/5.
Both review findings are fixed and resolved.

## Risks

- The guard only handles startup before React renders root content.
Existing React boundaries handle later rendering errors.
- A slow startup can show the message after 30 seconds. A later
successful render removes it; the page does not reload by itself.
- The fallback uses native HTML when the app stylesheet is unavailable.
- The worker changes only its offline navigation response. It returns
static HTML with a reload button and `Cache-Control: no-store`. Its
cache allowlist, private-response protections, task state, provider
prompts, and grading rules stay unchanged.
- This does not establish or fix the unknown cause of the historical
intermittent Vite stall.

## Model Used

OpenAI GPT-6 through Codex. The session exposes the GPT-6 family but not
an exact served model ID or context window size. Used reasoning, code
editing, shell tools, and browser testing. No subagents were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 19:36:45 -05:00
DottaandPaperclip 8781f06a87 feat(connections): enable MCP aggregators by default (#13964)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - Connections let those agents use external services with explicit
access rules.
> - Zapier, Arcade, Composio Connect, and Executor already have setup
and runtime support.
> - Their experimental switch still blocks discovery and setup by
default.
> - This pull request removes those gates and the Settings toggle.
> - Users can connect these providers without enabling an experiment.

## Linked Issues or Issue Description

Refs #13755. Refs #13941.

**What existing behavior does this improve?**
Apps browsing, inline setup, and agent connection search for the four
MCP aggregators.

**Current behavior**
An instance must enable the MCP aggregators experiment before users or
agents can start setup.

**Proposed behavior**
All four providers are available by default on local and managed
instances. Old stored and managed values still parse but cannot disable
them.

## What Changed

- Remove the aggregator gates from Apps, inline setup, server setup, and
agent search.
- Remove the Settings toggle and its UI hook.
- Retain the old setting key only for upgrade compatibility. Normalize
it to true and ignore managed overrides, as Apps already does.
- Replace opt-in fixtures with default-on coverage. Test old false
values, all four setup flows, provider choice, and the removed toggle.
- Update current connector guidance and remove the opt-in from the
runner acceptance fixture.

## Verification

- 306 focused tests passed across eight files: shared remote MCP
contracts; server remote MCP lifecycle, aggregator fallback, settings
normalization, and managed overlay; UI Apps browsing, setup, and
experimental settings.
- Server and UI TypeScript checks passed.
- UI token gates and `git diff --check` passed.
- The full local suite was not run, per the maintainer's instruction.
All 54 CI checks passed; two checks were skipped. One unrelated
workspace-preview readiness timeout passed on one failed-shard retry.
- The setup fixtures use simulated MCP responses. This change does not
claim new live provider acceptance.

## Risks

- Existing instances now show all four providers, even if the old flag
was false. This is intentional.
- External provider choice, credentials, company isolation, agent
grants, and tool policies still apply. Showing a connector does not
authorize an external account.
- No data migration is required. The compatibility key keeps old managed
configuration documents valid.
- Historical Zapier live acceptance remains incomplete in the existing
evidence report. The maintainer explicitly requested the default-on
rollout for all four existing providers; the report records that scoped
exception.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, repository
tools, and test execution. The context window size is not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 17:31:32 -05:00
DottaandPaperclip aa8fc86331 feat(connections): prefer native apps and ask users to choose external providers (#13941)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents governed access to external services.
> - Native connections should remain the first choice for a supported
app.
> - Other apps may be available through an external MCP provider.
> - The user must know which external provider handles the connection
and choose it before setup.
> - This pull request adds ranked alternatives and server-authored
instructions to connection search.
> - Agents can follow the returned instructions while Paperclip
validates saved choices and access.

## Linked Issues or Issue Description

Related: #13879, which fixed inline MCP provider setup. This PR adds
discovery and provider selection on top of that work.

**Subsystem affected**

Cross-cutting: shared connection contracts, server search and intent
services, native runtime, CLI, inline setup UI, and evals.

**Problem or motivation**

An agent cannot offer a clear external-provider choice when Paperclip
has no native connection for an app. Adding provider-specific branches
to the core prompt would make those instructions harder to maintain.

**Proposed solution**

Prefer a native connection. Otherwise return verified alternatives in
Composio, Arcade, Executor, Zapier order. Include an external-service
disclosure, a question with None, and the next instruction in the search
result. Validate the saved human choice before creating a selected
fallback setup card. Reuse existing provider accounts and verify
underlying app access separately.

**Alternatives considered**

Do not silently choose a provider. Do not claim that broad execution
tools prove support for every app. Reuse existing questions and
connection intents rather than add another connection model.

**Roadmap alignment**

Extends the existing MCP Tool Gateway & Apps and Agent evals & feedback
capabilities. The MCP aggregators experiment remains the gate. No
duplicate provider-routing PR was found in the public search.

## What Changed

- Add a dated support index and authorized cached-tool evidence for
external routes.
- Return provider questions and next-step instructions from
`connections_search`.
- Preserve pending choices and declines across continuation. Validate
company, task, agent, human, app, and current route eligibility.
- Carry the selected app into new setup and account reuse, validate
explicit provider requests against persisted human messages, and
distinguish provider readiness from app authorization.
- Sync native, MCP, REST, and CLI contracts. Keep core agent
instructions provider-neutral.
- Add production-component Storybooks, focused database tests, and three
real-agent browser eval cases.
- Record the plan, observed failures, fixes, passing evidence, and
acceptance limits.

## Verification

- Latest head `586f0e6cd`: 54 checks passed, 2 skipped; Greptile 5/5 and
all review threads resolved.

- After rebasing on master `18dac1e1e`: 64 focused shared, validator,
route-contract, and database tests passed; server typecheck passed.
- Embedded-browser test drive on the rebased head: native Jira card,
HubSpot external-provider question, Arcade account reuse, one actual MCP
read against a local synthetic fixture, reload persistence, and None
preventing further calls. A real OpenAI-backed agent performed discovery
and continuation.
- UX observation: the agent initially combined mutually exclusive
request fields; the server rejected it and the agent recovered without
changing access. This extra retry remains visible in the transcript.
- After rebase: 23 focused eval grader/catalog tests, affected
TypeScript checks, token gates, production UI build, and Storybook build
passed. Full local tests are intentionally excluded at the maintainer's
request.
- Before rebase: four browser/real-agent attempts passed: native Jira,
None, and reuse of the second provider on two Codex profiles.
- Browser evals used an isolated deterministic MCP fixture through the
real Paperclip gateway. They do not prove production compatibility with
all four providers.
- Review `Apps / Connections / Provider choice` in Storybook. Choose
Arcade, continue through Access, and verify the app name,
external-service disclosure, and URL configuration.
- The detailed verification report is
`doc/connections/2026-09-23-aggregator-routing-verification.md`.

## Risks

- The public support index is finite and can age. Account capability and
app authorization still require verification after selection.
- Existing installed-tool permissions remain in effect. Provider choice
is not a new execution permission boundary. An early Mini attempt
skipped search; clearer provider-neutral instructions made the targeted
rerun pass. This is not a measured reliability rate.
- Explicit requests skip provider confirmation only when a clear
persisted human message or saved provider choice supports them. Other
phrasing falls back to confirmation; the agent query alone is not
consent. These routes do not add tool permissions.
- No database migration or legacy Composio broker is added.
Real-provider acceptance remains separate from fixture proof.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser tools. The exact deployment variant and
context-window size are not exposed in this session. Product evals
separately used the repository's primary Codex and Codex Mini profiles;
those agents supplied test behavior, not independent provider
compatibility proof.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 16:36:49 -05:00
DottaandPaperclip b0155a681a feat(slack): connect Paperclip conversations and scheduled messages (#13920)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Slack conversations use the same tasks and agents as the Paperclip
board.
> - A board reply must reach that Slack conversation and let the agent
continue the work.
> - An assigned agent also needs its Slack tools during normal tasks and
scheduled routines.
> - Both paths must keep the linked user's authority, delivery rules,
and conversation history.
> - This pull request adds those paths and reduces setup friction for
Slack bots.

## Linked Issues or Issue Description

**Subsystem affected**

Server orchestration, Slack connector tools, shared contracts, and chat
setup UI.

**Problem or motivation**

Replies entered in Paperclip did not provide a complete round trip to
the linked Slack thread. Slack and board wakeups could select different
model sessions for the same task. Agents also lacked their assigned
Slack tools outside Slack-origin work, which prevented a routine from
sending its responsible user a briefing. Inviting a bot could leave the
new channel disabled.

**Proposed solution**

Mirror human board messages with author attribution and route the agent
result to the same thread. Use the same session key across both entry
points. Supply Slack tools to the connection's assigned agent in normal
tasks and routines, using the current responsible user's verified link.
Enable newly invited channels while preserving explicit disabled
choices. Add a browser-agent setup prompt to the Slack wizard.

**Alternatives considered**

A separate Slack scheduler or task dispatcher would duplicate existing
Paperclip workflows. Reusing the connection owner's identity would grant
the wrong authority. Replaying old channel history could start
unintended work. This change uses ordinary task wakeups, routine
dispatch, and Slack's original invitation mention event instead.

**Roadmap alignment**

Extends the shipped Scheduled Routines and governed Apps capabilities.
It does not add a separate task lifecycle. Related work: #13828 and
#13809. Related test stabilization: #13877. The existing plugin
Slack-control proposals are separate from this built-in connector
change.

## What Changed

- Queue human Paperclip messages for the original Slack thread with
display-name attribution and stable delivery identities. Require the
author’s current linked Slack identity and recheck access before
delivering messages or agent replies.
- Apply pause, dependency, cancellation, and closed-workspace guards
before explicit Board sends request work and again when the durable
outbox dispatches it.
- Route agent results back to Slack and preserve model-session
continuity, including replies that reopen completed tasks.
- Resolve assigned Slack connections for normal agent tasks and
routines. Recheck the responsible user's link, membership, and
permissions at execution.
- Add `slack_open_dm` for the responsible user's bot DM and request the
`im:write` scope.
- Enable newly discovered invited channels. Keep explicit OFF choices
and normal admission and deduplication rules.
- Add a copyable Slack setup prompt for a computer-use agent, with
Storybook coverage. Share the prompt-button component with GitHub.
- Update Slack tool documentation and runtime instructions.
- Stabilize the mobile project browser test by waiting for the final
canonical route before editing, preserving all persistence assertions.

## Verification

- Live staging: invited the bot after the first mention. The channel
became enabled and the bot answered that original mention.
- Live staging: a normal Paperclip reply appeared in Slack with author
attribution. The agent completed the calculation and replied once in the
original thread and in Paperclip.
- Live staging: a codeword entered in Slack was recalled from Paperclip.
A following Slack calculation used the result from the Paperclip turn.
Run metadata confirmed the same model session for both entry points.
- Live staging: a scheduled routine used `slack_open_dm` and
`slack_post_message` to deliver one DM. The existing app was reinstalled
with `im:write`. The test routine was paused after verification.
- Before the master merge: 397 focused feature tests passed. The
continuity fix passed all 76 issue comment/update route tests and six
focused route/integration cases. Typecheck, build, and token gates
passed.
- Review fixes: 47 focused integration cases passed, covering link
revocation/replacement, private membership removal, guarded outbox
dispatch, concurrent workers, lost scheduler responses, a real
one-connection pool, exact reply provenance, and attachment retries. All
99 issue-comment route tests and the Slack catalog browser test passed.
- Full local typecheck, production build, token gates, and
module-boundary checks passed. The full local test command passed 25,915
tests before a 15-second timeout in
`issue-thread-interaction-routes.test.ts`; that entire suite passed on
isolated rerun (81 tests). Remaining serialized coverage is provided by
the current-head CI shards.
- An unchanged Cursor adapter test hit its 10-second limit in CI; all
five tests in that file passed on a local rerun in 3.11 seconds, and the
failed CI shard passed on its single retry.
- The preview-server readiness test passed a local rerun (28 tests). The
mobile-project readiness fix passed three repetitions of both browser
tests (6/6).
- Final commit `64ac0d9897f4353375996f1b1b38e5040bdeb0a0`: all CI gates
passed, including all eight browser shards, all server/chat suites,
typecheck, build, runner checks, and security checks. Greptile reviewed
this exact commit at 5/5; all review threads are resolved.

## Risks

- Human messages on a Slack-linked task now publish to its Slack thread.
The task banner states this behavior. Incoming Slack messages and
internal agent bookkeeping must not echo back.
- Normal tasks and routines can now use the assigned bot. Authority
remains bound to the current responsible user's link; it does not fall
back to the connection owner. Revocation, private-context limits, and
queued-write checks still apply.
- Existing Slack apps need `im:write` and a reinstall to open DMs. Other
existing capabilities remain available without that scope.
- New invited channels default to enabled. Explicit disabled choices
remain disabled. Channels created by bot tools still require a person to
enable responses.
- No database migration or new provider credentials are required.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, GitHub CLI,
and browser tools. The runtime does not expose a more specific model
build identifier or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 10:51:45 -05:00
Devin FoleyandPaperclip 6681c71b40 fix(ci): stabilize chat startup and close the initial live-update gap (#13895)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Browser tests check chat state across navigation, reload, and agent
runs.
> - Runtime tests check that a service is ready before Paperclip
publishes its address.
> - CI for #13891 failed in these paths, then passed attempt 3 with the
same code.
> - The browser traces stopped during development asset startup. A
separate sidebar assertion used an unstable focus path through the rich
editor.
> - This PR gives browser tests fresh built assets and a direct keyboard
path to the star button. It adds service-worker reload coverage under
CPU throttling.
> - A later browser failure exposed a real reload race: comments can
change between the first query and the first live subscription. The UI
now refreshes active queries when that subscription opens.
> - Runtime fixtures now have separate registry state and better failure
evidence. The readiness deadline remains unchanged.
> - A later CI run exposed a wall-clock backoff assertion and three
authorization cases sharing one test lifecycle. The PR anchors the
assertion to transport time and separates the cases.

## Linked Issues or Issue Description

Refs #13891. Related runtime ownership and cleanup work: #11791, #11389,
#11278.

Evidence: [original CI run, attempt
3](https://github.com/paperclipai/paperclip/actions/runs/35910335089/attempts/3).
Attempt 3 passed both affected shards without code changes. This PR is
separate from the wake-payload change, which has since merged.

| Failure | Diagnosis and classification |
| --- | --- |
| Sidebar blank page and retry text missing after reload in CI | Both
traces show a blank document before React startup. Vite connects, but
the failing page makes no application API requests. The service worker
forwards the unbundled development module graph. The retry response is
already stored and its process-adapter run succeeded before reload. This
places the failure in browser bootstrap, not wake payload or reply
persistence. The exact reason the development module graph stopped is
not established by the retained trace. The harness now serves built
assets, and the new test covers startup and controlled reload at 4x CPU
throttling. |
| Star opacity remains zero on macOS | Reproduced on unchanged master
`f55759942b`. The old test clicked the rich editor, focused the star,
then used Tab and Shift+Tab. An instrumented baseline run captured the
sequence: Tab moved from the star to the next sidebar link, then the
editor bundle called `focus()` on its contenteditable before Shift+Tab.
That key reached the composer, where it is a work-mode shortcut. This
confirms a pending editor selection update stole focus; it was not the
browser skipping sidebar buttons. The assertion did not prove the star
retained focus. The test now moves the pointer away, focuses the
preceding sidebar link, presses Tab once, and asserts both actual focus
and opacity. This is test synchronization and keyboard traversal, not a
demonstrated CSS defect. |
| Runtime readiness exceeds 10 seconds | The old error only says `fetch
failed`. There is no child startup output in that failure, so it cannot
distinguish slow process startup from a refused or stalled probe. It
passed unchanged locally and took 3.416 seconds in attempt 3. Resource
contention is plausible but unproved. This PR does not claim a proven
historical runtime root cause: it isolates fixture registry/log files,
checks ports before spawn and after stop, checks live backends at
publication, and preserves transport errors, probe count, elapsed time,
and fixture startup timestamps for the next occurrence. |
| Later CI: recovery reply disappears after reload | Product
synchronization bug, distinct from the blank-page bootstrap failure.
Reproduced locally with a trace: the comments request started at
`21:56:07.443`, the server saved the reply at `.520`, and the first live
subscription started at `.541`. The reload fetched a successful run and
its complete log, but missed the comment event. The provider refreshed
after reconnects only. It now refreshes active queries on the first
connection too. A deterministic regression test fails before the fix. No
browser assertion was changed. |
| Later CI: Slack backoff and authorization tests | The 30-second
backoff assertion required more than 25 seconds to remain when it read
the saved action. CI spent 10.311 seconds in the test, exceeding that
five-second allowance. A local six-second read delay reproduces the
failure; the new transport-anchored lower and upper bounds pass the same
fault injection. The neighbouring authorization test ran three
independent fixtures in one test and timed out at 15 seconds. Each case
now has its own fixture cleanup and the normal per-test deadline, so
earlier cases do not remain active during later global worker sweeps. No
specific production slow call was established. |
| Self-hosted runner loses communication or shuts down | The original
lost-communication failure has no assertion. On final-head [attempt
1](https://github.com/paperclipai/paperclip/actions/runs/35925901613/attempts/1),
Build, Runner Vitest 1/2, and chat 2/3 ran on three separate fleet
instances. All received a runner shutdown signal at `22:04:55 UTC`,
within 35 milliseconds, then cancellation. Server shard 6/12 received
the same shutdown signal one minute later. All four were Spot
`m7i-flex.xlarge` instances in `us-east-1a`. No test assertion or build
error preceded those stops. This is infrastructure interruption; the
reason the fleet stopped the runners is not available in job logs. |

## What Changed

- Build the browser fixture UI into the static server's preferred
directory, `server/ui-dist`, and disable Vite middleware. Ignore these
generated assets. Start the source CLI directly from the repository
root, as required by the CLI invocation safety contract.
- Keep all existing chat assertions. Add first takeover and three
service-worker-controlled reloads under CPU throttling. Assert that the
page uses built module assets and that opening it creates no chat task.
- Use forward keyboard traversal from the Zeta link to its star. Check
focus before checking the reveal style.
- Give each runtime exposure test a temporary Paperclip home and restore
environment state after process cleanup.
- Check that reserved ports are free before spawn, serve the fixture
response before exposure, and become free after stop.
- Include the nested fetch error, probe count, and elapsed time in
readiness failures. Add a unit test for this diagnostic contract.
- Refresh active queries when the first live connection opens, covering
events missed during initial page loading. Keep reconnect toast
suppression unchanged.
- Measure Slack retry timing from the transport attempt and recovery
completion. This checks the full provider-requested delay without
spending a small wall-clock allowance on unrelated processing.
- Run each Slack authorization-revocation scenario as a separate test,
with cleanup between cases. All assertions remain.
- Document the browser fixture's build and serving mode. No configured
assertion deadline, readiness deadline, retry count, or skip was added.

## Verification

- Final-head [Linux CI, attempt
2](https://github.com/paperclipai/paperclip/actions/runs/35925901613/attempts/2):
**green**. All 52 check runs passed; two conditional checks were
skipped. The legacy Snyk status also passed. No pending or failed checks
remain.
- `pnpm -r typecheck` passed.
- `pnpm build` passed again after the final UI fix.
- Focused runtime suites: 38 passed, 3 existing platform skips.
- CLI invocation safety suite: 39 passed.
- Slack timing negative control: inserting a six-second delay before
reading the saved action fails the old assertion. All four revised
authorization/backoff cases pass with that same delay. The diagnostic
delay is not committed.
- Live update suites: 92 passed, including the new regression. UI
typecheck and token gates passed.
- Recovery browser negative control: the unchanged tests reproduced the
missing reply (1 failed, 9 passed). After the first-connection fix, all
six recovery paths passed twice (12 passed). Browser assertions and
deadlines are unchanged.
- Server typecheck passed after the Slack test adjustment.
- `GITHUB_WORKFLOW=PR pnpm test:run:general -- --group general-chat
--shard-index 0 --shard-count 3`: 342 passed. The other 682 tests belong
to the remaining shards; collection verified exact coverage.
- Final direct-source CLI launch: all five chat session tests passed.
- `PAPERCLIP_E2E_PORT=32993 PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm
exec playwright test --config tests/e2e/playwright.config.ts
tests/e2e/agent-chat-sessions.spec.ts --repeat-each=3`: 15 passed,
including nine controlled reloads at 4x CPU throttling.
- `GITHUB_WORKFLOW=PR pnpm test:run:general -- --group
general-server-without-chat --shard-index 10 --shard-count 12`: 55 files
passed, 867 tests passed, 21 existing skips. An earlier run could not
initialize PostgreSQL because this Mac exhausted its System V
shared-memory slots. After reclaiming the orphaned segment from this
task's stopped browser server, the full shard passed.
- On final commit `1f1fafc08d`, [browser
4/8](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400837947)
passed all 16 tests; [browser
8/8](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400838155)
passed all 23 tests; [server
11/12](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400838328)
passed 887 tests with one existing skip. The readiness lifecycle case
took 1.842 seconds. Chat 1/3 also passed. All four jobs interrupted by
runner shutdowns passed unchanged on their single rerun.
- Greptile reviewed `1f1fafc08d`: **5/5**, with no unresolved comments.
- Full local `pnpm test:run` was attempted and stopped after confirming
failures outside this patch: a sibling `skills` directory shadows
bundled Slack/AgentMail skills; macOS rejects rename of read-only
skill-cache directories (`EACCES`, also reproduced in an isolated
filesystem probe); and the host exhausts PostgreSQL System V
shared-memory slots. Focused reruns confirmed these limits. The complete
Linux CI run is the repository-wide verification; the full local run is
not green.
- Baseline evidence: the original sidebar test failed on unchanged
master; a development-mode run at 4x CPU throttling passed six selected
cases, so CPU pressure alone did not reproduce the CI bootstrap stall.

## Risks

- Default browser tests now exercise the shipped static UI. They no
longer implicitly cover Vite middleware or HMR; use the development
server for those checks.
- The new browser startup test uses Chromium CDP, matching the only
configured browser project.
- Opening a live subscription now causes one active-query refresh to
close the initial event gap. This adds startup API reads but no
recurring poll.
- Runtime behavior and deadlines are unchanged except for error details.
The historical readiness stall remains unconfirmed; a green rerun alone
cannot establish its cause.
- Local verification runs on macOS. The runtime lifecycle checks also
passed on Linux CI.

## Model Used

- OpenAI GPT-6 through Codex. The session identifies the model family as
GPT-6; an exact served model ID and context window size are not exposed.
Used reasoning, repository inspection, shell tools, code editing, and
test execution. No subagents were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#123` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — targeted suites passed;
full local host limits are listed above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 15:22:21 -07:00
DottaandPaperclip 24429024e7 feat: add Fireflies connector and summary-ready routines (#13890)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps gives agents governed access to external tools through stored
credentials.
> - Routines start work when an external service sends an event.
> - Fireflies provides meeting transcripts and summaries through an
official hosted MCP server.
> - This PR adds that connection and accepts signed meeting events
through the shared app webhook flow.
> - Agents can review completed meetings with the same permissions and
audit records as other work.

## Linked Issues or Issue Description

**Problem or motivation**

Operators need agents to read Fireflies meetings and start follow-up
work when a summary is ready. The Apps catalog lacks Fireflies. The
shared app webhook flow needs to accept its signed deliveries.

**Proposed solution**

Use the official Fireflies MCP endpoint with OAuth or a vaulted bearer
API key. Extend the existing Another app or script flow with signed
webhook support. Verify the raw-body signature and pass the JSON payload
as external data. Select Meeting Summarized in Fireflies. Deduplicate
identical signed deliveries, including setup deliveries.

**Alternatives considered**

A separate REST connector would duplicate the governed MCP path.
Polling, legacy V1 payloads, and automatic provider-side webhook
registration are outside this change.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps and Scheduled Routines
surfaces. It adds a provider to those systems. It does not introduce a
second integration framework.

**Additional context**

A GitHub search found no existing Fireflies issues or PRs. Provider
references and verification limits are in
`doc/connections/FIREFLIES.md`.

## What Changed

- Add the official Fireflies catalog definition, generated registry,
provider evidence, and branded artwork.
- Reuse Access → Connect, dynamic discovery, Permissions, vault storage,
policy, and audit behavior.
- Classify Fireflies sharing, movement, and access revocation as writes.
- Preserve Off and Ask first restrictions during OAuth reauthorization
and API-key replacement. New actions retain normal defaults.
- Add `app_webhook` authentication to the shared Another app or script
flow. Accept bearer tokens or raw-body HMAC-SHA256. Preserve earlier
`fireflies_hmac` triggers and revision snapshots for compatibility.
Existing text columns need no migration.
- Verify `X-Hub-Signature` or `X-Hub-Signature-256` against the exact
request body. Preserve generic event payloads and deduplicate identical
signed requests.
- Keep the routine wizard generic. Show one webhook URL and secret in
Another app or script. Keep all new app webhook event names
provider-neutral. Keep provider setup instructions in the connector
documentation.
- Pass generic webhook JSON to the task in an explicit external-data
block, capped at 16,384 characters. Keep strict meeting validation for
existing legacy Fireflies triggers.

## Verification

- Feature implementation commit `0882dc8a1`: all 54 CI checks passed;
two conditional Storybook checks skipped. This includes full tests,
typecheck, build, browser E2E, canary dry run, and security checks.
Greptile rated this commit 5/5; all review threads are resolved.
- Full local `pnpm -r typecheck`, `pnpm build`, and token gates passed
on the final code. Targeted connector, gateway, webhook, revision, and
UI suites passed during implementation. After the provider-neutral
follow-up, all 84 app-webhook and routine-service tests passed; the
final payload-to-task assertion also passed in the 72-test routine suite
and a clean-config rerun.
- The long local `pnpm test:run` invocation started before the final
edits and was stopped after the final-commit CI suites passed. It
reported one generic webhook test failure while those files were
changing; that test and the entire routine suite passed on the final
source, including a clean-config reproduction. The interrupted local run
is not counted as a full-suite pass.
- In the embedded browser, completed official OAuth consent and
discovered 20 live actions. Real meeting listing, transcript retrieval,
and summary/action-item retrieval succeeded as the selected agent.
Turning a live read Off blocked its test; catalog refresh preserved the
restriction.
- Embedded-browser Another app or script setup, back/save/resume, narrow
layout, and a signed synthetic Fireflies delivery succeeded. The UI
reported authentication passed without creating a task. Fixtures cover
signature tampering, malformed requests, ordinary app event names,
duplicate/setup deliveries, rotation, revisions, pause/archive, and
company isolation.
- Existing MCP browser suite: 8 passed and 2 provider-dependent cases
skipped. Branding checks passed; connector artwork and webhook setup
were checked at desktop/mobile widths and in light/dark modes.
- An unauthenticated POST to a correctly formatted public webhook URL
reached the staging tenant verifier through the existing Cloud gateway.
- A real Fireflies webhook delivery remains unverified. A staging
callback is available for the operator walkthrough. Live API-key
authorization, credential expiry, and a new meeting's summary completion
were not tested against the provider. Fixtures cover these protocol and
lifecycle paths where applicable.

- Storybook follow-up `c54174faa`: 27 production-component stories cover
every UI change, with a source-to-story map in the connector
documentation. Static Storybook build, UI typecheck, token gates, and
Playwright checks for all stories and the mobile footer pass. All PR
checks passed for this Storybook follow-up; Greptile reviewed
`c54174faa` at 5/5.

## Risks

- Fireflies may change its hosted MCP tools or OAuth behavior. Tool
discovery stays dynamic. Experimental search/fetch tools are not
required.
- Public webhook setup requires HTTPS and a separate signing secret.
Fireflies normally emits events for meetings owned by the configuring
account.
- Reauthorization touches shared MCP permission code. Regression tests
cover existing restrictions, new actions, connection removal, and other
gateway callers.
- Webhook receipt grants no tool access. The routine agent still needs
an authorized Fireflies connection.

- New generic triggers rely on provider event subscriptions. Without a
sender-supplied idempotency key, changed request bytes count as a new
event. Existing legacy Fireflies triggers retain summary-only filtering
and per-meeting deduplication.

## Model Used

OpenAI Codex, model `gpt-6-astra`. Used reasoning, repository editing,
code execution, and embedded-browser testing. The runtime did not expose
a context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 17:11:16 -05:00
DottaandPaperclip b648d8cdda fix(evals): support explicit Grok qualification workflows (#13878)
## Thinking Path

> - Paperclip manages AI agents and their provider connections.
> - Product E2E checks real tasks through the browser, server, and
runner.
> - Grok qualification needs separate API-key and subscription evidence.
> - Product subscription tests and direct Grok protocol evals need
explicit credential delivery.
> - This change supplies each credential only to its selected profile
and prepares the pinned binary.
> - Maintainer authorization and protected-environment gates remain
required.

## Linked Issues or Issue Description

Refs #13845, #13847, #13850, #13882.

The Grok feature branch has a manual subscription qualification profile.
The trusted master workflow must admit its selected credential and
prepare the same verified binary and artifact verifier as the API
profile. Direct protocol evals also need the selected xAI key and pinned
Grok binary. These prerequisites do not register or schedule the new
profiles on master.

## What Changed

- Deliver `GROK_AUTH_JSON` from the protected paid environment only when
the selected profile requests that credential.
- Install the checksum-verified Grok binary for the local subscription
profile.
- Prepare the pinned artifact verifier for the manual subscription
suite.
- Extend workflow security assertions to cover the new credential and
profile.
- Add the ACPX Grok credential mapping to the trusted-master catalog,
then deliver only the selected `XAI_API_KEY` to direct protocol cells
and install the target’s checksum-verified Grok binary before packaging.
- Allow a direct-protocol concurrency override from two cases up to the
existing configured ceiling; it can only lower concurrency.
- Document the Grok protocol workflow and its API-only credential
boundary.
- Render missing LLM usage and cost as Unavailable, and label partial
observations with coverage. Preserve raw records, grades, and actual
zero costs.
- Preserve measured campaign source metadata during report regeneration
instead of inheriting the renderer checkout or CI event; skip empty
legacy source records when recovering older provenance.

## Verification

- Latest commit `05d05801477104c8155977bbbe3e119a5241f960`: all 54
reported checks successful, two intentional skips, Greptile 5/5, and
zero unresolved review threads. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/35890978288).

- After merging current master, all 17 workflow security/image tests and
23 catalog/workflow policy tests passed. The trusted catalog also
generated all 39 pinned Grok cells with `XAI_API_KEY` and one case per
shard. The new policy tests execute the concurrency guard against valid,
out-of-range, and malformed values.
- The Grok branch separately passed 450 Product harness unit tests,
including private company credential staging, cleanup, and
token-fragment redaction.
- The fresh-login native subscription smoke passed three repetitions of
tool execution, session resume, restrictive permissions, and cleanup.
These are setup evidence; full subscription Product qualification
remains pending.
- All 72 focused report/billing/history/catalog tests and the Product
harness typecheck passed for the report-display change. The initial
sandbox run could not open the tsx IPC socket; the permitted rerun
passed. A zero-provider-call replay of the actual 16-cell Grok campaign
preserved all result records, grades, timing, and source provenance
while correcting missing usage labels.
- Review the thirteen-file diff. Provider credentials still enter only
the selected paid-test step; default-branch, numeric-actor, and
environment restrictions are unchanged.

## Risks

This admits a refreshable subscription credential to explicitly selected
trusted tests. Store it only in `runner-e2e-paid`, use a test login, and
remove it after qualification. Unselected profiles receive an empty
value. Pull requests cannot trigger the paid workflow. This PR changes
no fleet admission or actor allowlist.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 17:05:52 -05:00
DottaandPaperclip 60c7c9cd1a fix(runner-e2e): pass verified lock digest to Daytona image build (#13876)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Product E2E campaigns test the native runner in local and Daytona
environments.
> - Each campaign resolves one target lockfile and verifies its
downloaded artifact.
> - The Daytona image job did not pass that artifact digest to Docker.
> - Docker used an older default digest and stopped before any selected
task ran.
> - This pull request passes and validates the campaign digest at the
image build boundary.
> - The image keeps its checksum check and frozen package installation.

## Linked Issues or Issue Description

**What happened?**

The merged-master [qualification
campaign](https://github.com/paperclipai/paperclip/actions/runs/35863582409)
stopped in the Daytona image build. The resolved target lock digest was
`e0c928a494f90ddad3c00791e83f09315ee8c82df2a0418a809dcf93649a8ab3`.
Docker used its default digest,
`57b298aceebc48bb94ea0593347348256475da7b2fddb77025d9e57cc8759420`. The
checksum check rejected the mismatch. All 12 selected model cells were
skipped. This follows the image provenance work in #13814.

**Expected behavior**

The image build must check the same lockfile artifact that the campaign
restored and verified. A changed lockfile must still fail the checksum
check.

**Steps to reproduce**

1. Start a Product E2E campaign with a Daytona cell on master
`7944ed3d976d1a7cc26a2d0cee51f227f3542084`.
2. Resolve a target lockfile whose digest differs from the Dockerfile
default.
3. Observe the provider-pack image stage reject the lockfile before
model execution.

**Paperclip version or commit**

`7944ed3d976d1a7cc26a2d0cee51f227f3542084`.

**Deployment mode**

GitHub Actions Product E2E campaign with a Daytona image build.

**Install method**

Built from source with the campaign lockfile artifact.

**Agent adapter(s) involved**

Native Codex and ACPX Claude cells were selected. No model cell ran in
this failed campaign.

**Database mode**

Not involved. The failure occurs during image creation.

**Access context**

The authorized default-branch paid workflow. The build receives no
provider credentials.

## What Changed

- Read the image checksum from the existing target-lock job output.
- Require a 64-character lowercase hexadecimal digest before image
inspection or build.
- Pass the digest as the existing Docker build argument.
- Add regression checks and document the campaign checksum handoff.

## Verification

- The Daytona image regression fails with the original workflow and
passes with the fix.
- All six Daytona image contract tests pass.
- All 450 Product E2E unit tests pass.
- Product E2E typecheck passes.
- Actionlint passes for the changed workflow.
- A context-shaped resolution probe preserves the downloaded lockfile
bytes and digest.
- All latest-head CI checks passed on
`67d41fd9439b2a9a809ddb05765f8617585072c5`
([run](https://github.com/paperclipai/paperclip/actions/runs/35865739359)).
- Greptile gave 5/5 on this head; its test-scoping comment is addressed
and resolved.
- A hosted Daytona image rebuild and the three remote qualification
cells remain pending after merge.

## Risks

The campaign digest comes from the existing trusted target-lock job. The
restored artifact checks, Docker checksum check, frozen install, content
identity, image signing, and verification remain in place. The
standalone Docker default remains available. This change does not alter
task behavior, prompts, credentials, or dependency versions.

## Model Used

OpenAI `gpt-6-astra` through Codex performed diagnosis and review with
code execution tools. OpenAI `gpt-5.6-luna` assisted with investigation,
implementation, and verification. Context window limits are not exposed
in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 08:45:52 -05:00
DottaandPaperclip 106d89314f fix(evals): keep raw diagnostics out of public history (#13850)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Product E2E campaigns retain evidence from live provider runs.
> - The public publisher copied per-attempt files based on their
extension.
> - Credential redaction does not remove hidden reasoning or provider
session IDs.
> - This pull request keeps those raw diagnostics out of public evidence
bundles.
> - Public reports retain normalized grades and declared fixture
screenshots.

## Linked Issues or Issue Description

**What happened?**

The Product history publisher accepted all JSON, log, Markdown, and text
files under an attempt directory. API snapshots and process logs can
contain provider reasoning even after credential redaction.

**Expected behavior**

Keep raw attempt diagnostics in retained Actions artifacts. Publish
normalized results and declared fixture screenshots. Preserve the
original grades and attempt evidence.

**Steps to reproduce**

Place a credential-redacted API snapshot with a reasoning event under an
attempt's snapshots directory. The previous public path check admitted
it because it ended in .json.

## What Changed

- Remove extension-based admission for per-attempt text files in both S3
and Pages staging.
- Require PNG paths to appear in the existing screenshot declaration
allowlist.
- Preserve root normalized results, grading, billing, provenance, and
reviewed screenshot behavior.
- Add seven denial cases for raw diagnostics, renamed files, malformed
JSON, and false screenshot declarations.
- Update report copy and the publication security contract.

## Verification

- Focused history and report tests: 47 passed.
- Replayed the filter against a copy of retained Grok evidence: six
diagnostic files removed, one declared screenshot retained, zero
provider calls. Original artifacts remain unchanged.
- git diff --check passed.
- Product typecheck in the reused local dependency tree reports an
unrelated plugin type mismatch for organizationSwitcher in
server/src/services/plugin-capability-validator.ts. CI will verify a
fresh install.
- No local browser or Docker execution.

## Risks

Public reports no longer link raw per-attempt logs or API snapshots.
Those files remain in the original Actions artifact. This change affects
future publication only; it does not withdraw or rewrite existing
immutable public campaigns. Normalized result content and marked
screenshot review retain their existing boundaries. No workflow
authorization, provider credentials, application behavior, or database
changes.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, code editing, and tool
execution. The exact deployment model ID and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 21:31:39 -05:00
DottaandPaperclip ac854a9b59 ci: build Daytona eval images on the authorized fleet (#13847)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Product E2E campaigns test the browser, server, and runner together.
> - The trusted workflow selects an authorized execution fleet.
> - Daytona image builds still use a fixed GitHub-hosted runner.
> - This pull request applies the existing fleet selection to image
builds.
> - Campaign builds and tests then use the configured EC2 fleet.

## Linked Issues or Issue Description

Refs: https://github.com/paperclipai/paperclip/pull/13845

**What existing behavior does this improve?**

The location of Daytona Product E2E image builds.

**Current behavior**

The image job uses ubuntu-latest even when RUNNER_E2E_AWS_ENABLED
selects EC2 for the rest of the campaign.

**Proposed behavior**

Use the existing authorized runner output for the image job. Preserve
the GitHub-hosted fallback when the EC2 switch is off.

## What Changed

- Route Daytona image builds through the existing authorized runner
selection.
- Assert that the image job depends on authorization and receives no
provider credentials.
- Document the build routing and credential boundary.

## Verification

- Product E2E workflow security: 11 tests passed.
- git diff --check passed.
- Reviewed actor checks, workflow triggers, target commit selection,
signing, and package permissions. They are unchanged.
- Required CI checks pass on the latest head. The first CI attempt had
failures in unchanged chat tests; one diagnostic rerun passed. Both
attempts remain in Actions.
- Greptile reviewed the latest head at 5/5 with no inline findings.
- Live EC2 image verification follows after this trusted workflow change
is merged. No local Docker execution.

## Risks

EC2 image builds depend on the fleet having working Docker and
sufficient disk space. Existing authorization, image digest checks,
signing, and the GitHub-hosted fallback remain in place. No provider
secrets are added to the image job. No application or database changes.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, code editing, and tool
execution. The exact deployment model ID and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 21:30:09 -05:00
DottaandPaperclip 950ccb8eef ci: enable Grok qualification in the trusted paid workflow (#13845)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Product E2E tests verify tasks through the browser, server, and
runner.
> - These paid tests use a trusted workflow from master and an isolated
target commit.
> - The Grok target branch selects XAI_API_KEY, but the trusted workflow
does not deliver that credential.
> - Its artifact test also needs the pinned Python verifier before
execution.
> - This pull request adds both bindings inside the existing paid
boundary.
> - The tests can then run on the configured EC2 fleet without laptop
Docker.

## Linked Issues or Issue Description

Related evaluation infrastructure:
https://github.com/paperclipai/paperclip/pull/11297. No duplicate Grok
paid-workflow change was found.

**What existing behavior does this improve?**

Branch-targeted Grok Product E2E qualification in the existing paid
workflow.

**Current behavior**

Grok cells cannot receive their selected API credential. The Grok
build-revise case also misses the artifact-verifier setup step.

**Proposed behavior**

Deliver XAI_API_KEY only when the selected matrix credential is
XAI_API_KEY. Prepare the existing pinned verifier for the Grok
qualification suite.

**Reason and benefit**

Run the controller, browser, runner and artifact checks on the EC2
fleet. Preserve default-branch workflow authorization and protected
environment secret access.

## What Changed

- Bind the selected XAI credential only in the paid test step.
- Install the checksum-verified Grok binary for local cells before
provider access.
- Include Grok qualification in the existing pinned artifact-verifier
preparation.
- Add an optional max_parallel input that can only lower the configured
campaign concurrency. Use 1 for the Grok test key.
- Extend security assertions and document setup.

## Verification

- Ran the Product E2E workflow-security tests: 11 passed.
- Checked the diff for whitespace errors.
- Reviewed credential selection, setup ordering, numeric actor gates,
target commit pinning, and trusted report checkout.
- Live Grok execution follows after this workflow is available on
master. This PR does not claim completed Grok qualification.

## Risks

The paid test step can use the selected XAI credential and incur
provider charges. The credential remains in runner-e2e-paid and is
absent from setup, build, and reporting jobs. The default-branch gate
and existing environment restrictions remain in place. No database
migration or product behavior changes.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, code editing, and tool
execution. The exact deployment model ID and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 20:17:25 -05:00