106 Commits
Author SHA1 Message Date
DottaandPaperclip 22a3ea3414 Invite assistants from Connections with scoped browser and device consent (#14933)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Public MCP lets people use their organization from an external
assistant.
> - Operators need a visible control for this experimental access.
> - Hosted users should select an organization once and then approve its
permissions.
> - This pull request adds the setting, invitation-first setup, and
browser or device consent.
> - Connections provides a copyable invitation with public instructions
that grant no access.
> - Users reach browser consent from their assistant and return to
inspect or revoke access.

## Linked Issues or Issue Description

Builds on merged foundation #14846. This PR now targets master. Related
settings convention: #13905.

**Current behavior**

The preview uses an environment variable to enable MCP. Hosted consent
repeats organization selection. Assistant access has no entry in
Connections, so users must already know the endpoint and how to reach
consent.

**Proposed behavior**

An administrator enables Settings → Experimental → Assistant connections
(MCP). A hosted connection shows the selected organization and its icon,
then asks for permissions. Requested write access starts checked when
the user’s role permits it; the user can opt out before connecting.
Direct instance connections show an organization picker with the first
available organization selected. The selection stays fixed across
refetches and still requires an explicit Connect action. Connections
includes Assistant Connection (MCP). Its setup page explains the
canonical endpoint, client configuration, browser authentication, and
connected access. It connects as the current person and does not select
or impersonate an agent.

**Reason and benefit**

Operators manage access with the other experiments. Users select one
organization, and both the UI and server enforce that choice.

**Breaking changes**

The old enable variable has no effect. Preview operators must enable the
setting once. Apply the additive consent-request migration before
deploying the tenant, then deploy the compatible Cloud broker. Existing
direct requests and grants keep their behavior.

## What Changed

- Simplify OAuth and device consent: show the Paperclip logo beside
“Connect {client} to Paperclip”, fall back to “your assistant”, and show
the identifying origin plus its favicon below, with the callback URL
also visible when different. Remove the hosted-organization creation
action. Default to the first available organization without silently
changing it on refetch; preserve company restrictions and write
opt-outs. Keep the button row contained on narrow screens.
- Make Copy invitation the primary action, using the shared animated
AgentSetupPrompt and a collapsed manual setup section with icon-labeled
line tabs. Remove redundant link actions, copy-status text, the extra
first-prompt well and revocation explanation from the setup page. Serve
shared version-aware HTML and Markdown instructions without private
organization data.
- Support guarded Client ID Metadata Documents alongside dynamic
registration, and include authorization response issuer identification.
- Add RFC 8628 device authorization with separately hashed codes,
expiry, shared request quotas, persistent polling backoff and atomic
redemption. Reuse human consent, role checks, scoped grants, audit and
revocation.
- Add CLI device login and a local stdio bridge. Store credentials
separately with private permissions and serialize rotating refreshes.
- Add device consent stories and five cold-start paid Product E2E cases
with independent grant, configuration and durable-work assertions.

- Add `enablePublicMcp` to the settings validator, normalizer, feature
catalog, and toggle UI. Check it live for OAuth, tools, subscriptions,
and event delivery. Keep connection management and revocation available
while disabled.
- Default the MCP origin to the existing auth public URL, with strict
validation and an explicit override.
- Persist the optional OAuth `company_id` restriction. Describe only
that company and reject approval for any other company, even if the
person belongs to both. Keep active-membership and role checks.
- Show the Paperclip icon and a large organization icon during consent.
Return the saved company logo through the company-scoped request
response and reuse the standard fallback icon. Use the requested concise
permission labels: “Read all of your Paperclip data” and “Allow write
access and creating tasks as me”. Use concise permission copy, retain a
compact client and callback-origin disclosure, and remove the footer
link.
- Default requested write access on for eligible roles. Preserve opt-out
across organization changes and refetch, reset defaults for a new
request, and submit read-only access when the request or role does not
allow writes. Align the shared checkbox with its label.
- Use organization wording in consent, management, settings, and
walkthroughs. Keep the organization fixed for hosted requests and retain
direct-instance choice.
- Add an Assistant Connection (MCP) card to the Connectors catalog, a
setup page in the app shell, and a return link from Experimental
settings. Include Codex, Claude Code, OpenCode, and generic remote MCP
instructions.
- Read the live gate and canonical server URL through authenticated
setup metadata. Show only the current person’s grants for the selected
organization, refresh after consent, and support revocation. Surface
catalog status failures with an explicit retry action; do not present
them as an empty connection list. Opening setup grants no authority.
- Start the eight guided chapters in Connections. Keep presenter notes
and chapter controls around real product pages in the app shell. Explain
the terminal, consent, delegation, retrieval, and revocation handoffs.
Mark conversation examples as illustrative. Cover first use, client
setup, connected, loading, and error states. Keep the existing consent
and management stories.
- Keep the paid-eval setup and browser helper aligned with the setting
and consent button.

## Verification

- Warm-standby integration fix `ec64ea05e`: public MCP ingress now
follows the Cloud claim guard; MCP and discovery paths return 503
instead of SPA HTML while unclaimed. Event polling checks the in-memory
claim before reading the persisted experimental setting. All 97 focused
OAuth/Cloud tests and server typecheck pass, including new request and
timer regressions for idle-before-claim and resume-after-claim behavior.
Fresh review is 5/5 with no unresolved threads, and all security scans
pass on this final head. All browser shards, typecheck, build, canary
installation and other test groups passed on the first attempt. The
unchanged Cursor sandbox default-command test timed out at 10 seconds;
the exact test passed locally without edits in 587 ms. The single
failed-job retry passed, with the original failure retained in workflow
37500711895. All 54 final-head checks pass on
`ec64ea05e9a03e2179d4e2f84c2de03761f7ce26` (two optional Storybook jobs
are intentionally skipped).

- Final master integration `8457828fc`: merged foundation #14846 and
current master, preserving the invitation changes and all 33 files from
the two newer upstream changes. No migration renumbering was required.
All 95 focused OAuth/Cloud integration tests, full recursive typecheck
and token gates pass. All CI gates passed on that integration head;
review identified the warm-standby issue fixed above.

- Security-review fix `8c1d0b696`: commit shared global/per-source
admission before outbound CIMD work, preserve failed-attempt receipts,
and validate resource/scope before fetching. Added migration
`0305_chubby_vin_gonzales.sql` and six concurrent/adversarial regression
cases. All 69 OAuth/metadata tests, 26 migration checks, full recursive
typecheck and production build pass. The security scanner passed that
commit. Follow-up `87f9658e7` limits only actual cache-miss fetches; 18
authorization requests sharing one proxy across two service instances
use just two fetches. All 70 OAuth/metadata tests and server typecheck
pass after that refinement. Final follow-up `8ebeae84c` reports
admission-storage failures as retryable HTTP 503 instead of invalid
client metadata. Its regression proves no outbound request before
admission and successful retry after storage recovers. All 71
OAuth/metadata tests and server typecheck pass. Final-head security
scanning passes; Greptile is 5/5 with no unresolved findings. CI passed
all browser shards, typecheck, build, token gates and canary
installation. One unchanged adapter-utils bridge test raced a
response-file write (expected a JSON error, received the safe
file-changed error). The exact test passed locally without edits. The
single failed-job retry passed; the original failure is retained in
workflow 37490609192. All 54 checks now pass on final head
`8ebeae84ca77c0cf7ac12c2006f0f8743fe50e0b`, with security scan and fresh
Greptile 5/5 and no unresolved threads. Foundation #14846 subsequently
merged as `e34abee670069cca84afb2efb86041bce7dccbec`; the final
integration above now targets master.

- Integration with current master: preserved the new Connections source
filters and pagination, kept all eval suites, and regenerated the
consent/device snapshots as migrations 0303/0304. All four MCP migration
SQL hashes are unchanged from the staging versions. Full recursive
typecheck and production build, 132 focused UI tests (including catalog
filtering), 89 server authorization/settings tests, 26 migration tests,
120 eval calibration tests and token gates pass. Review follow-up
`4820ce74c` also keeps active assistant grants in Installed, with
pending/error recovery and revocation/company-isolation coverage. All 76
setup/catalog tests, UI typecheck and token gates pass after that fix.
The unchanged signoff browser test timed out waiting for a heartbeat in
CI at `4820ce74c`; the exact test passed locally without code changes,
and the preceding CI head passed that shard. That same unchanged test
failed at the reviewer stage in the next CI run. All five signoff tests
passed three times locally (15/15), without test changes. All eight
browser shards pass at final head `8ebeae84c`; no browser-test edits or
failed-browser-job retries were needed.

- Setup-page refinement at `9ab009178`: all 17 focused setup/consent
tests pass, along with UI typecheck, production build, Storybook build
and token gates. Browser exercised the shared prompt preview and client
tab switching, and the updated InvitationCopied Storybook interaction
checks its clipboard fixture. All final-head CI checks pass at
`9ab009178`, with no unresolved review findings. Deployed successfully
to Butter in
https://github.com/paperclipai/paperclip-cloud/actions/runs/37475189524.
Verified the actual page, tab switching and line styling, removed
actions/copy, and successful native copy/paste of the complete Butter
invitation into a local-only test field. The existing Claude grant was
left intact.
- Consent follow-up at `dc8e9fd11`: all 10 consent tests and token gates
pass. UI typecheck and production build passed again at `4e4d5e4d9`;
Storybook build and eval-helper typecheck passed for `28101cf91`.
Follow-ups let the primary button wrap on narrow screens, preserve a
distinct callback URL, and use only bundled icons to avoid pre-consent
requests to client-selected sites. Browser-verified the real consent
component in desktop and 320px mobile stories, including default
selection, write access and preserved opt-out. Updated E2E
heading/default-selection helpers. All CI checks passed at `dc8e9fd11`,
with review 5/5 and no unresolved threads. The Butter preview
publication needed a retry because npm initially accepted the DB package
before making it visible; the retry succeeded and `dc8e9fd11` deployed.
Verified a fresh, unapproved native Codex CIMD request on Butter:
default organization/write selection, known-client heading and icon,
distinct callback origin, and removed creation action. No grant was
approved for this UI check. Prior paid runs below retain their exact
source provenance; this UI-only follow-up did not rerun paid
qualification.
- Source-pinned paid matrix at
`2992ef2710f47230e7f484c709c6ba02524f884c`: **15/15 passed**, five cases
each on GPT-5.4 Mini, Claude Haiku and Sonnet. Campaign
`local-2026-10-06T02-41-14-462Z`. Covers cold start, existing config,
unavailable host, denied consent and reconnect/later retrieval, with
independent configuration/grant/task/run/document assertions. Original
failures, transcripts, source fingerprints and billing remain retained.
- Final instruction follow-up `cda8178af`: **3/3 cold starts passed** on
Mini, Haiku and Sonnet. Campaign `local-2026-10-06T02-58-30-041Z`.
Latest `0637b9f1c` shares that same guidance across HTML, Markdown and
manual UI after review; generated Markdown is verified byte-identical to
the paid-evaluated version. Shared build, server/UI typechecks, token
gates and 63 auth/metadata tests passed again. Every CI gate passed at
prior HEAD `0637b9f1c`, with review 5/5 and no unresolved threads.
- Other focused checks: 11 CLI credential/refresh-lock tests, 120 eval
calibration tests, server/UI/eval typechecks, token gates and Storybook
build passed. Full recursive typecheck and production build passed
during implementation; CI also passed them at `2992ef271`.
- Local full-suite limitations: a large-file Git streaming test times
out on this Mac, and broader CLI/route runs hit DB hook timeouts. Fresh
MCP reruns passed, and the corresponding CI groups passed. No claim that
the local full suite is green.
- Actual clients: Codex 0.153.4 and Claude Code 2.1.245 reach CIMD
consent; device CLI reaches verification/consent. New grants await human
approval. Existing local OpenCode retrieved a saved result in a fresh
conversation through its previously approved grant.
- Fresh OpenCode 1.18.17 on Butter: started with no MCP config, received
the exact copied invitation, read public setup, configured its server
and started PKCE consent. Its shell command timed out; background retry
reached the client's own callback deadline while approval remained
pending. Latest instructions cover that handoff. **No completed Butter
read/delegation/result retrieval is claimed.**
- Cloud companion
https://github.com/paperclipai/paperclip-cloud/pull/672 passes
checks/review and deployed. Anonymous setup and device-protocol routing
verified. Core `2992ef271` deployed successfully and the actual Claude
web flow now reaches consent. Its extra JWT-bearer metadata is filtered
to implemented grants; unsupported token grants remain rejected. Final
`0637b9f1c` deployed successfully to Butter in
https://github.com/paperclipai/paperclip-cloud/actions/runs/37409195300;
live HTML and Markdown both contain the final guidance. The superseded
instruction-only build was canceled before deployment. This is a
core-only staging preview; private Cloud plugins are omitted. ChatGPT
web is signed out, so browser connector use is unverified.
- Screenshot gallery begins at Butter's dashboard and distinguishes real
setup/pending consent from local reuse and fixtures. It records the
timeout finding. New persistent access needs human confirmation before
the remaining actual-client acceptance work.
- Manual path: Connectors → Assistant Connection (MCP) → Copy invitation
→ paste into assistant → configure and start authorization → sign in and
approve → verify `paperclip_connection` → delegate → retrieve the saved
report later.
- Plan and instructions: `doc/plans/2026-10-05-assistant-invitations.md`
and `doc/public-mcp.md`.

## Risks

- Apply additive, replay-safe migration `0304_curvy_shadow_king.sql`
before using device authorization. The public setup link carries no
credential. Device codes and tokens stay private; neither sharing
instructions nor installing a plugin authorizes access.
- Apply additive migration `0305_chubby_vin_gonzales.sql` before
deploying the shared metadata admission gate. It retains at most 60
short-lived, hashed-source receipts per instance and rejects excess
attempts with 429.
- CIMD metadata fetching is a new external-input boundary. It requires
HTTPS, exact client ID and redirect validation, bounded responses and
guarded DNS/network access. Client names remain self-reported.
- Device support is per-instance. The central Cloud broker retains its
existing grant support. Host installation and tool reload capabilities
vary by client; instructions describe manual settings and restart
requirements.

- Consent names the registered client in its heading and displays its
identifying origin below. Known-origin icons are bundled; all other
origins show a neutral site icon without contacting client-selected
sites. Client names are self-reported; the callback origin is the
recipient check. The Cloud chooser also displays the original client and
receiving origin before tenant handoff.

- A user who accepts the preselected write permission can create tasks
and comments. Task creation and comments can start or wake agents and
use execution budget; the consent label uses the concise wording
explicitly requested by the maintainer. Scope requests, role checks, and
the final Connect action still apply.
- Migration `0303_supreme_garia.sql` adds one nullable UUID column with
`IF NOT EXISTS`. Requests without a company restriction keep the
direct-instance picker. The binding stays recorded if its company is
deleted; consent then fails closed.
- Deploy tenant support before the Cloud broker sends `company_id`.
Unknown or inaccessible organizations must never fall back to a
different company.
- The setting defaults off. Disabling access does not cancel work
already delegated. Existing tokens and unexpired subscriptions can
resume when enabled again; revocation remains separate.
- The catalog entry is visible for discovery while the feature is off.
Setup instructions, OAuth, and tool execution remain gated. No access is
granted by viewing the entry.
- Assistant sign-in starts in the external client so it owns PKCE and
callback state. Client command syntax can change and links to official
setup documentation are included.
- An authenticated instance and valid public URL are required. Hosting,
paid execution, and store publication remain separate rollout steps.

## Model Used

OpenAI GPT-6 in Codex, with tool use and code execution. The exact
serving model version and context window are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks pass;
unrelated local full-suite timeouts are explicitly recorded above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 12:35:16 -05:00
DottaandPaperclip e34abee670 feat(mcp): connect assistants to a team with user OAuth (#14846)
## Thinking Path

> - Paperclip gives teams durable tasks, agent execution, budgets, and
approvals.
> - People also use assistants in Codex, Claude, and other MCP clients.
> - Those assistants need a scoped connection that preserves the
person’s permissions and attribution.
> - Delegating a task must not turn the assistant into the assigned
agent.
> - This PR adds opt-in user OAuth, ten first-party tools, browser
consent, and workflow packages.
> - Paid product evals verify the resulting tasks, documents,
attribution, retries, and access boundaries.
> - The team keeps working after the assistant conversation ends.

## Linked Issues or Issue Description

**Problem or motivation**

A person cannot connect an external assistant to an existing team
through browser consent and safely delegate durable work as themselves.

**Proposed solution**

Expose an opt-in `/mcp/paperclip` endpoint with individually described
first-party operations. Bind every connection to a person, client,
company, resource, and scopes. Reuse domain authorization and
scheduling. Package shared team-review, delegation, and follow-up
workflows for OpenAI/Codex and Claude.

**Alternatives considered**

Related PRs #9393 and #12549 cover earlier remote MCP and board-operator
approaches. This change uses user OAuth and a bounded public catalog. It
does not expose a generic executor, operator administration, static
shared board credentials, or external agent execution. Registry listing
work in #9851 is a separate distribution step.

**Roadmap alignment**

This maintainer-requested implementation extends the governed MCP
gateway, activity attribution, durable work products, and hosted
deployment direction in `ROADMAP.md`. It implements the first release of
the saved design plan; external agent participation and granted
third-party tools remain later releases.

## What Changed

- Add MCP 2.0 discovery and task status/comment/document Events on the
same authenticated endpoint. Persist subscriptions and delivery
receipts, verify HTTPS callbacks, sign Standard Webhooks, encrypt
callback material, recheck permissions/Cloud membership, and bound
retries/expiry. Older MCP clients keep their existing tools.
- Add discovery, dynamic client registration, S256 PKCE, resource
validation, rotating refresh tokens, revocation, and company consent.
Store credentials as hashes and recheck membership at execution.
- Add tools for connection identity, agents/projects, task
search/read/create, human comments, documents/deliverables, and
pending-approval links. Preserve current domain permissions and
scheduling.
- Add durable mutation receipts across reconnects. Matching retries
replay results; uncertain outcomes keep the same request ID and require
inspection.
- Add consent and connection-management pages, OAuth log redaction,
shared plugin workflows, and separate OpenAI/Codex and Claude package
outputs.
- Add eight paid Product E2E cases across three models, independent
durable-state grading, usage evidence, cleanup, and report integration.
Add task-document guidance and regenerate the runner capability
inventories.
- Add migrations 0301 and 0302, the dated implementation plan, result
notes, and direct-client setup instructions in `doc/public-mcp.md`.

## Verification

- Merge integration `e180b1948`: resolved conflicts with current master,
preserved both eval registries, regenerated capability catalogs, and
regenerated migrations as 0301/0302 while keeping the original
replay-safe SQL byte-identical. Local migration safety/snapshot tests
(26), MCP/OAuth tests (38), redaction/OpenAPI tests (71), and eval
catalog/grading tests (198) pass. Token and capability gates pass. Full
recursive typecheck passed. Fresh Greptile review is 5/5 with no
unresolved findings. CI is green on this exact head (55 successes, two
intentional skips, one neutral result): one unchanged Cursor sandbox
test timed out at 10 seconds, then passed locally in 856 ms. A single
retry of that failed shard and the aggregate workflow passed. Merge
remains blocked on the repository code-owner approval rule.

Earlier checks passed at `6aa0962d4fb715f2190bb7bb22efacab2e58495d`: 55
successes, two intentional skips and one neutral result. [The earlier CI
run](https://github.com/paperclipai/paperclip/actions/runs/36901592350)
includes all test shards, browser tests, typecheck, build and canary dry
run. Greptile was 5/5 on that commit with no unresolved review threads.
GitHub still requires code-owner review under the repository merge
rules; passing checks do not bypass that approval. Paid source
fingerprints remain separate below and in the dated result note.

- Paid Events qualification passes **3/3**: GPT-5.4 Mini, Claude Haiku
4.5 and Claude Sonnet 4.6. Each uses a real public HTTPS callback,
signature verification and report retrieval in a fresh conversation. A
final Mini regression passes after the quota/status fixes. All evidence
validates. Bounded tunnel startup retries occur before provider calls
and remain visible; failed earlier attempts retain their original
grades.
- The earlier complete seven-case matrix passes **21/21**, with a
separate **3/3** delegation regression. Two preceding matrices also
passed 21/21 each. A complete 24-cell matrix including Events has not
been run. [The dated
results](doc/plans/2026-10-01-public-mcp-paid-eval-results.md) retain
exact source fingerprints, failures, model IDs and partial costs.
- Node 24: repository-wide `pnpm -r typecheck` and `pnpm build` pass
after merging master. Server typecheck passes after the final
quota/status changes. Eval typecheck and all 892 eval-support tests
pass.
- All 33 real MCP/OAuth tests pass. The preceding combined MCP,
redaction, private-address and DNS-rebinding run passed 129 tests; two
later MCP regressions cover quota reuse and unchanged-status
suppression. All 28 adjacent issue-tree/stale-lock route tests pass. CI
then found a null checkout result in the existing concurrent-workspace
path; logging now uses optional status access. All 12 closed-workspace
tests and all 33 MCP tests pass after that correction. The exact-start
event calibration exposed a timestamp gap; scanning now includes the
subscription start, with all 33 MCP tests and server typecheck passing.
These two narrow corrections follow the paid regression.
- A real Core → Cloud → Core authority round trip passes OAuth, MCP 2.0
subscription/delivery, current membership loss, unsubscribe, legacy SDK
tools, refresh and revocation. Its callback transport is a fixture with
independent HMAC verification. The paid Events campaigns separately
prove public HTTPS delivery.
- Earlier component qualification passed UI 7,117 tests, CLI 502, shared
832, skills catalog 20, database 160 and OpenAPI 10. Token gates, module
boundaries, migration order and plugin regeneration passed. CI covers
general/serialized suites, eight browser shards, runner checks,
typecheck, build and canary dry run.
- **Local full-suite limitation:** the earlier monolithic run was not
clean. It encountered overlapping schema rebuilding, Mac database
shared-memory limits and isolated CLI/fixture failures. Targeted reruns
passed. The existing >32 MiB Git filename stress test still hit its
300-second Mac timeout. The additional serialized sweep stopped after 62
passing suites once CI passed. Original failures and partial logs
remain; this PR does not claim a wholly green local monolithic run.
- Local Codex CLI and Claude Code OAuth login and MCP SDK
interoperability were verified. Public-store installation, actual
ChatGPT Work Cloud Events UI, staging HTTPS client behavior and hosted
newcomer provisioning remain release gates.

Enablement is moving to **Settings → Experimental → Assistant
connections (MCP)** in the stacked follow-up
[#14933](https://github.com/paperclipai/paperclip/pull/14933). Merge
both for the intended setup experience. This foundation branch alone
still uses `PAPERCLIP_PUBLIC_MCP_ENABLED=true`. After deployment, set
`PAPERCLIP_PUBLIC_URL` to the authenticated instance's HTTPS origin, and
connect to `/mcp/paperclip`. Select a team and allow writes in browser
consent. Configure an available agent and budget, then delegate and
retrieve results later. For Events, rescan the deployed plugin catalog
in ChatGPT Work Cloud; the host supplies its webhook credentials when
the user asks to watch a task. See [the setup
runbook](doc/public-mcp.md).

## Risks

- Events are at-least-once and may arrive out of order. No replay cursor
is advertised. Clients must refresh finite subscriptions, read current
state and avoid comment feedback loops. Callback material uses the
instance secrets master key; hosted subscriptions require the updated
Cloud broker and are bounded to five minutes/the access proof expiry.
- ChatGPT Work Cloud/dot event UI, plugin rescan and a hosted staging
subscription remain deployment gates. Local signed-webhook and paid
model evidence does not claim those surfaces have been exercised.
- Disabled by default. Merging adds schema and opt-in code; it does not
deploy a public endpoint, publish a store listing, create a team, or
start paid agents.
- Migrations 0301 and 0302 are additive and idempotent. Their SQL is
unchanged from the earlier preview numbers, so hash-aware upgrade
reconciliation preserves prior staging applications. Normal instance
upgrades must apply it before enabling MCP.
- Task creation and comments can schedule paid agent work. Consent and
tool descriptions disclose that effect. Revocation blocks future calls
but does not undo delegated work.
- Public deployments need edge rate limits and credential-safe logging.
Internal dispatch is restricted to the closed catalog and carries a
request-local verified actor.
- Hosted onboarding requires the companion Cloud broker, encryption-key
configuration, and tenant rollout. Self-hosted direct connections can
use this PR alone.
- Store acceptance and agent-mode participation are not claimed.
Checked-in plugin endpoints are development defaults; rebuild packages
for a real deployment before installation.

## Model Used

OpenAI GPT-6 in Codex, with reasoning, tool use, and code execution. A
more specific serving version and context-window size were not exposed
by the session. Paid eval models: `gpt-5.4-mini-2026-03-17`,
`claude-haiku-4-5-20251001`, and `claude-sonnet-4-6`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted/component checks;
full local-run limitations are recorded above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 11:48:53 -05:00
DottaandPaperclip 16b7db35ff Shorten planning skills and measure task decomposition (#15296)
## Thinking Path

> - Paperclip manages work for AI agents.
> - Planning guidance helps agents choose owners and dependencies.
> - The runtime skill favors few tasks, but the catalog skill requires a
child-task breakdown.
> - Both add repeated process instructions that can distract from the
requested outcome.
> - This change keeps the ownership and dependency rules and removes the
required matrix and repeated checklist.
> - A bounded Product E2E comparison measures saved outcomes and task
handoffs before qualification.

## Linked Issues or Issue Description

Refs #11057. Related measurement work: #15218.

**What existing behavior does this improve?**
Planning and delegation through the runtime plan-to-tasks and bundled
task-planning skills.

**Current behavior**
The two skills contain about 1,900 words and conflicting guidance on
whether plans require child tasks.

**Proposed behavior**
Keep cohesive work with one owner. Split only for a real owner, parallel
output, dependency, independent review, or follow-up lifecycle. Preserve
existing authorization and planning mechanics.

## What Changed

- Shorten both skills to about 400 words combined. Preserve their keys
and installed-version behavior.
- Remove the duplicate operational-skill pointer and regenerate affected
source metadata.
- Add twelve explicit Product E2E cells: four scenarios with current,
short and disabled planning skills.
- Use the current task composer and actual create-response ID; calibrate
public skill APIs and browser creation without providers.
- Eliminate an observed collision in chat-test company prefixes with a
per-suite sequence.
- Grade saved documents, exact author/run attribution, child count,
prerequisite execution order, review boundaries and completion handoffs.
- Retain current skill bytes and report source, selections, run
accounting and failures.

## Verification

- `pnpm test:e2e:runner:typecheck`: pass.
- `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks
pass.
- `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve
local Codex cells.
- Archived current skills match master
`72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly.
- Corrected fixture: three real public-API/database calibrations pass
with zero provider runs; all 35 evaluator checks and Product E2E
typecheck pass.
- Setup campaign
[37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253)
was canceled after source review found unsupported bundled edits and
automatic core reinstallation. Its paid-cell step was skipped: zero
provider runs, no behavioral grade.
- The next setup
[37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799)
failed before task creation on the old title-field selector: zero actual
runs, original FAIL retained, cleanup passed. A real browser/API
calibration of the new helper passes with paused non-provider agents and
zero runs.
- Full local typecheck/build pass. Full local tests retain one unchanged
five-minute Git streaming timeout (also fails isolated), 9,591 passes
and 5,796 skips. CI's chat failure was a proven random fixture-prefix
collision; five affected cases pass after the test-only repair.
- Paid behavior comparison and new-head CI/review remain pending. This
PR remains a draft.

## Risks

- The shorter text may change delegation decisions. Live outcomes are
not yet qualified.
- The initial comparison uses one profile and one attempt per cell. It
cannot establish cross-model reliability or cost trends.
- Disabled means unassigned company-owned copies; the company library
remains discoverable. This does not qualify global removal, automatic
accepted-plan wiring changes, or installed-copy migration.
- Skill availability does not prove a model read or cognitively used it.
- No provider/tool protocol, permission, timeout or runtime lifecycle
behavior changes in production.

## Model Used

OpenAI Codex (GPT-6), with repository inspection, code editing and tool
use. The exact backend model ID and context-window size are not exposed
in this session. The declared eval model is native Codex `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 09:36:01 -05:00
DottaandPaperclip 72ff3a9f27 Measure native tool context and expand bounded workflow evals (#15218)
## Thinking Path

> - Paperclip manages persistent agents and their assigned work.
> - Agents receive both fixed instructions and tool definitions.
> - Moving a procedure into a tool description still adds model context.
> - We need to measure the complete delivered catalog and test real
outcomes.
> - The tested reductions saved little space and introduced failing
outcomes.
> - This PR keeps measurement, bounded eval coverage and the original
evidence.
> - Production prompts, tools and runtime behavior stay unchanged.

## Linked Issues or Issue Description

Refs: #15151, #14961, #14948, #14985.

This adds the measurement and eval coverage needed to assess further
native
instruction changes. The attempted hiring and dependency reduction
failed
qualification and is excluded from the final diff.

## What Changed

- Measure the actual standard-mode tool authority, including all 39
tools and their input schemas. Capture scripted native start, resume and
continuation payloads and the OpenCode MCP declaration list.
- Add OpenCode to the two explicit-only local hiring/reuse and
delegation/feedback stories. Preserve the original task requests and
independent oracles.
- Apply one attempt per selected story and explicit company and
lead-agent budget stops.
- Select managed hiring credentials from the requested profile,
including OpenRouter.
- Run the existing Node test files under Node instead of collecting them
as Vitest suites.
- Retain sanitized comparison reports, original failed grades, source
hashes and evidence gaps.

## Verification

Final source: `2e7cef78eef7cdfe02265e0dcb03e855b8e50bd8`. Local
verification passes:

- Full repository `pnpm -r typecheck` and `pnpm build`.
- Eval-support unit tests: 1,252 Vitest checks and 128 Node checks.
- Eval TypeScript check and six final-source measurement tests. All 36
normalized
components across nine scripted deliveries and the OpenCode MCP catalog
match
the baseline exactly; the complete standing projection is 49,200 bytes.

Fresh review of this exact head is [5/5 with no remaining
findings](https://github.com/paperclipai/paperclip/pull/15218#issuecomment-5996723170);
all three review threads are resolved.
[Final-head
CI](https://github.com/paperclipai/paperclip/actions/runs/37354539126)
passes all 47 jobs, including repository typecheck, build, tests, runner
checks,
browser shards and the canary check. One initial annotation-test timeout
is
retained in attempt 1; its seven-test suite passed locally, and the
affected
CI lane passed on one targeted retry. No source or paid eval rerun was
needed.

Both ready-transition security scans passed with zero annotations. The
PR is
out of draft and conflict-free. GitHub still requires code-owner
approval for
the `package.json` test-script change; its requested reviewers are
already set.

A byte-for-byte comparison against master context
`a65ca0950834a85bb93bcc4b4042ecacdebfef53` confirms no production
changes under
`packages/` or production `server/` paths. The only server addition is a
measurement test. No new paid rerun is needed to compare unchanged
production
bytes. This does not claim that existing product defects have been
fixed.

The rejected corrected experiment had baseline **5 PASS / 1 FAIL** and
candidate
**3 PASS / 3 FAIL**, including **two newly failing pairs**. The later
readiness
experiment had candidate **3 PASS / 3 FAIL** and baseline **3 PASS / 2
FAIL / one
setup cell without a behavioral grade**. Its five comparable pairs had
two new
failures, two new passes and one unchanged pass. The missing baseline
Codex
cell never reached its provider step because Docker setup timed out.

The observed missing behaviors include parent continuation, waiting for
the
latest child revision, revised ZIP delivery and OpenCode credential
persistence.
Those failures remain failures. Source review and passing CI do not
regrade them.
The original reduction saved 460 bytes; its first repair saved only 125
bytes,
and the larger unqualified runtime repair increased the full projection.
None
of those production changes is shipped here.

Read the
[report](https://github.com/paperclipai/paperclip/blob/2e7cef78eef7cdfe02265e0dcb03e855b8e50bd8/doc/plans/2026-10-05-native-procedure-guidance.md)
and its linked sanitized receipts for exact
sources, original campaign links, pair-level results and evidence
limitations.

## Risks

- Full-catalog bytes are not model tokens, invoices, private vendor
prompts, lazy-loading behavior or truncation proof.
- The two stories are explicit-only and do not prove general coding
quality or arbitrary resume behavior. Single trials do not establish
causation or performance trends.
- The retained OpenCode candidate credential guard failed. Cleanup
removed the original provider database, so the precise persistence
mechanism remains unknown. The guard is unchanged.
- ACPX provider-execution IDs lack proven host-call mapping. No
extra-work or feedback-consumption claim is inferred by matching names,
order or counts.
- PostgreSQL cannot start locally while the host's shared-memory slots
are exhausted. Hosted CI must supply the full database checks; the full
local database suite is not claimed green.

## Model Used

OpenAI Codex, based on GPT-6, with repository tools and code execution.
The exact deployment model ID and context-window limit are not exposed
in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 20:03:35 -05:00
DottaandPaperclip 4857799a88 feat(connections): deliver saved instructions to authorized agent turns (#15216)
Persist optional connection instructions and deliver authorized snapshots to agent execution prompts. Keep provider templates with each app definition, preserve edits and opt-outs, and replace sessions when guidance or access changes.

Use shared production settings across setup and Permissions, with source visibility in agent Instructions. Add the initial memory-provider defaults and managed Honcho workspace configuration. Include migration 0298 and regression coverage for generic providers, runtime delivery, authorization, and catalog regeneration.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-05 18:16:00 -05:00
DottaandPaperclip b43073d11f feat(connections): sync and group accounts managed by aggregators (#15254)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents governed access to external tools.
> - Aggregator gateways can expose accounts that users already connected
upstream.
> - The Apps catalog did not show those accounts or their current
provider status.
> - Separate cards and setup tasks also made account ownership unclear.
> - This pull request discovers upstream accounts and groups them under
one app card.
> - Users can find connected apps while each provider keeps control of
its accounts.

## Linked Issues or Issue Description

**Subsystem affected**

Connections across the database, shared contracts, server, and board UI.

**Problem or motivation**

Users cannot see which apps are connected through a saved aggregator
gateway. Native and upstream accounts need one app card. Discovery must
preserve company, user, gateway, and credential boundaries.

**Proposed solution**

Sync account metadata from Composio, Arcade, and supported Executor
gateways. Keep upstream account management in each provider. Use source
chips and search to browse the catalog. Preserve native setup and the
gateway's existing access policy.

**Alternatives considered**

Creating a local executable connection for each upstream account would
duplicate authorization state. Using an agent task for routine Composio
setup would add an unnecessary step. The board now calls the saved
gateway directly for that setup.

**Roadmap alignment**

This extends the shipped Connected Apps and MCP Tool Gateway features in
ROADMAP.md. The duplicate search found no open PR for managed account
discovery.

Related work: Refs #13755, Refs #13941, Refs #14725, Refs #13855. Open
PR #12906 covers adjacent toolkit routing work.

## What Changed

- Add provider-neutral discovery, sync, and refresh APIs. Preserve the
Composio API paths.
- Cache observations by company, saved gateway, viewing user, and
credential version. Retain stale observations after failed or incomplete
scans.
- Add optional Arcade account sync credentials in the vault. Discover
Executor accounts through its supported inventory interface.
- Group native and upstream accounts in one app card. Imported account
menus open their provider. Gateway menus own refresh and sync setup.
- Add Paperclip, Composio, Arcade, Installed, and All chips. Show 50
catalog entries per page. Keep connected accounts above discovery. Keep
explicit provider searches scoped.
- Simplify Composio app setup and refresh its connected app list on the
gateway Permissions page.
- Add a compact agent access card and task creation defaults for
connection setup. Preserve explicit blocks and approval policies.
- Add two replay-safe migrations, service and UI tests, Storybook
journeys, and acceptance stories.

## Verification

- Passed the repository typecheck, full build, token gates, and
migration ordering check.
- Passed the focused provider adapter, connection interaction, and
catalog tests after rebasing onto master.
- Passed all nine database sync and migration replay tests using a
disposable database on the test-drive PostgreSQL cluster. Removed that
database after the run.
- Verified Arcade cursor pagination against its official Go SDK and
passed all eight adapter tests, including short and incomplete pages.
- Passed all 45 interaction tests after making the exact requested tools
and their Allowed/Ask first permissions visible before granting access.
Verified the compact card in Storybook.
- Passed the complete UI suite on the final code: 683 files and 7,432
tests, including the corrected Composio destination assertions. Passed
130 focused tests for the UUID, management-link, and health-status
corrections.
- Passed 22 Composio setup/sync tests, 23 connection-intent service
tests, and the connection migration test in separate disposable
databases. Database startup alone was substituted; the suites exercised
their real SQL and services.
- Passed all 10 OpenAPI route checks and the full-stack
connection-intent browser test, including scoped consent, agent
continuation, and task completion.
- The local full runner encountered embedded PostgreSQL startup failures
on this loaded macOS host. The earlier in-flight run also held the
pre-fix Arcade transform; a fresh run of the final provider suite
passes. The final-head CI is queued during GitHub’s active Actions
incident: https://www.githubstatus.com/. The previous run also lost
several runners simultaneously; its real catalog assertion failures are
fixed and the fresh complete UI suite passes.
- Tested the real test-drive server in the embedded browser with a live
Composio gateway. Detected Airtable and Circleback. Verified refresh
progress, account rows, source chips, search scope, and 50-entry
pagination.
- Arcade and Executor coverage uses provider fixtures. Live credentials
were unavailable.
- Storybook builds successfully and includes grouped native/provider
accounts, stale and unavailable discovery, optional Arcade setup, and
mobile states. The acceptance document records the simulated and live
coverage separately.

- Greptile reviewed final commit
`217b024c27b5933e773ce9419c4e92b1032042c6` at 5/5. All six review
threads are resolved, security scans pass, and the PR has no merge
conflicts. The outstanding remote checks are `ci / Select trusted
runner` and `review`, queued by GitHub. They need to complete before
merge.

## Risks

- Provider response changes can break inventory discovery. Failed scans
retain observations and show stale status.
- Composio scans only the supported catalog and can take time. Large
inventories run in the background with progress and a bounded lease.
- Arcade requires a project API key and user ID when the gateway cannot
supply them. This key is used only for discovery.
- Executor discovery depends on the server's exposed inventory tools.
Unsupported servers report unavailable discovery.
- Cached account rows do not grant access or create executable
connections. Gateway policies still govern tool use. Account deletion
and per-app authorization remain upstream.
- The migrations add tables and one nullable column. Replay preserves
existing rows and company-scoped foreign keys.

## Model Used

OpenAI Codex, based on GPT-6. The session does not expose a more
specific serving model ID or context limit. Used reasoning, repository
tools, code execution, and browser verification.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 14:57:04 -05:00
DottaandPaperclip a386a59998 Reduce repeated native completion guidance and preserve final replies (#15151)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native agents receive task constraints and completion tools from
Paperclip.
> - Completion tools already define the procedure for reporting a
result.
> - Repeated procedure text adds instructions to each full task turn.
> - The final reply must still explain a blocker and link a saved
document.
> - This pull request removes repeated procedure text and keeps these
visible outcome requirements explicit.
> - A document receipt supplies the exact link, and stricter evals check
the persisted reply and browser navigation.

## Linked Issues or Issue Description

Refs: #14961. Related: #14948 and #15007.

**What happened?**

Native task envelopes repeat completion procedure text. A reduced
envelope needs explicit final-reply requirements. The `write_document`
receipt also lacks a canonical document link.

**Expected behavior**

Keep the completion tools as the source of procedure details. Require
one accepted completion result before the final reply. A blocked reply
must explain the reason, owner and unblock action. A document reply must
contain a working link to the saved document.

**Steps to reproduce**

1. Run the native assigned-skill document case and native blocker case.
2. Inspect the run-attributed provider final and its persisted comment.
3. Check the blocker explanation or open the final reply's document
link.

## What Changed

- Remove repeated completion procedure text from the native task
constraints and backend instructions.
- Keep explicit blocker and document-link requirements in full task
turns.
- Return a company/task-scoped `documentHref` from `write_document`.
Preserve the link in the idempotent mutation receipt.
- Repeat canonical links for this run's current saved revisions in
accepted completion feedback. Give blocked providers final-response
guidance for the cause, owner and unblock action.
- Keep internal document/comment anchors when Markdown issue links load
cached issue details.
- Add a manual six-cell comparison suite with strict source, build,
default-instruction and budget admission.
- Capture eighteen shared runnerd RPC projections and six direct
OpenCode HTTP projections across start, resume and continuation phases,
using scripted local transports and no provider execution.
- Apply v3 checks only to the manual instruction comparison; preserve v2
checks for the existing native completion suite. Check the actual
persisted blocker reason and exact saved-document link. Click the
rendered document link and check the original content marker in the
classic document card or the new document tab.
- Forward exact OpenCode finishing calls through the controller. Wait
for acceptance, keep accepted feedback and concrete rejection text, and
reject malformed responses. Preserve ordinary dynamic-tool response
handling.
- Settle the completion decision and tool response before mapping a
racing idle/error/abort event or handling explicit close/interruption.
Reject a concurrent finishing call before controller admission.
- Add a provider-free regression through real runnerd, the OpenCode
proxy and a fake provider. Reject the first completion, accept the
corrected report in the same turn, and propose one result.
- Keep all original verdicts unchanged. Treat replay under new checks as
separate diagnostics.

## Verification

- `pnpm -r typecheck` and `pnpm build` pass locally.
- Native document-authority tests pass, including company/run
authorization and idempotent replay.
- Native runtime-context, backend and measurement tests pass.
- Final-answer calibration, protocol scoring, source-admission and
catalog tests pass. Wrong reasons, absent links and wrong link targets
fail.
- `pnpm test:e2e:runner:typecheck` passes. Discovery lists exactly six
single-attempt local cells with the declared models.
- Exported `prepareNativeInstructionPreflight` then
`verifyNativeInstructionPreflight` pass on this clean committed source.
They build locally and make zero provider calls.
- Corrective live confirmation is incomplete. Source 3a7349d passed both
Claude and both Codex cases. OpenCode saved the correct document but
omitted its final link; its blocker case was canceled before paid
execution. Preserve this failure. The e171282 confirmation was stopped
during build after fresh review found a completion-settlement race; it
executed zero providers. Source c3e0cb303 fixes that race. Two affected
OpenCode cases await fresh review and one bounded confirmation; earlier
results remain attributed to their original source.
- OpenCode proxy parsing, driver, factory and input tests: 81 pass
across retained focused runs, including six settlement races.
Evaluator/scoring/admission checks: 122 pass. The real proxy regression
passes. Fresh local prepare then verify passes with 18 shared and 6
direct scripted captures, fresh SDK/Rust builds and zero providers.
- The full local suite recorded two failures: a webhook timeout and a
Git-scan load count mismatch. Both files pass in isolation with
unchanged assertions/time budgets; preserve the original failure log.
Fresh c3e0cb303 CI and review are pending. This PR remains draft.

## Risks

- Final-answer wording can vary by provider. The checks cover the
declared release-access blocker and saved document fixture, not general
answer quality.
- A single trial does not establish general equivalence, cause, speed,
cost or live resume behavior.
- `documentHref` is an additive receipt field. It points to the current
saved document, not an immutable historical revision. Replaying an older
receipt does not fabricate a new link.
- The correction adds four production paths for document receipts,
accepted completion feedback and UI navigation, plus four OpenCode
controller/proxy paths, beyond the original three instruction paths.
Completion rejection must remain repairable; the production-boundary
regression covers it.
- Preserve the frozen comparison context for live measurement. A
merge-tree check against current master is clean. Do not relabel earlier
live results as results from a later source tree.

## Model Used

- OpenAI Codex, GPT-6 family. The exact serving model ID and context
window are unavailable in this session. Capabilities used: reasoning,
code editing, shell execution, test authoring and evidence review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 08:10:14 -05:00
DottaandPaperclip b17019e14d fix(agents): reduce default instructions and qualify stock harnesses (#14948)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its adapters supply task context and access to Paperclip skills and
tools.
> - The default hire manual and shared prompts also repeat general work
procedures.
> - Those procedures overlap with stock provider instructions and the
Paperclip skill.
> - Existing E2E fixtures supply a QA manual, so they do not qualify the
production default.
> - This pull request reduces the generic instructions and adds real
default-hire coverage.
> - The benefit is less competing guidance, with inspectable evidence
for preserved skills and task context.

## Linked Issues or Issue Description

Refs: #14920. That merged change preserves native Codex base
instructions. This PR covers the default manual, shared legacy prompts,
operational skill guidance, and the narrowly approved ACP
skill-discovery/session-environment repair for measured delivery and
credential-persistence failures.

**What existing behavior does this improve?**

New non-CEO hires without a custom bundle and legacy task/chat startup
and continuation prompts.

**Current behavior**

The shipped default manual contains 602 words. Generic task/chat prompts
and ordinary resume deltas repeat work procedures already available
through the harness and Paperclip skill.

**Proposed behavior**

The default manual contains only the eight-word company identity. Shared
startup prompts retain identity and connection guidance. Ordinary resume
deltas retain current work context without the generic execution
contract.

**Reason and benefit**

Let the stock harness guide general work. Keep Paperclip-specific
capabilities and independently test default hires, skills, ordered
comments, and chat restart.

**Breaking changes**

New default hires receive less guidance. Existing saved manuals,
explicit custom bundles, CEO templates, and specialized wake contracts
retain their behavior. The obsolete includeExecutionContract option
remains accepted for source compatibility.

## What Changed

- Reduce the default hire manual to one sentence.
- Reduce shared task/chat defaults and remove the generic
ordinary-resume contract.
- Keep connection guidance, auth, skills, custom prompts, and
specialized wake context.
- Add credential-free instruction-boundary gates and 26 explicit Product
E2E cells across eight legacy/native profiles, including two focused
Paperclip-storage cases.
- Capture public hire receipts before providers run, then grade
delivered prompts and independent task/chat outcomes.
- Add an early legacy skill API recipe for saving a task document,
checking the saved revision receipt and linking the document. Improve
stock task/heartbeat skill-selection metadata and show a clickable
Markdown UI-link example. Keep native tool completion separate.
- Advertise bounded routing descriptions and exact successfully staged
SKILL.md paths in legacy ACP Claude; keep full bodies on demand and
preserve remote path rebasing.
- Remove only the provider environment from copied persisted ACP session
records, while loading current run credentials and preserving all other
options/conversation state.
- Regenerate both capability metadata inventories and reject stale
manifests/inventories before provider admission.
- Publish the original reduction and focused skill-repair comparisons,
preserving all failures, automatic recovery, cost coverage and
limitations.

## Verification

**Behavioral qualification remains pending.** Original legacy ACP Claude
loses the issue document only in the reduced cohort beneath an unchanged
credential failure. A source-backed diagnosis finds that neither
ordinary assignment reads the staged operational skill, while the
runtime persists provider environment in session state. The new common
repairs expose skill metadata/path and omit persisted env; strict
document and credential guards stay intact. [Inspectable diagnosis and
retained
hashes](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-readiness.md).

Current repair head `de0965984ff3edf611ae6d0e7ca5c7d5ae3947bb`
incorporates master `569c7203aa24b95440682983ce7940ba1d4247bd` (merged
#14961/#15007). All 222 affected adapter tests, adapter-utils/E2E
typechecks, and final 96 variant/grader/retry calibrations pass. The
frozen historical comparator is
`c25697f4260b6f3adfea143c3ae9932e2f42986d`: 8,280 of 8,291 paths
identical, exactly two production instruction paths plus nine declared
unit expectations differ. The operational skill/discovery/environment
repairs, selected model/profile/task/core grader/auth/permissions/retry
policy are identical. Both actual launcher prepare→verify admissions
pass with zero providers. [Immutable manifest and exact
receipts](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-evidence/manifest.json).

One original legacy ACP Claude cell per variant is authorized, with
enforced single campaign attempts, 12-minute deadlines and company/agent
1,000-cent hard stops; every product recovery run/cost is counted.
Actual live outcomes are pending. Current normal CI has one failed
server shard and failed aggregate verify under diagnosis; other normal
gates including typecheck/build/Rust/all eight browser shards pass.
Fresh review completed successfully; the valid historical startup/resume
masking finding was fixed with per-invocation task/chat checks and
strict complete-snapshot capture, calibrated and resolved. Prior heads,
failures and campaigns below remain historical evidence, not checks on
this repair head.

- Prior head `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` is replayed on
merged hiring master `862a5758ba0e88a33232c1f1fa645e85c38a3113`. All 52
current-head checks pass with two intentional Storybook skips, including
repository typecheck/test/build and the browser shard. Fresh Greptile is
5/5 with zero unresolved review threads. Exact-head stock prerequisites
pass 599 assertions (598 TypeScript + 1 Rust), all six gates and
retained receipt verification, zero providers/source errors. Fingerprint
`a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`.
Combined catalog/hiring calibrations pass 67 assertions, E2E typecheck
and 26-cell stock discovery pass. Canonical contract/inventory checks
and the later issue-derived reference calibration are retained; that
reference-only follow-up is not live-qualified by earlier frozen runs.
- Prior full repository typecheck/build passed. The complete local
Vitest run executed 14,956 tests: 14,870 passed, 83 skipped, three
timing failures. All three affected files passed unchanged narrow
reruns; original failures remain retained. Current-head CI now passes
the full general checks; the original local failures remain retained.
- The original 24-pair default-manual/shared-prompt comparison has two
new overall classic Claude/OpenCode document-delivery failures plus an
additional legacy ACP Claude document loss beneath an unchanged
credential-guard failure (not closed by later runs), two newly passing
OpenCode ordered cases, seven unchanged failures and 13 unchanged
passes. Equal 15/24 totals do not establish behavioral equivalence.
[Complete original
report](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-stock-harness-live-comparison.md).
- The skill-only repair holds the eight-word manual/shared prompts and
merged #14920 fixed. All four matched profile configurations and 203
fixture/behavior files match. Candidate
`abd0b628ca642c09a54a4edc56a5227402f6686e` varies only the two skill
sources against baseline `bc83fe030234439ac51279502a28803958963e2e`.
[Candidate
workflow](https://github.com/paperclipai/paperclip/actions/runs/37060885547)
and [baseline
workflow](https://github.com/paperclipai/paperclip/actions/runs/37060888047)
each pass 571 exact-source prerequisites before providers; all eight
cells clean up successfully. Failed campaigns publish successfully and
remain failed.
- Repair pairs: Claude original Fail → Pass; Claude explicit Pass →
Pass; both OpenCode cases Fail → Fail. Explicit OpenCode's handoff
worsens beneath the unchanged failing UI-link grade: baseline gives a
clickable API URL, candidate gives a code-formatted path without an
anchor. The request's usable-link wording is narrower in the UI-only
oracle. [Complete repair report and safe
projection](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-legacy-document-skill-repair.md).
- The subsequent narrow stock metadata/link correction has two matched
Pass → Pass cases, zero new machine failures/passes and no pending
pairs. Both original-case handoff links remain deficient: candidate uses
a wrong PAP prefix, baseline supplies a bare prefix-less slug path; the
preserved original oracle only requires a durable document. Both
explicit clickable UI-link cases pass revision/content/link grading. All
four exact-source 587-check gates, single assignment runs and cleanup
pass. This does not establish fix causality because baseline also
succeeds. [Candidate
workflow](https://github.com/paperclipai/paperclip/actions/runs/37069547401)
freezes `fe9dc1e3c518825242ed889ab9c8352986f8c2ed`; [matched
baseline](https://github.com/paperclipai/paperclip/actions/runs/37069552374)
freezes `0d7ecfa96d72fba79b7f0a25052b42c0686c0488`. This is a skill-only
comparison with reduced manuals/shared prompts held constant, not a
repeat of the historical-manual comparison. Only original and clarified
explicit classic OpenCode cases are selected, two per variant/four
expected turns. 8,242 other tracked files and both profile hashes match;
protected workflows admit each exact source before credentials.
[Complete qualification
report](https://github.com/paperclipai/paperclip/blob/74d0d3d945f4c52d0814b5a845ab5bd09f33cd6b/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.md).
Candidate original loads Paperclip/reference before saving publicly;
baseline original loads it after writing locally, then saves publicly
within the same assignment. Reported cost totals are $0.0107824490
candidate / $0.0107909015 baseline, with unmetered runtime. The later
reference-only issue-derived link correction is provider-free calibrated
and **not live-qualified** by these frozen runs; no further paid runs.
- Retained tool calls show the repaired original OpenCode assignment
loads only its assigned output skill before writing locally. Operational
Paperclip is first loaded during automatic disposition recovery; its
early recipe is visible then, but it never saves the missing document.
Explicit candidate loads Paperclip and reads the new reference before
saving successfully. All nine actual runs are counted. Reported LLM
totals are $0.3802537209 baseline and $0.4918990161 candidate; local
runtime is unmetered.
- Initial setup, packaging, cancelled/missing-cell recovery, callback
test and relative-output attempts remain retained. No completed provider
failure was rerun. Frozen measurement branches are unchanged by later
canonical metadata maintenance.
- Run `pnpm test:e2e:runner:stock-harness`, `pnpm test:e2e:runner:unit`,
and `pnpm test:e2e:runner:typecheck`. Select `stock-harness` explicitly
for paid execution; it is excluded from `--all`.

Prior-head integration: `36aa4d81c49a1a8f6f04b1a068fae19aa901955f`
replays this PR on merged hiring #14985
(`862a5758ba0e88a33232c1f1fa645e85c38a3113`), preserving the four
explicit custom-CEO-bundle checks, minimal generic manual boundary, and
both suites. The combined fixture catalog and hiring calibrations pass
67 assertions; exact-head stock prerequisites pass 599 assertions (598
TypeScript + 1 Rust), all six gates and retained-receipt verification,
zero providers/source errors, fingerprint
`a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. E2E
typecheck and 26-cell stock discovery pass. Fresh current-head CI passes
all 52 checks with two intentional skips, and fresh Greptile is 5/5 with
zero unresolved review threads.

The prior source-plan browser failure is retained: a deterministic
process fixture replayed its last `fixture:plan` command on
`chat_task_completed`, writing revision 2 with identical body after the
approval handoff. This was not paid provider execution. Rebased
current-head CI passes the same assertion without an old-head retry or a
change to that browser fixture.

The merged hiring change was measured separately on immutable matched
unions, with this reduced/shared/operational context and native
completion guidance held constant. [Complete original two-profile
report](https://github.com/paperclipai/paperclip/blob/f0512647656be78e48abd8c22a3078db8bf6bcd2/doc/plans/2026-10-02-hiring-template-live-comparison.md):
[candidate](https://github.com/paperclipai/paperclip/actions/runs/37075466208)
/ [historical
baseline](https://github.com/paperclipai/paperclip/actions/runs/37075469463),
705 provider-free prerequisites each. Both pairs are unchanged Fail →
Fail on the exact-five count, with six core delivery checks passing all
four cells; 28 actual successful runs include eight automatic completion
wakes, zero retries, four successful cleanups. Source-read coverage is
uncomparable, actual model charges unknown. Separately versioned
provider-free accounting remains analytical work; original verdicts are
preserved. This does not rerun or qualify the completed default-manual
or native campaigns.

## Risks

- Legacy ACP Claude's additional delivery loss is not closed by any
later matched run and blocks the no-extra-failing-behavior merge
criterion. Legacy document delivery may have relied on the prior
manual/shared prompts. The early skill repair improves Claude in one
trial; the later OpenCode pairs pass in both variants and cannot
establish causality or robust recovery. Both original-case links remain
deficient beneath the storage-only grade. The later issue-derived
reference correction has only provider-free validation. Native
finish/block descriptions must not be supplied to legacy agents.
- The comparison holds merged native Codex fix #14920 constant; it
cannot measure that fix's before/after task performance.
- These bounded skill/context/chat workflows do not measure general
coding quality. Unrepresented providers remain unqualified.
- Saved manuals and old Codex sessions are not automatically migrated.
Codex through ACP still has a separate base-instruction follow-up.

## Model Used

OpenAI Codex, GPT-6 family as identified by this session. The exact
deployment ID and context-window size are not exposed. The assistant
used reasoning, repository tools, code execution, and delegated PR/eval
work.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (relevant suites and all
three unchanged narrow reruns pass; complete-run timing failures
retained in Verification)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green on the new repair head
(prior-head checks retained above)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
on the new repair head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 12:32:42 -05:00
DottaandPaperclip 78e0034498 fix(evals): account for hiring completion notifications (#15007)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Product E2E evals check real hiring and delegated task completion.
> - The hiring fixture requires three requested CEO turns and two coder
executions.
> - The server can also wake the CEO when each delegated task completes.
> - Two exact-five-run guards rejected these valid completion turns in
all four retained cells.
> - This pull request validates bounded completion turns in both guards.
> - The benefit is accurate workflow grading while all actual runs and
coverage failures remain visible.

## Linked Issues or Issue Description

Refs: #14985, #14948, #14961.

**What happened?**

The original hiring comparison reports Codex Fail → Fail and Claude Fail
→ Fail. Each cell has seven successful runs. The five requested work
turns are accompanied by two server task-completion notifications. All
six other delivery checks pass.

**Expected behavior**

Require exactly three distinct user-requested CEO turns and one coder
execution for each of two known tasks. Admit at most two strictly
attributed server completion turns, including one turn that batches both
tasks. Reject unknown, duplicate, failed, retried or extra-work runs.

**Steps to reproduce**

Inspect the retained four-cell report linked below. Each original result
fails `five-successful-turns`. The same exact count was also enforced by
the final chat-flow guard.

## What Changed

- Add one typed lifecycle helper shared by the hiring scorer and the
hiring-only final chat guard.
- Validate public run ledgers, company/user/account identity, request
attribution, task origins, completion deliveries, timing and replies.
- Keep exactly five required work turns; declare seven maximum total
turns for cost and timeout planning.
- Count all actual runs, including notification runs and unexpected
resets. Keep other chat count guards unchanged.
- Version the hiring grader as v3 (turn accounting v2) and include the
helper and chat guard in its definition digest.
- Keep source-read, exact coder-body and all six other delivery checks
unchanged.
- Add 144 focused helper/scorer/settlement calibrations and separately
versioned exact retained-input replay reports.
- Retry complete bracketed observations, await both owed callbacks and
attributed replies, and refresh the final guard consistently.
- Reject unrelated completion writes and failed mutation attempts using
exact canonical/native action IDs. Missing identity mapping is
uncomparable action coverage.

## Verification

- All 977 credential-free E2E support tests pass across 64 files,
including 144 focused lifecycle/action/scorer/settlement calibrations.
- E2E typecheck, ordinary plugin SDK and Runner TypeScript dependency
builds, capability contract/inventory checks and the existing two-cell
hiring discovery pass.
- [Executable replay
report](https://github.com/paperclipai/paperclip/blob/fed1729018cc100f5f4bbfb692777e49009c423b/doc/plans/2026-10-02-hiring-executable-accounting-replay.md)
pins current code revision `e4077ade1818d98b9862ae79ee1d49a007dcf9c1`,
v3 definition digest, exact source/input hashes and each original/new
check.
- The stricter replay verifies both Codex variants through both
executable guards. ACPX Claude action attribution remains
unresolved/uncomparable because provider execution IDs cannot be exactly
joined to native request IDs; guards fail closed. No notification writes
are observed. All six other outcomes and every original source/template
coverage check stay unchanged. Original files and Fail → Fail machine
verdicts remain preserved; zero providers are called.
- Full attempts remain uncomparable in both profiles. Historical Claude
also keeps its six-backtick exact-template mismatch. This grading repair
does not prove model-performance equivalence.
- The limited sidecar-v1 and initial executable-v2 passes checked
notification-created tasks but could miss unrelated document writes.
Those assessments remain preserved and do not prove harmless
notifications. The stricter v3 replay is separate.
- [Original measurement and separate
sidecar](https://github.com/paperclipai/paperclip/blob/8eb517ca1497687237163bdef4dfc4d3332ea916/doc/plans/2026-10-02-hiring-template-live-comparison.md)
retain 28 actual runs, eight automatic notifications, four successful
cleanups and unknown actual model charges. No models are rerun.
- The branch is replayed on master `59c07ede7`. Intervening master
changes are UI-only; eval source bytes and replay verdicts match. The
four-cell provider-free replay was repeated against the reachable code
revision.
- Initial-head normal CI retained browser failures in agent-run denial
feedback and touch-picker scroll position. Those browser paths and
imports were unchanged, but their cause was not established. The
necessary review-fix head passes both browser checks; no blind rerun was
requested.
- Local full repository typecheck/test/build were not repeated.
Exact-head normal CI passes the required repository gates, including
typecheck, tests, build and browser shards. Fresh Greptile review
completed on `fed1729018cc100f5f4bbfb692777e49009c423b` with 5/5 and
zero unresolved threads. An independent rerun of the 144 focused
helper/scorer/settlement tests passes on the unchanged head.


**Merge readiness:** This PR repairs the evaluator. Its positive and
negative calibrations pass, both guards reject missing action
attribution, current-head CI and review pass, and there are no merge
conflicts. The retained ACPX cells remain uncomparable because their
action IDs cannot be joined. That coverage limit remains a separate
follow-up; it does not require relaxing this grader or changing the old
results. No model calls, production instructions, carrier changes, or
historical regrades are part of this readiness update.

## Risks

- Missing or inconsistent public lifecycle evidence fails the bounded
helper. The focused calibrations reject plausible false positives and
malformed observations. Unmatched action IDs fail closed and are
reported as uncomparable rather than a model task regression.
- Source-read evidence remains incomplete. This PR does not change
provider event carriers or relax the coverage oracle.
- The versioned count check differs from original v1 results. Reports
retain both versions and exact input hashes.

## Model Used

OpenAI Codex, GPT-6 family as identified by this session. The exact
deployment ID and context-window size are not exposed. The assistant
used reasoning, repository tools, code execution and delegated
calibration work.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip normal CI gates are green (exact head
`fed1729018cc100f5f4bbfb692777e49009c423b`; fresh review tracked
separately below)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(completed exact-head review; zero unresolved threads)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 06:42:51 -05:00
DottaandPaperclip 862a5758ba fix(agents): reduce hiring templates to role descriptions (#14985)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for work.
> - New agents receive role instructions from onboarding, the hiring skill, or a team package.
> - These sources repeat harness procedures and impose generic work policies.
> - They can crowd out the task and the harness instructions.
> - This pull request reduces those sources to short role descriptions.
> - It preserves configuration, skills, authentication, reporting lines, and approval controls.
> - The benefit is less repeated instruction text with explicit coverage for the default hiring path.

## Linked Issues or Issue Description

Refs #3307. The CEO template can impose a fixed delegation route instead of letting the agent choose how to fulfill the request. This change removes that route. It does not implement autonomous goal selection.

Related work: #14920 preserves stock Codex base instructions. #14948 reduces the generic manual and shared runtime prompts. #14961 improves native completion-tool descriptions. This PR is separate from those changes.

## What Changed

- Select only the short CEO `AGENTS.md` for new default CEO bundles. Keep the three former companion files as compatibility assets.
- Reduce the first-agent chief-of-staff prompt and coder, QA, UX, and security role examples.
- Reduce seven bundled team role bodies. Preserve their role, reporting, and skill metadata. Regenerate the catalog.
- Make hiring examples optional. Replace the long generic role manual with short role drafting guidance. Preserve explicit requester instructions.
- Add configuration and import coverage for native and legacy managed bundles, custom instructions, first-agent rendering, and catalog contents.
- Add an explicit-only hiring eval that starts from the production CEO default and checks one coder hire, independently computed JSON output, saved instructions, and worker reuse.
- Include the full prompt comparison and a separate three-request drafting simulation. Neither is a live provider comparison.

Prompt differences: [before and after](doc/plans/2026-10-02-hiring-template-prompt-diff.md). The CEO default falls from 1,897 to 20 words. The coder example falls from 652 to 18 words. Word counts describe instruction size, not outcome quality or billing.

## Verification

- PASS: 99 focused server tests and eight shipped-catalog tests.
- PASS: catalog generation and validation for four shipped teams.
- PASS: hiring skill validation.
- PASS: `pnpm -r typecheck`.
- PASS: `pnpm build`.
- INCOMPLETE: the full local `pnpm test:run` was stopped before rebase. Its original log is retained. This is not a completed full-suite pass. The full current-head GitHub CI workflow passed: https://github.com/paperclipai/paperclip/actions/runs/37073372419.
- PASS: `pnpm test:e2e:runner:typecheck` and `pnpm test:e2e:runner:unit` (63 files / 843 tests).
- PASS: discovery for the two new hiring cells, 50 existing everyday cells, and the full 438-cell catalog.
- PASS after rebase: 99 server tests, 11 catalog tests, 62 selected E2E support tests, and the E2E typecheck.
- PASS: all current-head PR checks at `57dcee147ed0b2d2e3cc657cd9e50fb16bf9ec25`: 51 successful check runs, two intentional Storybook skips, and successful Snyk status. Fresh Greptile is 5/5 with zero unresolved threads.
- PENDING follow-up: matched live hiring runs on frozen integration refs. No live outcome-quality or non-regression result is claimed from the configuration checks or this merge.

The new suite has two local native cells: Codex and ACPX Claude. It expects five provider turns per cell. It compares source-derived bundles, so the historical long templates remain admissible. Missing successful source-read receipts make a pair uncomparable. They do not establish a behavior regression or equivalence.

## Risks

- New default roles have fewer prescribed procedures. Live checks must determine whether a removed instruction was needed for an outcome.
- Existing custom and saved bundles keep their contents. The retained companion assets avoid a source-file compatibility break.
- Specialized Summarizer, Reflection Coach, and Wiki Maintainer prompts remain unchanged. Their product contracts need separate review.
- The generic non-CEO fallback reduction is in #14948. This PR alone does not provide its eight-word fallback.
- Configuration tests and drafting simulations do not establish live outcome quality. QA, UX, security, and chief-of-staff hiring behavior remain outside the new two-cell comparison.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code editing, shell tools, and delegated verification. The runtime does not expose the exact deployment model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context to this change
- [x] I have specified the model used (with version and capability details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before requesting merge

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-02 17:56:42 -05:00
DottaandPaperclip 408f70e69f fix(runner): preserve stock Codex base instructions (#14920)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native Runner connects Paperclip tasks to Codex app-server.
> - Paperclip passed its runtime context as `baseInstructions`.
> - That field replaces the stock Codex base prompt.
> - This pull request sends the same Paperclip context as additive
developer instructions.
> - Codex keeps its stock prompt and still receives Paperclip task
instructions and tools.

## Linked Issues or Issue Description

**What happened?**

The native Codex driver and Rust provider sent Paperclip context through
`baseInstructions` on thread start and resume. Codex used this text in
place of its stock base instructions. Direct-chat resume also sent an
empty replacement base. The Runner Lab session path used the same
replacement field.

**Expected behavior**

Codex should retain its stock base prompt. Paperclip should add its
runtime context through `developerInstructions`. Other provider facades
should retain their current instruction handling.

**Steps to reproduce**

1. Create a native Codex session through Paperclip Runner.
2. Inspect the `thread/start` request in the native provider trace.
3. Resume the session and inspect `thread/resume`.
4. Before this fix, these paths set `baseInstructions`. After this fix,
the Codex paths set `developerInstructions` and omit `baseInstructions`.

**Paperclip version or commit**

Reproduced against master at `cad26c6bfb736039c8ed5743da650a44792a083c`.

**Deployment mode**

Built from source. Native Codex app-server and runnerd paths. A local
protocol probe used codex-cli 0.153.4 and a localhost Responses stub.

No duplicate fix or matching public issue was found in the GitHub
search.

## What Changed

- Send additive developer instructions on Codex start and resume in the
TypeScript driver, Rust provider, and Runner Lab session path.
- Carry the additive fragment through runnerd, including runtime asset
path mapping.
- Preserve existing instruction fields for other provider facades,
including OpenCode.
- Add start/resume/direct-chat regression coverage and check the actual
Rust provider request.
- Document the historical option and trace field names. Record progress
and follow-ups in the working checklist.

## Verification

- `pnpm -r typecheck` — passed.
- `pnpm build` — passed.
- Targeted Codex driver lifecycle, driver, and live-session Vitest
suites — 139 tests passed.
- `cargo test --manifest-path
packages/paperclip-runner/runner/Cargo.toml --locked -p
paperclip-runner-core --test codex_provider` — 91 passed, 2 ignored
subprocess helpers.
- Real app-server probe: a localhost Responses stub captured identical
14,732-character stock base instructions on fresh start and cold resume.
Both requests retained the Paperclip marker in developer input. Both
stub turns completed. No paid inference was used.
- Runnerd transport Vitest suite — 182 tests passed.
- The initial `pnpm test:run` attempt reported local dependency-loading,
embedded PostgreSQL startup, and macOS `/var` versus `/private/var` path
failures. It was stopped after those failures. Loading-suite reruns
passed 1,428 tests after the build; native interaction/finalization
reruns passed 38 tests. A seven-suite diagnostic rerun passed 463 tests
and isolated the remaining path and PostgreSQL setup failures.
- With `TMPDIR=/private/tmp`, workspace, gateway, interaction, and
attachment suites passed all 356 tests. The remaining environment-image
and native-session-resumption suites passed all 44 tests with the same
canonical temp path. All affected suites passed on rerun. The original
full local command was stopped after failures and is not claimed as
passing.
- All 55 PR checks passed at `83281439456181396f3707eecda5d2ebc90bd14d`.
Greptile scored 5/5 with no open review threads.
- No paid live campaign or Product E2E browser suite was run. This
change has protocol and regression coverage; it does not claim improved
task quality.

## Risks

- Stock Codex behavior may differ from behavior under the previous
Paperclip replacement prompt. Restoring that behavior is the intended
change.
- Existing Codex threads retain their saved replacement base prompt.
They need a provider session reset to receive the stock base. This PR
does not reset active sessions or alter recovery rules.
- The legacy `baseInstructions` option and trace field names remain for
compatibility. They now describe the additive Paperclip fragment for
Codex.
- The separate Codex-through-ACP dependency patch remains a follow-up in
the harness coverage checklist. This PR covers native app-server
execution.

## Model Used

OpenAI Codex, GPT-6. The exact runtime model variant and context window
are not exposed in this session. Used reasoning, repository inspection,
code editing, shell execution, and test tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 08:53:47 -05:00
DottaandPaperclip d72389bee2 feat: add Browser Use Cloud connector and live task browsers (#14627)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Apps gateway gives agents governed access to external tools.
> - Browser Use Cloud can run browser work, but a tool result alone does
not let a person watch or take over.
> - A task needs a durable browser session, a visible viewer, and
recorded costs.
> - This pull request adds a Browser Use Cloud v4 connection and
interactive browser tabs on tasks.
> - People can follow the work, interact with the page, and retain the
browser after the agent finishes.

## Linked Issues or Issue Description

**Problem or motivation**

Agents need governed access to Browser Use Cloud. People need to see and
interact with the same browser from the task. A browser must remain
available after a run finishes and appear at the correct point in the
task feed.

**Proposed solution**

Add a native REST connection for the v4 API. Bind each session to its
company, task, agent, and credential grant. Open its interactive viewer
in the task side panel. Record provider costs as financial events. Use
`browser-use-cloud` as the app and connector key. Keep its skill with
the connector and deliver it only with authorized connection tools.

**Alternatives considered**

A v3 MCP connection would expose tools without the v4 lifecycle
integration. An external viewer link would leave the task. A fixed
viewer size would prevent pages from responding to changes in the task
pane.

**Roadmap alignment**

This extends the governed Apps gateway and Connected Apps roadmap. It
uses the existing task, grant, secret, approval, and financial records.
The work was requested by the maintainer. A search found no duplicate
Browser Use connector PR or issue.

## What Changed

- Add the Browser Use Cloud app, brand asset, API-key connection, and
profile settings under the `browser-use-cloud` key.
- Bundle the `browser-use-cloud` skill with the connector. Keep it out
of global `skills/` discovery. Deliver it only with authorized task/run
connection tools. Remove retired connector skill keys from runtime
overlays and preserve unrelated browser skills.
- Expose seven v4 tools through the governed gateway and deliver them to
native and CLI agents.
- Persist sessions, browsers, runs, event and recovery cursors, shutdown
leases, and cumulative cost accounting. Recover uncertain paid starts
without replaying them.
- Enforce task ownership, credential grants, approvals, revoked access,
and budget limits.
- Add interactive task browser tabs and compact chronological feed
entries. Retain the viewer across tab switches and keep visible idle
browsers open.
- Add debounced automatic viewport fitting, standard size presets, and a
viewer ownership lease.
- Add lifecycle, authorization, accounting, viewport, UI, and Storybook
coverage.
- Add an idempotent database migration after the current master
migration. Preserve deployed migration hashes. Migrate pre-release Cloud
connection and financial keys without replacing grants, credentials, or
browser history.
- Document provider behavior, live acceptance results, and the lack of
documented passkey forwarding.

## Verification

- Full workspace typecheck and production build pass on the updated
branch.
- Token gates, brand asset validation, module boundaries, and migration
ordering pass.
- Cloud tests verify global skill exclusion, authorized task/run
delivery, unassigned agents, disabled connections, revocation, adapter
isolation, and secret exclusion. The existing AgentMail connector
assignment test also passes.
- Migration replay runs twice against existing browser work and
financial records. It preserves the records and avoids duplicate costs.
- The focused provider, app catalog, OpenAPI, connection gateway, and
migration regression suites pass. Recovery coverage includes lost
replies, process crashes, provider rejection, and browser arrival
acknowledgement.
- All 54 checks pass on `2974b5f03641ad0cea3c941d8c02579316fa8c92`,
including the full test matrix, browser E2E shards, build, typecheck,
security, and release canary. Two optional Storybook jobs are skipped.
- Greptile is 5/5 on the same commit, with zero unresolved review
threads. The corrected review uses the actual master-to-head diff.
- The local `pnpm test:run` started and was stopped after the full CI
matrix passed. It did not complete locally; the full-suite result above
comes from CI.
- Earlier live acceptance used an isolated company with a capped
provider credential. The agent opened paperclip.ing, the embedded viewer
accepted navigation, and the same browser stayed available after
completion and tab switches.
- The local Storybook build passes. Stories cover the panel, footer,
feed entries, settings, lifecycle failures, and viewport modes with an
offline viewer fixture.

## Risks

- Browser Use charges for hosted work. Provider caps and local budget
checks reduce exposure; reported costs can arrive after work completes.
- Viewer and CDP URLs grant access to the browser. The server validates
and restricts them. They are excluded from agent results and durable
event data.
- Runtime resizing of v4 agent browsers uses a provider option confirmed
by live testing but absent from its published agent schema. Resizing
during a click may invalidate coordinates. Fixed presets remain
available.
- Viewport ownership is process-local and resets on restart. The
lifecycle and accounting records remain in the database.
- The original intermittent embedded-viewer stall has not been fully
diagnosed. A bounded reconnect and active-session recovery cover the
observed failure paths.
- Live tests did not cover every revocation, approval, rate-limit, or
restart case. Deterministic integration tests cover those paths. Passkey
forwarding is not claimed.
- Unknown create outcomes keep the credential available for cleanup.
Run-list absence cannot prove a paid POST was rejected, so recovery
stays pending until it can identify provider work.

## Model Used

OpenAI Codex, GPT-6. Used reasoning, repository search, code execution,
browser interaction, and test tools. The exact serving model ID and
context-window size were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 07:13:16 -05:00
DottaandPaperclip 83076d7e7c feat: return completed handoffs to Agent Chat (#14408)
Return completed Agent Chat handoffs through a durable outbox and scope each generated update to its supplied tasks. Add recovery, browser delivery, result access, and calibrated quality coverage.

Validated with two consecutive ten-case Claude/Codex campaigns, all CI checks, and a 5/5 review.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 10:25:52 -05:00
DottaandPaperclip 24beb00575 feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification.

Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 08:56:21 -05:00
DottaandPaperclip 992f720262 fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English
(ASD-STE100). -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task descriptions, comments, continuation data, skills, and
execution rules enter several agent adapters.
> - The same source can be rendered by more than one automatic input
carrier.
> - Failed resumes can also rebuild input from stale or compact context.
> - This pull request gives each Paperclip-owned source one delivery
owner and preserves the required transport boundaries.
> - It adds deterministic adapter, interaction, runner, and browser
tests for these boundaries.
> - The benefit is more predictable context delivery with explicit
evidence for later live qualification.

## Linked Issues or Issue Description

Related: #13144 removes a duplicate environment payload and bounds wake
lists. Related: #11360 addresses Hermes resume behavior. This pull
request preserves compatible active-session formats while repairing
context ownership and stale question creation.

**What happened?**

Task descriptions and comments could enter more than one automatic
context block. Native transports could wrap a complete model input in a
second task envelope. Some legacy and gateway adapters could omit the
owned assignment on ordinary tasks or rebuild a failed resume with stale
compact context. A continuation could also request a question after
newer human comments had arrived.

**Expected behavior**

Each task or comment source has one automatic model-facing owner.
Distinct comment IDs and repeated wording remain distinct. Fresh
fallback attempts rebuild the required full context. A question request
is rejected when newer queued human direction makes it stale. Harness
access policy remains owned by execution configuration.

**Steps to reproduce**

1. Build a task with a description and current comments.
2. Capture the actual adapter or runner input.
3. Compare source ownership and task-envelope nesting.
4. Queue a human comment before a continuation requests a question.
5. Trigger a failed resume and inspect the fresh retry input.
6. Run the focused adapter, interaction, runner, and browser checks.

## What Changed

- Add shared prompt-section selection at the provider-attempt boundary.
- Deliver owned assignment context through native, legacy CLI, ACP,
gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and
Hermes paths.
- Rebuild full or compact context after resume recovery changes the
attempt. Add native and Claude ACP tests of actual recovery requests.
- Preserve custom templates, loaded instruction files, execution
policies, and older active-session formats.
- Record continuation source metadata and reject stale question creation
under the issue-row lock.
- Add explicit Product E2E context-integrity profiles, prerequisite
gates, credential-isolation checks, and report fixtures.
- Bypass service-worker forwarding for same-origin Vite development
modules. A real Chromium test fails with resource exhaustion before the
repair and passes after it. Production asset caching keeps its existing
policy.
- Add browser diagnostics and service-worker module-loading regressions.
- Add an explicit zero-retry eval option. The default retry behavior
remains unchanged. Each campaign records its effective policy.
- Remove the model-facing working-directory sentence from four prompt
builders. Existing workspace, sandbox, permission, and custom-template
configuration remains unchanged.
- Align the everyday workflow assertion with the current 47-entry
catalog.

Compared with current upstream master, the branch carries the
context-ownership implementation and its tests, the explicit
context-integrity catalog and evidence harness, and the focused browser
regression checks.

## Verification

**Merge assessment:** focused regression evidence supports merge. This
is not full completion of the original broad qualification matrix. The
maintainer has authorized merge after fresh verification of the master
integration.

- Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This
integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`.
All 14 conflicts are resolved. Cancellation checks, workspace
finalization, native Grok support, and both sets of tests are retained.
- Current-head Greptile: **5/5**, with no blocking findings. The review
names this exact commit. All **59 reported checks are terminal: 55
successful, 4 skipped, zero pending or failing**. This includes the full
root general and serialized suites, separate runner checks, typecheck,
build, canary, browser E2E, Docker, and security checks. The successful
legacy security status is included in that total.
- After integration: workspace typecheck and full build passed. Separate
runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust
tests, and 39 preparation checks**. Other passing checks include 621
Product E2E harness units, 376 focused shared/adapter tests, 160
real-database/API tests, 86 Hermes tests, 18 browser-support checks, and
Product E2E typechecking. The complete root suite passed in CI. The
duplicate local monolithic root run was stopped after that CI result; it
is not counted as a completed local pass.
- New native recovery coverage retains full assignment, completion
contract, and explicit skill selection after safe replacement, for old
and prepared input formats. Full native session test file: **136/136
passed**.
- New Claude ACP coverage captures actual fresh, resumed, and
missing-session fallback requests. It verifies one assignment copy,
comment order, identical text under distinct comment IDs, and full
fallback context. Full file: **33/33 passed**. Both affected TypeScript
checks passed.
- Existing deterministic tests cover source revisions, approval and
trust boundaries, completion validation, custom templates, compatible
sessions, standalone driver wrapping, and maintained adapter transport
requests.
- Provider-free browser support: **17/17 passed** after the master
merge. Service-worker unit tests: **33/33 passed**. The module-overload
regression failed before the repair and passed after it in real
Chromium.

### Fresh live comparisons

The new batch ran exactly four Product E2E attempts. **All four passed
on the first attempt; no retries.** Each has six terminal matchers plus
the existing browser lifecycle and invariant checks.

| Exact case ID | Control | Candidate |
|---|---|---|
| `core-compatibility.runner-codex.local.plan-revise-accept` | Passed |
Passed |
|
`local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume`
| Passed | Passed |

The plan case checks a revised canonical plan and revision-bound
approval before completion. The question case restarts the server before
submitting the answer, then verifies the continuation completes.

Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate
source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical
frozen definitions and provider versions: Codex `0.156.0` with
`gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with
`claude-sonnet-5`. The September 24 head added master browser recovery
and test-only changes. The September 28 head also integrates newer
master changes, including cancellation, workspace finalization, and
native Grok. These are frozen-source live results, not exact-head live
runs.

The candidate received one description copy where the control initially
received three. The submitted initial plan envelopes were 7,969 versus
19,097 characters. Question envelopes were 7,592 versus 18,919. These
are structural measurements, not whole-provider token or dollar savings.

### Earlier evidence and failed attempts

- The preceding fresh batch has four effective passing pairs: OpenCode
comment continuation and assigned skill, native Codex comment
continuation, and native Claude comment continuation. It retains **11
attempts: eight passed and three failed**.
- Original failures remain recorded: missing local PostgreSQL library
links before task creation; host-sleep cleanup after task/page checks
passed; and a Claude **control** session-open rejection before a model
turn. Setup was repaired identically on both worktrees. The permitted
unchanged infrastructure retries passed. The underlying Claude provider
startup error was not retained and remains unknown.
- Older R2 retains **17 passes and one failure** across 18 attempts,
including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode
blank-page failure led to the service-worker repair. R2 is historical
evidence: master changed the native fixed prompt and removed duplicate
wake environment data afterward.
- The September 24 CI run initially failed one unrelated preview
readiness test (`ECONNREFUSED` on its local fixture). Its test and
production code match master. Isolated local verification passed **28
tests, 3 skipped**. One unchanged CI retry passed the full shard: **831
passed, 1 skipped**, including all **31 preview-exposure tests**. The
aggregate CI gate passed afterward. The precise startup cause remains
unknown; a port race is a hypothesis, not a proved cause.

### Limits

The original wider profile/workflow matrix, repeated trials, and remote
Daytona qualification are incomplete. These results support a focused
merge recommendation, not statistical equivalence or universal harness
qualification. Some usage receipts are missing in both variants, so no
token or dollar savings are claimed. The $500 ceiling was preserved
using conservative allowances; failed attempts and unknown charges
remain in the ledger.

Reproduce the focused additions with `pnpm exec vitest run
packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm
--filter @paperclipai/paperclip-runner exec vitest run
src/native-session-runtime.test.ts`. Full checks use `pnpm -r
typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner
checks. Paid evals require the frozen definitions, profiles, and
credentials; do not use `--all` as a substitute for the selected cases.

## Risks

- Context placement changes can affect model behavior. Deterministic
checks cover the selected paths, but live qualification remains
incomplete.
- The stale-question guard can reject a request when queued human
comments arrived during the run. This is intended.
- New stored inputs and model envelopes retain compatibility readers for
older active sessions.
- Custom templates may intentionally repeat content.
- Removing a model-facing working-directory sentence does not change
filesystem, command, sandbox, or permission configuration.
- The worker bypass applies only to same-origin development module
paths. Cache-policy tests preserve private-response handling and
production asset caching. Mounted HTTP fixture changes remain test-only.
- This PR does not claim measured token savings or statistical
equivalence across every harness.

## Model Used

OpenAI Codex, exact model gpt-6-astra, with repository tools and code
execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The
serving context-window size is not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR using the required issue fields
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run the focused local checks and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect these changes
- [x] I have considered and documented risks above
- [x] All current-head Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
for the current head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:49:14 -05:00
DottaandPaperclip 96bf004a79 fix: use persisted state for lifecycle continuation and retry budgets (#13888)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its control plane decides when a task can continue, wait, stop, or
complete.
> - Legacy continuation could change when an agent changed its wording
without changing task state.
> - Shared attempt counts also let repair and infrastructure retries
affect each other's limits.
> - This pull request uses persisted state and separate, bounded
allowances for these decisions.
> - If automatic repair stops, the task explains what happened and
offers a guarded retry.
> - Paired tests and real-provider evaluations verify that Stop,
approvals, ownership, and spending limits remain authoritative.

## Linked Issues or Issue Description

Related work: Refs #13761, Refs #11126, Refs #13610. These cover
obsolete continuation dispatch and retry storms. Open and closed issues
and PRs were searched for related lifecycle, continuation, and retry
work.

**What happened?**
Legacy continuation depended on English wording and progress heuristics.
Repair, failure retry, and productive continuation could consume shared
counts. When bounded repair stopped, the task showed a technical
recovery message without a clear next action.

**Expected behavior**
Persisted disposition and owned execution paths determine the next
action. Missing disposition prompts bounded agent repair. Explicit work
mode determines planning mode. Narrative changes and raw activity counts
cannot replenish allowances. An exhausted repair shows a readable
notice. An explicit retry checks current controls and preserves the
assigned agent.

**Steps to reproduce**
Run `pnpm test:lifecycle-baseline`. The paired probes keep structured
state constant while varying completion, planning, blocker, and progress
prose. Run the explicit `lifecycle-baseline` and
`continuation-accounting` Product E2E suites for real-provider coverage.
In Storybook, open **Design previews / Recovery notice** to inspect the
production component's normal, pending, acknowledged, unavailable,
failure, and mobile states.

## What Changed

- Hide the image attachment button, icon, and drop/paste hint in answer
composers. Image paste and drop support remains available.
- Merge current master and retain both browser regression sets. Use a
production-stamped service worker in the offline recovery browser
fixture.
- Share one state-based legacy continuation decision across immediate,
delayed, and recovered dispatch. Bind bounded repairs to their source
run and episode.
- Remove title and description wording from work-mode authority. Agents
can still write requested plans in execution mode.
- Persist separate failure-retry and productive-continuation counters.
Disposition repair and resource waits cannot consume or reset those
allowances.
- Validate delayed repair identity, then recheck current gates before
provider dispatch. Fence native startup cancellation.
- Show **Agent needs attention**, a plain-language explanation, **Retry
agent**, and expandable details in both task interfaces. Report request
progress, acknowledgement, and errors inline.
- Store typed recovery notice metadata. Recognize older active notices
only through exact stored action and run IDs. Notice text never grants
retry authority.
- Use the existing recovery-action endpoint for retry. Recheck current
action, status, owner, agent availability, dependencies, active runs,
pending questions and confirmations, approvals, pause controls, and
budget. Duplicate requests do not wake twice.
- Add component, page, route, database, contract, and Storybook
coverage. Keep the scenario inventory and executable evals here.
Historical reports and snapshots live in the [commit-pinned
paperclip-evals
archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md).
- Preserve unsaved project fields while the same project URL changes to
its canonical alias. Do not reuse data across projects or companies.
This separate fix addresses the repeated repository-editor browser
failure without changing the browser test.
- Keep the development service worker from intercepting Vite module
reloads. Update the connection-intent browser fixture to record progress
and completion through the agent API.

## Verification

Merge preparation on September 25, commit
`c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`:

- Merged master `bd2030932` and resolved the browser test-list conflict
by keeping both sets of regressions.
- Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no
failures, skips, or missing selected evidence. Unit 423, runner 184,
database integration 397, grading 86.
- Browser support: 17/17 passed. The offline recovery test first failed
with an unstamped development worker, then passed with the production
stamp. Its assertions are unchanged.
- Focused interaction UI and offline fallback tests: 19/19 passed.
Verified the custom-answer composer in Storybook: no attachment controls
or hint; entering an answer enables Next.
- Recursive typecheck, production build, token gates, and diff checks
passed. The worktree is clean. No new real-provider campaign was run.
- Current CI and review: [Current PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243):
55 successful checks and two optional Storybook skips. Greptile scored
this exact commit 5/5. Hiding the question attachment controls is an
intentional UI change; paste/drop remains available.

Earlier recovery UI verification, commit
`21be0fec0e90e86b6d662b8ee4831847cd041cdb`:

- Recursive typecheck, production build, token gates, and diff checks
passed.
- Focused UI coverage: 338 tests passed across six suites (336 before
the interaction guard, with the two affected suites rerun at 149 passed
after it). Covers both task interfaces, the real page mutation,
pending/error acknowledgement, stale state, and unavailable controls.
- Recovery database integration: 352 tests passed before the interaction
guard. The complete recovery-action and mutation-route suites passed 181
tests after it. The two new pending question/confirmation regressions
failed before the fix and passed afterward, including
resolved-interaction controls. Shared validator suite: 31 passed. E2E
catalog suites: 34 passed.
- Browser inspection passed for light/dark themes, mobile layout,
expandable details, pending retry, acknowledgement, failure, and
disabled retry. Storybook renders the production component; its request
is simulated.
- The broad local run hit two chat callback-order wait failures and was
stopped after all CI unit/database/runner shards passed. Both local
failures passed when rerun without the competing full-suite process.
- CI exposed a repeated project-repository draft-loss race during
canonical redirects. A new unit regression failed before the fix; all
nine project-page tests now pass, including controls for other projects
and companies. Both unchanged repository browser tests passed against a
fresh local server. UI typecheck, production UI build, and token gates
passed after this fix.
- [Earlier PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798)
on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two
optional Storybook skips, and no failed or pending checks. The
repository browser shard passed with the production fix. Greptile is 5/5
on this exact commit with no unresolved review threads. The PR is
mergeable.

Historical, source-qualified lifecycle evidence:

- Lifecycle baseline: 1,074 assertions. Native session coverage: 447
tests. Product E2E support: 515 tests. Browser support: 11 tests. Full
earlier verification is retained in the archive.
- [Real-provider campaign: 8/8 passed, zero
retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html),
source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup
checks passed. This includes deliberately exhausted repair cases that
correctly remain blocked; it does not mean every task finished Done.
This campaign predates the recovery UI change.
- Archive migration verified all 16 original JSON files byte-for-byte
and all 24 checksum entries. App tests do not need private archive
access. [Archive PR
#27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged.

## Risks

- Agents that omit durable disposition receive at most two repair
attempts by default. Prose-only completion exposes missing state rather
than silently changing scheduling.
- A retry is an explicit board action. The server rechecks current
controls. A successful response confirms the task returned to To do; it
does not claim that the provider has already started.
- Existing notice metadata remains valid. Only older active notices with
matching structured evidence receive the new UI. Historical notices
without that evidence keep their existing rendering. No schema migration
is required.
- Old run records require conservative retry accounting. Tests cover old
counters, alternating retry lanes, restarts, and exhausted repairs.
- Historical snapshots require private `paperclip-evals` access. The app
index retains public campaign links. Live campaigns qualify specific
sources and scenarios; no new real-provider campaign has run for the
recovery UI commit.

> This fixes existing lifecycle and recovery behavior and does not
duplicate planned core work.

## Model Used

OpenAI GPT-6 through Codex assisted implementation, reasoning, code
execution, and review. The exact serving model ID and context window are
not exposed in this task. Historical real-provider evaluations used
Codex model `gpt-5.6-sol`, separately from the implementation assistant.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 15:28:11 -07:00
DottaandPaperclip aa8fc86331 feat(connections): prefer native apps and ask users to choose external providers (#13941)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents governed access to external services.
> - Native connections should remain the first choice for a supported
app.
> - Other apps may be available through an external MCP provider.
> - The user must know which external provider handles the connection
and choose it before setup.
> - This pull request adds ranked alternatives and server-authored
instructions to connection search.
> - Agents can follow the returned instructions while Paperclip
validates saved choices and access.

## Linked Issues or Issue Description

Related: #13879, which fixed inline MCP provider setup. This PR adds
discovery and provider selection on top of that work.

**Subsystem affected**

Cross-cutting: shared connection contracts, server search and intent
services, native runtime, CLI, inline setup UI, and evals.

**Problem or motivation**

An agent cannot offer a clear external-provider choice when Paperclip
has no native connection for an app. Adding provider-specific branches
to the core prompt would make those instructions harder to maintain.

**Proposed solution**

Prefer a native connection. Otherwise return verified alternatives in
Composio, Arcade, Executor, Zapier order. Include an external-service
disclosure, a question with None, and the next instruction in the search
result. Validate the saved human choice before creating a selected
fallback setup card. Reuse existing provider accounts and verify
underlying app access separately.

**Alternatives considered**

Do not silently choose a provider. Do not claim that broad execution
tools prove support for every app. Reuse existing questions and
connection intents rather than add another connection model.

**Roadmap alignment**

Extends the existing MCP Tool Gateway & Apps and Agent evals & feedback
capabilities. The MCP aggregators experiment remains the gate. No
duplicate provider-routing PR was found in the public search.

## What Changed

- Add a dated support index and authorized cached-tool evidence for
external routes.
- Return provider questions and next-step instructions from
`connections_search`.
- Preserve pending choices and declines across continuation. Validate
company, task, agent, human, app, and current route eligibility.
- Carry the selected app into new setup and account reuse, validate
explicit provider requests against persisted human messages, and
distinguish provider readiness from app authorization.
- Sync native, MCP, REST, and CLI contracts. Keep core agent
instructions provider-neutral.
- Add production-component Storybooks, focused database tests, and three
real-agent browser eval cases.
- Record the plan, observed failures, fixes, passing evidence, and
acceptance limits.

## Verification

- Latest head `586f0e6cd`: 54 checks passed, 2 skipped; Greptile 5/5 and
all review threads resolved.

- After rebasing on master `18dac1e1e`: 64 focused shared, validator,
route-contract, and database tests passed; server typecheck passed.
- Embedded-browser test drive on the rebased head: native Jira card,
HubSpot external-provider question, Arcade account reuse, one actual MCP
read against a local synthetic fixture, reload persistence, and None
preventing further calls. A real OpenAI-backed agent performed discovery
and continuation.
- UX observation: the agent initially combined mutually exclusive
request fields; the server rejected it and the agent recovered without
changing access. This extra retry remains visible in the transcript.
- After rebase: 23 focused eval grader/catalog tests, affected
TypeScript checks, token gates, production UI build, and Storybook build
passed. Full local tests are intentionally excluded at the maintainer's
request.
- Before rebase: four browser/real-agent attempts passed: native Jira,
None, and reuse of the second provider on two Codex profiles.
- Browser evals used an isolated deterministic MCP fixture through the
real Paperclip gateway. They do not prove production compatibility with
all four providers.
- Review `Apps / Connections / Provider choice` in Storybook. Choose
Arcade, continue through Access, and verify the app name,
external-service disclosure, and URL configuration.
- The detailed verification report is
`doc/connections/2026-09-23-aggregator-routing-verification.md`.

## Risks

- The public support index is finite and can age. Account capability and
app authorization still require verification after selection.
- Existing installed-tool permissions remain in effect. Provider choice
is not a new execution permission boundary. An early Mini attempt
skipped search; clearer provider-neutral instructions made the targeted
rerun pass. This is not a measured reliability rate.
- Explicit requests skip provider confirmation only when a clear
persisted human message or saved provider choice supports them. Other
phrasing falls back to confirmation; the agent query alone is not
consent. These routes do not add tool permissions.
- No database migration or legacy Composio broker is added.
Real-provider acceptance remains separate from fixture proof.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser tools. The exact deployment variant and
context-window size are not exposed in this session. Product evals
separately used the repository's primary Codex and Codex Mini profiles;
those agents supplied test behavior, not independent provider
compatibility proof.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 16:36:49 -05:00
DottaandPaperclip 24429024e7 feat: add Fireflies connector and summary-ready routines (#13890)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps gives agents governed access to external tools through stored
credentials.
> - Routines start work when an external service sends an event.
> - Fireflies provides meeting transcripts and summaries through an
official hosted MCP server.
> - This PR adds that connection and accepts signed meeting events
through the shared app webhook flow.
> - Agents can review completed meetings with the same permissions and
audit records as other work.

## Linked Issues or Issue Description

**Problem or motivation**

Operators need agents to read Fireflies meetings and start follow-up
work when a summary is ready. The Apps catalog lacks Fireflies. The
shared app webhook flow needs to accept its signed deliveries.

**Proposed solution**

Use the official Fireflies MCP endpoint with OAuth or a vaulted bearer
API key. Extend the existing Another app or script flow with signed
webhook support. Verify the raw-body signature and pass the JSON payload
as external data. Select Meeting Summarized in Fireflies. Deduplicate
identical signed deliveries, including setup deliveries.

**Alternatives considered**

A separate REST connector would duplicate the governed MCP path.
Polling, legacy V1 payloads, and automatic provider-side webhook
registration are outside this change.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps and Scheduled Routines
surfaces. It adds a provider to those systems. It does not introduce a
second integration framework.

**Additional context**

A GitHub search found no existing Fireflies issues or PRs. Provider
references and verification limits are in
`doc/connections/FIREFLIES.md`.

## What Changed

- Add the official Fireflies catalog definition, generated registry,
provider evidence, and branded artwork.
- Reuse Access → Connect, dynamic discovery, Permissions, vault storage,
policy, and audit behavior.
- Classify Fireflies sharing, movement, and access revocation as writes.
- Preserve Off and Ask first restrictions during OAuth reauthorization
and API-key replacement. New actions retain normal defaults.
- Add `app_webhook` authentication to the shared Another app or script
flow. Accept bearer tokens or raw-body HMAC-SHA256. Preserve earlier
`fireflies_hmac` triggers and revision snapshots for compatibility.
Existing text columns need no migration.
- Verify `X-Hub-Signature` or `X-Hub-Signature-256` against the exact
request body. Preserve generic event payloads and deduplicate identical
signed requests.
- Keep the routine wizard generic. Show one webhook URL and secret in
Another app or script. Keep all new app webhook event names
provider-neutral. Keep provider setup instructions in the connector
documentation.
- Pass generic webhook JSON to the task in an explicit external-data
block, capped at 16,384 characters. Keep strict meeting validation for
existing legacy Fireflies triggers.

## Verification

- Feature implementation commit `0882dc8a1`: all 54 CI checks passed;
two conditional Storybook checks skipped. This includes full tests,
typecheck, build, browser E2E, canary dry run, and security checks.
Greptile rated this commit 5/5; all review threads are resolved.
- Full local `pnpm -r typecheck`, `pnpm build`, and token gates passed
on the final code. Targeted connector, gateway, webhook, revision, and
UI suites passed during implementation. After the provider-neutral
follow-up, all 84 app-webhook and routine-service tests passed; the
final payload-to-task assertion also passed in the 72-test routine suite
and a clean-config rerun.
- The long local `pnpm test:run` invocation started before the final
edits and was stopped after the final-commit CI suites passed. It
reported one generic webhook test failure while those files were
changing; that test and the entire routine suite passed on the final
source, including a clean-config reproduction. The interrupted local run
is not counted as a full-suite pass.
- In the embedded browser, completed official OAuth consent and
discovered 20 live actions. Real meeting listing, transcript retrieval,
and summary/action-item retrieval succeeded as the selected agent.
Turning a live read Off blocked its test; catalog refresh preserved the
restriction.
- Embedded-browser Another app or script setup, back/save/resume, narrow
layout, and a signed synthetic Fireflies delivery succeeded. The UI
reported authentication passed without creating a task. Fixtures cover
signature tampering, malformed requests, ordinary app event names,
duplicate/setup deliveries, rotation, revisions, pause/archive, and
company isolation.
- Existing MCP browser suite: 8 passed and 2 provider-dependent cases
skipped. Branding checks passed; connector artwork and webhook setup
were checked at desktop/mobile widths and in light/dark modes.
- An unauthenticated POST to a correctly formatted public webhook URL
reached the staging tenant verifier through the existing Cloud gateway.
- A real Fireflies webhook delivery remains unverified. A staging
callback is available for the operator walkthrough. Live API-key
authorization, credential expiry, and a new meeting's summary completion
were not tested against the provider. Fixtures cover these protocol and
lifecycle paths where applicable.

- Storybook follow-up `c54174faa`: 27 production-component stories cover
every UI change, with a source-to-story map in the connector
documentation. Static Storybook build, UI typecheck, token gates, and
Playwright checks for all stories and the mobile footer pass. All PR
checks passed for this Storybook follow-up; Greptile reviewed
`c54174faa` at 5/5.

## Risks

- Fireflies may change its hosted MCP tools or OAuth behavior. Tool
discovery stays dynamic. Experimental search/fetch tools are not
required.
- Public webhook setup requires HTTPS and a separate signing secret.
Fireflies normally emits events for meetings owned by the configuring
account.
- Reauthorization touches shared MCP permission code. Regression tests
cover existing restrictions, new actions, connection removal, and other
gateway callers.
- Webhook receipt grants no tool access. The routine agent still needs
an authorized Fireflies connection.

- New generic triggers rely on provider event subscriptions. Without a
sender-supplied idempotency key, changed request bytes count as a new
event. Existing legacy Fireflies triggers retain summary-only filtering
and per-meeting deduplication.

## Model Used

OpenAI Codex, model `gpt-6-astra`. Used reasoning, repository editing,
code execution, and embedded-browser testing. The runtime did not expose
a context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 17:11:16 -05:00
DottaandPaperclip a10702a878 feat(slack): add governed tools for Slack-origin tasks (#13828)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors let people start and continue agent tasks from other
services.
> - A Slack conversation needs access to its surrounding discussion and
Slack collaboration tools.
> - The agent must use the linked requester's access and keep private
material within its permitted audience.
> - This pull request adds Slack tools through the existing connector
contribution and approval framework.
> - People can ask an invited bot to read a discussion, create follow-up
tasks, and collaborate in Slack.

## Linked Issues or Issue Description

**Subsystem affected**

Chat connectors, connector runtime, tool gateway, and connection
Settings/Access.

**Problem or motivation**

Slack-origin tasks can receive messages but cannot inspect the rest of a
channel or act through the originating bot. People must paste context or
configure a separate integration.

**Proposed solution**

Supply typed Slack tools and a bundled skill only to the originating
task and assigned agent. Resolve the linked requester on the server.
Check bot and requester access before reads and writes. Use existing
durable actions and approvals. Retrieved messages remain source
material.

**Alternatives considered**

Slack's user-OAuth MCP server does not replace the customer-created chat
bot. An unrestricted Web API proxy would not provide suitable permission
or publication boundaries.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps work with a provider
contribution. It does not add a task dispatcher or a separate Slack task
lifecycle. Related: #11144 covers generic per-user MCP grant execution;
this change binds Slack bot operations to chat-origin tasks.

## What Changed

- Add 39 typed Slack tools, a method/scope matrix, a bundled skill, and
shared native/HTTP execution.
- Bind tools to company, endpoint, task, run, assigned agent, and
admitted linked requester. Check membership and revocation on each call
and before queued writes.
- Add paginated reads, bounded history search, source links, messages,
file uploads, reactions, pins, bookmarks, topics, canvases, lists, and
approved channel operations.
- Restrict private-source publication, including automatic replies and
uploaded deliverables. Keep other people's bot DMs inaccessible.
- Reuse action receipts, idempotency, approvals, and reconciliation.
Suppress an identical explicit-send/final-reply duplicate. Return
governed results through their verified originating conversation.
- Add endpoint-bound personal search OAuth storage and lifecycle. Keep
native real-time search disabled until a runtime meets Slack's
transient-result requirements. Current runtimes use bounded history
search.
- Show capabilities, scope upgrades, and personal search authorization
in Settings/Access and Storybook. Document provider and runtime limits.

## Verification

- Current head `0eb21cba4`: CI checks pass and Greptile is 5/5 with no
unresolved findings. One unchanged rapid-callback timing test passed on
a single CI retry.
- Approval presentation regressions cover board-comment precedence and
exact Slack publication; the expanded database assertion passed in CI.
The local PostgreSQL startup probe later became unavailable, so that
final assertion was verified in CI. Slack setup and failed-run retry
browser tests also passed locally.

- Full workspace typecheck and build passed. Server typecheck/build
passed again after the approval routing fix.
- Broad local suites passed in separate groups: server 12,958 tests, UI
6,555, shared 770, skills catalog 20, and other workspace packages
2,652. CLI and serialized server checks passed after environment/timeout
retries. These are composite results, not one uninterrupted green
full-suite invocation.
- PostgreSQL authority regression covers admitted identity,
cross-company/task/agent rejection, recovery, retained-session
revocation, OAuth refresh/disconnect races, approval execution, exact
publication lineage, retries, uncertain sends, and duplicate
suppression.
- Gateway/response regressions cover separate-origin approval batches
and durable continuation. Focused provider, access, search, native
runtime, route, and AgentMail regressions pass.
- Storybook capability, missing-scope, OAuth configuration,
authorization, and disconnect states were inspected in the browser.
- Live staging: read a channel decision and full thread, create exactly
two assigned backlog tasks, add a reaction, paginate discovery to
exhaustion, and return bounded search matches with source links and
coverage.
- Live staging: create/edit/read a canvas and list, inspect the canvas
in Slack, post/edit one message, and create a channel only after
approval. New channels remain disabled for responses.
- Live staging: read a response-disabled channel from the requester's
DM; writes to that channel were denied. The test setting was restored.
- Final live retest passed: explicit file upload and exact content
read-back; approved deletion of only the disposable bot message;
continuation confirmation returned to the original Slack thread without
repeating the action.
- Optional OAuth, private multi-user boundaries, native RTS, and CLI
provider execution are not fully live-qualified. The staging agent
initially supplied malformed tool arguments; valid arguments succeeded,
and the tool/skill descriptions now emphasize UUID write keys.

## Risks

- Existing Slack apps must add scopes and reinstall for new
capabilities. Provider plans and document permissions can still restrict
operations.
- Instances need an independent `PAPERCLIP_TOOL_ACTION_SIGNING_SECRET`
for governed tool actions. The staging instance was configured with
explicit operator approval; fleet provisioning is a separate gap.
- Native RTS is not exposed on current transcript-retaining runtimes.
Bounded history scans are deliberately reported as incomplete. Inline
file reads support text/canvas content up to 256 KiB; other types return
metadata.
- Private document edits fail closed when the full audience cannot be
verified. Uncertain effects other than posts/uploads require inspection
instead of blind retries.
- Shared approval-delivery code now separates outcomes by source run to
preserve origin boundaries. No database migration is required.
- A separate completion-validator gap remains when the agent cites a
prior run's registered artifact during finalization. It asked for
registration again even though Slack delivery was confirmed. This change
does not add a connector-specific task-completion policy.

## Model Used

OpenAI GPT-6 through Codex, with repository tools, code execution, and
browser testing. The exact deployed model identifier and context-window
size were not exposed in the session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 15:29:07 -05:00
DottaandPaperclip 8813a50105 feat: run GitHub review bots through Paperclip agents (#13717)
## Thinking Path

> - Paperclip manages agent work as tasks and runs.
> - GitHub chat brings repository conversations into those tasks.
> - A review bot needs the assigned agent, its authority, and governed
provider tools.
> - The existing channel connection did not supply that review workflow
or a complete setup journey.
> - This pull request adds GitHub App setup, account access, event
prompts, task-bound review tools, and exact-commit checks.
> - Operators can inspect each review through the same task, run, and
activity systems.

## Linked Issues or Issue Description

**Subsystem affected**

GitHub chat, governed connection tools, task execution, shared/database
contracts, and connector setup UI.

**Problem or motivation**

Operators need a GitHub review bot that runs their assigned Paperclip
agent. Mentions and PR events must preserve task ownership and requester
authority. Provider publication must use the bot App identity and
enforce the configured permissions.

**Proposed solution**

Extend the existing GitHub chat connector with resumable App onboarding,
linked-member and sponsored-guest access, editable event prompts, and
governed review operations. Validate structured assessments on the
server and compute a stable Paperclip Review check for the exact head
commit.

**Alternatives considered**

A separate review scheduler would duplicate Paperclip execution and
permissions. Reusing personal GitHub credentials would change the bot
identity and credential boundary.

**Roadmap alignment**

This extends the existing Connected Apps and governed-tool
infrastructure. The project owner requested and approved this design.
Related PR #8645 imports external Codex review feedback; this change
runs an assigned Paperclip agent and publishes its results through the
existing chat connector.

## What Changed

- Include the current Paperclip instance origin in the copied setup
prompt. Storybook uses its configured Paperclip origin; callback
parameters and URL credentials are excluded.

- Add a Claude/Codex copy button in the real setup and Storybook opening
step. Its detailed prompt asks four setup questions and guides
embedded-browser setup, verification, and optional required checks.
Clipboard failure exposes selectable instructions.
- Add a tutorial that explains why App installation, review scheduling,
and required checks are separate choices.

- Add manifest registration, an existing-App path, separate installation
and repository selection, repository refresh, and explicit account
confirmation.
- Add low-trust agent guidance, effective capability verification,
member selection, and explicit restricted guests with a sponsor.
- Add configurable PR events, prompts, repository overrides, rating
thresholds, and separate formal-review permissions.
- Give the assigned agent governed App tools to read PRs, comment, begin
an assessment, submit findings, and optionally submit a formal review.
- Bind review history, root PR events, and inline replies to ordinary
tasks. Deduplicate deliveries/findings and reject stale publication.
- Link check Details to the underlying task on the current trusted
hostname, or to Reviews before task creation.
- Add schema migration 0283, API contracts, production UI, and 49
interactive Storybook states.
- Repair local lease recovery. Keep the Cloud Dockerfile identical to
master; no provider-pack layer or runtime-default environment variable
is added.
- Retry only rolled-back wake-admission transactions after transient
endpoint-lock contention. A deterministic held-lock regression proves
one accepted wake.

## Verification

- Current head: `7ba761fe007bb798400d3e62346fa964f607f0f8`, rebased on
master `d9b3a5653e41f2ee5a1345b97c86a238f7a5c8e9`. Dockerfile has zero
diff against master. Final workspace typecheck and build passed. The new
PostgreSQL migration regression passed and preserves existing relation
and constraint identities after replay.
- Greptile reviewed this exact head at 5/5. There are zero unresolved
review threads and no merge conflicts.
- All current-head checks are green: 54 passed and two conditional
Storybook jobs skipped. This includes complete server/workspace test
suites, build, typechecks, policy checks, Runner suites, browser suites,
and security status. One timing-sensitive callback-ordering test passed
in isolation and its CI shard passed one retry. The duplicate local
full-suite run was stopped after CI completed; it is not counted as a
local full-suite pass.
- Before the final Slack rebase and migration renumbering, 186 focused
GitHub tests, 14 native bootstrap cases, token gates, and Storybook
build passed. The final rebase retained the new Slack communication
guidance.
- The embedded-browser setup test copied the full detailed prompt,
including the configured Paperclip instance URL. Desktop and narrow
layouts were checked. Component tests cover successful copying and
clipboard failure with selectable text and retry.
- Live local and hosted GitHub acceptance evidence refers to application
revision `cb703ac959876a07ebf3d7a295847f9f351eb6fc`. Real agent tasks
exercised issue mentions, automatic PR reviews, inline findings,
repeated mentions, task continuation, and failing-to-passing checks
after a push. The Storybook agent generated, built, and browser-rendered
pages; missing acceptance text failed, matching text passed, and broken
JSX produced an incomplete result.
- Live cases also covered independently disabled push events, prompt
injection, duplicate signed deliveries, rapid pushes, stale-result
rejection, finding deduplication, and restart recovery. Formal reviews
were denied while disabled and published only after explicit enablement.
Check Details links pointed to the underlying task on the trusted
hostname.
- Those hosted native Claude runs used the provider-pack layer now
removed from this PR. They do not prove native Claude works on the
standard Cloud image. A replacement hosted native Codex run is not yet
verified: the disposable QA tenant has only an Anthropic AI connection.
No new staging or production deployment was made for the packaging
removal.
- Required-check merge enforcement could not be tested because the
private disposable repository's GitHub plan rejected the rules
configuration. Published success/failure/incomplete check states were
verified directly.

## Risks

- Latest master allocated migration 0282 to Slack. The GitHub migration
is regenerated as 0283 with replay-safe table/index/constraint creation;
a PostgreSQL regression verifies existing relations and constraints are
preserved. Existing preview tenants remain subject to the fleet
migration-history compatibility preflight; no bypass is introduced.

- Migration 0283 adds company-scoped configuration, registration,
review, and publication records. Existing connections retain their
behavior until reviews/tools are enabled.
- Signed webhooks and expiring registration state remain required.
Hosted installations also need the companion narrow Cloud gateway
exemptions.
- Agent assessments can be incomplete or wrong. The server enforces
coverage/result structure, current-head publication, rating policy, and
separate formal-review permission; it does not replace code-review
judgment.
- No Cloud image packaging changes are included. Remote native
ACPX/Claude and OpenCode retain their existing operator-supplied
provider-pack prerequisite. Native Codex and Codex with managed MCP
tools do not require that pack. Earlier staging deployment evidence
refers to its stated revision, not this packaging-removal head.
Production rollout and merging remain outside this change.

## Model Used

OpenAI GPT-6 through Codex, with repository, code execution, API, and
embedded-browser tools. The exact serving model ID and context-window
size were not exposed by the environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 14:41:19 -05:00
DottaandPaperclip e8c8ba3c19 feat(apps): add experimental MCP aggregator connectors (#13755)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its tool gateway applies company access rules and approval controls
to connected apps.
> - MCP aggregators expose many apps through one provider endpoint.
> - Each aggregator needs its own credential, catalog, grants, and
lifecycle in Paperclip.
> - This pull request adds independent Zapier, Arcade, Composio Connect,
and Executor setup with a common Access → Connect layout.
> - A default-off MCP aggregators flag lets operators opt in while we
complete provider acceptance tests.
> - Agents use the normal Paperclip permissions, Test screen, and
gateway after setup.

## Linked Issues or Issue Description

**Subsystem affected**

Apps, connection setup, shared contracts, and the remote MCP gateway.

**Problem or motivation**

Aggregator endpoints need clear provider setup and correct MCP sessions.
Generic setup does not explain each provider's authentication or broad
execution tools. Provider approval must preserve the original execution
instead of replaying a write.

**Proposed solution**

Add four separate connectors behind Settings → Experimental → MCP
aggregators. Start with human and agent access, then connect the
endpoint and read its tools. Enable tools by default. Use the existing
Permissions and Test screens after setup. Keep legacy Composio API-key
and child connections intact.

**Alternatives considered**

A shared connection for all providers would mix credentials and access
rules. Separate provider-specific permission and test screens would
duplicate existing controls. Vercel Connect is outside this change.

**Roadmap alignment**

Extends the existing MCP Tool Gateway & Apps capability and the
Connected Apps roadmap area. This work was requested and reviewed by the
maintainer.

Related work: #11894, #12630, #12632, #12634, and #12906 concern the
legacy Composio broker. #13102 also covers remote MCP pagination. This
change preserves the broker path and adds initialized sessions, response
matching, and provider resume handling alongside pagination.

## What Changed

- Add branded setup and interactive Storybooks for Zapier, Arcade,
Composio Connect, and Executor. Use the existing access controls and
normal action tests. Do not request a connection name or action choices
during setup.
- Add the default-off `enableMcpAggregators` flag to settings, managed
feature metadata, the catalog, and setup guards. Hidden connections keep
running. Legacy Composio connections remain unchanged.
- Reuse the vault, grants, policy, and catalog models. Support OAuth
discovery, bearer tokens, custom headers, and credential-bearing URLs.
Add no database tables or migrations.
- Initialize and retain Streamable HTTP sessions by connection and
effective credentials. Read paginated catalogs and match streaming
responses to request IDs.
- Classify unfamiliar aggregator tools as writes despite upstream
read-only hints; only exact reviewed read capabilities enter the
read-only allowlist. Legacy Composio child behavior is preserved.
- Preserve provider authorization links and execution IDs. Support
Executor approve/resume, decline, and cancel without automatic replay of
uncertain writes.
- Preserve Off and Ask first choices during refresh and reconnect. Allow
new tools and retire removed tools. Keep agent access updates atomic and
preserve an empty agent selection.
- Document connector UX rules, provider branding sources, and live
acceptance results.
- Stabilize the existing Sentry release fixture after its repeated CI
failure by reusing one module mock; production Sentry behavior is
unchanged.

## Verification

- Final head `d11781970`: [CI
run](https://github.com/paperclipai/paperclip/actions/runs/35633534900)
passed, including broad typecheck, test shards, build, and E2E. All 54
checks pass; 2 optional checks are skipped. Greptile is 5/5, Security
Scan passes, and all review threads are resolved.

- Passed 27 focused connector Vitest checks and 18 connector-only
Storybook browser checks before the flag change. All 85 stories rendered
at desktop and narrow widths.
- Passed 5 connector lifecycle/server checks and 7 selected flag checks
after adding the flag. The latter cover settings, managed defaults,
cached catalog visibility, and all four setup routes.
- Review fixes passed 13 risk/handoff/lifecycle checks, dedicated
session-expiration and transport regressions, 13 selected
connector/gateway CI cases, and 10 selected setup/reconnect UI cases. A
real Composio connection-list call also succeeded through the refreshed
UI on `9ab115f71`.
- UI and server TypeScript checks passed. UI build, Storybook build,
token gates, and diff whitespace checks passed during implementation.
- Real browser and real Paperclip agent tests passed for Arcade,
Composio, and Executor. Tested action permissions, denied agent access,
reconnect, disconnect, and isolation. Tested Arcade catalog
additions/removal and Executor provider approve/resume, decline, and
cancel.
- Zapier live acceptance is incomplete. Its dedicated provider server is
configured, but its credential-copy dialog returned an empty clipboard
through browser automation. No live Zapier action is claimed.
- The three isolated Sentry release cases pass after the CI fixture fix.
- Local verification is deliberately narrow at the maintainer's request.
The full local suite, recursive typecheck, and repository-wide build
were not run. CI provides the broader checks.

## Risks

- Shared MCP transport changes affect other remote MCP servers. Protocol
fixtures cover initialized sessions, streaming response matching,
pagination, and isolation.
- Broad execution tools remain broad permissions. The provider governs
actions inside those tools.
- Provider handoff links are retained briefly in memory. After a server
restart, a one-time link may require reopening the provider dashboard.
Paperclip does not replay the original call.
- Zapier remains unproven live. Custom-header imports and self-hosted
endpoints have fixture coverage rather than a separate live account for
every variant.
- Turning the experimental flag off hides setup; it does not revoke
existing credentials or stop existing connections.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, shell
execution, and browser automation. The exact runtime model ID and
context-window size are not exposed in this session. A separate
Anthropic-backed Paperclip agent performed live gateway acceptance
tasks.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 12:53:11 -05:00
Devin FoleyandPaperclip 6fe8e30625 feat(apps): add Railway connection and governed deployment tools (#13415)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps gives agents governed access to external resources.
> - Operators need to inspect Railway services, read logs, deploy code,
and run container commands.
> - Railway offers hosted MCP with OAuth, but broad remote actions hide
their internal operations.
> - This PR adds a branded connection and fixed direct operations
through the existing gateway.
> - Separate SSH keys enable container commands under the same grants
and policies.
> - Operators can require approval for an action and inspect the
resulting audit record.

## Linked Issues or Issue Description

**Subsystem affected**

Apps catalog, connection setup, gateway execution, and connection
documentation.

**Problem or motivation**

Agents need Railway access through Paperclip. Operators need to grant
and revoke that access, inspect available actions, and govern deployment
and container operations without giving agents provider credentials.

**Proposed solution**

Reuse hosted MCP OAuth, vault storage, catalog discovery, grants, and
the gateway. Probe the actual credential before enabling fixed GraphQL
operations. Use a dedicated grant-owned SSH key for bounded container
commands.

**Alternatives considered**

A catalog entry alone cannot execute the missing operations. The hosted
general agent has opaque internal effects. An unrestricted CLI runtime
can bypass action policy and inherit ambient credentials.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps path and the Connected
Apps direction in ROADMAP.md. It does not add a plugin or parallel
connection service.

Related PRs #311, #939, and #7861 concern hosting Paperclip on Railway.
They do not add this outbound Apps connection. The separate shared
agent-picker fix is #13414 and is not included here.

## What Changed

- Add the generated Railway catalog entry, official marks, provenance,
and OAuth setup guidance.
- Add fixed service/deployment status, bounded logs, and
redeploy/restart/rollback tools. Block source deployment until the
provider can atomically bind the approved repository and commit.
- Verify API access with an explicit workspace before exposing direct
tools.
- Add grant-owned SSH key setup and a bounded runner with host
verification, target checks, isolated state, and cleanup.
- Block the opaque hosted railway-agent and accept-deploy actions.
Preserve normal Allowed defaults and Ask-first policies for other
actions.
- Quarantine new or changed Railway schemas after initial discovery,
including reconnect.
- Add provider, lifecycle, gateway, SSH, UI, and browser fixtures.
Document setup, limitations, and the release checklist.

## Verification

- Security follow-up: removed the unsafe source-deployment mutation.
Direct calls and old active catalog entries are denied before any
upstream request, including normalized aliases. Refresh marks retired
entries disabled. All 386 focused Railway, catalog and gateway tests
passed, and server TypeScript checking passed. Full [GitHub
CI](https://github.com/paperclipai/paperclip/actions/runs/35139421144)
passed on d86530ab9, including typecheck, build, all tests, runner
checks, and browser tests. Superagent passed and confirmed the P2 fix.
Greptile reviewed the same commit at 5/5 with no findings.

- CI follow-up: fixed the missing Railway SSH operation in the OpenAPI
document, including its request schema, operator-only authentication,
and error responses. The failure reproduced locally before the fix; all
403 selected API, Railway, catalog, and artwork tests passed after it.
Synced current master and resolved the catalog/artwork conflicts.

- After rebase: 440 focused provider, lifecycle, gateway, catalog, and
container-panel tests passed. AppDetail and AppsConnect passed another
196 tests.
- Full typecheck, build, token gates, and the gallery browser check
passed after rebase.
- During implementation, full build and the gallery browser check
passed. Shared generic-MCP fixtures covered OAuth callback/state/issuer
binding and failure paths.
- Local live consent and tools/list succeeded. There were 44 active
hosted actions and two blocked actions. A workspace-bound API probe and
direct project/service/environment reads succeeded. The inspected
project had no deployed services. No provider mutation ran.
- Full GitHub CI passed on commit 303340f19, including all
server/workspace test groups, typecheck, build, runtime verification,
release dry run, and browser tests. The original local full-run attempt
was incomplete; the complete automated suite is now verified in CI.

Manual review: connect Railway, review the actual actions, install for
an agent, and run a resource read through the gateway. Choose Ask first
before testing a deployment mutation. Configure a dedicated key only
when container access is needed.

**Release qualification is still open.** Live agent gateway reads/logs,
rejected and approved deployment calls, refresh/revoke, public HTTPS
consent, and SSH enrollment/commands/cleanup need an authorized
disposable service. The passing API diagnostic does not replace those
tests. See doc/connections/RAILWAY.md and RAILWAY-REVIEW.md.

## Risks

Overall risk is medium. New runtime behavior is gated to Railway
connections, but the PR changes shared catalog, credential lifecycle,
and gateway code. A regression in those paths can affect other Apps
connections. The highest-impact operations are Railway deployments and
container commands.

- Provider consent can authorize an entire workspace. Catalog labels are
not local resource allowlists. Direct tools check target membership, and
provider permissions still apply.
- Shell commands have broad internal authority. Action policy cannot
approve each internal shell step. Timeouts close the local connection
but cannot guarantee remote child-process termination.
- Log and command output may contain application secrets that pattern
redaction cannot recognize.
- Source deployment is unavailable until the provider supports atomic
repository/commit binding. Existing deployments can still be redeployed,
restarted or rolled back.
- No database migration is required. Rollback can remove promotion and
direct dispatch while preserving connection data and the generic MCP
path.
- Live Railway qualification must still pass before release acceptance.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and browser testing.
An independent read-only security agent reviewed the local
implementation. The exact serving model ID and context window were not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 13:54:44 -07:00
Devin Foley d08abcba15 ci: cut PR wall clock from ~16 to ~6 minutes (#13521)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every pull request runs the Trusted PR CI workflow before merge
> - The test suites roughly tripled in six weeks, and shard balance did
not keep up, so PR runs crept from ~4 to ~17 minutes
> - Slow CI delays every merge and every contributor
> - This pull request rebalances the shards from fresh measurements,
splits the largest test files, reuses the Rust build cache in three more
jobs, and takes the policy job off the critical path
> - The benefit is a PR wall clock near 6 minutes with the same coverage

## Linked Issues or Issue Description

**What existing behavior does this improve?**

PR CI wall clock. A typical green run took 16-17 minutes. Two months ago
it took about 4 minutes.

**Subsystem affected**

The Trusted PR CI workflow (`.github/workflows/pr-trusted.yml`), the
shard-duration manifests, the vitest shard runner scripts, the
`paperclip-runner` package scripts, and the dry-run branch of
`release.sh`.

**Current behavior**

The shard-duration manifests were stale. The general-server manifest had
durations for ~400 of 649 suites. The e2e manifest was missing 14 of 29
specs. Stale median weights made shard steps range 417s-806s (server)
and 277s-745s (e2e). Three jobs each paid a ~3m40s cold cargo release
build. Every test lane waited ~60s for the policy job before it could
start.

**Proposed behavior**

All lanes finish in a narrow ~200-290s band. The manifests carry fresh
measured durations for every suite. The three largest test files are
split so no single file caps a shard. The Rust cache restore runs in
every job that builds the Runner binary. Test lanes start as soon as the
gate resolves.

**Reason and benefit**

Merges stop waiting on CI. The projected wall clock is ~6 minutes for
the same test coverage.

## What Changed

- Rebuild `scripts/general-server-shard-durations.json` (646 suites) and
`scripts/e2e-shard-durations.json` (all specs) from per-suite completion
timestamps in runs 35036001734 and 35024948947.
- Move the PR server lane to the release-verify shape:
`general-server-without-chat` across twelve duration-balanced shards,
plus the chat integration suite split by collected test location across
three dedicated lanes.
- Split `tests/e2e/chat-adapters-ui.spec.ts` into `-providers` and
`-messaging` specs, and `tests/e2e/agent-chat.spec.ts` into `-sessions`
and `-projects` specs. Each pair shares fixtures through a `.shared.ts`
module. Playwright collects the same test sets (39 and 20 tests).
- Raise e2e shards to eight and serialized shards to nine.
- Run the runner package's `check:all` as four matrix lanes:
`check:static`, `check:runner`, and two native vitest `--shard` halves.
The union is exactly `check:all`.
- Add the read-only Rust cache restore (toolchain pin, `save-if: false`)
to the Canary Dry Run, Build, and Typecheck jobs.
- Make release.sh preview publish payloads concurrently in batches of
eight during `--dry-run`. The real publish path stays strictly serial.
- Drop the policy-job lockfile artifact chain. Each lane installs with
`--frozen-lockfile` and falls back to an inline `--resolution-only`
regeneration. The policy job stays a required check through the `verify`
and `e2e` aggregates.
- Update the shard-count mirrors and workflow assertions in the
partition and gate tests.

## Verification

- `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs
scripts/__tests__/e2e-shard.test.mjs` — 30 pass.
- `node --test '.github/scripts/tests/'*.test.mjs` — 410 pass.
- `node --test scripts/__tests__/release-verify-workflow.test.mjs
scripts/cloud-source-verification.test.mjs
scripts/__tests__/release-dry-run-notes.test.mjs` — 42 pass.
- `playwright test --list` collects 39 tests across the chat-adapters
split and 20 across the agent-chat split, equal to the original files.
- A local vitest collection of the chat suite partitions 995 tests into
498/497 line shards.
- Projected shard weights: server 230s x12, chat ~143s x3, e2e 207-242s
x8, serialized ~216s x9.

## Risks

- The split spec files reorder tests relative to the original files.
Every describe seeds its own company, so the specs stay independent; a
hidden cross-describe dependency would surface as a deterministic
failure in one shard.
- The inline lockfile fallback changes install behavior for
manifest-changing and stacked PRs. The policy job still validates
resolution as a required check.
- `release.sh` changes are confined to the `--dry-run` preview branch.
The publish loop is untouched. `bash -n` passes and the release dry-run
tests pass.
- One PR now schedules ~44 fleet runners. If the RunsOn fleet caps
concurrency, queueing may absorb part of the gain; watch the first runs.

## Model Used

- Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking, with
tool use (shell, file edits) in Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-16 11:45:14 -07:00
DottaandPaperclip d351e08dee fix(ui): stop Recent Tasks storage feedback across tabs (#13402)
Version recent-task snapshots separately from activity, reject stale cache updates, and publish only on query or membership changes. Migrate history into an isolated storage namespace while preserving restart retries.

Add regression coverage for conflicting caches, migration, comments, unchanged writes, and storage implementations on Node 24 and Node 26.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-14 08:05:06 -05:00
DottaandPaperclip 422287eecd fix: preserve runner recovery, warm sessions, and task outcomes (#13338)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner connects task messages, provider execution, and
task outcomes.
> - First-time user tests exposed gaps in recovery, completion
permissions, message delivery, and Stop behavior.
> - These gaps left usable output hidden, completed work waiting for
bookkeeping, or safe work unable to continue.
> - This pull request fixes the shared lifecycle and receipt paths while
preserving process ownership and action checks.
> - Users can continue work with accurate task state and durable
messages.

## Linked Issues or Issue Description

**What happened?**

A stopped local Codex execution could remain blocked even after its
processes had stopped and its complete transcript proved that no
external action needed replay. Claude under Conservative permissions
could fail to call task completion tools. Recovery could reuse an
assistant item ID and overwrite prior output. A delivered comment could
remain marked uncertain after navigation. Stop could look like Pause or
a new recovery incident. Workspace contention could look like
cancellation. A direct reply reopening Done could enter a clarification
loop.

**Expected behavior**

Recover automatically only with verified termination and complete action
receipts. Preserve answers and messages. Keep task completion available
under Conservative permissions without broad tool access. Show crashes
as Blocked, actual human decisions as In Review, and ordinary workspace
contention as waiting. Stop the current response and allow a new
direction.

**Steps to reproduce**

1. Create ordinary response tasks with local Codex and Claude Code, then
send follow-up messages through the task composer.
2. Interrupt a disposable local Codex runner during text-only work.
Verify automatic continuation and retained output.
3. Stop a response, send a new request, answer a clarification, and
reopen completed work with another message.
4. Navigate or reload while a comment submission is pending. Confirm the
exact persisted request receipt settles it without removing newer draft
text.
5. Run two tasks in a shared Daytona workspace. Confirm waiting does not
appear as failure.

**Paperclip version or commit**

Initial acceptance baseline: `c9021c6721f91e2c74bd9fee9d3fd41c999d17b7`.
Current integration base: `6cef9743c`. Both operator-interruption and
workspace-waiting guards are preserved; native restart and legacy
permission rules remain documented.

**Deployment mode**

An isolated source-built test-drive instance, with real local Codex and
Claude Code providers and disposable Daytona environments.

Related work: #13314, #13316, #13327, #13344, #13239, #13254, #13163.
This PR addresses additional failures from ordinary task journeys,
including controller restart handoff and repeated warm sandbox setup.
Historical task status reconciliation is excluded.

## What Changed

- Persist runner ownership immediately at spawn and resume an explicitly
adopted runner even when the controller crashed before the first driver
checkpoint. Detach the controller safely across graceful restarts,
including session startup. Prevent an old finalizer from suspending or
signaling an adopted runner. Checkpoint idle warm sessions before
shutdown. Preserve the same run and queued follow-up messages.
- Scope saved legacy queue successor checks to the queue owner while
preserving ordinary task locks, operator identity, assignment gates, and
exactly-once delivery.
- Preserve managed Codex credential files when an old session is
detached for restart; normal owned cleanup still copies refreshed auth
back and removes the scoped copy.
- Reuse the bound warm shared sandbox and fully verify an existing
staged provider pack before using it. This avoids repeated uploads when
the pack is already valid.
- Add a narrow local Codex replacement path with stopped-process proof,
a closed transcript inventory, exact completion receipts, and
fresh-session lineage. Preserve no-replay holds when evidence is
incomplete. Recovery may clear only the same run's recorded Blocked
status version; manual re-blocking and dependency changes invalidate
that receipt, while queued comments do not. Later blocks stop scheduled,
queued, and final dispatch; queued/final checks re-read dependencies
even when the task status stays In Progress.
- Permit only task delivery and human-input tools through the isolated
Claude runner's exact task bridge.
- Scope assistant item identity to the provider turn and ignore only
authority-free Codex skill-change notifications during startup.
- Reconcile composer submissions by client request ID across response
loss, navigation, and reload. Retain text typed during delivery.
- Keep acknowledged run-only Stop neutral and show workspace contention
as waiting. Project exhausted native failures as Blocked.
- Restore the guarded task-page retry action for failed legacy runs,
including the server-supported explicit new-attempt path for stopped
conversation adapters. Preserve native/process recovery holds and avoid
promising Retry while a decision or execution gate hides it.
- Refresh delivered artifacts and handle direct user replies that reopen
completed work without a clarification loop.
- Check the embedded PostgreSQL PID, data directory, and actual port
before connecting or migrating.
- Document accepted behavior and add focused regressions at lifecycle,
route, transcript, and UI boundaries.

## Verification

- Final head `fece606ac2` passes the complete GitHub CI matrix: **34
green checks, two expected Storybook skips, no failures or pending
checks**, including `ci / verify`, `ci / e2e`, full runner verification,
typecheck, build, every server/workspace shard, and all browser shards.
[CI
run](https://github.com/paperclipai/paperclip/actions/runs/34727183287).
Greptile is **5/5 with no open findings**. The final two commits only
refine test fixtures; both affected suites pass 24/24 locally and in CI,
with server typecheck green.
- Complete local Vitest coverage uses the canonical groups/shards: all
635 general server suites, all 145 serialized suites, and all workspace
packages. The aggregate began on `0a8001c18` while the final queue fix
arrived: 23,903 passed, five failed, 87 skipped. The five
port/socket/timing failures passed unchanged in follow-ups (60 tests in
the exposure/file suites and 412 tests covering the serialized failures
and unrun tails). The final queue/operator-identity suites separately
passed 52/52. This is aggregate coverage plus explicit reruns, not a
pristine single-command final-head run.
- After integration with current master,
queue/operator-identity/continuation suites passed 162/162 and affected
UI suites passed 140/140. ACP Stop/continuation and legacy
task/Inbox/message browser suites passed 9/9, including both task
recovery Retry and thread Try again, automatic saved-message delivery,
exactly one new run, Done, and retained output after reload. The default
process Stop/Pause/Resume browser case passed (the native-provider case
is opt-in and skipped by default). The complete Board attachment/receipt
browser suite passed 11/11 on a disposable instance, covering both
composers, exact receipts after lost responses, no replay, bound
attachments, and newer drafts after reload.
- Blocking-intent regressions cover pre-existing Blocked, a mismatched
run/cause, an explicit manual re-block, changed dependencies, a queued
comment after failure, and a block arriving between scheduling and
provider dispatch. The negative cases reproduced before the fix. All 478
affected executor/recovery/dispatch tests passed; both database suites
ran separately after availability-probe skips in the first combined
command. The final late-dependency check passed all 143 affected
recovery/dispatch tests (zero skips) after two new negative cases
reproduced the bug.
- Focused runtime regressions cover awaited runner ownership
publication, authenticated adoption before the first checkpoint,
old-finalizer detachment, idle and busy warm-session shutdown, rejected
checkpoint propagation, provider-pack verification, and managed-Codex
credential preservation. Four managed credential detachment cases
reproduced the bug before the fix; normal owned cleanup still succeeds
exactly once.
- Live local Claude: SIGKILL 2.6 seconds into startup recovered the same
run automatically in 53 seconds, then a normal follow-up completed in 24
seconds. SIGTERM 2.5 seconds into startup preserved the same run (54
seconds) and its queued follow-up (21 seconds). Answers remained visible
and the task reached Done.
- Live Claude Daytona: a warm follow-up retained its sandbox and fell
from 121 seconds to 44 seconds. A separate cold turn took 127 seconds;
after controller shutdown and checkpointing, its follow-up completed in
33 seconds with the same sandbox, workspace, native session, and runner.
Both answers remained visible and the task was Done.
- Other live journeys covered task completion and follow-up with local
and Daytona Codex, local Codex crash recovery, Stop then new direction,
clarification response, live artifact refresh, and shared-workspace
waiting.
- Validation limits: the opt-in native composer Stop/Pause→subtree
Resume fixture exposes terminal/result ordering and subtree-cancellation
attribution bugs that can leave a child task blocked; that new finding
is assigned to a separate follow-up and is not claimed fixed here.
Default CI skips this optional native-provider fixture. Managed-Codex
credential handoff and the queue-agent integration use automated
regression evidence. Cold custom provider-pack uploads still add startup
latency.

## Risks

- Automatic replacement remains deliberately narrow: local Codex,
verified stopped identities, unchanged retained state, and a complete
text/completion-only turn. Unknown actions, partial history, or changed
ownership remain blocked.
- Claude completion permission handling changes an upstream package
patch. The exact isolated task bridge must remain pinned; unrelated
tools keep their existing permissions.
- New task failure projection changes user-visible status. No historical
status backfill or database migration is included.
- This is a broad lifecycle fix across server and UI. Live proof covers
graceful local Claude restart during startup and idle Claude Daytona
session recovery across controller shutdown. Live abrupt SIGKILL during
local Claude startup also recovered the same run. Unknown ownership or
missing action evidence still blocks reuse. Cold custom provider-pack
uploads still add startup latency; this change avoids unnecessary repeat
uploads.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, code execution, browser
automation, and tool use. The exact hosted model ID and context window
are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 19:41:15 -05:00
DottaandPaperclip 8d6232e7b0 feat: reuse provider sign-in across AI connection workflows (#13248)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users connect provider accounts during onboarding and agent setup.
> - They should reuse and manage those accounts through the existing
Connectors interface.
> - A second login wizard would diverge from the established provider
workflows.
> - This pull request composes the existing sign-in components into
Connections and agent configuration.
> - Users can select accounts without changing their agent's harness or
model.

## Linked Issues or Issue Description

**Problem or motivation**
AI credentials are configured separately from Connections. Agents cannot
consistently reuse a responsible user's account or a permitted shared
account.

**Proposed solution**
Manage AI accounts with the existing Connections grants and permissions.
Keep model and harness selection independent from credential selection.
Preserve legacy authentication until validated adoption.

**Alternatives considered**
A separate credential registry would duplicate ownership and access
policy. Automatic fallback would risk using the wrong account.

**Roadmap alignment**
This extends the shipped Apps, multi-user, secrets, and agent-runtime
capabilities. The maintainer requested the feature and reviewed the UI.
Related groundwork: #11899 (connection permissions), #10910 (connection
wizard), #11692 (Claude subscription profiles), and #11854 (Codex
account rotation).

## What Changed

- Add compact AI-account management to the existing Connectors pages.
- Reuse AgentProviderConnection, AdapterLoginPanel, AdapterLoginChrome,
and authentication controllers.
- Add the shared connection picker to agent setup/settings and task
requests.
- Preserve onboarding's sequence and reuse existing accounts.
- Add local-login recovery, retry, cancellation, and React StrictMode
handling.
- Add interactive Storybook scenarios, design-guide examples, and app
acceptance checks.

This is part 2 of the AI Connections change. The runtime foundation in
#13247 is merged. This PR now targets master.

## Verification

- Updated against master `47ded8bf9`, including the landed runtime
foundation and upstream task-search changes.
- Full workspace typecheck, production build, Storybook build, and token
gates passed on the integrated branch. Final local-login changes passed
59 focused tests; new-agent and inbox regression suites passed 63 tests.
- Browser checks verified automatic local Claude account detection,
resumable Codex login commands, retry, focus restoration, and
desktop/phone layouts. Commands create their isolated directory before
invoking the CLI.
- All CI test, browser, build, packaging, and runner jobs passed on
final head `dd17d3211931dd70aaa6ea619d83a7f9966dd18e`. The fresh
Greptile review is 5/5, the security scan passed, and there are no
unresolved review threads. The final CI aggregate gates passed.
- Local general-server coverage passed 11,804 tests; three
port-collision failures passed in an isolated 25-test rerun. All 6,111
UI tests passed. CLI coverage passed 484 tests; its remaining doctor
test requires port 3199, which is occupied by an unrelated report server
on this Mac. The complete CLI suite passed in CI.
- Live browser testing verified Codex API-key reconnect inside a task
card on desktop and phone. Real provider runs resumed and completed with
unchanged connection/grant identity and agent routing.
- Tested opening, cancelling, reopening, and completing connection
creation. A regression confirms Connect another account cannot submit
the new-agent form or copy provider keys into agent settings.
- Added shared inline repair, automatic local sign-in checks, and
responsive connection dialogs. Standalone Daytona installation ignores
workspace configuration and suppresses dependency scripts. Its
standalone build also passed with CI's exact pnpm 9.15.4.
- Destructive live tests are excluded by default. Explicit opt-in, local
deployment checks, and matching disposable fixture identities are
required before any mutation.

## Risks

- Local Codex/Grok creation requires the connection-specific terminal
login command.
- Browser sign-in uses the existing supported-environment controllers.
- This update verifies live local Claude detection and Codex API-key
task repair. New subscription authorization/refresh and
independent-human/native-runner isolation were not reverified in this
update.
- No agent automatically adopts managed Connections.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact runtime model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 16:51:26 -05:00
DottaandPaperclip 4d317274ce feat(channels): add experimental iMessage Photon (#13299)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Channels connect external conversations to company tasks and agent
execution.
> - Slack, Discord, and AgentMail already provide durable delivery and
access controls.
> - People also need to reach an agent from Apple Messages and send
photos.
> - Photon provides shared Pro DMs, dedicated numbers, and authenticated
event recovery.
> - This pull request connects Photon to the existing channel services.
> - People can message an agent while Paperclip retains task ownership
and approval authority.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: channel services, shared contracts, database constraints,
Apps, and agent Channels UI.

**Problem or motivation**

Paperclip has no iMessage channel. A person cannot use Apple Messages to
start a task, send a photo, or answer an agent's pending question.

**Proposed solution**

Add experimental **iMessage Photon** with Pro-compatible shared DMs or a
dedicated Photon Cloud number per agent channel. Reuse channel
admission, identity links, task generations, publication, and
interaction continuation. Keep groups disabled for shared allocation.
Dedicated lines support groups that an operator explicitly enables.
Require a fresh linked message and a published agent response before
setup completes.

**Alternatives considered**

Shared allocation has no owned phone number, so it reserves one project
and allows DMs only. Dedicated allocation reserves one stable number.
Local Mac access needs a separate deployment model. The upstream Photon
Chat SDK adapter does not persist the poll mappings and send receipts
required here. This change uses the lower-level SDK without adding
another agent runtime.

**Roadmap alignment**

This extends Connected Apps and agent communication through the existing
channel subsystem. It does not add a parallel tool connection or agent
loop. GitHub searches for Photon and iMessage found no matching provider
implementation.

**Additional context**

This ships behind the existing experimental channel gate. Dedicated-line
release qualification remains incomplete. Real Photon Pro DMs passed
task/reply, native poll, text answers, confirmation rejection, media,
restart, pause, reconnect, revocation, and removal tests. An
operator-supplied iPhone camera HEIC also passed the full round trip.
Dedicated groups remain unqualified. See [the verification
record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) and [the
implementation plan](doc/plans/2026-09-11-imessage-photon.md).

## What Changed

- Add the provider catalog entry, shared setup contracts, and a forward
migration. A global partial index reserves the dedicated number or
shared project until its endpoint is archived.
- Add Cloud project inspection, vaulted project credentials,
selected-line token renewal, and a leased receiver. Persist checkpoint
updates under the receiver lease. Shared project replay accepts sparse
increasing sequences only after a complete recovery barrier.
- Connect DMs and enabled groups to existing task generations, sender
authorization, ordered delivery, and publication services. Keep each
iMessage conversation on its task after completion; only explicit `/new`
or `/close` releases the binding. Publish committed inbound comments
live and label their human bubbles “Sent from iMessage” in both
task-chat renderers.
- Persist immutable text/file send identities, upload receipts, poll
IDs, option IDs, per-person drafts, and canonical interaction
continuation proofs.
- Add source-bound file recovery, bounded HEIC/HEIF conversion, JPEG
previews, and related Live Photo companion video retention.
- Add the three-step setup flow and channel management surfaces with
official branding. Preserve the experimental gate and existing
pause/disconnect behavior.
- Add interactive production-component Storybooks for setup, access,
recovery, and ongoing conversations. Add provider, integration, catalog,
and browser regression coverage. Document setup, recovery, supported
boundaries, and qualification gaps.

## Verification

- Live Photon Pro, SDK 2.1.0: linked iPhone messages create a task and
receive native Codex replies in Apple Messages. Unlinked senders cannot
start work.
- Three real follow-ups each reopened the same completed task. Incoming
bubbles appeared on its open page without reload and showed “Sent from
iMessage.” The third follow-up ran after restarting the server on
`4d7222110`; the agent correctly repeated its previous reply from before
the restart.
- Native polls after restart, sequential text drafts, required-field
correction, explicit submission, approval rejection with a required
reason, and native continuation passed against Photon.
- PNG, text documents, synthetic HEIC, and a real iPhone camera HEIC
passed in both directions. The camera photo produced a 3024×4032 JPEG
preview. The native agent described it and returned the received HEIC
byte-for-byte.
- Pause/resume, reconnect, identity revocation, removal, `/status`,
`/new`, `/close`, and stale answers after close passed live. Messages
suppressed by pause did not become work on resume. Removal stopped
intake and removed credential bindings.
- All 304 focused tests passed on `4d7222110`. These cover Photon
unit/integration behavior, both task-chat renderers, live comment
hydration, completed-task continuity after restart, enabled groups,
duplicate delivery, and explicit reset/close. The selected Teams
completion-boundary regression also passed. Full workspace
typecheck/build and token gates passed for the conversation fix; the
final UI changes passed their affected typecheck/build and tests.
- All 26 new Photon Storybook Playwright cases passed in light and dark
themes, including the complete shared-DM setup journey and 390px mobile
follow-ups. UI typecheck and the Storybook build passed. These stories
use simulated Photon responses and do not replace the live evidence
above.
- The full chat-adapters browser suite previously passed all 39 cases.
Migration checks passed, and migration 0275 applied to the isolated live
instance with the earlier Photon migration already applied.
- The local full Vitest run was previously interrupted by the host's
embedded-Postgres shared-memory limit; it is not a full-suite pass. All
30 applicable CI checks passed on preceding head `7a5419cac`, with two
skipped checks and Greptile 5/5. Head `24f8e1aae` adds an explicit
required-story discovery guard to the 26 passing Storybook cases.
Greptile rates this final head 5/5 with no unresolved review threads.
All 30 applicable CI checks passed, with two optional checks skipped.
- A repeated live send key suppressed the duplicate but returned gRPC 6
/ SDK `internalError` without an original receipt. Paperclip keeps
unknown delivery unresolved. This provider behavior is covered by a
regression test.
- See [the verification
record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) for package
versions, redacted live evidence, deterministic coverage, and remaining
qualification gaps.

## Risks

- Dedicated group qualification remains unrun; groups are disabled for
the approved Pro scope. Real iPhone camera HEIC passed transport,
preview generation, agent inspection, and return. Keep the channel
experimental; the dedicated-line release matrix remains incomplete.
- Shared recovery and attachment aliases were verified against the live
gateway. Duplicate writes currently return an error without the original
receipt; unresolved sends require operator resolution. The
implementation fails visibly on invalid replay ordering, a reset cursor,
or changed identity.
- The HEIF converter passed on macOS arm64 and in Linux CI. Windows HEIF
binaries have not been executed in this work. Linux musl has no packaged
converter. Unsupported conversion retains the original and reports the
missing preview.
- The migration adds a global reservation across companies for Photon
numbers and shared projects. Paused and revoked endpoints keep that
reservation until removal.
- Integration touches shared channel services. Existing provider browser
coverage passes; broad repository verification is recorded above.
- `pnpm-lock.yaml` is intentionally excluded under repository policy.
The repository bot owns lockfile updates. The additional Superagent
supply-chain scan is neutral/inconclusive because these new dependencies
are not yet in the committed lockfile. Its security scan passed; all
required CI checks pass.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository inspection, code
execution, browser testing, and tool use. The exact served model
identifier and context-window size are not exposed in this session. No
sub-agents were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 15:23:50 -05:00
DottaandPaperclip ab15aff390 feat: add experimental persistent agent chat (#13284)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Conversations must use the same tasks, controls, and execution
history.
> - Users need an ongoing chat with an agent without managing task
properties.
> - Agents should clarify and plan work, then hand execution to assigned
project tasks.
> - This pull request combines the reviewed Agent Chat stack for one
squash merge.
> - The benefit is persistent conversation with normal task governance
and shared UI.

## Linked Issues or Issue Description

**Subsystem affected**

Task lifecycle, agent runtime tools, shared task UI, and browser/paid
runner tests.

**Problem or motivation**

Users need one persistent conversation with each agent. A separate chat
store or renderer would duplicate task behavior and bypass existing
controls.

**Proposed solution**

Use a task-backed chat per company, user, and agent. Reuse the task
composer and transcript. Clarify and plan in chat, then create assigned
project tasks with the relevant plan. Keep Agent Chat behind its own
disabled-by-default experimental setting.

**Roadmap alignment**

This implements the task-backed direction in [CEO
Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat).
Related proposals: #2504 and #9693. Related request: #7981. The
maintainer requested one squash merge of the complete stack.

Consolidates the reviewed runtime
[#13281](https://github.com/paperclipai/paperclip/pull/13281), backend
[#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI
[#13283](https://github.com/paperclipai/paperclip/pull/13283) layers
with this PR's E2E coverage. All four layers passed CI and received
Greptile 5/5 before consolidation. This PR targets master and includes
the complete feature.

## What Changed

- Add personal canonical chat tasks with ordinary company visibility,
immutable identity, idempotent first sends, and an idle waiting state.
- Process `/new` in queue order. Preserve history, release a chat pause,
and fence old provider context and delayed writes.
- Keep chat lifecycle rules across recovery, finalization, assignment,
task lists, and rollups.
- Support research and plan revision in chat. Hand plans to ordinary
assigned project tasks before execution starts. Reject new chat
subtasks.
- Add repository-aware project creation and discovery tools, including
multiple repository IDs and GitHub URLs, authorization, idempotency, and
durable project-created cards.
- Reuse task UI components for chat, with starred/recent agent
navigation and a separate `enableAgentChat` experimental flag.
- Add deterministic browser tests and 24 paid chat cells across four
Codex/Claude profiles, with validated reports and screenshots.
- Integrate current master recovery, controller lease, queued-message,
and task UI changes. Gate chat interruption and deferred promotion on
ownership/feature policy. Guarantee lease renewal and active controls
are stopped even if teardown fails.
- Preserve master's migration 0273 and generate chat migration 0274 with
idempotent replay for development databases.

## Verification

- Prior exact heads of all four PRs passed Linux CI, including build,
typecheck, general/serialized tests, and browser E2E. Each had Greptile
5/5 and no unresolved findings.
- Integrated local verification passed: full repository typecheck and
production build, Storybook build, token gates, 340 focused UI tests,
all 20 deterministic chat browser tests, two migration replay tests, 88
focused chat/queue/native/controller tests, and provider/session
regressions including real lease expiry. These include the three
lifecycle regressions for the final admission/teardown fixes; server
typecheck also passes. Current head
`1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no
unresolved findings and passing security scans. All final-head CI gates
passed: build, full Runner verification, typecheck/release registry,
canary, all general/serialized test shards, and all browser E2E shards
([CI
run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)).
Local PostgreSQL startup contention required serialized retries; skipped
fixtures do not count as passing coverage.
- The earlier paid campaign passed all 24 chat cells and retained 32
screenshots:
[report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat).
It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior
evidence, not a paid run of this integrated head.
- Manual check: enable Agent Chat in Experimental settings, open an
agent, clarify and revise a plan, then hand off to an assigned project
task. Stop a reply, send `/new`, and verify fresh context with retained
history. Disable the setting and verify agent shortcuts/new chat turns
are blocked.

## Risks

- Queue/session integration can affect retries and delayed writes. Tests
cover ownership, cancellation, reset boundaries, idle recovery, and
ordinary task behavior.
- Migration 0274 adds conversation fields and constraints. Replay is
idempotent and preserves existing development chat history.
- This combines the previously reviewed stack at the maintainer's
request. Agent Chat remains off by default and is separate from
Conference Room.

## Model Used

OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code
execution, browser tools, and parallel review. The exact context-window
size is not exposed in this session. Codex and Claude also ran as test
subjects in the linked paid campaign.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 08:56:04 -05:00
DottaandPaperclip 889947c238 feat: add experimental native chat connectors (#13038)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - People also ask agents for work in their existing chat tools.
> - Each external conversation needs one task and a current authorized
source.
> - Retries, Stop, and provider failures must not duplicate work or
expose private data.
> - The first chat PR establishes the opt-in provider and data
contracts.
> - This PR adds experimental channel integration and its durable
control plane.
> - Users can request work from connected channels and inspect delivery
in Paperclip.

## Linked Issues or Issue Description

Refs #13100 and #13092. This is the second of exactly two chat PRs.
Foundation #13100 is merged and changed 143 files. Runner prerequisite
#13092 is also merged. This PR changes 400 files against master, below
the 500-file review limit. It contains no wireframe images or HTML
galleries.

## What Changed

- Add native Slack, GitHub, Microsoft Teams, Telegram, and Discord chat
connections. Keep chat disabled unless the operator enables experimental
chat connectors. Preserve the production GitHub tool connection and its
normal setup path.
- Bind each provider bot identity to one immutable Paperclip agent. Bind
each admitted external conversation to one task. Paperclip owns tasks,
runs, permissions, and audit records.
- Add durable admission, per-conversation queues, questions, task
controls, progress, final replies, images, files, and delivery receipts.
Board comments remain internal unless explicitly sent to the channel.
- Check current identity, provider reach, resource access, credentials,
runtime generation, and exact source before provider effects. Keep
private responses private. Never send raw reasoning, private logs,
credentials, or tool arguments.
- Hold uncertain sends for explicit audited resolution. Make Board
Send-to-channel atomic and idempotent. Keep reconnect and setup
credentials in Paperclip secret storage.
- Preserve current native-runner authority across retries, lost
acknowledgements, and recovery. Keep immutable input and completion
contracts separate from newer user input. Receipt reconciliation cannot
launch a provider.
- Reconcile chat close/new ordering and provider-effect lock order.
Audit resource access changes in the same transaction. Submit only the
selected resource from each UI toggle so stale pages cannot undo
unrelated access changes.
- Drain Codex stdout before certifying process exit. Bound the drain
with the existing shutdown grace. Preserve observed terminal authority
without treating an undrained process as successful or reusable.
- Incorporate master `018ca5da` with its ACP Stop, mobile task layout,
runner packaging, and official lock changes. Preserve dedicated
chat-answer continuations in both directions when ordinary queued
comments are adopted after Stop.
- Fence late adapter readiness behind an earlier Stop for the same run.
Preserve verified cleanup for registered adapters. Handle single Stop,
agent pause, duplicate Stops, and failure release without creating a
false cancellation receipt.
- Incorporate master's `6dd48cad4` wake-queue extraction. Preserve exact
failed-chat retry authorization and lineage, retired question-source
suppression, and the block on generic recovery that would discard the
admitted source. Fresh deferred input retains its separate promotion
path.
- Incorporate master `2a05b5ed3` and its queue-admission extraction,
simplified transaction ports, and separate runner CI job. Preserve exact
durable receipts, actor separation, and dedicated-answer isolation
through the new module. A failed receipt insert rolls back the
accompanying deferred-wake merge.

## Verification

Current head: `afe19299d06253cb628eb398e91d1200ea9f412a`, incorporating
master `2a05b5ed3457ea33efd6895520447d1d97fe98d8`. The conflicts are
resolved. This successor fixes two test-harness boundaries exposed by
CI: per-case route-module preparation and actual durable-save completion
before intentional runner termination. Production code and all existing
test/turn deadlines are unchanged. [Exact-head Greptile
review](https://github.com/paperclipai/paperclip/pull/13038#issuecomment-5587250594)
is **5/5**, completed September 10 at 13:20:55 UTC, with no actionable
findings or open review threads. [Fresh exact-head
CI](https://github.com/paperclipai/paperclip/actions/runs/34481724341)
passes **all 24 jobs**, including Build and both required aggregates.
Normal exact-head guarded merge was attempted and rejected by the
remaining branch approval policy: CODEOWNER review is required and no
human approval is present. Normal **squash auto-merge is enabled** as of
September 10 at 13:36:26 UTC. Requested CODEOWNERS have been notified;
no approval bypass or self-approval was used. Earlier-head results below
remain historical evidence, not qualification of this successor.

- Final exact-head Linux evidence: 995/995 chat integration cases; 36/36
agent-skills routes; 35/35 runner live-session cases, including real
process kill/resume; 1948 runner Vitest cases with three existing
benchmark/platform guards; 870/870 API-authority cases; and 104 browser
cases with four existing optional skips. Rust, conformance/replay, full
repository build, typecheck, canary, all server/workspace shards, and
both required aggregates pass with normal CI concurrency. Earlier failed
attempts remain recorded below.

- Latest test-only qualification: 141/141
route/permissions/authentication cases pass in separate cold forks, with
plain server types and independent review clear. The real-runner suite
passes 35/35, with plain runner types and independent review clear. A
controlled premature-save acknowledgement fails as expected; matching
ownership/effect/process evidence, rejected saves, real turn outcome,
test abort, and pre-kill liveness are covered. No local reproduction of
the original CI scheduling failure is claimed. The preceding [CI
run](https://github.com/paperclipai/paperclip/actions/runs/34479680858)
passes 21/24 jobs, including all 995 Linux chat cases and browser
aggregate (104 passed, four existing optional skips); only Build, the
skills serialized shard, and the required verification aggregate fail.
Its exact-head Greptile review was 5/5. Both failed job logs are
retained.

- Final fixture qualification: all eight focused Discord cases and all
995 chat integration cases pass. The exact modal statement/PID is
observed before taking the real connection lock; the test then proves
its actual blocking relationship before mutation. Original SQL
execution, provider behavior, negative assertions, and 1s/15s timeouts
remain unchanged. Independent review is clear and test/production hashes
remain frozen. The preceding [CI
attempt](https://github.com/paperclipai/paperclip/actions/runs/34477184777)
passed 22 jobs, including Build/runner, typecheck, canary, all other
test shards, and browser aggregate (104 passed, four existing optional
skips); the two fixture failures and failed verification aggregate
remain recorded, not relabeled as a pass.

- Current queue-module composition: 308/308 recovery/batching/queue/Stop
tests; 995/995 full chat integration; 89/89 module tests, including real
PostgreSQL receipt-insert rollback; 24/24 workflow/module-boundary
tests; plain server and UI types. All four actual local process/ACP
browser paths pass in 1.4 minutes. Fresh databases, no skips or retries,
stable reviewed source hashes. The initial boundary failure is retained;
its no-op service wrapper was removed without changing recovery context
or weakening the check. An exploratory standalone test-directory
typecheck fails because its new upstream transformation config is not a
standalone typechecking project; standard CI/build does not invoke it,
and no configuration was weakened to suppress those diagnostics.

- The preceding head `e02a63d462ce5d47433b0aeb632bb6fd20aab1ba` passed
[all 24 CI
jobs](https://github.com/paperclipai/paperclip/actions/runs/34436462958)
and exact-head Greptile review at 5/5. Required CODEOWNER review
prevented its normal merge before master advanced again.

- Final extracted-module composition: 307/307 recovery, batching, queue
and Stop-control tests; 995/995 full chat integration; 49/49 module
tests including eight PostgreSQL adapter cases; and 19/19 issue-update
tests. Plain server types pass. All four actual local process/ACP
browser paths pass in 1.3 minutes. Fresh databases, no skips or retries
in these cohorts, frozen source hashes, and independent review clear.

- The preceding head `3e4e1c1c` passes [all PR CI
jobs](https://github.com/paperclipai/paperclip/actions/runs/34415826820),
including Build and required `ci / verify` and `ci / e2e`. Both the
original Rust failure and the previously load-sensitive lineage fixture
pass with unchanged Linux concurrency. Master advanced afterward and
required this reconciliation.
- Final master composition: 448/448 focused UI tests, 186/186 adapter
tests, 24/24 queue/control tests, and 11/11 packaging tests. Plain UI,
server, shared, and adapter types pass. Token gates and diff checks
pass. Independent server and UI reviews are clear.
- Stop-registration regression: both real-service cases fail against
exact `a95` source and pass with the fix. The full corrected
recovery/control suite passes 265/265. Duplicate-owner and failed-Stop
controls also pass. Plain server types pass. The readiness barrier
prevents provider startup without adding an acknowledgment to an already
terminal run.
- Final qualification strengthens terminal-field equality and repeats
both affected cases successfully on a fresh database. All four actual
local process/ACP browser paths pass again in 1.3 minutes, without skips
or retries. The final screenshot shows Cancelled, a paused subtree,
retained input, and no error toast.
- Two new actual-service regressions fail before the merge fix. They
prove that queued-comment adoption could consume a dedicated chat answer
or add unrelated input to that answer. The fixed four-case cohort
passes, including ordinary upstream continuation and adapter Stop
controls. Full recovery passes 257/257. All four actual local
process/ACP Stop browser flows pass in 1.4 minutes, without skips or
retries, on a fresh database.
- The unchanged runner artifact was qualified with 171/171 transport
tests, 870/870 API-authority tests, conformance 1/1, and replay 11/11.
Six controlled reader tests prove the exit/drain repair. Its local
serial Rust workspace passed 546 top-level cases plus two invoked
helpers; the later passing Linux CI supplies default-concurrency
evidence.
- Prior exact-source full chat integration passes 995/995. Settings
regressions cover concurrent stale pages, 501 destinations, pending
state, rejected updates, and explicit retry. These deterministic tests
do not prove live provider behavior.
- Retained failed attempts and their causes are in the [qualification
log](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-chat-queue-and-webhook-repair.md).
The first merge adapter run timed out while macOS slept for 290 seconds.
Its unchanged repeat passed with a temporary sleep guard. No assertion,
deadline, or CI gate was weakened.

Review commands include `pnpm --filter @paperclipai/server exec vitest
run src/__tests__/heartbeat-process-recovery.test.ts
src/__tests__/issue-queued-comments-routes.test.ts` and `pnpm exec
playwright test --config tests/e2e/playwright.config.ts
tests/e2e/acp-stop-continuation.spec.ts`. Database suites require fresh
disposable databases. See the [browser
runbook](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-04-chat-adapters-browser-e2e-runbook.md)
for provider setup and separate live acceptance steps.

## Risks

- This remains experimental. Deterministic tests and bounded live
evidence do not establish every provider feature, tenant, permission
layout, or media shape. Teams work-tenant qualification is still open.
- Failed and uncertain provider effects remain visible and can require
operator action. A transport receipt does not prove recipient
visibility.
- Native controller and runner artifacts must remain compatible.
Preserve lease ownership, terminal authority, source binding, and
quarantine during future changes.
- Access and audit rows commit together, but activity notifications
remain best-effort. This is not a new durable event outbox.
- The PR operation does not deploy a live server, replace its runner, or
change provider permissions. Remaining live qualification is documented
in the [temporary
handoff](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-open-qualification-followups.md).

## Model Used

OpenAI Codex assisted with implementation, tool execution, testing, and
review. The work records `gpt-6-astra` assistance. The environment does
not report a context-window size. No private reasoning traces are
included.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-10 10:06:45 -05:00
DottaandPaperclip 3b550c80fa fix(codex): correct startup trust, history reads, and resume usage (#13110)
## Thinking Path

> - Paperclip runs Codex locally and in remote sandboxes.
> - The runner must preserve startup configuration and session identity.
> - Missing project trust can disable repository configuration.
> - Full-history requests use deprecated provider fields.
> - Resume usage describes old work and must not become new run usage.
> - This change corrects startup trust, state reads, and usage
classification.

## Linked Issues or Issue Description

**What happened?**

Normal Codex runs could show repository-trust and history-deprecation
warnings.
Resume could report the preceding turn's token snapshot as a late-turn
warning.
The historical last-usage value could also be attributed to the new run.

**Expected behavior**

Trust the server-selected startup root in isolated configuration. Read
lightweight
provider state and paginated evidence. Use historical cumulative usage
as a
baseline without a new charge or user-facing warning.

**Steps to reproduce**

1. Start a native Codex task in a selected repository.
2. Finish the turn and resume the provider thread.
3. Inspect provider notices, history requests, and per-run usage.
4. Repeat startup and cold resume inside a Daytona sandbox.

**Paperclip version or commit**

Codex CLI 0.153.4 is the pinned runtime and reproduced baseline.
Replayed onto master at 6abeb6733. Related authority work: Refs #13092.
This PR retains its startup cleanup and protocol-integrity checks.

**Deployment mode**

Local source checkout and disposable Daytona sandbox.

## What Changed

- Classify the exact historical resume usage event before the generic
stale-turn warning.
- Persist cumulative usage baselines across recovery of the same run.
- Use excludeTurns on resume and lightweight thread reads.
- Page turn metadata and selected turn items with cursor and identity
validation.
- Reject unsupported or incomplete history instead of guessing that
execution is idle.
- Trust the startup execution root on its host, including Git worktree
trust keys.
- Start Codex in that root and retain the selected sandbox profile on
later turns.
- Keep unrelated isolated configuration and Codex's separate hook trust
policy.
- Add Rust, TypeScript, accounting, native integration, and local
run-log documentation.

## Verification

- Codex and native-transport TypeScript: 333 passed before PR replay.
- Adjacent OpenCode/ACPX driver and accounting tests: 49 passed.
- Rust library, serialized: 226 passed. Native Codex integration: 72
passed, 1 ignored, plus two pagination regressions.
- Repository typecheck and build passed. All repository test groups have
passing coverage after fixture and resource retests; the initial
monolithic command was not clean.
- Fresh real Codex native browser tasks returned correct answers without
the three targeted notices. Answers persisted after refresh and restart.
- Real same-thread TypeScript driver tests passed locally and in
Daytona, including cold resume, configuration, skills, and an approved
harmless hook.
- Local usage summed to 64,607 tokens. Daytona usage summed to 42,737
tokens. Each sum matched its final session total exactly.
- See doc/plans/2026-09-09-codex-integration-acceptance.md for the scope
and limits of the live tests.
- After replay onto current master and review fixes: 334 Codex, backend,
and live-session tests passed, including checkpoint serialization and
real-runner process restart. TypeScript checks passed.
- The native Codex integration run passed 83 tests; the large lineage
test passed separately with the release runner (its debug build exceeded
the test deadline).
- All GitHub checks passed on the final PR head. Greptile is 5/5 with no
unresolved review threads. CI regenerates the lockfile for the added
TOML dependency, per repository policy.
- The first server shard hit a timing-dependent duplicate-key failure in
the unchanged artifact-document concurrency test. Its focused 11-test
suite passed locally. One CI retry on the same head passed all 103 files
and 1,405 tests (2 skipped): [retry
result](https://github.com/paperclipai/paperclip/actions/runs/34398832930/job/102631274667).

## Risks

- Trust applies only to the server-selected startup root and isolated
configuration. Sandbox and tool permissions remain authoritative.
- Codex still requires approval of individual hook hashes. This change
does not bypass that policy.
- Providers without the required history APIs fail explicitly.
- Daytona acceptance used the production TypeScript driver. Remote
Paperclip UI and remote Rust execution were not tested.
- No new public API, database state, recovery policy, or UI control is
included.

## Model Used

OpenAI Codex, GPT-6 (`gpt-6-astra`). Used for reasoning, code edits,
tool use,
and test execution. The exact context-window limit is not exposed in
this
session. Real-provider acceptance used Codex CLI 0.153.4 with
`gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-09 15:35:18 -05:00
DottaandPaperclip 6abeb67334 feat: add opt-in chat provider and data foundation (#13100)
Add dormant provider contracts, qualified patched adapters, tenant-scoped persistence and lifecycle ownership without activating chat routes. Preserve the experimental integration as dependent PR #13038.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-09 13:49:12 -05:00
DottaandPaperclip 35fdc0c66b fix: make task recovery durable and preserve current requests (#13075)
Make task recovery durable and preserve the latest user request across native and legacy continuations. Keep routine recovery quiet and prevent replay when action outcomes are uncertain.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-09 09:14:25 -05:00
DottaandPaperclip 5bddff0920 feat(runner): add guarded API search and call fallback (#13003)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - The new runner gives agents dedicated tools for common tasks.
> - Some API operations and parameters have no dedicated tool.
> - Agents need a controlled way to find and use those operations.
> - This pull request adds API search and calls through the real server
routes.
> - Existing tools remain the preferred path. The new tools are disabled
by default.
> - Paired tests measure correctness, tool choice, cost and time.

## Linked Issues or Issue Description

**Subsystem affected**

Paperclip Runner contracts, production tool authority and the server API
catalog.

**Problem or motivation**

The runner cannot use much of the API described by the old Paperclip
skill. A generic HTTP client would also let agents bypass runner control
rules.

**Proposed solution**

Add `search_api` and `call_api`. Resolve calls from the mounted API
catalog. Use server-held, run-bound credentials. Preserve route checks
and runner lifecycle rules. Keep the tools disabled until an operator
enables selected companies.

**Alternatives considered**

A dedicated tool for every endpoint would add a large initial prompt. An
unrestricted HTTP tool would weaken authorization and replay controls.

**Roadmap alignment**

This extends the native runner tooling. The repository owner requested
this design and implementation. The roadmap and related open PRs were
checked. No duplicate API escape-hatch PR was found.

## What Changed

- Register two compact fallback tools in canonical contracts and
provider projections.
- Build deterministic API discovery from OpenAPI, mounted experimental
routes and the old skill reference.
- Execute bounded JSON, text, file and download requests through
authenticated HTTP routes.
- Recheck active runs, company access and work modes. Block runner
lifecycle, scheduling, credential and approval bypasses. Keep routine
annotation collaboration available.
- Retain mutation receipts. Report uncertain outcomes without blindly
repeating writes.
- Add a company rollout gate and a durable eval worker with complete
cost accounting checks.
- Record child-task creation in the activity log with the agent and run.
- Add contract, authorization, file, replay and real runnerd/PRP/HTTP
tests.
- Document rollout gates and paid coverage limits. The companion eval
repository retains immutable attempts and reports.

## Verification

- Final app commit `da58370524c3626a744eec20164397c5fb6ba9ef`: all 32
checks passed; the unrelated Storybook visual check was skipped.
Greptile 5/5; no unresolved review threads.

- Full Linux build and recursive typecheck passed. Repository tests were
run by project and serialized shard; all 143 serialized server suites
passed.
- Runner TypeScript: 1,599 passed, two skipped. Rust release: 451
passing test reports. Conformance and replay parity passed. The required
API check passed 837 tests, including runnerd → PRP → authority → real
HTTP.
- Bindings cannot enable API tools without the explicit deployment flag.
Unit and real-authority tests prove the default-off boundary.
- The standalone API check builds and stages its own binary. It passed
after existing staged and debug binaries were removed from the test
container.
- UI and CLI tests passed. Initial environment failures (missing jq,
Docker overlay file identity, and parallel linker memory pressure) and
focused passing reruns are retained. The macOS full runner suite has
platform-specific failures; Linux is the qualified full-check platform.
- Eval harness: 27 tests passed; existing CI discovery ran 86 tests with
two unrelated skips. Credential export rejection is tested against the
actual report command.
- Luna and OpenRouter Sonnet each passed 60 common-workflow runs: ten
workflows, three repetitions per arm, zero unnecessary API fallback.
- Sonnet passed 11 selected capability/contract cases after fixes.
Gemini passed three smoke cases. DeepSeek exceeded the 120-second limit
and remains unqualified.
- Luna's two cost flags received focused follow-up. The original flags
and a later n=1 latency flag remain visible. Sonnet had no cost or
latency increase above 20%.
- The catalog contains 785 entries; 58 were exercised across all stages.
Most operation probes remain unrun and some need additional fixtures.
Authored probes do not establish successful coverage.
- Total conservative accounted cost: $9.875960. Active paid-campaign
time: 88.16/90 minutes. No missing accounting. Later security and
harness fixes have provider-free verification; no paid validation is
claimed for those revisions.
- Inspect the [qualification
report](https://github.com/paperclipai/paperclip-evals/blob/codex/seach-call-api-tools/evals/runner-api-tools/reports/2026-09-07-production/READINESS.md)
and [verification
record](https://github.com/paperclipai/paperclip-evals/blob/codex/seach-call-api-tools/evals/runner-api-tools/reports/2026-09-07-production/verification.json).

## Risks

- This is a broad authenticated API surface. Keep the default-off gate
until an operator selects initial rollout companies.
- Paid coverage is incomplete. Small regression samples do not prove all
workflows are unchanged.
- A timeout or server failure can follow a committed mutation. The
result reports an unknown outcome and requires state inspection.
- The new definitions add prompt tokens. The report retains cost flags
and cache variation.
- No database migration is required.
- Repository rules require code-owner approval before merge. Technical
CI and automated review are complete.

## Model Used

OpenAI Codex based on GPT-6 assisted with code, tests and review. The
exact serving model ID and context window are not exposed in this
session. It used reasoning, tool calls and code execution.

Eval models: `gpt-5.6-luna` with low reasoning,
`openrouter/anthropic/claude-sonnet-5`,
`openrouter/google/gemini-3.8-flash`, and
`openrouter/deepseek/deepseek-v4-flash-0731`. Attempts retain runtime
versions, model identity, usage and source provenance.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-07 14:14:43 -05:00
DottaandPaperclip a7e6b818e9 feat(apps): add Paperclip Cloud managed OAuth connector (#12600)
## Thinking Path

> - Paperclip lets operators give governed tools to AI agents.
> - Connected Apps already support provider OAuth and personal
connection grants.
> - Some providers require one stable callback and do not support
dynamic client registration.
> - Self-hosted Paperclip instances can run at private or changeable
origins.
> - Paperclip Cloud can provide the stable callback while each instance
keeps its durable provider credentials.
> - This pull request adds the instance side of that managed OAuth
protocol and keeps customer-created clients available.
> - The benefit is a safe path to one-click Workspace connections for
hosted and enrolled self-hosted instances.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting. This change updates the server, Apps UI, shared app
definitions, and connection documentation.

**Problem or motivation**

Some OAuth providers require a pre-registered callback and
provider-owned client. An arbitrary self-hosted Paperclip origin cannot
use that client callback directly. Paperclip ID must also stay limited
to product identity instead of resource authorization.

**Proposed solution**

Use the existing Paperclip Cloud application as the fixed callback
broker. Enroll each instance to an exact origin and separate Ed25519 and
X25519 keys. Bind every request and sealed envelope to the instance,
environment, user, company, provider, profile, and exact scope set.
Store durable provider credentials only in the originating instance
vault.

**Alternatives considered**

Customer-created OAuth clients remain available as the independent
fallback. A generic redirect relay was rejected because it would allow
caller-selected destinations and scopes. Paperclip ID was rejected as
the broker because it is the identity boundary. A new service was
rejected because the existing Cloud application already owns customer
login and the public callback origin.

**Roadmap alignment**

This work extends the shipped MCP Tool Gateway and Apps milestone. It
also supports the Connected Apps and Cloud deployments roadmap items.

Companion Cloud implementation:
https://github.com/paperclipai/paperclip-cloud/pull/312

The duplicate search found no related open Paperclip PR or issue.

## What Changed

- Add a `paperclip_cloud_connector` client with signed requests, exact
profile and scope bindings, and X25519-sealed credential handling.
- Add explicit self-hosted enrollment with owner-only instance key
storage and exact HTTPS origins.
- Route managed Google Workspace setup through Paperclip Cloud and
preserve customer-created OAuth clients.
- Keep broker claims retryable until the local vault transaction
commits.
- Keep managed Google per-profile removal local-only to avoid
client-wide provider revocation.
- Add setup status to the Connections page and retain the Paperclip ID
names as compatibility aliases.
- Document the trust boundaries, enrollment, callback, refresh, removal,
and rollout flows.

## Verification

- `pnpm -r typecheck`
- `pnpm --filter @paperclipai/shared exec vitest run
src/app-definitions.test.ts`
- `pnpm --filter @paperclipai/server exec vitest run
src/services/paperclip-cloud-connector.test.ts
src/services/paperclip-cloud-connector-enrollment.test.ts`
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/tool-access-service.test.ts -t 'brokered Gmail
OAuth|brokered OAuth state'`
- `pnpm --filter @paperclipai/ui exec vitest run
src/pages/apps/Connections.test.tsx`
- `pnpm check:token-gates`
- `pnpm build`
- The full stable test runner also reproduced existing macOS workspace,
skill-discovery, and listener fixture failures outside the changed
paths. GitHub Linux CI is the authoritative full-suite result.

## Risks

- The managed flow depends on
https://github.com/paperclipai/paperclip-cloud/pull/312. Real provider
profiles stay disabled until Cloud deploys that protocol and the
provider approves the managed client.
- A Cloud outage blocks new authorization and refresh. Existing access
tokens continue to work until expiry.
- Managed Google profile removal only deletes the local grant. This
avoids invalidating the user's other profiles that share the managed
Google client.
- Legacy `paperclip_id_connector` records require a reconnect after
their current access tokens expire. Old Paperclip ID keys and refresh
tokens are not sent to Paperclip Cloud.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI GPT-5.6 (Codex). Agentic coding, tool use, code execution, and
subagents were enabled. The context-window size is not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-31 14:34:46 -05:00
Dotta d387cc0ff0 feat(connections): add managed external MCP connectors (#12346)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connection intents need secure provider implementations to complete
setup.
> - Some providers use managed OAuth or external credential brokers.
> - Those tokens must stay out of durable Paperclip state and fail
closed when refresh fails.
> - This pull request adds managed connector backends and the required
storage contract.
> - The benefit is safer provider setup with governed credential
lifecycles.

## Linked Issues or Issue Description

Refs #11965

This is stack 8 of 11. It depends on stack 7 and replaces another
reviewable part of #11965.

## What Changed

- Add managed Google Workspace and external connector backends.
- Add Vercel Connect support without storing provider bearer tokens.
- Add replay-safe migration 0232 and its generated snapshot.
- Fail closed and clear stale token bindings when organization OAuth
refresh needs reauthorization.

## Verification

- `pnpm --filter @paperclipai/server typecheck`
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/tool-access-service.test.ts`
- Result: 194 tests passed.
- `pnpm --filter @paperclipai/db check:migrations`
- `pnpm build`
- `pnpm exec vitest run --project @paperclipai/server
server/src/services/remote-url-credentials.test.ts` (5 passed, including
URL userinfo vault extraction)

## Risks

- Broker metadata errors can block provider setup.
- OAuth refresh failure disables the shared organization connection
until reauthorization.
- Migration 0232 is generated, ordered after 0231, and safe to replay.

> I checked `ROADMAP.md`. This stack continues the existing app
connection work from #11965 and does not duplicate another planned item.

## Model Used

OpenAI Codex, GPT-5. The runtime model ID and context window were not
exposed. The model used reasoning, tool use, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have linked the public source pull request with `Refs #`
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-29 12:08:34 -05:00
DottaandPaperclip b3343dbd64 feat(connections): add self-serve intent runtime (#12345)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents need a governed way to request app connections during issue
work.
> - The catalog now describes the available providers and setup methods.
> - A request must become a durable, company-scoped intent before an
operator acts on it.
> - This pull request adds that intent runtime across server, agent,
CLI, and shared contracts.
> - The benefit is a safe bridge from agent need to operator-approved
setup.

## Linked Issues or Issue Description

Refs #11965

This is stack 7 of 11. It depends on stack 6 and replaces another
reviewable part of #11965.

## What Changed

- Add connection intent types, validation, service logic, and routes.
- Add agent runtime tools and CLI support for connection requests.
- Add issue-thread interaction support for connection intents.
- Add runtime, route, adapter, and contract tests.
- Hold the final resolved-continuation row lock through asynchronous
adapter preparation until an actual process spawn, so parking or
reassignment cannot cross that boundary.
- Report Hermes Gateway's first remote run request through the shared
dispatch hook so the resolved-intent lock is released at the true
dispatch boundary.
- Revalidate the addressed user's live non-viewer membership and
connection-management authority for every intent mutation, including
OAuth completion.

## Verification

- `pnpm --filter @paperclipai/server typecheck`
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/tool-access-service.test.ts`
- Result: 176 tests passed.
- `pnpm build`
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/heartbeat-stale-queue-invalidation.test.ts` (32 passed;
includes non-process dispatch lock-release coverage)
- `pnpm exec vitest run --project @paperclipai/server
server/src/__tests__/connection-intents-service.test.ts -t
"addressed-user mutation"` (1 passed)
- `pnpm exec vitest run --project @paperclipai/server
server/src/__tests__/tool-access-service.test.ts -t "binds OAuth
callback completion to the initiating board session"` (1 passed)
- `pnpm --filter @paperclipai/hermes-paperclip-adapter test --
src/gateway/server/execute.test.ts` (23 passed; includes dispatch-hook
ordering and exactly-once coverage)
- `pnpm --filter @paperclipai/hermes-paperclip-adapter typecheck`

## Risks

- A malformed intent could create an unusable operator request.
- Validators and company checks reject invalid or cross-company
requests.
- The final continuation gate holds the issue row lock through adapter
preparation until process or remote dispatch; later operator changes use
the normal active-run interruption path.
- The change does not add a database migration.

> I checked `ROADMAP.md`. This stack continues the existing app
connection work from #11965 and does not duplicate another planned item.

## Model Used

OpenAI Codex, GPT-5. The runtime model ID and context window were not
exposed. The model used reasoning, tool use, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have linked the public source pull request with `Refs #`
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-29 12:08:34 -05:00
DottaandPaperclip 3db2e6bdd2 feat(mcp) [split 8/8]: add e2e coverage and operator docs (#9563)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Governed MCP access spans contracts, runtime enforcement, adapters,
UI surfaces, and operator verification
> - The parity reference PR #9534 is too large for effective automated
or human review
> - The feature therefore needs a linear stack whose individual diffs
stay below the 100-file review limit
> - This pull request is split 8/8 and focuses on end-to-end coverage,
operator docs, evals, and release notes
> - The benefit is a standalone, testable review boundary while
preserving byte-for-byte parity at the top of the stack

## Linked Issues or Issue Description

- Related parity reference: #9534
- Problem: The complete stack needs discoverable browser scenarios,
operator guidance, threat modeling, eval coverage, and a parity proof
before merge.
- Proposed solution: Adds MCP user-story and Smoke Lab e2e suites,
docs/evals/release notes, the skill update, and the root e2e driver
script registration.
- Alternatives considered: keeping #9534 as one 403-file review, or
rewriting the feature to manufacture seams; both were rejected in favor
of path extraction plus compile-driven boundary moves.
- Roadmap alignment: this advances the existing governed MCP/tool-access
work already represented by #9534; it does not introduce a separate
roadmap initiative.
- Stack position: base branch is `pap10341-split/07-ui-apps-activation`.
- Merge policy: merge bottom-up, in order, only after the complete
eight-PR stack has been reviewed and the top-of-stack parity gate
remains empty.
- Requested review: QA for flag audit and e2e/browser acceptance;
Greptile on every PR.

## What Changed

- Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release
notes, the skill update, and the root e2e driver script registration.
- Keeps this PR below 100 changed files and independently typecheckable.
- Preserves the final tree from #9534 when combined with the other seven
stack levels.

## Verification

- `pnpm typecheck`
- `node --check scripts/e2e-mcp-user-stories.mjs`
- `pnpm exec playwright test --config tests/e2e/playwright.config.ts
--list` — 43 tests discovered
- `git diff pap10341-split/08-e2e-docs
6b40e3876d9297105d4ec306e47e46d351c86172` — empty (0 bytes)

## Risks

- Browser suites depend on runtime services and environment setup; this
PR validates discovery locally while QA owns full flag-on/flag-off
execution.
- Stack risk: merging out of order can expose incomplete layers;
mitigate by following the documented bottom-up merge policy.
- Parity risk: later edits to an intermediate branch can drift from
#9534; mitigate by re-running the empty top-of-stack diff before merge.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context
window; medium reasoning with repository, shell, Git, GitHub CLI, and
code-execution tools enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] Internal references are omitted except the execution-plan link
explicitly required for this coordinated split stack
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge


## Stack Coordination

- Internal execution plan:
[PAP-13874](/PAP/issues/PAP-13874#document-plan)
- Parity reference: #9534
- Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563
- Merge bottom-up only after full-stack review and an empty parity diff
at #9563.

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-14 15:48:57 -05:00
DottaandPaperclip 2dbaf4a7fa External object references across issue surfaces (#8512)
## Thinking Path

> - Paperclip is the open source control plane people use to coordinate
AI agents, issues, approvals, comments, and work products.
> - The involved subsystem is issue context: markdown links, issue
properties, related work, lists, filters, inbox/sidebar status, and
plugin-provided external context.
> - The gap is that URLs to external systems currently remain mostly
plain links, so humans and agents must manually open them to understand
status, identity, and liveness.
> - This matters because external work objects such as GitHub issues and
pull requests are part of the operational state of a Paperclip company.
> - The implementation keeps core provider-neutral: shared contracts,
storage, sync, routes, and UI surfaces live in core while providers can
contribute detection and status resolution.
> - This pull request adds the external object reference foundation,
GitHub provider support, issue-surface rendering, filters,
sidebar/list/inbox signals, and test/story coverage.
> - The benefit is that linked external work becomes inspectable
Paperclip context without hardcoding every provider directly into the
UI.

## Linked Issues or Issue Description

No public GitHub issue exists for this work.

Feature request:

- Problem: URLs in Paperclip issues, comments, documents, and related
surfaces do not expose provider status or object identity inline.
- Proposed behavior: detect supported external object URLs, persist
normalized references, refresh provider status, and render concise
status-aware links across issue surfaces.
- Users affected: board users, agents, and maintainers who triage issues
containing external work links.
- Acceptance: external object references are company-scoped,
provider-extensible, visible in key issue surfaces, filterable where
relevant, and covered by focused shared/server/UI tests.

Related PR search:

- No open duplicate PRs found for `external object references`.
- Closed related prior attempt: #4556.

## What Changed

- Added shared external-object contracts, validators, status/liveness
helpers, and plugin protocol declarations.
- Added database schema and additive migrations for external objects,
source mentions, and display metadata.
- Added server services/routes for detecting, syncing, summarizing,
refreshing, and resolving external objects across issues, documents,
comments, projects, and plugins.
- Added a GitHub external-object provider plus plugin SDK authoring
docs.
- Wired UI presentation across markdown links, comments, issue chat,
documents, properties, related work, issue rows, filters, inbox/sidebar
badges, and Storybook stories.
- Rebasing cleanup: moved the branch onto current `master`, repaired
stale worktree provision config, hardened environment-sensitive
tests/mocks, and removed committed screenshot artifacts from the PR
branch to keep the reviewable file set below tool limits.

## Verification

- `pnpm exec vitest run packages/shared/src/external-objects.test.ts
server/src/__tests__/external-object-routes.test.ts
server/src/__tests__/external-objects-service.test.ts
ui/src/components/ExternalObjectPill.test.tsx
ui/src/lib/external-objects.test.ts` passed after rebasing: 5 files, 56
tests.
- Historical branch verification before this PR creation included `pnpm
test:run`, `pnpm -r typecheck`, and `pnpm build`; this PR body does not
claim those were rerun after the final rebase.

## Risks

- Medium: this adds a new cross-surface sync path on
issue/document/comment writes. The implementation uses safe sync
wrappers so external-object failures warn instead of blocking core
mutations.
- Medium: the migrations introduce new tables and indexes. They are
additive and company-scoped.
- Medium: provider-specific URL parsing can miss or misclassify edge
cases. Shared canonicalization tests and provider tests cover current
GitHub shapes.
- Low: UI badge/filter behavior could add visual noise for object-heavy
issues; component tests and Storybook stories cover the intended
surfaces.

> Roadmap checked: `ROADMAP.md` references the plugin system as the
current extension path and does not list a duplicate core feature.
Related long-range docs discuss external references, work products,
preview URLs, and plugin extension points; this PR implements the scoped
external-object reference foundation.

## Model Used

OpenAI Codex, GPT-5 coding-agent runtime, with shell and GitHub CLI tool
use. Reasoning mode: medium. Exact deployed runtime model ID and context
window were not exposed in the environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-06-23 08:27:19 -05:00
DottaandPaperclip dbebf30c89 Add low-trust review containment (#7530)
## Thinking Path

> - Paperclip is a control plane for AI-agent companies, so execution
policy and trust boundaries are part of the product's safety contract.
> - Low-trust review work needs narrower authority than normal
same-company agents because hostile PRs, comments, attachments, and
generated output can carry prompt-injection payloads.
> - The current V1 shape gives trusted workers broad company context,
which is useful for normal execution but too permissive for a reviewer
assigned to hostile content.
> - This branch adds a `low_trust_review` preset, source-trust tagging,
route-level containment, and quarantine handling so low-trust output
does not automatically flow into higher-trust wake context.
> - The branch has been rebased onto current `origin/master`, and the
low-trust migration was renumbered to `0097_low_trust_source_trust.sql`
to avoid collisions with existing `0091` through `0096` migrations.
> - Greptile feedback was addressed by tightening low-trust detection,
preserving project-level trust policy checks, fixing issue-kind
promotion lookup, removing duplicate post-lease isolation assertion,
documenting fail-closed source-trust behavior, bounding ancestry checks,
enforcing runtime issue context for CEOs, awaiting accepted-plan monitor
authorization, and making low-trust issue source-trust tagging atomic.
> - The benefit is a first production slice of deny-by-default review
containment with regression coverage for the main control-plane pivot
surfaces.

Fixes #7531.

## What Changed

- Added shared trust-policy types and validators, plus
database/source-trust fields for issues, comments, documents, and work
products.
- Implemented server enforcement for low-trust issue scope, agent
self-view redaction, secret/plugin/runtime denial paths, promotion
checks, and quarantined continuation/wake context.
- Added focused low-trust regression tests for resolver behavior, source
trust, route authorization, heartbeat preflight ordering, runtime
containment, and quarantine redaction.
- Added board UI affordances for selecting/reviewing the low-trust
preset and surfacing source-trust badges in relevant issue views.
- Added `doc/LOW-TRUST-PRESETS.md`, updated
`doc/SPEC-implementation.md`, and committed the low-trust review
contract plan under `doc/plans/`.
- Rebasing note: the original `0097_low_trust_source_trust.sql`
migration was renamed to `0097_low_trust_source_trust.sql`; the SQL uses
`ADD COLUMN IF NOT EXISTS` so users who already applied the old-numbered
migration are not broken by the renumbered migration.

## Verification

- Rebased branch onto current `origin/master` and force-pushed with
lease to `origin/PAP-10211-low-trust-agent` at head `2719f31e3`.
- Confirmed the PR diff does not include `pnpm-lock.yaml` or
`.github/workflows` changes.
- Resolved upstream UI/comment conflicts by preserving deleted-comment
tombstone behavior and low-trust source-trust badges/metadata.
- Renumbered the low-trust source-trust migration to
`0097_low_trust_source_trust.sql`; the SQL uses `ADD COLUMN IF NOT
EXISTS` so users who already applied an old-numbered copy are not
broken.
- `pnpm exec vitest run ui/src/lib/issue-chat-messages.test.ts
server/src/__tests__/heartbeat-workspace-session.test.ts`
- `pnpm exec vitest run server/src/__tests__/source-trust.test.ts
server/src/__tests__/workspace-runtime-service-authz.test.ts
ui/src/lib/trust-policy-ui.test.ts
ui/src/components/TrustPresetSection.test.tsx`
- `pnpm run typecheck:build-gaps`
- `git diff --check`
- GitHub checks pass on head `2719f31e3`: build, typecheck/release
registry, general tests, serialized server suites, e2e, canary, verify,
policy/review, Socket, and Snyk.
- Greptile Review passes with Confidence Score 5/5 and zero unresolved
Greptile review threads.
- No design screenshots/images were added because the task explicitly
says not to add them unless they are specifically part of the work.

## Risks

- Medium risk: this touches shared trust-policy contracts, server
authorization paths, heartbeat context generation, migration metadata,
and UI preset controls.
- Low-trust containment is intentionally deny-by-default; legitimate
future review workflows may need explicit allowlisted exceptions.
- Plugin/runtime/security surfaces are broad, so regression tests cover
the current known routes but future integrations must route through the
same containment layer.
- The PR is ready for review; GitHub checks are green and Greptile is
5/5.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex, GPT-5 coding agent, tool-enabled shell and GitHub CLI
workflow.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] UI changes are covered by focused tests; no screenshots were added
per task instruction not to add design images unless specifically
required
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-06-05 16:48:02 -05:00
Aron PrinsandDevin Foley 70b1a9109d Improve CLI API parity coverage (#6626)
## Thinking Path

> - Paperclip is a control plane for AI-agent companies, with the CLI
acting as a scriptable operator and agent interface to that control
plane.
> - The REST API surface has grown across companies, agents, issues,
routines, plugins, auth, workspaces, secrets, and operational inspection
commands.
> - The CLI had drifted from that API surface: some commands were
missing, some command shapes differed from docs/reference material, and
several edge cases only failed during end-to-end local-source testing.
> - The local development runbook requires these tests to be disposable
and isolated from a real `~/.paperclip`, `~/.codex`, or `~/.claude`
installation.
> - This pull request adds broad CLI/API parity coverage, fixes the
actionable bugs found during that pass, and records the reproducible
test log under `doc/logs`.
> - The benefit is a more complete, scriptable CLI surface with
regression coverage for the command families exercised by the parity
run.

## What Changed

- Added or expanded CLI command coverage for access/auth, companies,
agents, projects, goals, issues and subresources, routines, plugins,
workspaces, activity/run/cost/dashboard inspection, assets, skills,
secrets, tokens, prompt/wake flows, and local setup helpers.
- Fixed CLI/API parity bugs found during the run, including context
profile patching, issue interaction optional payloads, malformed
tree-hold errors, environment duplicate handling, configure
invalid-section exit codes, worktree pnpm invocation, token agent ID
resolution, plugin tool worker lookup, and routine webhook secret
cleanup.
- Added missing CLI wrappers and route coverage for health/access,
invite resolution URL forwarding, join status normalization, secret
lifecycle commands, LLM docs routes, available-skill isolation, positive
board-claim coverage, and interactive `connect` prompt-flow tests.
- Added a schema-backed `/api/openapi.json` route sufficient for CLI
parity and `paperclipai openapi --json` smoke coverage.
- Added `doc/logs/2026-05-24-cli-api-parity-e2e-log.md` with the
detailed living test/bug log and renamed the log directory from
`doc/bugs` to `doc/logs`.
- Added `doc/plans/2026-05-23-cli-api-parity.md` and the OpenAPI parity
reference used during the pass.

OpenAPI note: this PR intentionally does not try to subsume
`feature/openapi-spec`. The OpenAPI implementation here is schema-backed
and better than the earlier route-inventory stub, but
`feature/openapi-spec` is the fuller/better OpenAPI branch because it
includes exact mounted-route coverage tests and additional current route
coverage. That branch should stay as its own PR and can supersede this
OpenAPI route implementation.

## Verification

Targeted automated checks run:

- `pnpm exec vitest run server/src/__tests__/openapi-routes.test.ts`
- `pnpm exec vitest run server/src/__tests__/board-claim.test.ts`
- `pnpm exec vitest run cli/src/__tests__/connect.test.ts`
- `pnpm exec vitest run cli/src/__tests__/agent-lifecycle.test.ts`
- `pnpm exec vitest run server/src/__tests__/plugin-database.test.ts`
- `pnpm exec vitest run server/src/__tests__/routines-service.test.ts`
- `pnpm --dir cli typecheck`
- `pnpm --dir server typecheck`

Manual/local E2E verification:

- Ran the full disposable local-source CLI/API parity pass with isolated
`PAPERCLIP_HOME`, `PAPERCLIP_CONFIG`, `PAPERCLIP_CONTEXT`,
`PAPERCLIP_AUTH_STORE`, `CODEX_HOME`, and `CLAUDE_HOME` under
`tmp/cli-api-parity`.
- Verified `DATABASE_URL` and `DATABASE_MIGRATION_URL` stayed unset for
the scratch server.
- Verified live health and schema-backed OpenAPI responses on
non-default port `3197`.
- Revoked created board/agent tokens and cleaned up temporary plugins,
secrets, non-default environments, and project workspaces.
- See `doc/logs/2026-05-24-cli-api-parity-e2e-log.md` for the full
command-by-command reproduction log.

Not run:

- Full `pnpm test`, `pnpm test:run`, or `pnpm build` were not run after
the entire branch because the branch is broad and the parity pass used
focused test/typecheck verification plus live isolated CLI reruns.

## Risks

- This is a broad PR and touches many CLI command modules, so review
surface is high. The changes are grouped around one theme, but a split
may be easier if maintainers prefer narrower PRs.
- The OpenAPI route in this PR is not the final/best OpenAPI
implementation. `feature/openapi-spec` has stronger exact-route coverage
and should remain the source for the dedicated OpenAPI PR.
- The living log is intentionally detailed and large. It is useful for
reproducibility but adds documentation weight.
- No UI changes are intended; screenshots are not applicable.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex, GPT-5-based coding agent in Codex desktop. Exact served
model/context-window identifier was not exposed in the local app. Work
used shell/Git/GitHub CLI tooling, local source inspection, targeted
test execution, and live isolated Paperclip CLI/API smoke testing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] If this change affects the UI, I have included before/after
screenshots
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Devin Foley <devin@devinfoley.com>
2026-06-02 17:13:29 -07:00
DottaandPaperclip 9eac727cf1 [codex] Add skills CLI and catalog management (#6782)
## Thinking Path

> - Paperclip orchestrates AI agents for zero-human companies through
company-scoped control-plane workflows.
> - Agents need reusable, inspectable skills that can be installed,
reset, audited, exported, and assigned without bespoke local setup.
> - The existing skill truth model needed cleanup so bundled skills,
optional catalog skills, runtime skills, and adapter-provided skills
have clear provenance.
> - Operators also need a practical CLI and board UI for discovering and
managing company skills.
> - This pull request adds the skills CLI, packaged skills catalog,
company skills APIs, and catalog-aware board UI.
> - The benefit is a more reusable Paperclip company setup where skills
are portable, auditable, and easier for operators and agents to manage.

## What Changed

- Added `paperclipai skills` CLI commands and coverage for catalog
listing, installing, resetting, and inspecting company skills.
- Added a packaged `@paperclipai/skills-catalog` workspace with bundled
and optional skill content plus validation/build tests.
- Added shared company-skill types and validators used across CLI,
server, and UI contracts.
- Added server catalog APIs/services for company skill catalog
operations, reset semantics, audit behavior, and portability provenance.
- Updated adapter skill handling so runtime/catalog provenance remains
explicit across local adapters.
- Added board UI support for browsing and managing catalog-backed
company skills.
- Updated docs for the skills CLI/catalog flow and the company skills
Paperclip skill reference.
- Rebased the branch onto current `paperclipai/paperclip:master`; no
`pnpm-lock.yaml`, `.github/workflows`, or migration files are included
in the final PR diff.

## Verification

- Passed: `pnpm run preflight:workspace-links && pnpm exec vitest run
cli/src/__tests__/skills.test.ts
packages/skills-catalog/src/catalog-builder.test.ts
packages/skills-catalog/src/shipped-catalog.test.ts
packages/shared/src/validators/company-skill.test.ts
packages/adapter-utils/src/server-utils.test.ts
packages/plugins/create-paperclip-plugin/src/entrypoints.test.ts
server/src/__tests__/company-skills-catalog-service.test.ts
server/src/__tests__/company-skills-routes.test.ts
server/src/__tests__/company-portability.test.ts`.
- Passed: `pnpm exec vitest run
server/src/__tests__/workspace-runtime.test.ts -t "default
branch|origin/master|symbolic-ref"`.
- Attempted: full `server/src/__tests__/workspace-runtime.test.ts`. Four
provisioning tests failed while seeding an isolated worktree database
from the local Paperclip instance because the local plugin schema dump
contains a duplicate-column foreign key
(`plugin_content_machine_18a7bc327b.content_case_signals`). The
default-branch tests touched by the rebase conflict passed in the
focused run above.
- Checked final diff: no `pnpm-lock.yaml`, no `.github/workflows`, and
no migration-file changes relative to `master`.

## Risks

- Medium: this is a broad skills/catalog change touching CLI, server
APIs, shared contracts, adapter skill sync, and UI.
- Catalog validation and reset semantics need careful reviewer attention
because they affect reusable company setup and portability.
- No database migrations are included in this PR, so there is no
migration ordering/idempotency risk in the final diff.
- No lockfile is included by design; dependency resolution will be
handled by the repository lockfile workflow.

## Model Used

- OpenAI Codex coding agent based on GPT-5, running in Paperclip via the
`codex_local` adapter with shell, git, GitHub CLI, and code-editing tool
access. Exact hosted model build/context-window metadata is not exposed
in this runtime.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have run targeted tests locally and documented the local
workspace-runtime seed failure above
- [x] I have added or updated tests where applicable
- [x] If this change affects the UI, screenshots were intentionally
omitted per PAP-10124 instructions; UI behavior is covered by tests and
reviewer inspection
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-05-28 07:33:51 -10:00
Emad Ibrahim afb73ba553 Scale issue kanban board for high-volume columns (#5309)
## Thinking Path

> - Paperclip is a control plane for autonomous AI-agent companies, and
the board UI needs to keep operator visibility clear as company work
scales.
> - The involved subsystem is the Issues page board mode, specifically
the Kanban rendering path for issue status columns.
> - The current board keeps the classic Kanban model, but high-volume
columns can become tall, slow, and hard to scan when hundreds of issues
are loaded.
> - We explored alternatives and chose the conservative Scaled Kanban
direction: preserve status lanes and drag/drop, but bound visible cards
and collapse low-signal lanes.
> - This pull request adds UI-only density controls and high-volume
defaults rather than introducing schema or API changes.
> - The benefit is a board that remains usable with large issue
inventories while keeping active workflow lanes visible.

## What Changed

- Added scaled Kanban behavior with compact cards, collapsed cold-lane
rails, per-column visible-card limits, and per-column "show more" reveal
controls.
- Added persisted board density preferences to the Issues page view
state, scoped through the existing company-specific localStorage path.
- Added board toolbar controls for compact cards, collapsed cold lanes,
cards-per-column page size (`10`, `25`, `50`), and density reset.
- Added a design spec and implementation plan under `doc/plans/`.
- Added focused Vitest coverage for `KanbanBoard` and `IssuesList`
high-volume board behavior.

## Verification

- `pnpm exec vitest run ui/src/components/IssuesList.test.tsx
ui/src/components/KanbanBoard.test.tsx` — pass, 35 tests.
- `pnpm -r typecheck` — pass.
- `pnpm build` — pass before the upstream merge; not rerun after
docs/assets cleanup.
- `curl -fsS http://127.0.0.1:3100/api/health` — pass against restarted
local dev server after applying pending migration
`0078_white_darwin.sql`.
- `pnpm test:run` — previously failed in unrelated Cursor remote-sandbox
server tests:
- `server/src/__tests__/cursor-local-adapter-environment.test.ts`:
expected probe status `pass`, received `fail`.
- `server/src/__tests__/cursor-local-execute.test.ts`: two remote
sandbox execution cases exited `127` instead of `0`.

Local dev server for manual UI inspection: `http://127.0.0.1:3100`.

Screenshots were captured for review and attached in the PR thread
rather than committed to source.

## Risks

- Low schema/API risk: this is UI-only and uses the existing issue data
path.
- Board users may need to notice the new density controls if they want
to override high-volume defaults.
- Collapsed cold lanes remain valid drop targets, so status moves can
happen without expanding the destination lane.
- Very large remote columns can still hit the existing 200-item
per-column query cap; this PR improves rendering, not server-side board
pagination.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex coding agent based on GPT-5, with repository tool use,
shell execution, local test/build execution, and inline implementation
planning. No subagents were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] If this change affects the UI, I have included before/after
screenshots
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-05-15 10:53:09 -05:00
DottaandPaperclip 0096b56a1c [codex] Add LLM Wiki plugin host support (#5597)
## Thinking Path

> - Paperclip orchestrates AI agents for zero-human companies.
> - The plugin system needs host contracts and runtime support before
large plugins can integrate cleanly.
> - The source branch mixed the LLM Wiki package with supporting
host/runtime work, managed plugin skills, root-level storage spaces, and
a bookmarks reference plugin.
> - [PAP-9173](/PAP/issues/PAP-9173) asked for the current branch to be
split by file boundary: plugin package separately from everything else.
> - [PAP-9188](/PAP/issues/PAP-9188) clarified that LLM Wiki may have
plugin-local spaces, but Paperclip core should not reorganize top-level
local storage into spaces.
> - Follow-up review clarified that the bookmarks example should not
ship in this PR either.
> - This pull request contains the
non-`packages/plugins/plugin-llm-wiki/` host/runtime work, keeps runtime
state under the selected Paperclip instance root, and no longer includes
the bookmarks example.

## What Changed

- Added/updated plugin host contracts, SDK types, worker RPC plumbing,
managed plugin skill support, and related server tests.
- Removed the bookmarks example plugin package and its
bundled-example/workspace references.
- Removed the root-level local spaces CLI/migration surface and restored
instance-root runtime defaults for config, db, logs, storage, secrets,
workspaces, projects, and adapter homes.
- Replaced shared root `space-paths` helpers with `home-paths` helpers
for core runtime storage.
- Tightened stranded recovery unique-conflict detection so concurrent
recovery scans reuse the raced recovery issue when Postgres errors are
wrapped.
- Kept `packages/plugins/plugin-llm-wiki/` out of this PR diff;
plugin-local spaces remain in the stacked plugin-only PR.

## Verification

- `pnpm exec vitest run cli/src/__tests__/data-dir.test.ts
cli/src/__tests__/home-paths.test.ts cli/src/__tests__/onboard.test.ts
packages/shared/src/home-paths.test.ts
packages/db/src/runtime-config.test.ts
server/src/__tests__/agent-instructions-service.test.ts
server/src/__tests__/claude-local-execute.test.ts
server/src/__tests__/codex-local-execute.test.ts`
- `pnpm exec vitest run packages/db/src/runtime-config.test.ts`
- `pnpm exec vitest run
server/src/__tests__/plugin-routes-authz.test.ts`
- `pnpm --filter @paperclipai/server typecheck`
- `pnpm exec vitest run
server/src/__tests__/heartbeat-process-recovery.test.ts -t "reuses the
raced stranded recovery issue"` skipped locally because embedded
Postgres did not initialize on this macOS temp host; the code path was
typechecked and is covered by Linux CI.
- Boundary check: no core references remain for `PAPERCLIP_SPACE_ID`,
`spaces migrate-default`, `@paperclipai/shared/space-paths`,
`registerSpacesCommands`, or the removed bookmarks example.
- Previous PR head `4f23e034` had green GitHub checks: `verify`, all
four serialized server shards, `e2e`, `Canary Dry Run`, `policy`, Snyk,
and `Greptile Review`. Current head `582f466d` is re-running checks
after the bookmarks deletion.

## Risks

- Plugin host changes touch shared runtime paths, so regressions would
most likely appear in adapter startup, plugin loading, or local dev path
defaults.
- Removing the bookmarks example also removes one demonstration of
plugin database namespaces plus local-folder persistence; remaining
plugin examples still cover bundled example discovery and plugin host
flows.
- The plugin package itself is intentionally deferred to the stacked
plugin-only PR, where LLM Wiki plugin-local spaces live.
- Existing installs that tested the transient root-level spaces CLI
should stop using it; this PR intentionally removes that unsupported
migration surface before merge.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI GPT-5 Codex via Codex CLI, tool use and local code execution
enabled; context window not exposed.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have run tests locally and they pass, except where noted above
for host-specific embedded Postgres initialization
- [x] I have added or updated tests where applicable
- [x] If this change affects the UI, I have included before/after
screenshots
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] I will address all Greptile and reviewer comments before
requesting merge

Stacked follow-up: PR #5592 contains only
`packages/plugins/plugin-llm-wiki/` and targets this branch.

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-05-10 07:34:12 -05:00
778e775c35 Add secrets provider vaults and remote import (#5429)
## Thinking Path

> - Paperclip orchestrates AI-agent companies and needs secrets handling
to work across local development, hosted operators, and governed agent
execution.
> - The affected subsystem is the company-scoped secrets control plane:
database schema, server services/routes, CLI workflows, and the Secrets
settings UI.
> - The gap was that secrets were local-only and operators could not
manage provider vaults or import existing remote references without
exposing plaintext.
> - This branch adds provider vault configuration plus an AWS Secrets
Manager remote-import path while preserving company boundaries, binding
context, and audit trails.
> - I kept the PR to a single branch PR, removed unrelated
lockfile/package drift, rebased the full branch onto the current
`public-gh/master`, and addressed fresh Greptile findings.
> - The benefit is a reviewable implementation of provider-backed
secrets with focused tests covering provider selection, import
conflicts, deleted secret reuse, rotation guards, and AWS signing
behavior.

## What Changed

- Added provider vault support for company secrets, including provider
config storage, default vault handling, health checks, binding usage,
access events, and remote import preview/commit.
- Added an AWS Secrets Manager provider using SigV4 request signing,
bounded request timeouts, namespace guardrails, cached runtime
credential resolution, and external-reference linking without plaintext
reads.
- Added Secrets UI surfaces for vault management and remote import, plus
CLI/API documentation for setup and operations.
- Stabilized routine webhook secret binding paths and SSH
environment-driver fixture bindings discovered during verification.
- Addressed Greptile and CI findings: no lockfile/package drift,
monotonic migration metadata, disabled-vault default races, soft-deleted
secret hiding/recreate behavior, remove behavior with disabled vaults,
soft-deleted external-reference re-import, non-active rotation guards,
managed-secret soft deletion through PATCH, and per-call AWS SDK
credential client churn.
- Rebased this branch onto `public-gh/master` at `0e1a5828` and
force-pushed with lease to keep this as the single PR for the branch.

## Verification

- `git fetch public-gh master`
- `git rebase public-gh/master`
- `git diff --name-only public-gh/master...HEAD | grep
'^pnpm-lock\.yaml$' || true` confirmed `pnpm-lock.yaml` is not in the PR
diff.
- Confirmed migration ordering: master ends at `0081_optimal_dormammu`;
this PR adds `0082_dry_vision` and
`0083_company_secret_provider_configs`.
- Inspected migrations for repeat safety: new tables/indexes use `IF NOT
EXISTS`; foreign keys are guarded by `DO $$ ... IF NOT EXISTS`; column
additions use `ADD COLUMN IF NOT EXISTS`.
- `pnpm -r typecheck` passed before the Greptile follow-up commits.
- `pnpm test:run` ran the full stable Vitest path before the Greptile
follow-up commits; it completed with 3 timing-related failures under
parallel load: `codex-local-execute.test.ts`,
`cursor-local-execute.test.ts`, and `environment-service.test.ts`.
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/codex-local-execute.test.ts
src/__tests__/cursor-local-execute.test.ts
src/__tests__/environment-service.test.ts` passed on targeted rerun
(`24/24`).
- `pnpm build` passed before the Greptile follow-up commits. Vite
reported existing chunk-size/dynamic-import warnings.
- After Greptile follow-up commits: `pnpm --filter @paperclipai/server
exec vitest run src/__tests__/secrets-service.test.ts` passed (`26/26`).
- After Greptile follow-up commits: `pnpm --filter @paperclipai/server
exec vitest run src/__tests__/aws-secrets-manager-provider.test.ts
src/__tests__/secrets-service.test.ts` passed (`39/39`).
- After Greptile follow-up commits: `pnpm --filter @paperclipai/server
typecheck` passed.
- Captured Storybook screenshots from `ui/storybook-static` for visual
review.
- Latest PR checks on `5ca3a5cf`: `policy`, serialized server suites
1/4-4/4, `Canary Dry Run`, `e2e`, `security/snyk`, and `Greptile Review`
pass; aggregate `verify` is still registering the completed child
checks.
- Greptile review loop continued through the latest requested pass; all
Greptile review threads are resolved and the latest `Greptile Review`
check on `5ca3a5cf` passed with 0 comments added.

## Screenshots

Before: the provider-vault and remote-import surfaces did not exist on
`master`; these are after-state screenshots from the Storybook fixtures.

![Secrets
inventory](https://raw.githubusercontent.com/paperclipai/paperclip/PAP-2339-secrets-make-a-plan/doc/pr/5429/secrets-inventory.png)

![Secret binding
picker](https://raw.githubusercontent.com/paperclipai/paperclip/PAP-2339-secrets-make-a-plan/doc/pr/5429/secret-binding-picker.png)

![Environment editor with
secrets](https://raw.githubusercontent.com/paperclipai/paperclip/PAP-2339-secrets-make-a-plan/doc/pr/5429/env-editor-with-secrets.png)

## Risks

- Migration risk: this adds new secret provider tables and extends
existing secret rows. The migrations were checked for monotonic ordering
and idempotent guards, but reviewers should still inspect upgrade
behavior carefully.
- Provider risk: AWS support uses direct SigV4 requests. Automated tests
cover signing, request timeouts, vault-config selection, namespace
guardrails, pending-version archival, sanitized provider errors, and
service-level cleanup paths. A real-vault AWS smoke test remains
deployment validation for an operator with AWS credentials rather than
an unverified merge blocker in this local branch.
- UI risk: the Secrets page and import dialog are large new surfaces;
screenshots are included above for reviewer inspection.
- Verification risk: the full local stable test command hit
parallel-load timing failures, although the exact failed files passed
when rerun directly.
- Operational risk: remote import intentionally avoids plaintext reads;
operators must understand that imported external references resolve at
runtime and may fail if AWS permissions change.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex, GPT-5 coding agent with local shell/tool use in the
Paperclip worktree. Exact context-window size was not exposed by the
runtime.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] If this change affects the UI, I have included before/after
screenshots
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-09 18:22:17 -05:00
DottaandPaperclip 7f893ac4ec [codex] Harden execution reliability and heartbeat tooling (#3679)
## Thinking Path

> - Paperclip orchestrates AI agents for zero-human companies
> - Reliable execution depends on heartbeat routing, issue lifecycle
semantics, telemetry, and a fast enough local verification loop to keep
regressions visible
> - The remaining commits on this branch were mostly server/runtime
correctness fixes plus test and documentation follow-ups in that area
> - Those changes are logically separate from the UI-focused
issue-detail and workspace/navigation branches even when they touch
overlapping issue APIs
> - This pull request groups the execution reliability, heartbeat,
telemetry, and tooling changes into one standalone branch
> - The benefit is a focused review of the control-plane correctness
work, including the follow-up fix that restored the implicit
comment-reopen helpers after branch splitting

## What Changed

- Hardened issue/heartbeat execution behavior, including self-review
stage skipping, deferred mention wakes during active execution, stranded
execution recovery, active-run scoping, assignee resolution, and
blocked-to-todo wake resumption
- Reduced noisy polling/logging overhead by trimming issue run payloads,
compacting persisted run logs, silencing high-volume request logs, and
capping heartbeat-run queries in dashboard/inbox surfaces
- Expanded telemetry and status semantics with adapter/model fields on
task completion plus clearer status guidance in docs/onboarding material
- Updated test infrastructure and verification defaults with faster
route-test module isolation, cheaper default `pnpm test`, e2e isolation
from local state, and repo verification follow-ups
- Included docs/release housekeeping from the branch and added a small
follow-up commit restoring the implicit comment-reopen helpers that were
dropped during branch reconstruction

## Verification

- `pnpm vitest run
server/src/__tests__/issue-comment-reopen-routes.test.ts
server/src/__tests__/issue-telemetry-routes.test.ts`
- `pnpm vitest run server/src/__tests__/http-log-policy.test.ts
server/src/__tests__/heartbeat-run-log.test.ts
server/src/__tests__/health.test.ts`
- `server/src/__tests__/activity-service.test.ts`,
`server/src/__tests__/heartbeat-comment-wake-batching.test.ts`, and
`server/src/__tests__/heartbeat-process-recovery.test.ts` were attempted
on this host but the embedded Postgres harness reported
init-script/data-dir problems and skipped or failed to start, so they
are noted as environment-limited

## Risks

- Medium: this branch changes core issue/heartbeat routing and
reopen/wakeup behavior, so regressions would affect agent execution flow
rather than isolated UI polish
- Because it also updates verification infrastructure, reviewers should
pay attention to whether the new tests are asserting the right failure
modes and not just reshaping harness behavior

## Model Used

- OpenAI Codex coding agent (GPT-5-class runtime in Codex CLI; exact
deployed model ID is not exposed in this environment), reasoning
enabled, tool use and local code execution enabled

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] If this change affects the UI, I have included before/after
screenshots
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-04-14 13:34:52 -05:00
Dotta 8bdf4081ee chore: improve worktree tooling and security docs 2026-04-10 22:26:30 -05:00
dotta b1e9215375 docs: add browser process cleanup plan 2026-04-09 06:14:12 -05:00
dotta 5758aba91e docs: add agent-os follow-up plan 2026-04-09 06:14:12 -05:00
dotta 482dac7097 docs: add agent-os technical report 2026-04-09 06:14:12 -05:00
dotta 0937f07c79 Remove standalone issue recovery plan doc 2026-04-09 06:14:12 -05:00