mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 03:08:10 +02:00
5ee9e751fbb38787c055ee793c3412ecfc89da2e
18
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6681c71b40 |
fix(ci): stabilize chat startup and close the initial live-update gap (#13895)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Browser tests check chat state across navigation, reload, and agent runs. > - Runtime tests check that a service is ready before Paperclip publishes its address. > - CI for #13891 failed in these paths, then passed attempt 3 with the same code. > - The browser traces stopped during development asset startup. A separate sidebar assertion used an unstable focus path through the rich editor. > - This PR gives browser tests fresh built assets and a direct keyboard path to the star button. It adds service-worker reload coverage under CPU throttling. > - A later browser failure exposed a real reload race: comments can change between the first query and the first live subscription. The UI now refreshes active queries when that subscription opens. > - Runtime fixtures now have separate registry state and better failure evidence. The readiness deadline remains unchanged. > - A later CI run exposed a wall-clock backoff assertion and three authorization cases sharing one test lifecycle. The PR anchors the assertion to transport time and separates the cases. ## Linked Issues or Issue Description Refs #13891. Related runtime ownership and cleanup work: #11791, #11389, #11278. Evidence: [original CI run, attempt 3](https://github.com/paperclipai/paperclip/actions/runs/35910335089/attempts/3). Attempt 3 passed both affected shards without code changes. This PR is separate from the wake-payload change, which has since merged. | Failure | Diagnosis and classification | | --- | --- | | Sidebar blank page and retry text missing after reload in CI | Both traces show a blank document before React startup. Vite connects, but the failing page makes no application API requests. The service worker forwards the unbundled development module graph. The retry response is already stored and its process-adapter run succeeded before reload. This places the failure in browser bootstrap, not wake payload or reply persistence. The exact reason the development module graph stopped is not established by the retained trace. The harness now serves built assets, and the new test covers startup and controlled reload at 4x CPU throttling. | | Star opacity remains zero on macOS | Reproduced on unchanged master `f55759942b`. The old test clicked the rich editor, focused the star, then used Tab and Shift+Tab. An instrumented baseline run captured the sequence: Tab moved from the star to the next sidebar link, then the editor bundle called `focus()` on its contenteditable before Shift+Tab. That key reached the composer, where it is a work-mode shortcut. This confirms a pending editor selection update stole focus; it was not the browser skipping sidebar buttons. The assertion did not prove the star retained focus. The test now moves the pointer away, focuses the preceding sidebar link, presses Tab once, and asserts both actual focus and opacity. This is test synchronization and keyboard traversal, not a demonstrated CSS defect. | | Runtime readiness exceeds 10 seconds | The old error only says `fetch failed`. There is no child startup output in that failure, so it cannot distinguish slow process startup from a refused or stalled probe. It passed unchanged locally and took 3.416 seconds in attempt 3. Resource contention is plausible but unproved. This PR does not claim a proven historical runtime root cause: it isolates fixture registry/log files, checks ports before spawn and after stop, checks live backends at publication, and preserves transport errors, probe count, elapsed time, and fixture startup timestamps for the next occurrence. | | Later CI: recovery reply disappears after reload | Product synchronization bug, distinct from the blank-page bootstrap failure. Reproduced locally with a trace: the comments request started at `21:56:07.443`, the server saved the reply at `.520`, and the first live subscription started at `.541`. The reload fetched a successful run and its complete log, but missed the comment event. The provider refreshed after reconnects only. It now refreshes active queries on the first connection too. A deterministic regression test fails before the fix. No browser assertion was changed. | | Later CI: Slack backoff and authorization tests | The 30-second backoff assertion required more than 25 seconds to remain when it read the saved action. CI spent 10.311 seconds in the test, exceeding that five-second allowance. A local six-second read delay reproduces the failure; the new transport-anchored lower and upper bounds pass the same fault injection. The neighbouring authorization test ran three independent fixtures in one test and timed out at 15 seconds. Each case now has its own fixture cleanup and the normal per-test deadline, so earlier cases do not remain active during later global worker sweeps. No specific production slow call was established. | | Self-hosted runner loses communication or shuts down | The original lost-communication failure has no assertion. On final-head [attempt 1](https://github.com/paperclipai/paperclip/actions/runs/35925901613/attempts/1), Build, Runner Vitest 1/2, and chat 2/3 ran on three separate fleet instances. All received a runner shutdown signal at `22:04:55 UTC`, within 35 milliseconds, then cancellation. Server shard 6/12 received the same shutdown signal one minute later. All four were Spot `m7i-flex.xlarge` instances in `us-east-1a`. No test assertion or build error preceded those stops. This is infrastructure interruption; the reason the fleet stopped the runners is not available in job logs. | ## What Changed - Build the browser fixture UI into the static server's preferred directory, `server/ui-dist`, and disable Vite middleware. Ignore these generated assets. Start the source CLI directly from the repository root, as required by the CLI invocation safety contract. - Keep all existing chat assertions. Add first takeover and three service-worker-controlled reloads under CPU throttling. Assert that the page uses built module assets and that opening it creates no chat task. - Use forward keyboard traversal from the Zeta link to its star. Check focus before checking the reveal style. - Give each runtime exposure test a temporary Paperclip home and restore environment state after process cleanup. - Check that reserved ports are free before spawn, serve the fixture response before exposure, and become free after stop. - Include the nested fetch error, probe count, and elapsed time in readiness failures. Add a unit test for this diagnostic contract. - Refresh active queries when the first live connection opens, covering events missed during initial page loading. Keep reconnect toast suppression unchanged. - Measure Slack retry timing from the transport attempt and recovery completion. This checks the full provider-requested delay without spending a small wall-clock allowance on unrelated processing. - Run each Slack authorization-revocation scenario as a separate test, with cleanup between cases. All assertions remain. - Document the browser fixture's build and serving mode. No configured assertion deadline, readiness deadline, retry count, or skip was added. ## Verification - Final-head [Linux CI, attempt 2](https://github.com/paperclipai/paperclip/actions/runs/35925901613/attempts/2): **green**. All 52 check runs passed; two conditional checks were skipped. The legacy Snyk status also passed. No pending or failed checks remain. - `pnpm -r typecheck` passed. - `pnpm build` passed again after the final UI fix. - Focused runtime suites: 38 passed, 3 existing platform skips. - CLI invocation safety suite: 39 passed. - Slack timing negative control: inserting a six-second delay before reading the saved action fails the old assertion. All four revised authorization/backoff cases pass with that same delay. The diagnostic delay is not committed. - Live update suites: 92 passed, including the new regression. UI typecheck and token gates passed. - Recovery browser negative control: the unchanged tests reproduced the missing reply (1 failed, 9 passed). After the first-connection fix, all six recovery paths passed twice (12 passed). Browser assertions and deadlines are unchanged. - Server typecheck passed after the Slack test adjustment. - `GITHUB_WORKFLOW=PR pnpm test:run:general -- --group general-chat --shard-index 0 --shard-count 3`: 342 passed. The other 682 tests belong to the remaining shards; collection verified exact coverage. - Final direct-source CLI launch: all five chat session tests passed. - `PAPERCLIP_E2E_PORT=32993 PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/agent-chat-sessions.spec.ts --repeat-each=3`: 15 passed, including nine controlled reloads at 4x CPU throttling. - `GITHUB_WORKFLOW=PR pnpm test:run:general -- --group general-server-without-chat --shard-index 10 --shard-count 12`: 55 files passed, 867 tests passed, 21 existing skips. An earlier run could not initialize PostgreSQL because this Mac exhausted its System V shared-memory slots. After reclaiming the orphaned segment from this task's stopped browser server, the full shard passed. - On final commit `1f1fafc08d`, [browser 4/8](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400837947) passed all 16 tests; [browser 8/8](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400838155) passed all 23 tests; [server 11/12](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400838328) passed 887 tests with one existing skip. The readiness lifecycle case took 1.842 seconds. Chat 1/3 also passed. All four jobs interrupted by runner shutdowns passed unchanged on their single rerun. - Greptile reviewed `1f1fafc08d`: **5/5**, with no unresolved comments. - Full local `pnpm test:run` was attempted and stopped after confirming failures outside this patch: a sibling `skills` directory shadows bundled Slack/AgentMail skills; macOS rejects rename of read-only skill-cache directories (`EACCES`, also reproduced in an isolated filesystem probe); and the host exhausts PostgreSQL System V shared-memory slots. Focused reruns confirmed these limits. The complete Linux CI run is the repository-wide verification; the full local run is not green. - Baseline evidence: the original sidebar test failed on unchanged master; a development-mode run at 4x CPU throttling passed six selected cases, so CPU pressure alone did not reproduce the CI bootstrap stall. ## Risks - Default browser tests now exercise the shipped static UI. They no longer implicitly cover Vite middleware or HMR; use the development server for those checks. - The new browser startup test uses Chromium CDP, matching the only configured browser project. - Opening a live subscription now causes one active-query refresh to close the initial event gap. This adds startup API reads but no recurring poll. - Runtime behavior and deadlines are unchanged except for error details. The historical readiness stall remains unconfirmed; a green rerun alone cannot establish its cause. - Local verification runs on macOS. The runtime lifecycle checks also passed on Linux CI. ## Model Used - OpenAI GPT-6 through Codex. The session identifies the model family as GPT-6; an exact served model ID and context window size are not exposed. Used reasoning, repository inspection, shell tools, code editing, and test execution. No subagents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #123` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — targeted suites passed; full local host limits are listed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ab15aff390 |
feat: add experimental persistent agent chat (#13284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Conversations must use the same tasks, controls, and execution history. > - Users need an ongoing chat with an agent without managing task properties. > - Agents should clarify and plan work, then hand execution to assigned project tasks. > - This pull request combines the reviewed Agent Chat stack for one squash merge. > - The benefit is persistent conversation with normal task governance and shared UI. ## Linked Issues or Issue Description **Subsystem affected** Task lifecycle, agent runtime tools, shared task UI, and browser/paid runner tests. **Problem or motivation** Users need one persistent conversation with each agent. A separate chat store or renderer would duplicate task behavior and bypass existing controls. **Proposed solution** Use a task-backed chat per company, user, and agent. Reuse the task composer and transcript. Clarify and plan in chat, then create assigned project tasks with the relevant plan. Keep Agent Chat behind its own disabled-by-default experimental setting. **Roadmap alignment** This implements the task-backed direction in [CEO Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat). Related proposals: #2504 and #9693. Related request: #7981. The maintainer requested one squash merge of the complete stack. Consolidates the reviewed runtime [#13281](https://github.com/paperclipai/paperclip/pull/13281), backend [#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI [#13283](https://github.com/paperclipai/paperclip/pull/13283) layers with this PR's E2E coverage. All four layers passed CI and received Greptile 5/5 before consolidation. This PR targets master and includes the complete feature. ## What Changed - Add personal canonical chat tasks with ordinary company visibility, immutable identity, idempotent first sends, and an idle waiting state. - Process `/new` in queue order. Preserve history, release a chat pause, and fence old provider context and delayed writes. - Keep chat lifecycle rules across recovery, finalization, assignment, task lists, and rollups. - Support research and plan revision in chat. Hand plans to ordinary assigned project tasks before execution starts. Reject new chat subtasks. - Add repository-aware project creation and discovery tools, including multiple repository IDs and GitHub URLs, authorization, idempotency, and durable project-created cards. - Reuse task UI components for chat, with starred/recent agent navigation and a separate `enableAgentChat` experimental flag. - Add deterministic browser tests and 24 paid chat cells across four Codex/Claude profiles, with validated reports and screenshots. - Integrate current master recovery, controller lease, queued-message, and task UI changes. Gate chat interruption and deferred promotion on ownership/feature policy. Guarantee lease renewal and active controls are stopped even if teardown fails. - Preserve master's migration 0273 and generate chat migration 0274 with idempotent replay for development databases. ## Verification - Prior exact heads of all four PRs passed Linux CI, including build, typecheck, general/serialized tests, and browser E2E. Each had Greptile 5/5 and no unresolved findings. - Integrated local verification passed: full repository typecheck and production build, Storybook build, token gates, 340 focused UI tests, all 20 deterministic chat browser tests, two migration replay tests, 88 focused chat/queue/native/controller tests, and provider/session regressions including real lease expiry. These include the three lifecycle regressions for the final admission/teardown fixes; server typecheck also passes. Current head `1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no unresolved findings and passing security scans. All final-head CI gates passed: build, full Runner verification, typecheck/release registry, canary, all general/serialized test shards, and all browser E2E shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)). Local PostgreSQL startup contention required serialized retries; skipped fixtures do not count as passing coverage. - The earlier paid campaign passed all 24 chat cells and retained 32 screenshots: [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat). It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior evidence, not a paid run of this integrated head. - Manual check: enable Agent Chat in Experimental settings, open an agent, clarify and revise a plan, then hand off to an assigned project task. Stop a reply, send `/new`, and verify fresh context with retained history. Disable the setting and verify agent shortcuts/new chat turns are blocked. ## Risks - Queue/session integration can affect retries and delayed writes. Tests cover ownership, cancellation, reset boundaries, idle recovery, and ordinary task behavior. - Migration 0274 adds conversation fields and constraints. Replay is idempotent and preserves existing development chat history. - This combines the previously reviewed stack at the maintainer's request. Agent Chat remains off by default and is separate from Conference Room. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code execution, browser tools, and parallel review. The exact context-window size is not exposed in this session. Codex and Claude also ran as test subjects in the linked paid campaign. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9031516a7e |
fix: recover legacy Daytona startup failures from task and inbox (#13272)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Legacy conversation adapters can run in Daytona sandboxes. > - A server restart during provisioning can occur before the invocation event exists. > - Recovery then lacks the old adapter identity and leaves a hold that ordinary user retries cannot clear. > - A remote launch can also fail when its host relay looks for Node in the sandbox PATH. > - This pull request records the adapter at claim time and restores explicit user continuation after verified cleanup. > - Users can recover from the task or inbox while the failed run and uncertain action history remain intact. ## Linked Issues or Issue Description Refs #13237, #13239, #13254. Those changes cover recorded conversation runs, native user continuation, and explicit remote Stop. This change covers legacy failure before `adapter.invoke` and exact task/inbox Retry. Refs #9771 for overlapping generated-command quoting. This change also supplies the absolute host Node executable. Refs #13163 and #13264 for the separate native restart and retained-workspace work. **What happened?** A legacy Daytona run interrupted during provisioning became `process_lost` without an invocation event. Recovery preserved an execution hold, and Retry or a new task reply could not resume it. Cleanup could also run before the Daytona plugin was ready. On a macOS host, a subsequent ACP relay launch failed with `env: node: No such file or directory` because the remote launch environment did not contain the host Node path. **Expected behavior** An interrupted conversation can continue after its previous execution stops. Explicit Retry and new user replies should start a fresh turn with the task history. Cleanup failures must remain visible and recoverable. The host relay must use the host Node executable. **Steps to reproduce** 1. Use a legacy Claude adapter with a Daytona environment. 2. Interrupt the server after it acquires the sandbox lease and before it records `adapter.invoke`. 3. Restart and inspect the task hold. 4. Retry from the task or inbox, or send a new task reply. 5. Confirm the old sandbox has stopped and one new response arrives. **Paperclip version or commit** Reproduced from master at `3bafac12f796fbea02e609e1074a9639f872e9c4`. The branch is rebased on `51b0e01ea`, including #13261 and #13270. **Deployment mode** Built from source on macOS with a real Daytona sandbox and the legacy Claude ACP adapter. ## What Changed - Count new browser specs with the scheduler's median duration in the shard-balance check. This fixes a false policy failure after new specs arrive from both branches. The balance threshold is unchanged. - Persist server-owned adapter identity in the queued-to-running claim before provisioning starts. - Wait for provider plugin startup before restart cleanup. Keep failed cleanup leases as active ownership blockers. - Admit exact board retries and new user comments after verified termination. Retain the old run, task history, approvals, and unknown action outcomes. - Adopt repeated Retry requests. Permit one scoped cleanup attempt per explicit user Retry after the automatic limit, with an activity record. A later user Retry can recover after a transient provider failure; automatic attempts remain capped. - Resume replies deferred during cleanup, including historical legacy startup failures. - Launch the host ACP relay through the absolute host Node executable. - Add a task-level Retry button and return actionable blockers when retry admission is refused. - Add database regressions and three browser recovery journeys. Exclude installed third-party dependency skills from the shipped-skill audit. ## Verification - Current head: `d23c84181`, rebased on `51b0e01ea`. Conflict resolution retains the saved-message recovery, local stop receipts, and wait reasons from #13270 alongside exact legacy Retry support. - Real Daytona: interrupted the server after lease acquisition and before adapter invocation. Restart cleanup confirmed provider termination. Task Retry cleared a seeded historical hold and a real Claude agent returned `Recovery verified.` in the task. Removed the disposable sandbox and environment after testing. - All three browser recovery journeys passed again after the final rebase. Task Retry, Inbox Retry, and a new reply each produced one fresh successor, completed the task, preserved the failed run, and retained the answer after reload. - All 29 e2e/server shard-partition tests passed. The balance check now uses the scheduler's median fallback for unmeasured specs, with the same balance threshold. - Server typecheck passed after rebuilding the generated runner dependencies. The combined recovery/route run passed 136 of 137 tests. Its remaining route test timed out during the first cold module import at its explicit 10-second limit; an isolated rerun reproduced that timeout and passed the other 51 route cases. The complete CI suite passed on this head. The same route file passed all 52 cases in CI, including the first cold import in 7.5 seconds. - Before the final rebase, recursive typecheck, full build, UI token gates, 132 targeted server tests, and the complete [CI workflow](https://github.com/paperclipai/paperclip/actions/runs/34650004085) passed. The subsequent CI failure was the shard-balance accounting mismatch fixed here. - Greptile reviewed `d23c84181` at 5/5 with no outstanding actionable findings. The complete [current CI workflow](https://github.com/paperclipai/paperclip/actions/runs/34653327949) passed on attempt 2. All test, typecheck, build, and canary jobs passed on the first attempt. Docker setup timed out fetching BuildKit from Docker Hub; retrying that job and its dependent aggregate succeeded. ## Risks - Recovery admission changes executable authority. Company, task, agent, user, approvals, process ownership, and provider termination checks remain required. - Explicit continuation starts a fresh conversation with history. It does not certify unknown external action outcomes or rerun non-conversation adapters automatically. - Changing task status alone does not clear an execution hold. The task now offers an explicit Retry action. - Historical adapter claims and invocation events take precedence over current agent settings. Known process or webhook runs retain their hold. Pre-upgrade rows with no adapter evidence may receive only a new explicit user turn after termination proof; they do not become eligible for automatic replay. - No schema migration or sandbox-image change is required. This branch has not been deployed to production. ## Model Used OpenAI GPT-6 through Codex, with repository inspection, code execution, browser automation, and test execution. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e095b84dab |
feat(connections): connect services from native task feeds (#13058)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents use connections to reach external services. > - A fresh native task can have no service tools installed. > - The agent needs a way to discover services and ask the responsible person for access. > - This pull request brings the existing connection-intent flow into native task execution. > - The person can connect from the task, and the agent can continue with updated tools. ## Linked Issues or Issue Description **Subsystem affected** Native runner tool authority, connection intents, task interactions, and shared connection setup. **Problem or motivation** A task that needs an unconnected service cannot finish its work. Leaving the task to configure access also loses context. A resolved request must survive a restart and resume the correct agent once. **Proposed solution** Expose connection discovery and access requests as server-owned native tools. Render a durable task card and use the shared setup dialog. Persist outcome delivery and start a fresh provider session after access is ready. **Alternatives considered** Sending the person to the Connections page adds navigation and does not solve continuation. Polling for authorization consumes runs and can create duplicate requests. **Roadmap alignment** This extends the existing connection-intent runtime and setup experience. It reuses the shared access model and the native runner. Related: #12345, #12347. The service-slug fix in #12906 is related but separate. Companion evaluation PR: https://github.com/paperclipai/paperclip-evals/pull/21. ## What Changed - Expose `connections_search` and `connection_request` with server-bound company, task, agent, and responsible user. Preserve the legacy entry points. - Discover catalog services and authorized custom connections. Check installation, identity, health, and executable permissions before reporting ready. - Keep pending cards through ordinary messages. Reuse requests and retire stale ownership. Put Connect at the right of Not now. - Reuse the shared setup flow in a task dialog. Keep access additive and default to the requesting agent. Recover from cancelled or blocked OAuth windows with a new-tab fallback. - Persist outcome delivery with an idempotent wake key. Resume in a fresh session and recheck ownership before dispatch. - Add native browser fixtures, offline Storybook states, server contracts, and evaluation fixtures. Update guidance and documentation. ## Verification - `pnpm build`: passed after replaying the change on current master. - `pnpm -r typecheck`: passed. - `pnpm check:token-gates`: passed. - `pnpm --filter @paperclipai/ui build-storybook`: passed. - New continuation-policy regression cases: 16 passed. - Docker-backed PostgreSQL regressions passed for requester-only OAuth access, assignment-only expiry, terminal expiry, and credential-free setup metadata. - Shared setup and task-card UI tests: 121 passed, including configured MCP reconnect URL recovery and preserving user edits across refetch. - Storybook browser checks: all 119 passed on the latest reconnect fix. - `pnpm test:run`: 4,734 tests passed in the first server group, but embedded PostgreSQL startup failures and resulting cleanup errors prevented a complete local pass. All Linux CI lanes passed on the latest reviewed commit. One external-object route test returned an unexplained 500 on the first run; it passed twice locally and the failed shard passed on retry without code changes. - Earlier feature-checkout evidence: three deterministic native browser journeys passed, including restart delivery and an actual fixture tool result. Legacy scripted coverage also passed. All 59 added stories were inspected in light and dark themes. - Live Notion testing recorded successful provider reads. The manual test used a local-trusted instance. It does not prove authenticated/cloud deployment or every provider journey. - Native browser rerun reached the embedded PostgreSQL startup limit before bootstrap, so the latest checkout’s full native browser journey remains unverified. Both OAuth page/task regression cases passed against isolated Docker-backed PostgreSQL 17. They verify no premature task access, requester-only completion, additive retries, and reconnect preservation. - Applied both new migrations twice to isolated PostgreSQL 17. Foreign keys remained intact, duplicate active delivery keys were rejected, and failed delivery records did not block retries. Reviewer path: start a fresh test drive, enable the native runner, use an agent that can perform work directly, and ask it to summarize a Notion page. Connect from the card, then verify the resumed provider call and source-linked answer. The default test-drive CEO is instructed to delegate, so it can introduce an unrelated hiring step. ## Risks - Two additive migrations create durable deliveries and a partial unique wake index. They are idempotent. The wake index can require a maintenance window on large tables because migrations run in a transaction. - OAuth and continuation cross asynchronous boundaries. Tests cover ownership changes, retries, additive access, and restart delivery; live provider behavior still varies. - The latest requester-scope fix has not yet been exercised through live OAuth. GitHub, API-key, authenticated-user, and all recovery journeys are not claimed as verified. ## Model Used OpenAI GPT-6-based Codex assisted with implementation, tests, and review using tools and code execution. The runtime does not expose the exact model version, context window, or reasoning setting. Live evaluation used `gpt-5.6-luna`; manual native testing used `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used and disclosed unavailable runtime details - [x] I have checked ROADMAP.md and confirmed this extends existing connection work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [ ] I have run all required tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation - [x] I have considered and documented risks - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e3eed3a3ae |
Keep browser startup explicitly opt-in (#12435)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The CLI can save a configuration and start the local server in one command. > - The onboarding path set a browser-open environment variable without an explicit user request. > - Headless test servers use the same onboarding path. > - Each server restart could therefore open a system browser. > - This pull request removes the implicit browser-open request and fixes test servers to disable it explicitly. > - The benefit is predictable foreground and test startup without unsolicited browser windows. ## Linked Issues or Issue Description This is stack 11 of 11. It depends on stack 10. **What happened?** `paperclipai onboard --yes --run` set `PAPERCLIP_OPEN_ON_LISTEN=true`. Headless server users, including browser test runners, opened the system browser on each server restart. **Expected behavior** Server startup must not open a browser unless the caller explicitly sets `PAPERCLIP_OPEN_ON_LISTEN=true`. **Steps to reproduce** 1. Run `paperclipai onboard --yes --run` from a clean source checkout. 2. Wait for the server to listen. 3. Observe that the default system browser opens. **Paperclip version or commit** Reproduced on `dbf052577` plus the dependent stack. **Deployment mode** Local dev from source. ## What Changed - Stop onboarding from setting `PAPERCLIP_OPEN_ON_LISTEN=true` for foreground startup. - Set `PAPERCLIP_OPEN_ON_LISTEN=false` in E2E and issue-detail performance test servers as defense in depth. - Preserve the existing explicit environment opt-in in the server. ## Verification - `pnpm exec vitest run cli/src/__tests__/onboard.test.ts` — 10 tests passed. - `pnpm --filter paperclipai typecheck` — passed. - `pnpm -r typecheck` — passed on the stacked head. - `pnpm build` — passed on the stacked head. - Playwright was not run locally by request. ## Risks - Low risk. The only behavior change removes an unsolicited side effect. - A caller that wants browser startup can still set `PAPERCLIP_OPEN_ON_LISTEN=true` explicitly. - No database or migration change exists in this layer. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5. The exact deployment suffix and context window are not exposed. The model used reasoning, repository tools, code execution, Git, and GitHub API access. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
20ccf3f476 |
feat(apps): add connection grants and delegated identities (#12341)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - External tools need explicit identity and access boundaries. > - Shared connection credentials cannot represent every user-scoped use case. > - Grants must stay company-scoped and support safe delegation. > - This pull request adds connection grants, identity rules, and their database contract. > - The benefit is durable control over which identity an agent may use. ## Linked Issues or Issue Description Refs #11965 This is stack 3 of 11. It depends on stack 2 and replaces another reviewable part of #11965. ## What Changed - Add company and user connection grants. - Add delegated identity and membership rules. - Synchronize database, shared, server, and UI contracts. - Register the grant-member replacement route in the OpenAPI surface in the same layer that mounts it. - Add migration 0231 with replay-safe guards and coverage. ## Verification - `pnpm -r typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/openapi-routes.test.ts` (5 passed) - `pnpm --filter @paperclipai/db check:migrations` - `pnpm build` ## Risks - Incorrect grant selection could expose the wrong credential scope. - The service enforces company and subject boundaries before credential use. - Migration 0231 is generated, ordered after 0230, and safe to replay. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked a public issue or pull request with `Refs #` - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1c366a9059 |
fix(server): reject invalid agent credentials instead of downgrading to the local user actor (#11589)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server authenticates each agent request in `actorMiddleware` before it attributes chat comments > - When an agent bearer token failed verification, the middleware called `next()` with no error and the request continued without an agent actor > - The request then fell back to the local user actor, so the server stored agent replies as user comments > - The task chat UI renders user comments in blue bubbles, so agent messages appeared as blue user bubbles > - This pull request rejects invalid agent credentials with 401 instead of a silent downgrade > - The benefit is that agent messages keep agent attribution, and broken credentials fail loudly with a clear retry message ## Linked Issues or Issue Description **What happened?** A user cancelled an onboarding question card. The agent posted a follow-up reply. The reply appeared in a blue bubble, which the UI reserves for human messages. The agent run held an expired local agent JWT. The auth middleware could not verify the token, called `next()` without an actor, and the request fell back to the local user identity. The server stored the agent comment as a user comment. **Expected behavior** Agent messages always render as agent bubbles. A request with invalid agent credentials must fail with 401 so the adapter can refresh credentials and retry. It must not post content under a human identity. **Steps to reproduce** 1. Start a local Paperclip instance. 2. Give an agent run an expired or malformed agent JWT. 3. Let the agent post an issue comment through the API bridge. 4. Before this change: the comment is stored with the local user identity and renders as a blue bubble. After this change: the request fails with 401 and a message that tells the caller to obtain fresh credentials. ## What Changed - `server/src/middleware/auth.ts`: a bearer token that fails verification now produces a 401 `unauthorized` error instead of a silent fall-through to the anonymous/local-user actor. - The 401 message states the cause: expired token, unverifiable token, empty bearer token, missing agent record, agent record in another company, terminated agent, or agent pending approval. - The API-key path now also rejects an agent record whose company does not match the key. - `packages/adapter-utils/src/execution-target.ts`: the bridge proxy now writes a `comment id: <id>` marker to the run log for each posted issue comment, so misattributed comments can be traced to a run. - `ui/src/components/task-chat/task-chat-adapter.test.ts`: a regression test asserts that a recovered `local-board` comment with a derived agent author renders as an agent bubble, not a user bubble. - `server/src/__tests__/agent-auth-middleware.test.ts` and `packages/adapter-utils/src/execution-target-sandbox.test.ts`: new tests cover each rejection path and the log marker. ## Verification - Run `pnpm vitest run src/__tests__/agent-auth-middleware.test.ts` in `server/` — 14 tests pass. - Run `pnpm vitest run execution-target-sandbox` at the repo root — 44 tests pass. - Run `pnpm vitest run src/components/task-chat/task-chat-adapter.test.ts` in `ui/` — 4 tests pass. - Manual check: post an issue comment with an expired agent JWT; the API returns 401 with a retry message and no comment is stored. ## Risks - Behavioral shift: requests that previously continued as anonymous or local-user actors after a failed agent-token verification now receive 401. Any caller that relied on the silent downgrade must refresh its credentials. This is the intended fix, and the adapters already handle 401 with a credential refresh. - No schema or migration changes. Low risk otherwise. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude (Anthropic), model ID `claude-fable-5`, via Claude Code with extended thinking and tool use (agent harness with shell, file, and git tools). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9c1f8e7887 |
feat(decisions): add first-class propose mode (#10010)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents can currently perform many mutations directly, while humans often need a durable review point before cross-issue or destructive actions occur > - Existing approvals and issue-thread interactions do not provide a standalone, reusable object for presenting options, collecting typed inputs, detecting stale targets, and auditing effect execution > - The control plane therefore needs a first-class propose mode that separates an agent's recommendation from the governed mutation it may cause > - This pull request adds Decisions v1 across the database, shared contracts, server execution and telemetry, agent skill guidance, and operator UI > - The benefit is that agents can propose multi-option actions safely while operators get explicit provenance, fail-closed execution, per-effect results, and a focused attention workflow ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting: `packages/db`, `packages/shared`, `server`, and `ui`. ### Problem or motivation Agents need a governed way to propose consequential work without immediately mutating issues, especially when one choice can affect several issue trees. Existing approvals and issue-thread interactions do not provide a standalone object with typed options, target snapshots, effect-level authorization, expiration, execution outcomes, and reusable attention-feed presentation. ### Proposed solution Add first-class Decisions that store options and typed inputs, surface open proposals in the operator attention feed, validate target freshness and the origin-agent/operator authorization intersection at decision time, execute a bounded set of auditable effects, and retain terminal outcomes. Decisions v1 supports comments, status and assignee changes, follow-up issue creation, blocker resolution, and issue-tree cancellation, plus bundle grouping, expiration/dismissal, rule-key telemetry, and agent-facing API guidance. ### Alternatives considered - Extend approvals with arbitrary effects: rejected because approvals represent governed yes/no actions and would become an unsafe generic mutation envelope. - Model every proposal as an issue-thread interaction: rejected because decisions can span several targets and need independent lifecycle, telemetry, idempotency, and effect results. - Let agents perform the mutation and ask for retrospective review: rejected because it removes the pre-execution governance boundary this feature is meant to provide. ### Roadmap alignment Aligns with `ROADMAP.md` sections **Agent Reviews and Approvals**, **Enforced Outcomes**, **MCP Tool Gateway & Apps (governed tool access)**, and **Activity History** by making explicit decisions, authorization gates, auditable execution, and terminal outcomes first-class control-plane objects. ### Additional context This does not replace existing approvals or issue-thread interactions, and it does not add an unrestricted generic mutation effect. ## What Changed - Added company-scoped decision, option, target, and effect-execution schema plus migration and shared TypeScript/Zod contracts. - Added decision routes and services for propose, list/get, decide, dismiss, cancel, target freshness checks, authorization intersection, idempotency, activity logging, and execution auditing. - Added rule-key decision telemetry and attention-feed metadata so open decisions are visible and measurable. - Added agent skill documentation for proposing and resolving decisions through the Paperclip API. - Added the Decisions UI: API client, query keys, inline attention resolver, bundle grouping, target-issue strip, terminal history, destructive confirmation, and per-effect result rendering. - Added server service coverage, DecisionCard state tests, and Storybook stories for the supported visual states. ## Verification - `pnpm -r typecheck` — passed. - `pnpm test:run` — 2,876 passed, 1 skipped, with one unrelated cross-suite cleanup-order failure in `heartbeat-responsible-user-invariant.test.ts`; the failing file passes in isolation (`6/6`). - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-responsible-user-invariant.test.ts` — passed. - `pnpm --filter @paperclipai/ui exec vitest run src/components/DecisionCard.test.tsx` — passed (`9/9`). - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/authz-existence-oracle-guard.test.ts src/__tests__/openapi-routes.test.ts` — passed (`5/5`). - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/decisions-service.test.ts` — passed (`16/16`). - `pnpm --filter paperclipai exec vitest run src/__tests__/company-import-export-e2e.test.ts` — passed (`1/1`). - `pnpm --filter @paperclipai/server typecheck` and `pnpm --filter paperclipai typecheck` — passed. - `pnpm build` — passed. - Rebased-head focused suite — passed (`6` files, `88` tests): shared decision contracts, Decisions service, OpenAPI routes, startup feedback export, DecisionCard states, and attention helpers. The follow-up stale-secondary-target regression passes in the DecisionCard suite (`10/10`). - Rebased-head scoped typechecks — passed for `@paperclipai/shared`, `@paperclipai/db`, `@paperclipai/server`, and `@paperclipai/ui`. - Rebased-head migration numbering and safety checks — passed after renumbering the additive migration to `0193` and making it replay-safe for environments that applied the earlier feature-branch number. - `pnpm check:token-gates` — passed with all gates clean. - GitHub PR workflow and Greptile review for `1f9f7645882d05dfdd9c99377c03a1f53f20e8be` — running after the stale-secondary-target fix and PR metadata refresh on July 27, 2026. - `pnpm --filter @paperclipai/ui build-storybook` exposes an existing Storybook version mismatch (`storybook` 10.4.6 vs `@storybook/addon-docs` 10.5.0); Decisions stories were validated with the docs addon temporarily disabled and the tracked config remains unchanged. ## Risks - **Migration:** Adds replay-safe migration `0193`; migration numbering and safety checks pass. The new tables and indexes are additive. - **Authorization:** Effect execution intersects the proposing agent's permissions with the responsible user context and fails closed; mistakes could reject a valid proposal rather than silently over-authorize it. - **Concurrency:** Target snapshots and idempotency keys protect against stale or duplicate execution, but reviewers should focus on mixed-effect partial outcomes and retry behavior. - **UI:** Decisions are integrated into the existing attention feed rather than a separate navigation surface, reducing routing risk but increasing the importance of attention-item metadata compatibility. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI using `gpt-5.6-sol` for final PR preparation, review fixes, and verification; repository tools and code execution were enabled, and context-window size is not exposed in this runtime. - Anthropic Claude Opus 4.8 with 1M context assisted with the Decisions UI implementation, as recorded in the relevant commits. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
3db2e6bdd2 |
feat(mcp) [split 8/8]: add e2e coverage and operator docs (#9563)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Governed MCP access spans contracts, runtime enforcement, adapters, UI surfaces, and operator verification > - The parity reference PR #9534 is too large for effective automated or human review > - The feature therefore needs a linear stack whose individual diffs stay below the 100-file review limit > - This pull request is split 8/8 and focuses on end-to-end coverage, operator docs, evals, and release notes > - The benefit is a standalone, testable review boundary while preserving byte-for-byte parity at the top of the stack ## Linked Issues or Issue Description - Related parity reference: #9534 - Problem: The complete stack needs discoverable browser scenarios, operator guidance, threat modeling, eval coverage, and a parity proof before merge. - Proposed solution: Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release notes, the skill update, and the root e2e driver script registration. - Alternatives considered: keeping #9534 as one 403-file review, or rewriting the feature to manufacture seams; both were rejected in favor of path extraction plus compile-driven boundary moves. - Roadmap alignment: this advances the existing governed MCP/tool-access work already represented by #9534; it does not introduce a separate roadmap initiative. - Stack position: base branch is `pap10341-split/07-ui-apps-activation`. - Merge policy: merge bottom-up, in order, only after the complete eight-PR stack has been reviewed and the top-of-stack parity gate remains empty. - Requested review: QA for flag audit and e2e/browser acceptance; Greptile on every PR. ## What Changed - Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release notes, the skill update, and the root e2e driver script registration. - Keeps this PR below 100 changed files and independently typecheckable. - Preserves the final tree from #9534 when combined with the other seven stack levels. ## Verification - `pnpm typecheck` - `node --check scripts/e2e-mcp-user-stories.mjs` - `pnpm exec playwright test --config tests/e2e/playwright.config.ts --list` — 43 tests discovered - `git diff pap10341-split/08-e2e-docs 6b40e3876d9297105d4ec306e47e46d351c86172` — empty (0 bytes) ## Risks - Browser suites depend on runtime services and environment setup; this PR validates discovery locally while QA owns full flag-on/flag-off execution. - Stack risk: merging out of order can expose incomplete layers; mitigate by following the documented bottom-up merge policy. - Parity risk: later edits to an intermediate branch can drift from #9534; mitigate by re-running the empty top-of-stack diff before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, Git, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] Internal references are omitted except the execution-plan link explicitly required for this coordinated split stack - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack Coordination - Internal execution plan: [PAP-13874](/PAP/issues/PAP-13874#document-plan) - Parity reference: #9534 - Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563 - Merge bottom-up only after full-stack review and an empty parity diff at #9563. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f42a66bf04 |
test: port pipelines tutorial e2e (#9149)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Pipelines are a control-plane workflow surface where operators need confidence that setup, intake, board movement, review, and learning flows keep working. > - The tutorial flow is broad enough that regressions can slip through unit tests when UI state, route behavior, and workflow fixtures drift apart. > - End-to-end coverage gives reviewers a realistic smoke path through the pipelines tutorial experience. > - This pull request ports the pipelines tutorial flow into Playwright and updates the variable flow covered by the scenario. > - The benefit is stronger browser-level regression coverage for a high-value onboarding/workflow path. ## Linked Issues or Issue Description No public GitHub issue exists. Inline feature request: **Subsystem affected** Cross-cutting: `tests/e2e`, server startup, and the pipelines UI workflow. **Problem or motivation** The pipelines tutorial flow lacked focused browser-level regression coverage for setup, intake, board movement, detail/review surfaces, and learning flow behavior. This left a broad user-facing workflow dependent on narrower unit coverage. **Proposed solution** Add a Playwright spec that boots a throwaway Paperclip instance and exercises the tutorial flow across setup, intake, board movement, item detail, review queue, learnings, agent fan-out, drift acknowledgement gates, child-terminal gates, and stale approvals. **Alternatives considered** Unit tests alone are faster but do not validate the integrated browser workflow. A release-smoke-only path would be broader than necessary for this tutorial-specific coverage. **Roadmap alignment** This supports the roadmap theme of easier onboarding and stronger first-run confidence without adding a new core feature surface. **Additional context** The local host used for this PR is missing Chromium’s `libatk-1.0.so.0` dependency, so the full browser run needs CI or a machine with Playwright system dependencies installed. ## What Changed - Added a Playwright e2e spec covering the pipelines tutorial flow. - Covered agent fan-out, drift acknowledgement gates, child-terminal gates, stale approvals, setup/intake, board movement, item detail, review queue, and learnings. - Added Playwright config support needed by the new tutorial flow. - Updated the tutorial variable flow assertions in the ported spec. ## Verification - Attempted `/srv/paperclip/home/paperclipai/paperclip/node_modules/.bin/playwright test --config tests/e2e/playwright.config.ts tests/e2e/pipelines-tutorial-flow.spec.ts` - Local result: the throwaway Paperclip server booted and `covers agent fan-out, drift acknowledgement gates, child-terminal gates, and stale approvals` passed. - Local blocker: the second Chromium test failed before executing because this host is missing the Playwright system library `libatk-1.0.so.0`. CI should run this on an image with browser dependencies installed. ## Risks Medium risk for CI duration/flakiness because this adds browser-level coverage over a broad workflow. The test is intentionally scoped to one spec file and uses the dedicated e2e config/server path. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5.5 coding agent with repository tool use and local shell execution. Context window was not surfaced by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6f9801a46b |
feat(ui): NUX rework behind enableConferenceRoomChat experimental flag — capsule onboarding, conference-room chat, unified composer (#8000)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The first-run experience (onboarding wizard) and the chat surfaces
(conference-room/board chat, task threads, composers) are the product's
front door — they decide whether a new operator understands "hire
agents, give them work, review results" in the first five minutes
> - Today those surfaces feel ticket-y and form-like: the wizard is a
static multi-step form that ends in an anticlimactic "Launch" screen,
the task composer and board chat behave differently from each other, and
agent-feed issue quicklooks misbehave (multiple flyouts open at once,
cards jump on hover)
> - We wanted to iterate toward a conversational, team-centric NUX — but
without risking the workflows of everyone already running Paperclip
> - This PR reworks the NUX behind a new default-OFF
`enableConferenceRoomChat` experimental flag: a capsule-motif onboarding
wizard that builds your team as you answer, a conference-room chat
surface, one shared ChatComposer across surfaces, brand-accurate status
chips, and feed-quicklook fixes — with the pre-existing UI
fork-and-frozen as `*Classic` components that flag-OFF users keep
> - The benefit is a complete, testable modern NUX that anyone can opt
into from Settings → Experimental, with zero default behavior change and
a clean path to either graduate or drop the experiment
## Linked Issues or Issue Description
No pre-existing GitHub issue — feature description per
`feature_request.yml`:
- **Problem / motivation:** Paperclip's onboarding wizard and chat
surfaces grew up as separate ticket-centric forms. New users get a
form-filling experience rather than the feeling of standing up a team;
the board chat and task threads use different composers with different
affordances; the agent feed's issue quicklook can stack multiple
popovers and shifts cards on hover.
- **Proposed solution:** A coherent NUX experiment behind one
experimental flag (`enableConferenceRoomChat`, Settings → Experimental,
default OFF): capsule onboarding wizard with an evolving team capsule,
conference-room chat, unified `ChatComposer`, team-centric copy, brand
status chips, quicklook single-flight fix. Flag-OFF users get the exact
pre-experiment UI via frozen `*Classic` forks, verified by an on/off
parity test matrix.
- **Alternatives considered:** (a) incremental unflagged restyling —
rejected: the changes interlock across surfaces and would drip risk into
every release; (b) a separate app shell / route for the new NUX —
rejected: too much divergence, the flag + classic-fork pattern keeps the
diff reviewable and reversible.
- **Roadmap alignment:** `ROADMAP.md` lists **CEO Chat** ("a
lighter-weight way to talk to leadership agents... should still resolve
to real work objects"). This experiment is groundwork in that direction
(conference-room chat resolves to issues/tasks via the same composer
used in task threads) and does not change the core task-and-comments
model.
Related PRs found in the dedup search (same area, none duplicate this
work — they target the classic wizard, which this PR intentionally
leaves intact and mergeable):
- #5385 — Coach-driven onboarding: conversational entry +
agent-companies package import
- #5378 — Onboarding wizard: reusable adapter picker + probe card
- #6636 — ui(onboarding): friendly error surface + retry for the wizard
- #7005 — fix(onboarding): explicitly await first-task wake
- #2616 — fix: restore workspace directory config in onboarding wizard
## What Changed
- **Experimental flag plumbing** — `enableConferenceRoomChat` in shared
types/validators, server instance-settings service + API, Settings →
Experimental card with explicit enable/disable copy
- **Onboarding wizard** — classic wizard forked and frozen
(`OnboardingWizardClassic`); flag-ON variant is a 5-step capsule wizard
with a persistent evolving `AgentCapsule` (gradient/glow motif),
team-centric reframed copy, and a typing-dots intro (hardened with
fake-timer tests)
- **Conference-room chat** — flag-ON board-chat surface with agent
bubble name/icon headers and copy/vote/timestamp action rows
(`AgentBubbleActionRow`)
- **Unified composer** — shared `ChatComposer` adopted across surfaces;
translucent surface + scroll-mask removal; "Agent mode"/"Plan mode"
relabels; no-assignee confirmation `AlertDialog` (new
`ui/alert-dialog.tsx` primitive); `@task` reference picker +
linkification in mentions
- **Agent feed** — single-flight issue-quicklook store (one popover at a
time), flyouts open to the left, removed hover translate-y jitter
- **Status chips** — brand-accurate task status chips behind the flag
(light/dark, 1px borders per paperclip.ing/brand)
- **Tests** — flag on/off parity matrix across IssueDetail,
NewIssueDialog, Sidebar, wizard, gate components; component tests for
all new pieces
- **Merge with `master`** — one conflict in
`ui/src/components/IssueChatThread.tsx`, resolved by keeping master's
new `AssigneeChip`/`HandoffWakeRow`/`RunStatusBadge` components inside
the flag-gated metadata-row chrome (details in commit `21a5642a`);
post-merge fixes: vitest 4 mock typing in `MarkdownEditor.test.tsx`,
flag hook made safe for provider-less mounts (master's new isolated
component tests)
- **Branch hygiene** — internal design wireframes/mockups stripped
before the PR (they live in the Paperclip issue threads)
- No user-facing documentation changes required: the flag is
intentionally experimental and self-described in the Settings card; no
existing docs reference the affected surfaces
## Verification
- `pnpm run typecheck` — green across the workspace (ui, server, shared,
plugins)
- Full UI suite (`vitest run` in `ui/`, clean worktree at this HEAD):
**1593/1595 passing, 223/224 files** — the 2 remaining failures are in
`src/components/artifacts/ArtifactCard.test.tsx` and **fail identically
on pristine `origin/master`** (pre-existing upstream, unrelated to this
branch)
- Full server suite (`vitest run` in `server/`, same clean worktree):
results in PR checks; flag plumbing covered by instance-settings tests
- Targeted post-merge resolution check: `IssueChatThread`,
`IssueChatThreadSystemNotice`, `IssueDetail`, `Sidebar`,
`ConferenceRoomChatGate`, `OnboardingWizardVariant`, `NewIssueDialog`,
`InstanceExperimentalSettings`, `MarkdownEditor` — 172/172 passing
- Manual walkthrough: flag OFF (default) → onboarding wizard, task
thread, board chat, composer all render the classic UI; flag ON via
Settings → Experimental → capsule wizard, conference-room chat, unified
composer, status chips active
- Screenshots: see below
**Flag on/off screenshots** (committed on this branch under
`screenshots/PR-8000-*`):
| Surface | Flag OFF (classic, default) | Flag ON (experimental) |
| --- | --- | --- |
| Settings → Experimental | 
| 
|
| Task thread | 
| 
|
| Home / nav | 
| 
|
| Conference Room (flag-ON only surface) | — | 
|
Capsule onboarding wizard walkthrough screenshots (flag ON) are attached
to the Paperclip design/implementation threads; the wizard requires a
fresh instance so it is captured via the e2e harness
(`tests/e2e/nux-phase4-screenshots.spec.ts`).
## Risks
- **Large surface, but gated:** all new behavior sits behind
`enableConferenceRoomChat`, default OFF; flag-OFF rendering is locked by
frozen `*Classic` forks plus an on/off parity test suite
- **Classic forks are frozen at the fork point (`e3aada1d`):** master
features added to the live thread component after that point (assignee
handoff chips, run status badge, composer mention coach) render in the
flag-ON path; the flag-OFF task thread keeps the fork-point behavior
until the experiment graduates (forks deleted) or is dropped (forks
restored as canonical). Called out for reviewer attention.
- **Merge-conflict resolution in `IssueChatThread.tsx`** (commit
`21a5642a`) deserves reviewer eyes: master's new handoff/run-status
components were kept; the base toast-style no-assignee flow remains
replaced by the AlertDialog flow introduced on this branch
- Schema/server changes are additive (one optional boolean instance
setting); no migrations of existing data
## Model Used
- Claude (Anthropic) via Claude Code running in the Paperclip agent
harness (agent: ClaudeCoder)
- Branch implemented across multiple agent sessions on Claude Opus-class
models with extended thinking + tool use (file edits, shell, Playwright
screenshots); merge/PR session model ID as reported by the harness:
`claude-fable-5` (Claude Code CLI)
- All code was agent-authored and board-reviewed through Paperclip issue
threads (plans, wireframes, confirmations) before merging
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] If this change affects the UI, I have included before/after
screenshots
- [x] I have updated relevant documentation to reflect my changes (none
required — experimental flag, self-documenting Settings card; noted
above)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (run 3 on `8af3041a`: all 16
gates SUCCESS, incl. e2e and all 4 serialized-suite shards)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(re-review verdict: Confidence 5/5, “Safe to merge”; all 4 round-1
findings fixed + confirmed resolved; both summary notes addressed in
`8af3041a`)
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
524e18b060 |
ci: use runner Chrome for headless workflows (#6967)
## Thinking Path > - Paperclip relies on CI browser suites to protect control-plane workflows, so a stalled browser bootstrap is a release blocker even when app code is unchanged. > - The failing signal on [PAPA-457](/PAP/issues/PAPA-457) was specific to the PR e2e lane timing out before tests started, which pointed at environment setup rather than assertions. > - The first shell-only Chromium attempt reduced download size, but the GitHub Actions log showed Playwright still hanging inside its install step after the headless shell download finished. > - That means the real problem is the Playwright browser-install path itself on the hosted Ubuntu runner, not just the size of the downloaded artifact. > - GitHub's Ubuntu runners already ship Google Chrome, and Playwright can target that binary through the `chrome` channel without downloading its own Chromium bundle. > - The safer workflow fix is therefore to remove the Playwright install step from the affected headless jobs and make the Playwright configs optionally use runner Chrome only when CI opts into it. > - This keeps local defaults unchanged, removes the failing browser-download dependency from CI, and preserves headless coverage for PR, standalone e2e, and release-smoke workflows. ## What Changed - Updated `.github/workflows/pr.yml`, `.github/workflows/e2e.yml`, and `.github/workflows/release-smoke.yml` to stop downloading Playwright browsers and instead verify the runner's preinstalled `google-chrome`. - Passed `PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome` into the headless PR, standalone e2e, and release-smoke test steps so those jobs explicitly use runner Chrome. - Updated `tests/e2e/playwright.config.ts` and `tests/release-smoke/playwright.config.ts` to honor `PAPERCLIP_PLAYWRIGHT_CHANNEL` while keeping the default local/browser-bundle behavior unchanged when the env var is absent. ## Verification - Investigated the failed PR run log and confirmed the prior `Install Playwright` step stalled after `chromium-headless-shell` reached 100% download. - `PLAYWRIGHT_BROWSERS_PATH="$(mktemp -d)" PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome PAPERCLIP_E2E_SKIP_LLM=true pnpm run test:e2e` Result: `7 passed (21.1s)` with an empty temporary Playwright browser cache, proving the e2e suite runs without any Playwright browser download when the `chrome` channel is selected. - `git diff --check` ## Risks - This assumes GitHub's Ubuntu runner continues to ship `google-chrome`; if that image contract changes, these workflows would need a dedicated Chrome install step. - The `chrome` channel can differ slightly from Playwright-managed Chromium, so the config gate is intentionally env-scoped to CI workflows that need the hosted-runner path. ## Model Used - OpenAI Codex, GPT-5-based coding agent running through Paperclip's `codex_local` adapter with tool use, shell execution, and repository editing enabled. The exact internal snapshot/version string is not exposed in-session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [ ] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b9a80dcf22 |
feat: implement multi-user access and invite flows (#3784)
## Thinking Path > - Paperclip is the control plane for autonomous AI companies. > - V1 needs to stay local-first while also supporting shared, authenticated deployments. > - Human operators need real identities, company membership, invite flows, profile surfaces, and company-scoped access controls. > - Agents and operators also need the existing issue, inbox, workspace, approval, and plugin flows to keep working under those authenticated boundaries. > - This branch accumulated the multi-user implementation, follow-up QA fixes, workspace/runtime refinements, invite UX improvements, release-branch conflict resolution, and review hardening. > - This pull request consolidates that branch onto the current `master` branch as a single reviewable PR. > - The benefit is a complete multi-user implementation path with tests and docs carried forward without dropping existing branch work. ## What Changed - Added authenticated human-user access surfaces: auth/session routes, company user directory, profile settings, company access/member management, join requests, and invite management. - Added invite creation, invite landing, onboarding, logo/branding, invite grants, deduped join requests, and authenticated multi-user E2E coverage. - Tightened company-scoped and instance-admin authorization across board, plugin, adapter, access, issue, and workspace routes. - Added profile-image URL validation hardening, avatar preservation on name-only profile updates, and join-request uniqueness migration cleanup for pending human requests. - Added an atomic member role/status/grants update path so Company Access saves no longer leave partially updated permissions. - Improved issue chat, inbox, assignee identity rendering, sidebar/account/company navigation, workspace routing, and execution workspace reuse behavior for multi-user operation. - Added and updated server/UI tests covering auth, invites, membership, issue workspace inheritance, plugin authz, inbox/chat behavior, and multi-user flows. - Merged current `public-gh/master` into this branch, resolved all conflicts, and verified no `pnpm-lock.yaml` change is included in this PR diff. ## Verification - `pnpm exec vitest run server/src/__tests__/issues-service.test.ts ui/src/components/IssueChatThread.test.tsx ui/src/pages/Inbox.test.tsx` - `pnpm run preflight:workspace-links && pnpm exec vitest run server/src/__tests__/plugin-routes-authz.test.ts` - `pnpm exec vitest run server/src/__tests__/plugin-routes-authz.test.ts server/src/__tests__/workspace-runtime-service-authz.test.ts server/src/__tests__/access-validators.test.ts` - `pnpm exec vitest run server/src/__tests__/authz-company-access.test.ts server/src/__tests__/routines-routes.test.ts server/src/__tests__/sidebar-preferences-routes.test.ts server/src/__tests__/approval-routes-idempotency.test.ts server/src/__tests__/openclaw-invite-prompt-route.test.ts server/src/__tests__/agent-cross-tenant-authz-routes.test.ts server/src/__tests__/routines-e2e.test.ts` - `pnpm exec vitest run server/src/__tests__/auth-routes.test.ts ui/src/pages/CompanyAccess.test.tsx` - `pnpm --filter @paperclipai/shared typecheck && pnpm --filter @paperclipai/db typecheck && pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/shared typecheck && pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm db:generate` - `npx playwright test --config tests/e2e/playwright.config.ts --list` - Confirmed branch has no uncommitted changes and is `0` commits behind `public-gh/master` before PR creation. - Confirmed no `pnpm-lock.yaml` change is staged or present in the PR diff. ## Risks - High review surface area: this PR contains the accumulated multi-user branch plus follow-up fixes, so reviewers should focus especially on company-boundary enforcement and authenticated-vs-local deployment behavior. - UI behavior changed across invites, inbox, issue chat, access settings, and sidebar navigation; no browser screenshots are included in this branch-consolidation PR. - Plugin install, upgrade, and lifecycle/config mutations now require instance-admin access, which is intentional but may change expectations for non-admin board users. - A join-request dedupe migration rejects duplicate pending human requests before creating unique indexes; deployments with unusual historical duplicates should review the migration behavior. - Company member role/status/grant saves now use a new combined endpoint; older separate endpoints remain for compatibility. - Full production build was not run locally in this heartbeat; CI should cover the full matrix. ## Model Used - OpenAI Codex coding agent, GPT-5-based model, CLI/tool-use environment. Exact deployed model identifier and context window were not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge Note on screenshots: this is a branch-consolidation PR for an already-developed multi-user branch, and no browser screenshots were captured during this heartbeat. --------- Co-authored-by: dotta <dotta@example.com> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
7f893ac4ec |
[codex] Harden execution reliability and heartbeat tooling (#3679)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies > - Reliable execution depends on heartbeat routing, issue lifecycle semantics, telemetry, and a fast enough local verification loop to keep regressions visible > - The remaining commits on this branch were mostly server/runtime correctness fixes plus test and documentation follow-ups in that area > - Those changes are logically separate from the UI-focused issue-detail and workspace/navigation branches even when they touch overlapping issue APIs > - This pull request groups the execution reliability, heartbeat, telemetry, and tooling changes into one standalone branch > - The benefit is a focused review of the control-plane correctness work, including the follow-up fix that restored the implicit comment-reopen helpers after branch splitting ## What Changed - Hardened issue/heartbeat execution behavior, including self-review stage skipping, deferred mention wakes during active execution, stranded execution recovery, active-run scoping, assignee resolution, and blocked-to-todo wake resumption - Reduced noisy polling/logging overhead by trimming issue run payloads, compacting persisted run logs, silencing high-volume request logs, and capping heartbeat-run queries in dashboard/inbox surfaces - Expanded telemetry and status semantics with adapter/model fields on task completion plus clearer status guidance in docs/onboarding material - Updated test infrastructure and verification defaults with faster route-test module isolation, cheaper default `pnpm test`, e2e isolation from local state, and repo verification follow-ups - Included docs/release housekeeping from the branch and added a small follow-up commit restoring the implicit comment-reopen helpers that were dropped during branch reconstruction ## Verification - `pnpm vitest run server/src/__tests__/issue-comment-reopen-routes.test.ts server/src/__tests__/issue-telemetry-routes.test.ts` - `pnpm vitest run server/src/__tests__/http-log-policy.test.ts server/src/__tests__/heartbeat-run-log.test.ts server/src/__tests__/health.test.ts` - `server/src/__tests__/activity-service.test.ts`, `server/src/__tests__/heartbeat-comment-wake-batching.test.ts`, and `server/src/__tests__/heartbeat-process-recovery.test.ts` were attempted on this host but the embedded Postgres harness reported init-script/data-dir problems and skipped or failed to start, so they are noted as environment-limited ## Risks - Medium: this branch changes core issue/heartbeat routing and reopen/wakeup behavior, so regressions would affect agent execution flow rather than isolated UI polish - Because it also updates verification infrastructure, reviewers should pay attention to whether the new tests are asserting the right failure modes and not just reshaping harness behavior ## Model Used - OpenAI Codex coding agent (GPT-5-class runtime in Codex CLI; exact deployed model ID is not exposed in this environment), reasoning enabled, tool use and local code execution enabled ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
42b326bcc6 |
fix(e2e): harden signoff policy tests for authenticated deployments
Address QA review feedback on the signoff e2e suite (86b24a5e): - Use dedicated port 3199 with local_trusted mode to avoid reusing the dev server in authenticated mode (fixes 403 errors) - Add proper agent authentication via API keys + heartbeat run IDs - Fix non-participant test to actually verify access control rejection - Add afterAll cleanup (dispose contexts, revoke keys, delete agents) - Reviewers/approvers PATCH without checkout to preserve in_review state Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
652fa8223e |
fix: invert reuseExistingServer and remove CI="" workaround
The playwright.config.ts had `reuseExistingServer: !!process.env.CI` which meant CI would reuse (expect) an existing server while local dev would start one. This is backwards — in CI Playwright should manage the server, and in local dev you likely already have one running. Flip to `!process.env.CI` and remove the `CI: ""` env override from the workflow. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
2c05c2c0ac | test: harden onboarding route coverage | ||
|
|
ccd501ea02 |
feat: add Playwright e2e tests for onboarding wizard flow
Scaffolds end-to-end testing with Playwright for the onboarding wizard. Runs in skip_llm mode by default (UI-only, no LLM costs). Set PAPERCLIP_E2E_SKIP_LLM=false for full heartbeat verification. - tests/e2e/playwright.config.ts: Playwright config with webServer - tests/e2e/onboarding.spec.ts: 4-step wizard flow test - .github/workflows/e2e.yml: manual workflow_dispatch CI workflow - package.json: test:e2e and test:e2e:headed scripts Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |