mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 16:11:46 +02:00
codex/native-completion-current-master-candidate
11
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d08abcba15 |
ci: cut PR wall clock from ~16 to ~6 minutes (#13521)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Every pull request runs the Trusted PR CI workflow before merge > - The test suites roughly tripled in six weeks, and shard balance did not keep up, so PR runs crept from ~4 to ~17 minutes > - Slow CI delays every merge and every contributor > - This pull request rebalances the shards from fresh measurements, splits the largest test files, reuses the Rust build cache in three more jobs, and takes the policy job off the critical path > - The benefit is a PR wall clock near 6 minutes with the same coverage ## Linked Issues or Issue Description **What existing behavior does this improve?** PR CI wall clock. A typical green run took 16-17 minutes. Two months ago it took about 4 minutes. **Subsystem affected** The Trusted PR CI workflow (`.github/workflows/pr-trusted.yml`), the shard-duration manifests, the vitest shard runner scripts, the `paperclip-runner` package scripts, and the dry-run branch of `release.sh`. **Current behavior** The shard-duration manifests were stale. The general-server manifest had durations for ~400 of 649 suites. The e2e manifest was missing 14 of 29 specs. Stale median weights made shard steps range 417s-806s (server) and 277s-745s (e2e). Three jobs each paid a ~3m40s cold cargo release build. Every test lane waited ~60s for the policy job before it could start. **Proposed behavior** All lanes finish in a narrow ~200-290s band. The manifests carry fresh measured durations for every suite. The three largest test files are split so no single file caps a shard. The Rust cache restore runs in every job that builds the Runner binary. Test lanes start as soon as the gate resolves. **Reason and benefit** Merges stop waiting on CI. The projected wall clock is ~6 minutes for the same test coverage. ## What Changed - Rebuild `scripts/general-server-shard-durations.json` (646 suites) and `scripts/e2e-shard-durations.json` (all specs) from per-suite completion timestamps in runs 35036001734 and 35024948947. - Move the PR server lane to the release-verify shape: `general-server-without-chat` across twelve duration-balanced shards, plus the chat integration suite split by collected test location across three dedicated lanes. - Split `tests/e2e/chat-adapters-ui.spec.ts` into `-providers` and `-messaging` specs, and `tests/e2e/agent-chat.spec.ts` into `-sessions` and `-projects` specs. Each pair shares fixtures through a `.shared.ts` module. Playwright collects the same test sets (39 and 20 tests). - Raise e2e shards to eight and serialized shards to nine. - Run the runner package's `check:all` as four matrix lanes: `check:static`, `check:runner`, and two native vitest `--shard` halves. The union is exactly `check:all`. - Add the read-only Rust cache restore (toolchain pin, `save-if: false`) to the Canary Dry Run, Build, and Typecheck jobs. - Make release.sh preview publish payloads concurrently in batches of eight during `--dry-run`. The real publish path stays strictly serial. - Drop the policy-job lockfile artifact chain. Each lane installs with `--frozen-lockfile` and falls back to an inline `--resolution-only` regeneration. The policy job stays a required check through the `verify` and `e2e` aggregates. - Update the shard-count mirrors and workflow assertions in the partition and gate tests. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs scripts/__tests__/e2e-shard.test.mjs` — 30 pass. - `node --test '.github/scripts/tests/'*.test.mjs` — 410 pass. - `node --test scripts/__tests__/release-verify-workflow.test.mjs scripts/cloud-source-verification.test.mjs scripts/__tests__/release-dry-run-notes.test.mjs` — 42 pass. - `playwright test --list` collects 39 tests across the chat-adapters split and 20 across the agent-chat split, equal to the original files. - A local vitest collection of the chat suite partitions 995 tests into 498/497 line shards. - Projected shard weights: server 230s x12, chat ~143s x3, e2e 207-242s x8, serialized ~216s x9. ## Risks - The split spec files reorder tests relative to the original files. Every describe seeds its own company, so the specs stay independent; a hidden cross-describe dependency would surface as a deterministic failure in one shard. - The inline lockfile fallback changes install behavior for manifest-changing and stacked PRs. The policy job still validates resolution as a required check. - `release.sh` changes are confined to the `--dry-run` preview branch. The publish loop is untouched. `bash -n` passes and the release dry-run tests pass. - One PR now schedules ~44 fleet runners. If the RunsOn fleet caps concurrency, queueing may absorb part of the gain; watch the first runs. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking, with tool use (shell, file edits) in Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
3bafac12f7 |
refactor: remove automatic productivity reviews (#13263)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its recovery loop keeps assigned work moving after execution failures. > - Productivity review used run counts, comment counts, and elapsed time to create management tasks. > - Infrastructure failures could satisfy those rules and create more tasks without evidence that the source work needed management review. > - This pull request removes that detector and its continuation holds. > - Bounded recovery, budgets, explicit blockers, and normal review stages remain in place. > - Existing task records stay readable and unchanged. ## Linked Issues or Issue Description Refs #5897. That request describes unwanted automatic productivity reviews and asks to preserve existing tasks. This change retires the feature instead of adding another configuration switch. Related prior approaches: Refs #9191, Refs #12489. Those changes excluded infrastructure failures or bounded review creation. This removal replaces the detector rather than tuning its thresholds. ## What Changed - Delete the scheduled detector, automatic task creation, evidence refresh, and productivity continuation holds. - Remove computed productivity fields, special attention items, badges, and Storybook fixtures. - Retain historical origin values, decision compatibility, and recovery recursion exclusions. Add no migration and change no existing task data. - Update the execution contract. Replace feature tests with regressions for legacy task reads, ordinary attention, and bounded continuation in the presence of an old review. ## Verification - Targeted attention, issue-route, startup, and UI tests: 4 files and 101 tests passed. - Updated issue-route and UI tests: 2 files and 61 tests passed. - Bounded continuation regression: 2 cases passed, including a legacy review plus pre-dispatch cancellation churn. - `pnpm check:token-gates`: all four gates passed. - `git diff --check`: passed. - `pnpm build-storybook`: passed. - Greptile: 5/5 on `a5a612eea`, with no actionable findings. - Scheduler and historical recovery regressions: 2 files and 28 tests passed. - Repository `pnpm -r typecheck` and `pnpm build`: passed. - The complete `pnpm test:run` suite passed across the CI server, serialized-server, and workspace shards on `a5a612eea`. Stopped the duplicate local monolithic run after the full CI suite passed; no completed local full-suite result is claimed. The targeted local suites above passed. - CI serialized shard 5 initially hit a 10-second timeout in the first interaction-route test. The complete file passed locally (78 tests), then the single CI rerun passed. - All CI gates are green, including the build and end-to-end suites. - A local merge check against current `master` (`ce09ea40b`) completed without conflicts. ## Risks - API responses no longer include the computed `productivityReview` field. Consumers must stop using it. - The scheduler no longer creates management work from elapsed time, run counts, or missing comments. This is the intended behavior change. - Existing review tasks and explicit dependencies remain in place. Historical origins still prevent recursive recovery treatment. No task cleanup or data migration occurs. - The native review handoff repair is separate from this removal. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
42961b6ef1 |
fix(ci): split release chat verification into test shards (#13198)
Split release chat verification into three validated test-line shards and balance other server suites across five runners using the measured native Runner integration cost. Retire each chat case's fixtures after assertions, preserve complete test coverage, and exercise the real shard CLI in PR tests. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
889947c238 |
feat: add experimental native chat connectors (#13038)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People also ask agents for work in their existing chat tools. > - Each external conversation needs one task and a current authorized source. > - Retries, Stop, and provider failures must not duplicate work or expose private data. > - The first chat PR establishes the opt-in provider and data contracts. > - This PR adds experimental channel integration and its durable control plane. > - Users can request work from connected channels and inspect delivery in Paperclip. ## Linked Issues or Issue Description Refs #13100 and #13092. This is the second of exactly two chat PRs. Foundation #13100 is merged and changed 143 files. Runner prerequisite #13092 is also merged. This PR changes 400 files against master, below the 500-file review limit. It contains no wireframe images or HTML galleries. ## What Changed - Add native Slack, GitHub, Microsoft Teams, Telegram, and Discord chat connections. Keep chat disabled unless the operator enables experimental chat connectors. Preserve the production GitHub tool connection and its normal setup path. - Bind each provider bot identity to one immutable Paperclip agent. Bind each admitted external conversation to one task. Paperclip owns tasks, runs, permissions, and audit records. - Add durable admission, per-conversation queues, questions, task controls, progress, final replies, images, files, and delivery receipts. Board comments remain internal unless explicitly sent to the channel. - Check current identity, provider reach, resource access, credentials, runtime generation, and exact source before provider effects. Keep private responses private. Never send raw reasoning, private logs, credentials, or tool arguments. - Hold uncertain sends for explicit audited resolution. Make Board Send-to-channel atomic and idempotent. Keep reconnect and setup credentials in Paperclip secret storage. - Preserve current native-runner authority across retries, lost acknowledgements, and recovery. Keep immutable input and completion contracts separate from newer user input. Receipt reconciliation cannot launch a provider. - Reconcile chat close/new ordering and provider-effect lock order. Audit resource access changes in the same transaction. Submit only the selected resource from each UI toggle so stale pages cannot undo unrelated access changes. - Drain Codex stdout before certifying process exit. Bound the drain with the existing shutdown grace. Preserve observed terminal authority without treating an undrained process as successful or reusable. - Incorporate master `018ca5da` with its ACP Stop, mobile task layout, runner packaging, and official lock changes. Preserve dedicated chat-answer continuations in both directions when ordinary queued comments are adopted after Stop. - Fence late adapter readiness behind an earlier Stop for the same run. Preserve verified cleanup for registered adapters. Handle single Stop, agent pause, duplicate Stops, and failure release without creating a false cancellation receipt. - Incorporate master's `6dd48cad4` wake-queue extraction. Preserve exact failed-chat retry authorization and lineage, retired question-source suppression, and the block on generic recovery that would discard the admitted source. Fresh deferred input retains its separate promotion path. - Incorporate master `2a05b5ed3` and its queue-admission extraction, simplified transaction ports, and separate runner CI job. Preserve exact durable receipts, actor separation, and dedicated-answer isolation through the new module. A failed receipt insert rolls back the accompanying deferred-wake merge. ## Verification Current head: `afe19299d06253cb628eb398e91d1200ea9f412a`, incorporating master `2a05b5ed3457ea33efd6895520447d1d97fe98d8`. The conflicts are resolved. This successor fixes two test-harness boundaries exposed by CI: per-case route-module preparation and actual durable-save completion before intentional runner termination. Production code and all existing test/turn deadlines are unchanged. [Exact-head Greptile review](https://github.com/paperclipai/paperclip/pull/13038#issuecomment-5587250594) is **5/5**, completed September 10 at 13:20:55 UTC, with no actionable findings or open review threads. [Fresh exact-head CI](https://github.com/paperclipai/paperclip/actions/runs/34481724341) passes **all 24 jobs**, including Build and both required aggregates. Normal exact-head guarded merge was attempted and rejected by the remaining branch approval policy: CODEOWNER review is required and no human approval is present. Normal **squash auto-merge is enabled** as of September 10 at 13:36:26 UTC. Requested CODEOWNERS have been notified; no approval bypass or self-approval was used. Earlier-head results below remain historical evidence, not qualification of this successor. - Final exact-head Linux evidence: 995/995 chat integration cases; 36/36 agent-skills routes; 35/35 runner live-session cases, including real process kill/resume; 1948 runner Vitest cases with three existing benchmark/platform guards; 870/870 API-authority cases; and 104 browser cases with four existing optional skips. Rust, conformance/replay, full repository build, typecheck, canary, all server/workspace shards, and both required aggregates pass with normal CI concurrency. Earlier failed attempts remain recorded below. - Latest test-only qualification: 141/141 route/permissions/authentication cases pass in separate cold forks, with plain server types and independent review clear. The real-runner suite passes 35/35, with plain runner types and independent review clear. A controlled premature-save acknowledgement fails as expected; matching ownership/effect/process evidence, rejected saves, real turn outcome, test abort, and pre-kill liveness are covered. No local reproduction of the original CI scheduling failure is claimed. The preceding [CI run](https://github.com/paperclipai/paperclip/actions/runs/34479680858) passes 21/24 jobs, including all 995 Linux chat cases and browser aggregate (104 passed, four existing optional skips); only Build, the skills serialized shard, and the required verification aggregate fail. Its exact-head Greptile review was 5/5. Both failed job logs are retained. - Final fixture qualification: all eight focused Discord cases and all 995 chat integration cases pass. The exact modal statement/PID is observed before taking the real connection lock; the test then proves its actual blocking relationship before mutation. Original SQL execution, provider behavior, negative assertions, and 1s/15s timeouts remain unchanged. Independent review is clear and test/production hashes remain frozen. The preceding [CI attempt](https://github.com/paperclipai/paperclip/actions/runs/34477184777) passed 22 jobs, including Build/runner, typecheck, canary, all other test shards, and browser aggregate (104 passed, four existing optional skips); the two fixture failures and failed verification aggregate remain recorded, not relabeled as a pass. - Current queue-module composition: 308/308 recovery/batching/queue/Stop tests; 995/995 full chat integration; 89/89 module tests, including real PostgreSQL receipt-insert rollback; 24/24 workflow/module-boundary tests; plain server and UI types. All four actual local process/ACP browser paths pass in 1.4 minutes. Fresh databases, no skips or retries, stable reviewed source hashes. The initial boundary failure is retained; its no-op service wrapper was removed without changing recovery context or weakening the check. An exploratory standalone test-directory typecheck fails because its new upstream transformation config is not a standalone typechecking project; standard CI/build does not invoke it, and no configuration was weakened to suppress those diagnostics. - The preceding head `e02a63d462ce5d47433b0aeb632bb6fd20aab1ba` passed [all 24 CI jobs](https://github.com/paperclipai/paperclip/actions/runs/34436462958) and exact-head Greptile review at 5/5. Required CODEOWNER review prevented its normal merge before master advanced again. - Final extracted-module composition: 307/307 recovery, batching, queue and Stop-control tests; 995/995 full chat integration; 49/49 module tests including eight PostgreSQL adapter cases; and 19/19 issue-update tests. Plain server types pass. All four actual local process/ACP browser paths pass in 1.3 minutes. Fresh databases, no skips or retries in these cohorts, frozen source hashes, and independent review clear. - The preceding head `3e4e1c1c` passes [all PR CI jobs](https://github.com/paperclipai/paperclip/actions/runs/34415826820), including Build and required `ci / verify` and `ci / e2e`. Both the original Rust failure and the previously load-sensitive lineage fixture pass with unchanged Linux concurrency. Master advanced afterward and required this reconciliation. - Final master composition: 448/448 focused UI tests, 186/186 adapter tests, 24/24 queue/control tests, and 11/11 packaging tests. Plain UI, server, shared, and adapter types pass. Token gates and diff checks pass. Independent server and UI reviews are clear. - Stop-registration regression: both real-service cases fail against exact `a95` source and pass with the fix. The full corrected recovery/control suite passes 265/265. Duplicate-owner and failed-Stop controls also pass. Plain server types pass. The readiness barrier prevents provider startup without adding an acknowledgment to an already terminal run. - Final qualification strengthens terminal-field equality and repeats both affected cases successfully on a fresh database. All four actual local process/ACP browser paths pass again in 1.3 minutes, without skips or retries. The final screenshot shows Cancelled, a paused subtree, retained input, and no error toast. - Two new actual-service regressions fail before the merge fix. They prove that queued-comment adoption could consume a dedicated chat answer or add unrelated input to that answer. The fixed four-case cohort passes, including ordinary upstream continuation and adapter Stop controls. Full recovery passes 257/257. All four actual local process/ACP Stop browser flows pass in 1.4 minutes, without skips or retries, on a fresh database. - The unchanged runner artifact was qualified with 171/171 transport tests, 870/870 API-authority tests, conformance 1/1, and replay 11/11. Six controlled reader tests prove the exit/drain repair. Its local serial Rust workspace passed 546 top-level cases plus two invoked helpers; the later passing Linux CI supplies default-concurrency evidence. - Prior exact-source full chat integration passes 995/995. Settings regressions cover concurrent stale pages, 501 destinations, pending state, rejected updates, and explicit retry. These deterministic tests do not prove live provider behavior. - Retained failed attempts and their causes are in the [qualification log](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-chat-queue-and-webhook-repair.md). The first merge adapter run timed out while macOS slept for 290 seconds. Its unchanged repeat passed with a temporary sleep guard. No assertion, deadline, or CI gate was weakened. Review commands include `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-process-recovery.test.ts src/__tests__/issue-queued-comments-routes.test.ts` and `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/acp-stop-continuation.spec.ts`. Database suites require fresh disposable databases. See the [browser runbook](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-04-chat-adapters-browser-e2e-runbook.md) for provider setup and separate live acceptance steps. ## Risks - This remains experimental. Deterministic tests and bounded live evidence do not establish every provider feature, tenant, permission layout, or media shape. Teams work-tenant qualification is still open. - Failed and uncertain provider effects remain visible and can require operator action. A transport receipt does not prove recipient visibility. - Native controller and runner artifacts must remain compatible. Preserve lease ownership, terminal authority, source binding, and quarantine during future changes. - Access and audit rows commit together, but activity notifications remain best-effort. This is not a new durable event outbox. - The PR operation does not deploy a live server, replace its runner, or change provider permissions. Remaining live qualification is documented in the [temporary handoff](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-open-qualification-followups.md). ## Model Used OpenAI Codex assisted with implementation, tool execution, testing, and review. The work records `gpt-6-astra` assistance. The environment does not report a context-window size. No private reasoning traces are included. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a20a4944ec |
feat: add Grok device login to the sandbox login panel (#12469)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip uses adapters to connect agents and model providers to its control plane > - The sandbox login panel supports displayed-code login for selected adapters > - Grok users need the same login path and a private credential home for later runs > - This pull request adds Grok support to the shared device-login path and preserves the existing Codex path > - The benefit is one secure login flow for both adapters with company-scoped credential storage ## Linked Issues or Issue Description **Agent or provider** Grok Local needs displayed-code login support in the sandbox login panel. **Why this adapter is useful** This change lets users sign in to Grok from the sandbox login panel. It also gives later Grok runs access to the stored credential. **How the agent is invoked** The Grok local adapter uses its login command through the shared displayed-code login flow. Later runs receive the managed home through `GROK_HOME`. **Additional context** The change uses adapter-scoped login lifecycle handling. It stores the credential in a company-scoped directory with mode `0700`, and it stores the credential file with mode `0600`. ## What Changed - Rename the shared device-login modules to adapter-neutral names. - Scope the shared login lifecycle to a closed adapter set. - Return the device-login URL that the provider prints. - Add the Grok prompt parser, login command, capability, and login panel entry. - Store the Grok credential in a private, company-scoped home directory. - Pass `GROK_HOME` to later Grok runs. - Add tests for the Grok adapter, the Daytona sandbox provider, the server login path, and the user interface. ## Verification - Run `pnpm vitest run packages/adapters/grok-local/src/server/adapter-auth-promotion.test.ts`. - Run the Grok adapter package suite. - Run the Daytona sandbox provider suite. - Run the server device-login suites. - Run the user interface suite. - Confirm the full CI suite passes. ## Risks The change extends shared login lifecycle code to another adapter. A regression could affect Codex login. The credential path uses explicit `chmod` calls to keep the directory at mode `0700` and the file at mode `0600`. ## Model Used OpenAI Codex, GPT-5. The runtime used tool calls and code review support. The runtime did not provide a context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
64b7dce0ad |
refactor(adapter-utils): replace the process-wide byte ledger with route-local byte bounds (#12465)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The adapter layer carries sandbox requests to host processes. > - The HTTP/2 bridge used one process-wide byte ledger for all routes. > - One busy route could exhaust that shared budget and move another route to file transport. > - This pull request gives each host retention site a fixed byte bound and limits concurrent HTTP/2 streams. > - The benefit is local protection: one route cannot consume the byte budget of another route. ## Linked Issues or Issue Description **What happened?** The HTTP/2 bridge used one aggregate byte ledger for retained bytes across all routes. A busy route could exhaust the shared budget and force an unrelated route to use file transport. **Expected behavior** Each route should protect its own retained bytes. A reset on one HTTP/2 stream should cancel only that stream's host forward. **Steps to reproduce** 1. Start the HTTP/2 bridge with multiple sandbox routes. 2. Send enough retained data through one route to reach the aggregate byte limit. 3. Send a request through a sibling route. 4. Observe that the sibling route can fall back to file transport because the first route used the shared ledger. **Paperclip version or commit** `47639e227e78e3c5e0dd1a3c0e2d792fe86895a3` **Deployment mode** Built from source with the adapter-utils and server test suites. ## What Changed - Bound each host retention site with a fixed local byte limit. - Limited concurrent live HTTP/2 streams with one built-in stream limit. - Bound each host forward and response-body read to its own HTTP/2 stream lifetime. - Removed the process-wide byte ledger, its environment override, its metrics, and its file-transport fallbacks. - Added tests for the stream limit, host body budget, and sibling-stream cancellation. ## Verification - Run `pnpm vitest run --project adapter-utils`. - Confirm that 996 adapter-utils tests pass. - Confirm that `test_live_forward_work_never_passes_the_stream_limit` passes. - Confirm that `test_the_host_body_budget_matches_the_stream_limit` passes. - Confirm that the sibling-stream cancellation test passes. - Run `pnpm tsc --noEmit`. - Confirm that all pull request checks pass. ## Risks The bridge no longer uses a process-wide byte ledger. A local bound or stream limit that is too low can reject or delay valid work. The tests cover the new limits and stream cancellation behavior. ## Model Used OpenAI GPT-5 Codex. Runtime model ID: GPT-5. The model used code execution and repository tools. The runtime does not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d1573244b5 |
refactor: disambiguate the Telemetry and Observability data paths (#12128)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip records first-party events, OpenTelemetry data, and local run-log events > - The code and documents used one term for these three data paths > - This naming made the required review level unclear > - This pull request names each data path in the module names, documents, and code comments > - The benefit is a clear review rule without a runtime change ## Linked Issues or Issue Description **Issue type** Unclear or confusing. **Where is the issue?** `packages/shared/src/telemetry/README.md`, `doc/observability.md`, `doc/run-log-events.md`, and the duplex instrumentation modules. **What's wrong?** The repository used Telemetry for first-party events, OpenTelemetry data, and local run-log events. This usage made the data path and review level unclear. **Suggested fix** Use Telemetry only for Paperclip first-party events. Use Observability for OpenTelemetry data. Use the run log for rows in `heartbeat_run_events`. Related public pull requests: #8476 and #9672. ## What Changed - Rename the duplex instrumentation modules and identifiers from `Telemetry` to `Observability`. - Move the Observability and run-log contracts out of the Telemetry README. - Add `doc/observability.md` and `doc/run-log-events.md` as the canonical documents. - Add a file-path review rule to `AGENTS.md`. - Correct the remaining code comments that name the wrong data path. - Keep all event names, payloads, database records, spans, configuration keys, environment variables, and runtime paths unchanged. ## Verification - `npx vitest run packages/shared/src/telemetry/readme-contract.test.ts` passes. - `npx vitest run packages/adapter-utils/src/published-exports.test.ts` passes. - `npx vitest run packages/adapter-utils/src/acpx-engine/startup-timing.test.ts` passes with 42 tests. - `pnpm --filter @paperclipai/adapter-utils typecheck` passes. - `pnpm --filter server typecheck` passes. - The old module name does not remain in TypeScript or JSON files, except for the intentional publication guard. - CI and Greptile checks remain pending after PR creation. ## Risks - The old duplex module subpath no longer has a compatibility shim. The board accepted this intentional hard break. - The new duplex module subpath stays blocked from package publication. - The change has no runtime effect. The main risk is an incorrect document or module reference. ## Model Used OpenAI GPT-5 Codex, exact model ID `gpt-5`, with tool use and code review support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR with the documentation issue fields - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0dfa0fb988 |
ci: refresh general-server shard duration manifest (#12075)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The PR verify workflow gates every pull request; its slowest check sets the feedback time for all contributors > - The general-server test lane splits its vitest suites across five runners with a duration-weighted partition (`scripts/general-server-shard.mjs`) > - The partition reads a duration manifest that was sampled on 2026-08-04, when the lane had 279 suites and 946s of serial time > - The lane has since grown to 405 suites and 1274s; 126 suites had no recorded duration and one suite grew from 37s to 123s > - The stale weights made the partition uneven: in the fully green actions run 32708351172, "General tests (server (2/5))" ran 364s and was the slowest check in the whole run, while sibling shards ran 292-330s > - This pull request refreshes the manifest with per-suite durations measured from that same run > - The benefit is a level five-shard split (255s ±1s of predicted suite time per shard), which removes ~50s from the slowest PR check ## Linked Issues or Issue Description **Describe the current behavior** In the fully green PR actions run [32708351172](https://github.com/paperclipai/paperclip/actions/runs/32708351172) (2026-08-24), the check "General tests (server (2/5))" completed in 364s. Its test step ran 315s while sibling shards ran 241-276s. It was the slowest check in the run. **Describe the improvement** The duration manifest `scripts/general-server-shard-durations.json` is stale. It holds 279 suites sampled on 2026-08-04, but the lane now has 405 suites. The 126 unknown suites fall back to the median weight (~1.3s), and `server/src/__tests__/workspace-runtime.test.ts` grew from 37.4s to 123.3s. The partition therefore predicts a level split but produces an uneven one. Refreshing the manifest restores the level split without any code change. **Expected impact** All five server shards level at ~255s of predicted suite time (~310s job time). The slowest PR check drops from 364s to about 317s, so the PR critical path improves by roughly 50s. ## What Changed - Regenerated `scripts/general-server-shard-durations.json` from actions run 32708351172 (2026-08-24): 405 suites, 1274s total serial time (was 279 suites, 946s from 2026-08-04) - Updated the `$comment` field to name the new sample run and date - No code changes; the partition logic in `scripts/general-server-shard.mjs` is untouched ## Verification - Parsed all five "General tests (server (n/5))" job logs from run 32708351172 with the consecutive-completion-timestamp method described in the manifest `$comment`; asserted that the parsed suite set equals the exact file list that `run-vitest-stable.mjs` collects (405/405, no misses, no extras) - Ran `node scripts/run-vitest-stable.mjs --mode general --group general-server --shard-index N --shard-count 5 --dry-run` for N=0..4 with the new manifest: each shard predicts 255s (±1s) of suite time, and the five shards form a complete, non-overlapping cover of all 405 suites - Ran `node --test ./scripts/__tests__/run-vitest-stable-shard.test.mjs`: 13/13 pass ## Risks - Low risk. The change is data-only. Wrong weights cannot break correctness: the partition always covers every suite exactly once, so the worst case of a bad weight is an uneven shard, which is the current state. ## Model Used - Claude (Anthropic), model ID `claude-fable-5`, agentic coding session with tool use (Claude Code / Claude Agent SDK) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable (no code change; existing partition tests pass) - [x] I have updated relevant documentation to reflect my changes (manifest `$comment` updated) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Related prior work: #11528 (balanced the serialized server shards by recorded duration), #10923 (split serialized tests into five shards), #11156 (split workspaces-a into two shards). Co-authored-by: Claude <noreply@paperclip.ing> |
||
|
|
00a24d7e8f |
ci: split general-server tests into five shards with refreshed durations (#10925)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The PR workflow runs the server vitest suite across sharded runners because the suite is pinned to one worker. > - In successful PR run 30930345729 (2026-08-04), shard `server (3/4)` took 311 seconds of wall time and was the slowest check in the run. > - The suite has grown to about 946 seconds of serial vitest time, but the duration manifest was last sampled on 2026-08-01 at about 882 seconds. > - This pull request refreshes the per-suite duration manifest from that run's logs and splits the lane into five shards. > - The benefit is a shorter PR critical path: each shard carries about 196 seconds of suite time, level with the other lanes. ## Linked Issues or Issue Description Refs #10663 (previous split of this lane into four shards). Related: #10923 splits the separate serialized-suites lane into five shards. Both PRs touch `.github/workflows/pr.yml` in different matrix blocks; whichever merges second needs a trivial rebase. **What existing behavior does this improve?** The `general-server` vitest lane runs in four shards with a duration manifest sampled on 2026-08-01. **Current behavior** In PR run 30930345729, shard 3/4 ran for 311 seconds (273 seconds in the test step) and was the longest check in the run. The suite now totals about 946 seconds of serial vitest time. **Proposed behavior** Run the same suite set in five shards, balanced with a per-suite duration manifest refreshed from that run's shard logs (279 suites measured by diffing consecutive completion timestamps). **Reason and benefit** The refreshed LPT partition balances at about 196 seconds of suite time per shard (about 240 seconds per job), level with the other PR lanes. No test coverage is lost. **Breaking changes** None. The change only alters the CI partition size and the duration manifest. ## What Changed - Bump the `general-server` shard matrix in `.github/workflows/pr.yml` from four to five shards. - Refresh `scripts/general-server-shard-durations.json` from the 2026-08-04 run's shard logs. - Update `SHARD_COUNT` in `scripts/__tests__/run-vitest-stable-shard.test.mjs` to five. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs` — 9/9 pass, including the complete non-overlapping partition proof and the duration-balance check. - `node --test scripts/__tests__/release-verify-workflow.test.mjs` — 2/2 pass. - `node --test scripts/__tests__/e2e-shard.test.mjs` — 7/7 pass. - A 5-way dry-run partition covers all suites exactly once with equal projected weights. ## Risks - Low risk. The change only alters CI partition size and duration weights; the suite set is unchanged. - One more runner is used per PR run for this lane. - Stale duration weights degrade gracefully: suites missing from the manifest get the median weight. ## Model Used - Claude (Anthropic), Claude Code CLI, model ID `claude-fable-5`, extended thinking with tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (workflow comments explain the new shard math) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Claude <claude@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8540ce2973 |
ci: shard general-server tests 4 ways (#10663)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Pull request CI must give contributors fast and stable feedback. > - The `general-server` Vitest lane runs many single-worker server suites. > - A recent completed PR run showed this lane as the slowest completed check. > - Three shards still left one runner with the largest share of work. > - This pull request splits that lane into four duration-balanced shards. > - The benefit is a shorter critical path for the same server test coverage. ## Linked Issues or Issue Description No public GitHub issue exists for this CI maintenance change. **Pre-submission checklist** - I confirmed this improves existing behavior. It does not add a new command, endpoint, or concept. - I searched open public issues and pull requests for related CI sharding work. **What existing behavior does this improve?** The pull request workflow's `general-server` Vitest lane. **Subsystem affected** Cross-cutting. This affects GitHub Actions CI and the Vitest shard duration manifest. **Current behavior** The `general-server` lane uses three shards. The server suites now total about 880 seconds of serial Vitest wall time. The slowest shard was about 313 seconds in the measured run. **Proposed behavior** The `general-server` lane uses four shards. Each shard receives about 220 seconds of predicted suite weight from the refreshed duration manifest. **Reason and benefit** The slowest PR check controls how soon a reviewer can trust the PR. Four balanced shards reduce the slowest `general-server` shard while keeping the same suite selection rules. **Breaking changes** None. This only changes CI partitioning and duration data for existing test suites. **Additional context** Related public searches found no exact open issue or pull request for this `general-server` sharding change. ## What Changed - Split the `general-server` CI matrix from three shards to four shards. - Refreshed `scripts/general-server-shard-durations.json` with wall-time weights from a recent completed PR run. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs` - `git diff --check origin/master...HEAD` - Dry-ran the four `general-server` shards locally during implementation. The partition covers 300 unique suites with about 220.56 seconds of predicted weight per shard. - Ran a local sensitive-data scan before push. It found only test filenames that contain words such as `secret` or `token`, not credential values. ## Risks Low risk. The main risk is that the duration manifest becomes stale as suite costs move. Missing suites fall back to the median weight, so the lane still runs if the manifest is incomplete. ## Model Used OpenAI Codex, GPT-5, with tool use and local command execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ce7dedf33d |
perf(ci): balance general-server test shards by recorded suite duration (#9516)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its PR CI runs the general-server vitest lane pinned to `maxWorkers=1` and sharded across 3 runners (introduced in #8360) > - Suites were assigned to shards round-robin by sorted file index, so shard test time was unbalanced: a recent PR run split 73s / 153s / 115s, and the heaviest shard made "General tests (server 2/3)" the slowest check in the whole workflow at 314s wall > - The slowest shard sets the lane's wall time, so unbalanced partitions waste the other two runners and stretch the PR critical path > - This pull request replaces the round-robin assignment with a deterministic longest-processing-time partition weighted by a checked-in per-suite duration manifest > - The benefit is near-even shard weights (projected 113s / 113s / 113s with the current manifest), taking roughly 40s off the PR critical path with no reduction in coverage ## Linked Issues or Issue Description - Refs #8360 (introduced the 3-way general-server sharding this PR rebalances) - No public issue exists. Problem: the general-server test lane's round-robin shard assignment ignores per-suite duration, so one shard can carry multiple 30s+ suites while another finishes in half the time; the slowest shard alone determines the check's wall time. ## What Changed - `scripts/general-server-shard.mjs` (new): manifest loader and deterministic LPT (longest-processing-time) partitioner; suites missing from the manifest get the median recorded weight, and a missing or malformed manifest degrades to uniform weights so the lane never fails on stale data - `scripts/general-server-shard-durations.json` (new): per-suite duration manifest sampled from a real PR run (240 suites); the `$comment` field documents how to regenerate it - `scripts/run-vitest-stable.mjs`: both shard-selection sites (run and `--dry-run`) now use the balanced partition instead of index round-robin - `scripts/__tests__/run-vitest-stable-shard.test.mjs`: 6 new tests covering skew-balance vs round-robin, determinism, median fallback for unlisted suites, malformed-manifest degradation, manifest coverage of the current suite set, and real-partition balance - `server/src/__tests__/heartbeat-issue-rewake-throttle.test.ts`: hardened the `afterEach` sweep — post-run bookkeeping (run-event records, follow-up wake scheduling) can still insert rows briefly after a run reaches a terminal status, and a late insert landing between the `agent_wakeup_requests` and `agents` deletes failed teardown with a foreign-key violation on the first CI attempt of this PR; the sweep now retries so a late background write cannot take down the shard - `release-verify.yml` shares the same runner script and inherits the balancing with no workflow change ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs` — 9/9 pass (run against current master) - `npx vitest run src/__tests__/heartbeat-issue-rewake-throttle.test.ts` — 6/6 pass against embedded Postgres with the hardened teardown - `node --test scripts/__tests__/release-verify-workflow.test.mjs` — 2/2 pass - `node scripts/run-vitest-stable.mjs --dry-run` with each shard flag shows every suite assigned exactly once across the 3 shards, with projected weights ~113s each ## Risks - Low risk: partition changes which runner executes which suite, not what runs; a completeness test asserts every suite is assigned to exactly one shard - The duration manifest will drift as suites are added/changed; unlisted suites get the median weight and a coverage test flags when the manifest covers less than half the suite set, so drift degrades balance gracefully rather than breaking the lane > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking enabled, agentic tool use (file edits, shell, test execution) via Claude Code ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude (Paperclip SWE) <noreply@paperclip.ing> |