mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 00:54:38 +02:00
codex/native-completion-master-final-baseline
10
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7a52dcdc74 |
fix: repair MCP validation and cancelled execution recovery (#14951)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The tool gateway gives agents access to connected services. Recovery controls what happens when a run stops. > - Generated tool names can exceed the provider limit after the MCP client adds its prefix. > - The same invalid definition can fail each automatic retry. A cancelled run can also hold saved messages without showing its cause. > - This pull request bounds tool names, stops configuration retries, and retains cancellation evidence. > - It shows the stopped run and admits saved input only after the existing safety checks pass. > - The benefit is a clear recovery path that preserves operator Stop and prevents duplicate message delivery. ## Linked Issues or Issue Description **What happened?** A long connected MCP tool name makes the provider reject the entire request. Automatic recovery repeats the invalid request. Separately, unexpected legacy cancellations can leave saved input behind a recovery hold. The notice does not identify the stopped run or its cause. **Expected behavior** Complete MCP names fit the provider limit. Tool-definition errors require configuration repair. Cancelled runs retain their source and reason. The recovery notice shows the cause and saved-message count. Verified unexpected cancellations can start a fresh turn through the existing admission checks. **Steps to reproduce** 1. Assign an App gallery connection with a long application key and tool name to a Claude agent. 2. Start a run. The provider rejects a name over 128 characters, including its MCP prefix. 3. For cancellation recovery, stop a legacy provider turn without an operator Stop request and send a user message while the recovery hold is active. 4. Inspect the recovery notice and the deferred message queue. **Paperclip version or commit** Rebased onto master at `cf8ad63c806685bfd7c48e3ed4a919d61a7c55f1`. **Deployment mode** Hosted or self-hosted server with legacy Claude or Codex execution. Related public work: - Refs #14017. That PR caps name segments. This PR preserves existing short names and uses stable hash aliases for long complete names. It also covers classification and recovery. - Refs #4510. That PR adds a cancellation-source column. This PR records bounded evidence in the existing run result, without a migration. - Refs #12552 and #4506. Those PRs suppress recovery after operator cancellation. This PR preserves operator intent and uses the existing continuation gates. ## What Changed - Bound gateway names with the full provider prefix in the 128-character budget. Retain the original upstream tool name for dispatch and permissions. - Classify invalid tool definitions as configuration failures before diagnostic redaction. Stop automatic retries and continuation attempts for that error code. - Persist cancellation source, expectedness, initiator, reason, and time. Preserve recorded Stop intent when adapter results arrive. Report unexpected started cancellations with closed diagnostic labels. - Show the run cause, saved-message count, and Inspect run link. Offer Continue for eligible unexpected cancellations. Require verified provider stop, empty tool inventory, ownership, and the existing pause, budget, approval, and dependency gates. Use the existing queue for single delivery. - Add regression coverage and update the execution, MCP gateway, and run-log documentation. ## Verification - `pnpm -r typecheck` and `pnpm build` passed. - `pnpm check:token-gates` passed. - Ran `pnpm test:run` and completed its workspace and serialized groups. Initial resource and timing failures passed on isolated reruns. All 149 serialized route suites passed. - Reran the changed server, adapter, and UI suites after the rebase. Coverage includes long-name upstream dispatch, configuration retry suppression, cancellation evidence retention, privacy labels, oversized run projection, and concurrent saved-message delivery. - `pnpm test:e2e tests/e2e/legacy-failure-continuation.spec.ts` passed all six browser scenarios. The recovery notice shows the run cause and inspection link, and each recovery entry point reaches one new response. - Added database-backed checks for active, removed, paused, unavailable, and disabled chat connections. The final continuation and recovery-notice suites passed 167 tests. Externally bound chats hide board Continue and show a usable next action. - All 55 GitHub checks passed on `42afbf1371dcaeb72646e3d8f65c19ff7cddf8de`. Two unrelated Storybook jobs were skipped by their normal conditions. Greptile reviewed that commit at 5/5 with no findings and no open review threads. ## Risks - Long tool names change to aliases. Existing short names stay compatible. The original connection and upstream name remain the dispatch authority. - Invalid tool definitions no longer get automatic retries. An operator must repair the configuration before a new attempt. - Continuation changes apply only to positively identified unexpected legacy cancellations with complete empty tool inventory. Operator Stop, unknown historical cancellations, outstanding tools, and unverified provider termination keep their holds. - No database migration. The added projection fields are optional. Cancellation reason and initiator IDs remain local run evidence; Sentry receives only closed source and initiator-type labels and expectedness. ## Model Used - OpenAI GPT-6 through Codex, with reasoning, repository editing, shell execution, and GitHub tool use. The runtime does not expose the exact model variant or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0b12ca9532 |
fix(server): retain context for unconfirmed adapter stops (#14639)
## Thinking Path > - Paperclip must keep ownership of work until termination is verified. > - A Stop request waits for the adapter and its cleanup to settle. > - After 60 seconds, an unconfirmed Stop raises an error. > - The error currently lacks run, adapter, and runtime context. > - This makes it difficult to investigate which stop path is stuck. > - This change adds bounded diagnostics while preserving termination checks. ## Linked Issues or Issue Description **What happened?** An adapter can remain unsettled after its Stop request. The resulting error says termination is unverified but does not identify the adapter or run in error monitoring. The optional Sentry setup does not capture request context, so the endpoint alone cannot fill the gap. **Expected behavior** Keep the Stop unconfirmed and preserve its live execution owner. When optional Sentry is enabled, attach enough bounded context to investigate the affected run. **Steps to reproduce** 1. Register an adapter execution control and abort its controller. 2. Leave its settlement promise pending. 3. Wait for the configured Stop timeout. 4. Observe that Stop still fails, but the event now includes the run UUID, built-in adapter, native/legacy runtime, timeout duration, and abort-requested flag. **Paperclip version or commit** Master commit `17780751551b3bc1c2521f7694026c34534c46c9`; reproduced with fake timers and mocked optional error monitoring. **Deployment mode** Server execution control, including local and hosted runs. Reporting remains opt-in. Searched related Stop PRs. #14523 and #14244 address Hermes cancellation contracts; this change only adds diagnostics to the shared unconfirmed-stop timeout. ## What Changed - Use a typed timeout error with the existing message, name, and timer stack. - Pass run/adapter/runtime identity from the cancellation owner. - Add an event-local, allowlisted Sentry context without changing the default fingerprint. - Rebuild the reported exception so arbitrary provider fields cannot be serialized. - Test timeout ownership, delayed settlement, privacy boundaries, and absence of context on unrelated events. - Document the additional opt-in fields. ## Verification - `pnpm exec vitest run server/src/services/adapter-execution-control.test.ts server/src/__tests__/sentry.test.ts`: 38 passed, five real-SDK checks skipped because the optional package is not installed. - `pnpm -r typecheck` passed; server typecheck passed again after the SDK test addition. - With audited optional `@sentry/node@10.71.0` installed only in local test dependencies, `PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1 pnpm exec vitest run server/src/__tests__/run-failure-sentry-real-sdk.test.ts server/src/services/adapter-execution-control.test.ts server/src/__tests__/sentry.test.ts`: all 45 tests passed. The real SDK uses an in-memory transport; no Sentry requests are sent. - The broad local `pnpm test:run` command did not complete in the available verification window and was stopped; no full local-suite pass is claimed. `pnpm build` passed. All sharded GitHub CI checks passed on the final PR head. - Tests use fake timers and a mocked Sentry package; no provider or monitoring requests. ## Risks This is diagnostic coverage, not a claim that the underlying stop delay is fixed. Unknown adapter/runtime values become `unknown`; malformed run identifiers become `null`. No stop reason, prompt, output, provider response, credentials, or arbitrary error properties are sent. Timeout, cancellation acknowledgement, live-owner retention, and retry behavior remain unchanged. No schema changes. ## Model Used OpenAI Codex (GPT-6), with reasoning, repository inspection, and command execution. The session does not expose a more specific model revision or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the targeted tests locally and they pass; full checks are in progress - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b3eb03fcba |
fix(sentry): preserve run failure stacks and diagnostic context (#14585)
Builds on merged #14575 and targets `master`. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators use optional Sentry reports to diagnose failed agent runs. > - The shared reporter currently converts each saved error message into a new exception. > - This loses the original stack and cause chain, and omits available adapter and provider diagnostics. > - This pull request selects, redacts, and bounds useful diagnostics before it sends them to Sentry. > - Operators can locate the failing operation and correlate upstream failures without copying arbitrary run data. ## Linked Issues or Issue Description Refs #14573. That change preserves ACP provider failure details in the run record. Refs #14575. That merged PR adds recorded exit codes and signals. This PR covers stacks, causes, and structured diagnostics without duplicating those fields. A thrown setup or adapter error currently appears in Sentry with the reporter's stack. A saved provider failure can contain useful details that never reach the Sentry event. Both cases use the same shared reporting path. ## What Changed - Pass caught setup and execution exceptions, and structured adapter error metadata, through successful terminal status transitions. - Snapshot a fixed set of execution, adapter, provider, and exception fields. Preserve up to four exceptions in the cause chain. - Remove registered secret values, declared runtime environment credentials, unknown inherited environment values and encoded credential forms, and credential patterns before truncation. Mark truncated fields and bound provider details and stacks. - Rebuild sanitized Sentry exceptions with the original stacks and causes. Use an adapter stack preview when available. Omit a fabricated reporting stack when the source has no stack. - Keep contexts local to each event. Preserve the optional DSN gate and existing error-code/adapter fingerprint. - Document the fields, limits, and omitted data. - Settle leftover chat fixture outbox rows only after assertions and worker shutdown, so subsequent tests cannot claim earlier cases’ pending actions or provider I/O. This fixes the CI shard contamination exposed during verification; production chat behavior and test timeouts are unchanged. ## Verification - `pnpm -r typecheck` passed. - `pnpm build` passed. - 86 focused diagnostic, Sentry, and startup tests passed after the rebase, including the real `@sentry/node@10.71.0` SDK with an in-memory transport. - Tests cover cause chains, HTTP status and request IDs, long provider details, secret redaction, failed secret resolution, cyclic causes, size limits, and context isolation. - Five targeted heartbeat integration cases passed, including thrown and returned errors containing an opaque environment-bound credential. The earlier full heartbeat integration run also passed its integration cases. - The focused suite also passed with an unknown inherited environment value set to `1`; reporting tests use a controlled environment and separately verify short-secret redaction. - Before the fixture cleanup, the full chat shard reproduced the CI Telegram timeout at `pending recovery before restart` (354 passed, 1 failed). After the cleanup, the same shard passed all 355 tests. No timeout or production behavior changed. - Latest-head CI (`0ae9e70df321d66dda025c3a7ba4787e169e0a4f`) passed typecheck, build, all 12 server shards, all 3 chat shards, workspace and serialized suites, browser shards, Runner checks, canary validation, the real Sentry SDK contract, and security checks. Required `ci / verify` and `ci / e2e` passed. - The branch has been rebased onto `master` after #14575 merged. All 55 applicable checks passed on this head (2 unrelated Storybook checks skipped). Greptile re-reviewed the current 11-file diff at 5/5 with no findings or unresolved review threads. - Local monolithic full-suite attempts were interrupted to apply fixes; full-suite success is not claimed from those runs. ## Risks - Error messages and stacks can contain credentials. The reporter uses existing redactors and the run's encrypted secret registry, reads only known fields, and skips capture if registered-secret resolution fails. - Unknown environment values remain private by default. Unrecognized short values can mask benign matches; known public settings are explicitly allowed. - Diagnostic text is bounded and can be truncated. Truncation is explicit. Data discarded upstream cannot be recovered. - These additional fields go to the operator's configured Sentry endpoint. Arbitrary request/response objects, headers, configuration, prompts, and stdout/stderr are not copied. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model ID, context window, and configured reasoning level are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e5bf9d49a5 |
fix(sentry): retain recorded process exit details (#14575)
## Thinking Path > - Paperclip manages agents and records their task runs. > - Operators can enable Sentry reports for terminal run failures. > - A failed adapter can leave only the generic message “Adapter failed.” > - The run already records a process exit code and signal, but the report omits them. > - This pull request carries those two values through a strict capture boundary. > - Operators can distinguish a nonzero exit from signal termination when the message is generic. ## Linked Issues or Issue Description **What happened?** A failed run can store `exitCode: 1` or `signal: "SIGTERM"` while its Sentry event contains only `adapter_failed` and “Adapter failed.” The existing reporter drops both recorded fields. This occurs on the current master reporting path. **Expected behavior** The opt-in report preserves bounded process exit evidence without exporting adapter output or changing run behavior. **Steps to reproduce** Enable the backend Sentry DSN and report a failed run whose message is “Adapter failed” and whose stored signal is `SIGTERM`. Before this change, the event has no signal field. After this change, `run_failure.signal` is `SIGTERM` and the existing fingerprint stays the same. Related: #12105 and #8222 describe missing adapter/HTTP failure details. #13152 changes terminal-result cleanup classification, and #12886 adds process-failure classification and runtime URL checks. None forwards these stored fields through the Sentry reporter. This change does not resolve those broader issues. ## What Changed - Forward the stored exit code and signal from the terminal run reporter. - Accept only signed 32-bit integer exit codes; use `null` for missing or malformed values. - Accept only the reporting host's Node signal constants; use `null` for missing values and `unknown` for unrecognized values. - Keep the added fields in event-local context, outside tags and fingerprints. - Cover database-backed reporting, malformed input, privacy, and isolation through the real Sentry SDK. - Document the fields and their limits. - Give the dedicated Sentry job the normal PR dependency-resolution fallback, with lifecycle scripts disabled on every install and the required real-SDK test retained. ## Verification - Before the change: 16 report-shape/exit-field assertions failed in the focused capture suite. - After the change: 91 focused Sentry, DSN, and database-backed reporting tests passed. The real SDK uses an in-memory transport. - Final real-SDK test also passed with malformed metadata; it verifies that arbitrary signal text is absent from captured events and unrelated errors inherit no run context. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm test:run`: exited nonzero after 718 general-server files: 14,085 tests passed, 14 failed, 86 skipped. Thirteen skill-service/cache failures reproduce on the unchanged base commit on this macOS host. One comment-wake test timed out; the complete 29-test suite passes on the unchanged base and in final-head Linux CI. Local isolated rechecks skipped because embedded PostgreSQL could not start; these are not counted as passes. The remaining local workspace/serialized lanes did not run after the failing first lane; all CI lanes passed. - Initial Sentry CI failed before tests with `ERR_PNPM_LOCKFILE_CONFIG_MISMATCH`. Its install step lacked the normal PR fallback. The repaired real-SDK job passed. Security review then requested disabling lifecycle scripts for resolved dependencies; every install now uses `--ignore-scripts`. A fresh isolated checkout passed the exact script-disabled fallback and real-SDK contract. Final-head real-SDK CI and the security scan passed. - GitHub CI: all 54 checks passed on `37e0a836e50660f7753d367bcf5a4959eaf89b90`, including required `ci / verify` and `ci / e2e`; two unrelated checks skipped. - Greptile: 5/5 on that commit. No unresolved review comments. - Merge status: conflict-free; required CODEOWNER approval for the workflow change is still pending. ## Risks Low risk: this only adds two validated fields to existing opt-in error reports. It changes no database schema, run status, retry, fingerprint, or suppression rule. Process output and adapter result payloads remain excluded. A recorded signal does not identify its sender or prove an out-of-memory kill. Missing exit evidence stays unknown; this change does not establish the cause of a historical generic adapter failure. ## Model Used OpenAI GPT-6 (Codex), with reasoning, terminal tools, and code execution. The context window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes:` / `Closes` / `Refs` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal ticket id or instance-derived details - [x] I have run focused tests locally and they pass; broader validation is recorded above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
92d4868e79 |
fix(server): isolate run errors and redact runtime capability headers (#13826)
## Thinking Path
> - Paperclip manages AI agents and their work.
> - Operators use Sentry to investigate failed runs and server errors.
> - Run reports attach a task ID, run ID, error code, and adapter
fingerprint.
> - The server skips Sentry's OpenTelemetry setup to preserve its
separate tracing and privacy settings.
> - Without an async context manager, a scope mutation can attach old
run data to later errors.
> - The HTTP logger also retains a runtime credential capability header.
> - This change isolates run metadata and redacts that header so
diagnostics identify failures without leaking credentials.
## Linked Issues or Issue Description
Refs #13446 and #13719.
**What happened?**
After a terminal run failure, an unrelated server exception can inherit
that run's tags, context, and fingerprint. Sentry then groups a database
error with an earlier adapter failure. The real SDK reproduces this with
the application's `skipOpenTelemetrySetup: true` setting. HTTP request
logs also retain the `x-paperclip-github-capability` header, which must
be treated as a credential.
**Expected behavior**
Run metadata belongs to the terminal run event. Later exceptions must
not inherit it. Every genuine error must still be captured. Runtime
capability headers must be redacted on success and failure logs.
**Steps to reproduce**
1. Initialize the optional Sentry SDK with the application's options and
an in-memory transport.
2. Capture a terminal run failure.
3. Capture an unrelated exception.
4. Inspect the second event. Before this fix, it contains the first
run's identity and fingerprint.
5. Send a request with a fixture runtime GitHub capability header.
Before this fix, HTTP logs retain the fixture value.
## What Changed
- Pass tags, context, and fingerprint directly to `captureException`
instead of mutating the ambient scope.
- Preserve the existing run fields, grouping keys, ordinary exception
capture, and privacy settings.
- Test two run identities interleaved with unrelated exceptions against
the real optional SDK.
- Update the capture contract tests and document event-local run
metadata.
- Redact the runtime GitHub capability header through the existing HTTP
logger policy. Test successful, denied, and failed requests.
- Add a dedicated GitHub-hosted CI check that installs the exact
optional SDK version declared in `server/package.json`. It fails if the
real-SDK regression would be skipped. The SDK stays outside the
workspace and production dependency graph.
## Verification
- The real-SDK regression failed before the fix because the unrelated
event contained `contexts.run_failure`.
- Five focused suites passed: 123 tests, including all optional SDK
tests. Suites: `run-failure-sentry-real-sdk.test.ts`,
`run-failure-sentry.test.ts`, `sentry.test.ts`,
`run-failure-report.test.ts`, and `http-log-redaction.test.ts`. A custom
in-memory transport prevented outbound Sentry delivery.
- All three new header-redaction cases failed before the policy fix and
passed afterward.
- The dedicated CI command passed locally with
`PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1` and the audited SDK available
through `NODE_PATH`.
- Server TypeScript check passed with a scratch configuration that
resolves this checkout's workspace packages. The existing dependency
links point to another checkout.
- `node scripts/check-module-boundaries.mjs` and `git diff --check`
passed.
- Gitleaks and a separate private-data scan passed before push.
- Full local workspace typecheck, test, and build were not run. The
machine has less than 2 GiB free and those commands include Rust builds.
Full PR CI must pass before merge.
- The dedicated real-SDK GitHub check passed with 1 test executed and no
skips: https://github.com/paperclipai/paperclip/actions/runs/35774449002
- Greptile reviewed
|
||
|
|
600e552d7b |
fix: attribute Sentry errors to the loaded source release (#13719)
Attribute optional server and browser Sentry events to their source build. Use validated build commits for Docker and source/npm artifacts, preserve explicit server release overrides, and keep cached browser bundles tied to the commit they loaded. Verify 127 focused tests, server/UI typechecks, Docker and source build stamps, all 53 CI checks, and Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
c276d3fdc3 |
feat(observability): report terminal run failures to Sentry (#13446)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server records the state of each agent run. > - Terminal run failures need clear error tracking for operators. > - The server did not report terminal `failed` or `timed_out` transitions to Sentry. > - This pull request reports each genuine terminal failure transition with safe diagnostic data. > - The benefit is faster diagnosis without changing run control flow or exposing credentials. ## Linked Issues or Issue Description **What happened?** The server wrote terminal run failures but did not report them to Sentry. Operators could not see these failures in error tracking. **Expected behavior** The server should report each genuine transition to `failed` or `timed_out` to Sentry. **Steps to reproduce** 1. Run an agent task that reaches a terminal failure state. 2. Inspect the Sentry events for the server. 3. Observe that the terminal run failure has no matching Sentry event. **Paperclip version or commit** The change targets the current `master` branch. No public GitHub issue or pull request covers this change. ## What Changed - Add `captureRunFailure()` as a fail-open Sentry entry point. - Add `reportRunFailure()` to filter status, resolve the adapter, redact text, and report the failure. - Call `reportRunFailure()` beside each of the eight terminal status writers. - Report six diagnostic values: the instance host, task identifier, run identifier, error message, error code, and agent adapter. - Group events by error code and agent adapter while keeping the redacted message in the event. - Report only genuine transitions and avoid duplicate finalization events. - Keep Sentry failures outside run control flow. ## Verification - `pnpm vitest run server/src/services/__tests__/run-failure-report.test.ts server/src/__tests__/run-failure-sentry.test.ts server/src/__tests__/native-session-resumption.test.ts` - `pnpm vitest run server/src/services/execution-control-reconciliation.test.ts` - `pnpm --filter @paperclipai/server typecheck` - The full continuous-integration suite must run on this pull request. ## Risks - The report path can add diagnostic events when Sentry is configured. - The report path returns without action when Sentry is not configured. - Redaction runs before length limits and before the event leaves the process. - The change has no migration and no schema change. ## Model Used OpenAI Codex, GPT-5, tool use and code review support. The exact context window and reasoning configuration are not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ed3559dd21 |
feat(server): split the Sentry DSN into front-end and backend variables (#12678)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip reports server and browser errors through optional Sentry monitoring > - One environment variable sends both error types to one Sentry project > - Operators need separate control for browser and server error data > - This pull request adds specific variables and keeps the existing variable as a fallback > - The benefit is separate monitoring without breaking current deployments ## Linked Issues or Issue Description **What existing behavior does this improve?** The Sentry configuration for server and browser monitoring uses one environment variable. **Subsystem affected** Cross-cutting (multiple of the above) **Current behavior** `SENTRY_DSN` supplies the server and browser clients. Both clients therefore report to the same Sentry project. **Proposed behavior** `SENTRY_DSN_FRONTEND` supplies the browser client. `SENTRY_DSN_BACKEND` supplies the server process. `SENTRY_DSN` remains a fallback for either component. **Reason and benefit** Operators can send browser and server errors to separate Sentry projects. Operators can also activate only one component. **Breaking changes** None. Existing deployments can continue to use `SENTRY_DSN`. ## What Changed - Add `resolveSentryDsns(env)` and use it in the server and browser configuration paths. - Add precedence, empty-string, fallback, and route tests. - Update the README, observability guide, and stale code comments. - Log one warning when the server uses the legacy fallback without exposing a DSN value. ## Verification - `pnpm vitest run --project server sentry-dsn` — 8 tests pass. - `pnpm vitest run --project server auth-routes` — 21 tests pass. - The earlier run of the three targeted suites passed 40 tests. - `tsc --noEmit` passes for the files in this diff. - All required GitHub Actions checks pass, including the full continuous-integration suite. ## Risks The main risk is an incorrect environment variable precedence rule. Unit tests cover specific values, empty strings, and legacy fallback behavior. The existing `SENTRY_DSN` path remains compatible. ## Model Used OpenAI Codex — GPT-5, current runtime, tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1de105c475 |
fix(observability): pin the Sentry browser SDK and gate the optional Sentry server peer on the exact version (#12270)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip uses separate server and browser packages for runtime
services and the board.
> - Sentry integrations need an exact SDK version and safe optional
loading.
> - A version range can select an SDK that the privacy tests did not
audit.
> - Missing peer metadata does not describe the optional server SDK
contract.
> - This pull request pins the browser SDK and gates the optional server
SDK on its exact version.
> - The benefit is a clear SDK contract with fail-open startup behavior.
## Linked Issues or Issue Description
**What happened?**
The browser package used the range ^10.71.0, so a lockfile refresh could
select a newer SDK. The server loaded @sentry/node dynamically but did
not declare its optional peer contract.
**Expected behavior**
The browser package must use the audited 10.71.0 version. The server
must load @sentry/node only when the installed peer matches 10.71.0. The
server must start when the optional peer is absent.
**Steps to reproduce**
1. Install the project dependencies.
2. Inspect the browser Sentry version and the server package metadata.
3. Start the server without installing @sentry/node.
4. Confirm that the server starts and that the dynamic Sentry bootstrap
does not load an unsupported peer version.
**Paperclip version or commit**
|
||
|
|
8f1e3cfe24 |
feat(observability): add opt-in Sentry error monitoring for the server and the browser (#12190)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server and the browser need clear error reports when an operator enables external monitoring. > - Paperclip already uses an opt-in OpenTelemetry pattern for server traces. > - Sentry can provide error reports for both runtime paths when the operator sets one data source name. > - This pull request adds one opt-in Sentry gate for the server and the browser. > - The benefit is faster diagnosis while the default setup sends no Sentry data. ## Linked Issues or Issue Description **What is improved?** Paperclip gains optional error monitoring for server and browser failures. **Subsystem affected** Cross-cutting (server, UI, and shared authentication data). **Current behavior** Paperclip has no built-in Sentry error capture for server failures or browser boundary failures. Operators must inspect local logs and browser tools. **Proposed behavior** When the operator sets `SENTRY_DSN`, the server and authenticated browser use the same Sentry project. When the variable is absent, both paths stay inactive. The server loads Sentry dynamically and fails open when the optional package is absent. **Reason and benefit** Operators can inspect runtime errors in one Sentry project. The default setup remains local and sends no monitoring data. **Breaking changes** None when `SENTRY_DSN` remains unset. Authenticated session responses add the optional `sentryDsn` field. **Additional context** The implementation uses built-in Sentry privacy options. It disables default HTTP context and breadcrumb integrations and keeps `sendDefaultPii` false. ## What Changed - Add an opt-in server Sentry gate with dynamic package loading and fail-open behavior. - Add the Sentry data source name to the authenticated session response. - Add an authenticated browser Sentry gate and React error boundary capture. - Add tests for server, browser, route, and application error paths. - Document activation, installation, privacy settings, capture behavior, and operator controls. ## Verification - Run `npx vitest run server/src/__tests__/sentry.test.ts`. - Run `npx vitest run ui/src/lib/sentry.test.ts`. - Run `npx vitest run server/src/__tests__/auth-routes.test.ts server/src/__tests__/shutdown.test.ts`. - Confirm that the full continuous integration suite passes on this pull request. - Leave `SENTRY_DSN` unset and confirm that the server and browser gates stay inactive. - Set `SENTRY_DSN` and install the optional Sentry packages before a manual capture check. ## Risks The operator controls the Sentry project and accepts the data risk when the operator enables the feature. Error objects can contain messages, stacks, or cause chains with private values. The default configuration sends no data because the feature stays off without `SENTRY_DSN`. A missing optional server package does not stop server boot. ## Model Used OpenAI Codex, GPT-5, with tool use, repository inspection, GitHub CLI operations, and code review support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |