Commit Graph
30 Commits
Author SHA1 Message Date
Devin FoleyandPaperclip 4b9a6000f7 Add bounded evidence for directory lock timeouts (#14787)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent files use directory locks during collection and cleanup.
> - A lock timeout can fail finalization after the model turn completes.
> - The timeout currently identifies no owner state or waiting
operation.
> - This pull request adds bounded evidence to the existing run failure
report.
> - Operators can distinguish a known local holder from a possible old
lock without changing lock safety.

## Linked Issues or Issue Description

**What happened?**

A directory lock timeout does not distinguish active local work from an
owner record left by an earlier process. The stored execution stage can
also precede the cleanup operation that failed.

**Expected behavior**

The failure report should identify the waiting operation and expose
bounded ownership clues. It must preserve the timeout and keep unknown
ownership protected.

**Steps to reproduce**

Hold a directory merge lock while a second caller reaches its
acquisition deadline. The regression tests exercise a live holder and an
older owner record with a live PID.

Related: #9667 proposes stale-lock recovery under a single-server
assumption. This change only adds evidence and does not adopt that
assumption. #14575 and #14665 add other run failure diagnostics.

## What Changed

- Record lock owner state, capped age and wait duration, same-process
and process-age comparisons, and whether this module holds the lock.
- Label agent-directory release, collection, checkpoint, and warm
handoff timeouts with a fixed operation code.
- Validate each field before the existing event-local Sentry report
accepts it. Exclude owner records, PIDs, paths, and absolute timestamps.
- Limit the extra diagnostic owner read to 100 ms with best-effort
abort; malformed JSON is `invalid` and unreadable owner records remain
`unknown`.
- Document the diagnostic limits and verify that contenders never
reclaim protected locks.

## Verification

- Focused lock, diagnostic, real Sentry SDK, and database-backed
agent-directory tests: 126 passed, including stalled-read and
malformed/missing/unreadable-owner regression coverage.
- Final revision `0691613dcc`: all 54 reported checks successful, with
two intentionally skipped Storybook checks. Greptile: 5/5, zero
unresolved review threads; no merge conflicts.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm test:run`: complete suite coverage ran with the existing
repository shard flags: four general-server shards, four serialized
shards, two general-workspaces-a shards, and general-workspaces-b. The
full run is not green because of the base failures below.
- The broad run found 13 failures in the unchanged macOS skill-cache
tests. All 13 reproduce on the clean base revision. Open PR #14290
covers that existing failure.
- Two unchanged CLI archive tests hit their five-second limits during
the broad run; all 17 tests in that file pass on recheck. A CLI auth
socket error also cleared on recheck (19 tests), and its full serialized
shard passed on rerun.

## Risks

This is a diagnostic change, not a stale-lock fix. Owner observations
can race with release. Wall-clock shifts can affect the age comparison.
A local-holder flag covers only this module instance. None of these
fields authorizes reclamation or proves a file save. Lock acquisition,
release, retries, task status, and recovery guards retain their current
behavior. No schema change or deployment action is required.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The exact model build and context window were not exposed to this agent.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass (focused checks pass;
existing base failures are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 18:35:12 -07:00
Devin FoleyandPaperclip 0ea6b10967 Record ACP activity and workspace restore failure evidence (#14665)
## Thinking Path

Paperclip records terminal run failures for operators. A timeout's own
log and cleanup output update the run's last-output timestamp, so that
timestamp can make a long-silent provider look active. Snapshot runtime
activity before finalization and include the saved workspace restore
classification to make the next failure actionable without copying tool
payloads.

## Linked Issues or Issue Description

**What existing behavior does this improve?**
Terminal run diagnostics in the existing opt-in Sentry integration.

**Current behavior**
Reports cannot distinguish runtime events from finalization logging and
omit the already-persisted workspace restore code. A later successful
run also does not establish that earlier workspace files were restored.

**Proposed behavior**
Record runtime-event age/count and pending-tool inventory at
finalization, before status reads and cleanup. Forward only finite
counts, a completeness boolean, and known restore codes through the
existing reporter.

**Reason and benefit**
Operators can distinguish a silent turn with unfinished tools from
recent runtime activity and see restore failures without retrieving
private run output. Neither signal certifies productive work or
successful recovery.

**Breaking changes**
None. Error grouping, execution deadlines, cancellation, recovery
policy, and the Sentry opt-in remain unchanged.

Related diagnostic work: #14573, #14575, #14639.

## What Changed

- Snapshot ACP activity before success/failure finalization, including
thrown relay failures.
- Forward bounded numeric/boolean evidence and shared workspace restore
codes; exclude commands, tool IDs, paths, and arbitrary result data.
- Document limitations and test silence, empty streams, timeout, cleanup
delay, incomplete tool inventory, and privacy.

## Verification

- `pnpm -r typecheck` passed after the final implementation.
- Changed suites: 252 tests passed; all 29 database reporter tests
subsequently passed after restoring the embedded-Postgres package
library symlinks. The migration test also passed (30 database cases
total).
- `pnpm build` passed during implementation. Final head
`fe78dba6f592b1abccac7cdbf341bd2e0b0d30cb` passed all 53 CI checks,
including complete test coverage, typecheck, build, and browser/runner
gates; two unrelated checks intentionally skipped.
- Full local test attempts initially hit missing embedded-Postgres
library symlinks; the dependency setup was repaired and database tests
passed. Duplicate local full-suite runs were stopped after full CI
completed. This PR does not claim a completed full local suite.
- Review regression: completed, failed, and cancelled tools are excluded
from the pending count; focused activity/timeout tests and adapter-utils
typecheck passed.
- Reviewed the diff for secrets, customer data, and internal references.

## Risks

Low risk, diagnostic-only. The existing tool inventory is incomplete for
some runtime events, so the report carries its completeness flag. Event
age is measured at finalization and does not prove useful work or
identify the underlying provider failure. No schema changes or new
capture gate.

## Model Used

OpenAI GPT-6 via Codex, with repository inspection, code execution, and
tests. Exact model build identifier is not exposed by this session.

## Checklist

- [x] Thinking path and model are specified
- [x] Checked ROADMAP.md; this is a maintenance correction, not planned
feature work
- [x] Searched for duplicate and related PRs
- [x] Described the issue using the enhancement template
- [x] No internal issue references, customer data, or private instance
links
- [x] Descriptive branch name
- [x] Focused regression tests pass
- [x] Added tests and updated documentation
- [x] Risks documented
- [x] Required validation and CI gates are green (full suite validated
in CI; local scope documented above)
- [x] Greptile is 5/5 with no unresolved findings
- [x] I will address review comments before requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 14:27:53 -07:00
DottaandPaperclip 0e5830887b perf: reveal task content sooner and parallelize issue reads (#14727)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - A task page must show saved replies quickly so a person can read the
work.
> - The title could appear while the conversation waited for unrelated
metadata and transcripts.
> - Thread requests also waited for enriched task details, while the
server read several independent fields in sequence.
> - This change shows saved content as soon as it is ready, starts
thread reads earlier, and runs independent server reads together.
> - Native event history takes priority over legacy log fallback, and
mentioned tasks load on intent.
> - Content-free timing spans make the remaining server delays visible
without recording task content.

## Linked Issues or Issue Description

**What happened?**

Task titles and properties appeared quickly, but saved conversation
content stayed hidden for several more seconds while metadata and run
transcripts loaded.

**Expected behavior**

Saved replies and the task description should be readable without
waiting for supporting history. Returning to a cached task should show
content within a frame or two.

**Steps to reproduce**

1. Open a task with saved comments and completed runs.
2. Delay the task activity and runs responses by five seconds in the
browser.
3. Observe whether saved content remains hidden until those responses
finish.
4. Navigate away and return to the task to check cached navigation.

**Paperclip version or commit**

The change was developed from `1b48e73e0` and rebased onto `44736c9c7`.

**Deployment mode**

Built from source, tested in an authenticated staging deployment and
with local response replay.

Related: #14667 overlaps the transcript reveal behavior and adds
separate retry UX. This PR also changes navigation prefetch, parent
metadata gates, native log fallback, server read scheduling, and timing
spans. #12647 proposes a separate SQL predicate optimization in the runs
service. #13597 and #13095 are earlier loading fixes.

## What Changed

- Start activity and runs when the task page mounts, alongside task
details and comments, using the route reference for shared query keys.
Hover/focus prefetch does not start full history reads.
- Reveal saved comments and descriptions while metadata and transcripts
load. Preserve strict waits for linked-comment navigation and tasks with
only runtime content.
- For settled native runs, fetch legacy logs only when event history is
empty or fails. Preserve live-log subscriptions for queued and running
native runs. Fetch mentioned-task details on hover or focus.
- Run independent issue-detail enrichment and run metadata reads in
parallel while preserving recovery dependencies.
- Add `Server-Timing` phases and opt-in OpenTelemetry spans, plus
regression tests and observability documentation.

## Verification

- **313 tests passed** across the initial seven focused component,
cache, timing, and scroll suites. After review fixes, **313 tests
passed** across five task-page, cache, prefetch, and live-transcript
suites (`pnpm exec vitest run` with `--maxWorkers=1`; these sets
overlap).
- `pnpm -r typecheck` and `pnpm build` passed locally before the final
UI-only review fixes. UI typecheck/build and `pnpm check:token-gates`
passed after those fixes. The final commit also passes full typecheck
and build in CI.
- The full local Vitest run was attempted. Several unrelated embedded
PostgreSQL fixtures failed to start, and parallel test workers hit
timeouts. Focused reruns passed. A later local full-suite rerun was
stopped after the complete CI suite passed; it is not claimed as a local
full-suite pass.
- Final commit `d1e147145`: **54 checks passed, 2 skipped**, including
all server/workspace test shards, all eight browser E2E shards, full
build/typecheck, Runner checks, release registry, and canary
clean-install verification. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36734642034).
Greptile **5/5** after two reviews; both findings fixed and all review
threads resolved. No merge conflicts.
- Live browser tests on an existing task with two saved replies: median
full reload to visible content fell from **1.92 s** (3 samples) to
**1.40 s** (5 samples). Cached return fell from **421 ms** (1 sample) to
**29 ms** (3 samples). These are observed samples, not a performance
guarantee.
- With activity and runs delayed by five seconds, saved content appeared
in **1.38 s** on desktop and **1.33 s** on mobile. The inspected comment
did not move when metadata arrived. Verified history expansion, task
properties, pending-input navigation, dashboard return, and mobile
layout.
- A separate local replay with fixed responses reduced visible-content
time from **5.12 s** to **2.15 s**. This isolates frontend behavior and
is not a live-server benchmark.

## Risks

Progressive history can change the thread after first paint. Existing
anchor behavior is retained and covered by tests and delayed-response
browser checks. Query aliases must stay aligned for invalidation.
Parallel reads can increase short bursts of database work; dependent
recovery operations remain ordered. Full reloads still depend on network
and task-detail latency.

No schema or authorization change. OpenTelemetry remains disabled
without an operator endpoint. The added spans use a closed set of phase
names and carry no task IDs, task content, or exception text.

I checked `ROADMAP.md`; this is a performance fix within the existing
task page.

## Model Used

OpenAI GPT-6 in Codex. The runtime does not expose a more specific model
ID or context-window size. The agent used reasoning, code editing,
terminal tools, and Chrome performance profiling.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 10:26:28 -05:00
Devin FoleyandPaperclip 0b12ca9532 fix(server): retain context for unconfirmed adapter stops (#14639)
## Thinking Path

> - Paperclip must keep ownership of work until termination is verified.
> - A Stop request waits for the adapter and its cleanup to settle.
> - After 60 seconds, an unconfirmed Stop raises an error.
> - The error currently lacks run, adapter, and runtime context.
> - This makes it difficult to investigate which stop path is stuck.
> - This change adds bounded diagnostics while preserving termination
checks.

## Linked Issues or Issue Description

**What happened?**

An adapter can remain unsettled after its Stop request. The resulting
error says termination is unverified but does not identify the adapter
or run in error monitoring. The optional Sentry setup does not capture
request context, so the endpoint alone cannot fill the gap.

**Expected behavior**

Keep the Stop unconfirmed and preserve its live execution owner. When
optional Sentry is enabled, attach enough bounded context to investigate
the affected run.

**Steps to reproduce**

1. Register an adapter execution control and abort its controller.
2. Leave its settlement promise pending.
3. Wait for the configured Stop timeout.
4. Observe that Stop still fails, but the event now includes the run
UUID, built-in adapter, native/legacy runtime, timeout duration, and
abort-requested flag.

**Paperclip version or commit**

Master commit `17780751551b3bc1c2521f7694026c34534c46c9`; reproduced
with fake timers and mocked optional error monitoring.

**Deployment mode**

Server execution control, including local and hosted runs. Reporting
remains opt-in.

Searched related Stop PRs. #14523 and #14244 address Hermes cancellation
contracts; this change only adds diagnostics to the shared
unconfirmed-stop timeout.

## What Changed

- Use a typed timeout error with the existing message, name, and timer
stack.
- Pass run/adapter/runtime identity from the cancellation owner.
- Add an event-local, allowlisted Sentry context without changing the
default fingerprint.
- Rebuild the reported exception so arbitrary provider fields cannot be
serialized.
- Test timeout ownership, delayed settlement, privacy boundaries, and
absence of context on unrelated events.
- Document the additional opt-in fields.

## Verification

- `pnpm exec vitest run
server/src/services/adapter-execution-control.test.ts
server/src/__tests__/sentry.test.ts`: 38 passed, five real-SDK checks
skipped because the optional package is not installed.
- `pnpm -r typecheck` passed; server typecheck passed again after the
SDK test addition.
- With audited optional `@sentry/node@10.71.0` installed only in local
test dependencies, `PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1 pnpm exec vitest
run server/src/__tests__/run-failure-sentry-real-sdk.test.ts
server/src/services/adapter-execution-control.test.ts
server/src/__tests__/sentry.test.ts`: all 45 tests passed. The real SDK
uses an in-memory transport; no Sentry requests are sent.
- The broad local `pnpm test:run` command did not complete in the
available verification window and was stopped; no full local-suite pass
is claimed. `pnpm build` passed. All sharded GitHub CI checks passed on
the final PR head.
- Tests use fake timers and a mocked Sentry package; no provider or
monitoring requests.

## Risks

This is diagnostic coverage, not a claim that the underlying stop delay
is fixed. Unknown adapter/runtime values become `unknown`; malformed run
identifiers become `null`. No stop reason, prompt, output, provider
response, credentials, or arbitrary error properties are sent. Timeout,
cancellation acknowledgement, live-owner retention, and retry behavior
remain unchanged. No schema changes.

## Model Used

OpenAI Codex (GPT-6), with reasoning, repository inspection, and command
execution. The session does not expose a more specific model revision or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run the targeted tests locally and they pass; full checks
are in progress
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 16:38:02 -07:00
DottaandPaperclip b3eb03fcba fix(sentry): preserve run failure stacks and diagnostic context (#14585)
Builds on merged #14575 and targets `master`.

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Operators use optional Sentry reports to diagnose failed agent runs.
> - The shared reporter currently converts each saved error message into
a new exception.
> - This loses the original stack and cause chain, and omits available
adapter and provider diagnostics.
> - This pull request selects, redacts, and bounds useful diagnostics
before it sends them to Sentry.
> - Operators can locate the failing operation and correlate upstream
failures without copying arbitrary run data.

## Linked Issues or Issue Description

Refs #14573. That change preserves ACP provider failure details in the
run record.
Refs #14575. That merged PR adds recorded exit codes and signals. This
PR covers stacks, causes, and structured diagnostics without duplicating
those fields.

A thrown setup or adapter error currently appears in Sentry with the
reporter's stack. A saved provider failure can contain useful details
that never reach the Sentry event. Both cases use the same shared
reporting path.

## What Changed

- Pass caught setup and execution exceptions, and structured adapter
error metadata, through successful terminal status transitions.
- Snapshot a fixed set of execution, adapter, provider, and exception
fields. Preserve up to four exceptions in the cause chain.
- Remove registered secret values, declared runtime environment
credentials, unknown inherited environment values and encoded credential
forms, and credential patterns before truncation. Mark truncated fields
and bound provider details and stacks.
- Rebuild sanitized Sentry exceptions with the original stacks and
causes. Use an adapter stack preview when available. Omit a fabricated
reporting stack when the source has no stack.
- Keep contexts local to each event. Preserve the optional DSN gate and
existing error-code/adapter fingerprint.
- Document the fields, limits, and omitted data.
- Settle leftover chat fixture outbox rows only after assertions and
worker shutdown, so subsequent tests cannot claim earlier cases’ pending
actions or provider I/O. This fixes the CI shard contamination exposed
during verification; production chat behavior and test timeouts are
unchanged.

## Verification

- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- 86 focused diagnostic, Sentry, and startup tests passed after the
rebase, including the real `@sentry/node@10.71.0` SDK with an in-memory
transport.
- Tests cover cause chains, HTTP status and request IDs, long provider
details, secret redaction, failed secret resolution, cyclic causes, size
limits, and context isolation.
- Five targeted heartbeat integration cases passed, including thrown and
returned errors containing an opaque environment-bound credential. The
earlier full heartbeat integration run also passed its integration
cases.
- The focused suite also passed with an unknown inherited environment
value set to `1`; reporting tests use a controlled environment and
separately verify short-secret redaction.
- Before the fixture cleanup, the full chat shard reproduced the CI
Telegram timeout at `pending recovery before restart` (354 passed, 1
failed). After the cleanup, the same shard passed all 355 tests. No
timeout or production behavior changed.
- Latest-head CI (`0ae9e70df321d66dda025c3a7ba4787e169e0a4f`) passed
typecheck, build, all 12 server shards, all 3 chat shards, workspace and
serialized suites, browser shards, Runner checks, canary validation, the
real Sentry SDK contract, and security checks. Required `ci / verify`
and `ci / e2e` passed.
- The branch has been rebased onto `master` after #14575 merged. All 55
applicable checks passed on this head (2 unrelated Storybook checks
skipped). Greptile re-reviewed the current 11-file diff at 5/5 with no
findings or unresolved review threads.
- Local monolithic full-suite attempts were interrupted to apply fixes;
full-suite success is not claimed from those runs.

## Risks

- Error messages and stacks can contain credentials. The reporter uses
existing redactors and the run's encrypted secret registry, reads only
known fields, and skips capture if registered-secret resolution fails.
- Unknown environment values remain private by default. Unrecognized
short values can mask benign matches; known public settings are
explicitly allowed.
- Diagnostic text is bounded and can be truncated. Truncation is
explicit. Data discarded upstream cannot be recovered.
- These additional fields go to the operator's configured Sentry
endpoint. Arbitrary request/response objects, headers, configuration,
prompts, and stdout/stderr are not copied.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model ID, context window, and configured reasoning level are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 12:54:12 -05:00
Devin FoleyandPaperclip e5bf9d49a5 fix(sentry): retain recorded process exit details (#14575)
## Thinking Path

> - Paperclip manages agents and records their task runs.
> - Operators can enable Sentry reports for terminal run failures.
> - A failed adapter can leave only the generic message “Adapter
failed.”
> - The run already records a process exit code and signal, but the
report omits them.
> - This pull request carries those two values through a strict capture
boundary.
> - Operators can distinguish a nonzero exit from signal termination
when the message is generic.

## Linked Issues or Issue Description

**What happened?**
A failed run can store `exitCode: 1` or `signal: "SIGTERM"` while its
Sentry event contains only `adapter_failed` and “Adapter failed.” The
existing reporter drops both recorded fields. This occurs on the current
master reporting path.

**Expected behavior**
The opt-in report preserves bounded process exit evidence without
exporting adapter output or changing run behavior.

**Steps to reproduce**
Enable the backend Sentry DSN and report a failed run whose message is
“Adapter failed” and whose stored signal is `SIGTERM`. Before this
change, the event has no signal field. After this change,
`run_failure.signal` is `SIGTERM` and the existing fingerprint stays the
same.

Related: #12105 and #8222 describe missing adapter/HTTP failure details.
#13152 changes terminal-result cleanup classification, and #12886 adds
process-failure classification and runtime URL checks. None forwards
these stored fields through the Sentry reporter. This change does not
resolve those broader issues.

## What Changed

- Forward the stored exit code and signal from the terminal run
reporter.
- Accept only signed 32-bit integer exit codes; use `null` for missing
or malformed values.
- Accept only the reporting host's Node signal constants; use `null` for
missing values and `unknown` for unrecognized values.
- Keep the added fields in event-local context, outside tags and
fingerprints.
- Cover database-backed reporting, malformed input, privacy, and
isolation through the real Sentry SDK.
- Document the fields and their limits.
- Give the dedicated Sentry job the normal PR dependency-resolution
fallback, with lifecycle scripts disabled on every install and the
required real-SDK test retained.

## Verification

- Before the change: 16 report-shape/exit-field assertions failed in the
focused capture suite.
- After the change: 91 focused Sentry, DSN, and database-backed
reporting tests passed. The real SDK uses an in-memory transport.
- Final real-SDK test also passed with malformed metadata; it verifies
that arbitrary signal text is absent from captured events and unrelated
errors inherit no run context.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm test:run`: exited nonzero after 718 general-server files: 14,085
tests passed, 14 failed, 86 skipped. Thirteen skill-service/cache
failures reproduce on the unchanged base commit on this macOS host. One
comment-wake test timed out; the complete 29-test suite passes on the
unchanged base and in final-head Linux CI. Local isolated rechecks
skipped because embedded PostgreSQL could not start; these are not
counted as passes. The remaining local workspace/serialized lanes did
not run after the failing first lane; all CI lanes passed.
- Initial Sentry CI failed before tests with
`ERR_PNPM_LOCKFILE_CONFIG_MISMATCH`. Its install step lacked the normal
PR fallback. The repaired real-SDK job passed. Security review then
requested disabling lifecycle scripts for resolved dependencies; every
install now uses `--ignore-scripts`. A fresh isolated checkout passed
the exact script-disabled fallback and real-SDK contract. Final-head
real-SDK CI and the security scan passed.

- GitHub CI: all 54 checks passed on
`37e0a836e50660f7753d367bcf5a4959eaf89b90`, including required `ci /
verify` and `ci / e2e`; two unrelated checks skipped.
- Greptile: 5/5 on that commit. No unresolved review comments.
- Merge status: conflict-free; required CODEOWNER approval for the
workflow change is still pending.

## Risks

Low risk: this only adds two validated fields to existing opt-in error
reports. It changes no database schema, run status, retry, fingerprint,
or suppression rule. Process output and adapter result payloads remain
excluded. A recorded signal does not identify its sender or prove an
out-of-memory kill. Missing exit evidence stays unknown; this change
does not establish the cause of a historical generic adapter failure.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, terminal tools, and code
execution. The context window size is not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes:` / `Closes`
/ `Refs` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
ticket id or instance-derived details
- [x] I have run focused tests locally and they pass; broader validation
is recorded above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 12:29:17 -05:00
Devin FoleyandPaperclip ad1f7e98ea fix: retain diagnostic reasons for native runner failures (#14481)
Retain bounded reasons for runner identity, harness recovery, and provider-pack read failures. Preserve existing ownership and cleanup proofs and compatibility with receipt-gated chat recovery.

Verified with executor, recovery, diagnostic privacy, typecheck, build, and full CI checks.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 16:16:19 -07:00
Devin FoleyandPaperclip faf72cb1d5 fix: preserve company context across browser hot reload (#14482)
## Thinking Path

> - Paperclip uses a shared company context for browser providers and
consumers.
> - Vite can load a new consumer module while an older provider is
mounted.
> - Recreating the context disconnects that consumer from the mounted
provider.
> - The consumer then reports a missing provider even though one is in
its React ancestry.
> - This change preserves the context object across development module
refreshes.
> - Bounded global error diagnostics distinguish development bundles and
otherwise context-free promise rejections.

## Linked Issues or Issue Description

**What happened?**
A refreshed company consumer can throw `useCompany must be used within a
CompanyProvider`. A real Vite and Chromium reproduction confirms that a
retained provider and a refreshed consumer can hold different context
objects. Global promise rejections also lack the bounded document state
already attached to React boundary errors.

**Expected behavior**
A refreshed consumer should read the mounted provider. Error reports
should identify the loaded bundle mode and bounded browser state while
preserving monitoring opt-in, sign-out, and privacy behavior.

**Steps to reproduce**
Run `pnpm test:e2e:browser-context`. The isolated Vite fixture renders
the real CompanyProvider, imports a new timestamped consumer module, and
renders that consumer below the retained provider. The test fails before
the context change and passes after it. The SDK regression invokes its
real unhandled-rejection handler with an undefined reason.

## What Changed

- Keep the React context object in Vite's per-module `hot.data`. Account
values stay in React.
- Add an isolated browser regression with mocked API responses and no
live instance, discovered by the existing Chrome CI shards.
- Add document-state diagnostics to global errors while preserving
earlier boundary snapshots.
- Tag events with development or production bundle mode and the type of
an unhandled rejected value.
- Document the new test command and diagnostic fields.

## Verification

- Company context, browser context, and Sentry suites: 59 passed.
- `pnpm test:e2e:browser-context`: passed in Chromium. The original
context code fails the reproduction.
- Real SDK tests preserve DSN/sign-out behavior and omit request
context, breadcrumbs, and private DOM data.
- `pnpm check:token-gates`: passed.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- Local `pnpm test:run` exited in the general-server phase: 13,819
passed, 86 skipped, 21 failures in unchanged filesystem and host-port
suites. Cache permission and long-path failures also reproduce on the
unmodified base; seven other failures involve local runtime port
ownership. This is not a full local pass.
- Full Linux CI and review passed at a0b5270662 (53 successful checks;
two conditional checks skipped). Three jobs interrupted by a runner
shutdown passed on rerun. Greptile is 5/5 with no unresolved comments.
The real browser regression also passed in the normal Chrome CI shard
(3.2 seconds).
- Code-owner approval remains required because the dedicated test
command changes `package.json`.
- Scanned the diff and PR text for credentials and private data before
pushing.

## Risks

The context cache applies only to development hot reload. It retains the
context object, not account state. Production continues to create an
ordinary React context. The diagnostic hook adds only bounded state and
type values, preserves boundary snapshots, and returns the original
event if enrichment fails. It does not suppress errors or restore raw
breadcrumbs. The global rejection diagnostics do not identify the
promise's originating operation by themselves.

## Model Used

OpenAI GPT-6 through Codex, with repository inspection, code editing,
and command execution. The runtime does not expose a more specific model
revision or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites and
Chromium regression; full local limitations documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 16:15:14 -07:00
Devin FoleyandPaperclip fff410dfe7 fix(ui): retain bounded context for browser render errors (#13904)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its browser error boundaries recover from failed renders and report
the exception when error monitoring is enabled.
> - The boundaries already receive the React component stack, but
discard it before reporting the error.
> - A minified DOM insertion error can therefore lack enough context to
identify the affected component.
> - Browser translation can replace text nodes that React still uses as
insertion anchors.
> - This pull request preserves bounded component names and browser
state for the next failure.
> - Maintainers can locate the failed component without collecting page
content or customer URLs.

## Linked Issues or Issue Description

Related: #13719 attributes browser errors to the loaded release. #13784
supplies the browser environment. This change adds context to
error-boundary reports.

**What happened?**

A browser error boundary reports a DOM `NotFoundError` with only a
minified JavaScript stack. The React component trace is logged to the
console, which production builds remove. Browser translation is a
plausible cause, but the report cannot identify the component or confirm
the translation marker.

**Expected behavior**

An error-boundary report includes a bounded component trace and limited
browser state. It excludes component props, text, HTML, element IDs,
arbitrary CSS classes, page URLs, and query strings.

**Steps to reproduce**

1. Render a component with conditional content before a text node.
2. Replace that text node with a translation element outside React.
3. Enable the preceding conditional content. React tries to insert
before the detached text node and throws `NotFoundError`.
4. The boundary shows its recovery UI, but the old report loses the
component trace.

**Paperclip version or commit**

Based on `6681c71b40`. The regression test reproduces the DOM mutation
with the installed React version.

**Deployment mode**

Built browser UI with optional Sentry monitoring enabled for the
signed-in session.

## What Changed

- Pass the component stack and boundary kind from both error boundaries.
- Keep at most 40 component names from at most 16 KiB of stack input.
Drop locations and unrecognized lines.
- Snapshot document readiness, visibility, and the browser translation
root-class marker before the asynchronous reporting queue runs.
- Attach diagnostics to that event only. Preserve the existing
monitoring gate, sign-out behavior, and original exception if
diagnostics fail.
- Preserve function names in production bundles. Test the actual Vite
production pipeline.
- Document the fields, privacy limits, and translation-marker
limitations.

## Verification

- Focused diagnostics, boundary, real Sentry SDK, and production-build
tests: 49 passed.
- The translation DOM-mutation test reproduces `NotFoundError`, verifies
the failed component trace, and keeps the recovery UI usable.
- The real SDK test checks emitted events, private fixture exclusion,
event isolation, and no capture after sign-out. Its transport stays
in-process.
- `pnpm check:token-gates`: passed.
- [Greptile
review](https://github.com/paperclipai/paperclip/pull/13904#issuecomment-5804220318):
5/5 on `92a2a23d9c`, with no review threads.
- `pnpm -r typecheck` and `pnpm build`: passed.
- Full CI unit and integration test matrix: passed, including all
server, chat, workspace, runner, and serialized suites.
- Local `pnpm test:run` was started, then stopped after the equivalent
full CI matrix passed. The unsharded local run was not completed and is
not claimed as a pass.
- `pnpm exec playwright test --config tests/e2e/playwright.config.ts
tests/e2e/project-repositories.spec.ts --repeat-each=3 --trace=on
--reporter=line`: six tests passed against the production build.
- Initial CI browser shard 8 timed out after the repository form
replaced an enabled Save button with a disabled one before the click.
The retained page snapshot shows the original repository selection. No
error-boundary fallback appeared. [CI attempt
2](https://github.com/paperclipai/paperclip/actions/runs/35929990870/attempts/2)
passed the failed jobs on the same commit. The full PR check matrix is
green. This confirms an intermittent failure, but does not establish the
cause of the first failure.
- Full browser builds with and without name preservation passed. The
initial JavaScript chunk grows from 1,571,919 to 1,668,453 gzip bytes
(+6.1%). Total JavaScript across all chunks grows by 219,610 gzip bytes
(+5.3%).

## Risks

- This adds diagnostics. It reproduces a translation failure mode but
does not identify or repair the specific application component from a
past report.
- The root-class marker is a hint. Other translation tools may omit it,
and its presence does not prove causation.
- Name preservation increases bundle size as measured above. No source
maps are published by this change.
- The error is still reported. DOM operations, browser translation, and
recovery behavior are unchanged.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code editing, terminal
tools, and test execution. The exact serving model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 16:23:42 -07:00
DottaandPaperclip 7944ed3d97 fix(runner): preserve hire runtime safety and first-activity timing (#13852)
## Thinking Path

> - Paperclip is the open source control plane for companies of AI
agents.
> - Native runner agents need governed tools, durable runtime state, and
useful execution evidence.
> - A first activity trace waited 53.467 seconds even though tool
activity took 6.274 seconds; provider input arrived before the server
API call executed.
> - Native agents also need a safe way to hire teammates without asking
the model to rebuild runtime configuration.
> - This pull request separates the observed ACP input-stream window
from the actual server `tool.execute` span and adds a server-owned
native hire contract.
> - The benefit is clearer latency evidence and safer native teammates
with existing approval, auth, and company boundaries preserved.

## Linked Issues or Issue Description

Related Daytona provenance work is in
[#13814](https://github.com/paperclipai/paperclip/pull/13814). No
duplicate public PR was found for this combined timing and native-hire
change.

**What existing behavior does this improve?**

Native runner agents can use governed tools and request hires. The
server did not expose a safe native hire operation that reused the
caller's validated runtime settings. First-activity traces also mixed
provider input timing with server tool execution timing.

**Current behavior**

A native hire must construct a separate runner configuration. Full
configuration copying could expose paths, instructions, secrets, or
sessions. Timing evidence could make a provider or MCP identity join
appear proven when the trace did not contain that join.

**Proposed behavior**

The native `hire_agent` operation accepts identity and persona inputs.
The server sends `adapterType: "paperclip_runner"` with
`inheritRuntimeFrom: "caller"`, then copies only validated provider,
model, permission, lifecycle, and bounded execution settings. It
inherits and validates the default environment, derives the managed AI
binding through existing normalization, preserves approval and
permissions, and creates fresh child instructions. Caller secrets,
paths, prompts, and sessions are excluded.

Provider events now include the optional boolean `inputUpdated`, with
Rust forwarding support. Timing evidence separately records the ACP
input-stream window and the actual server `tool.execute` activity. It
does not claim a provider or MCP join without matching evidence.

**Reason and benefit**

Native agents can hire teammates that start with the caller's approved
execution policy. Operators retain company boundaries, auth rules,
approval gates, and requalification. Reviewers can distinguish provider
streaming time from server API execution time when diagnosing
first-activity delays.

**Breaking changes**

None for existing hires or tool calls. `inheritRuntimeFrom` is optional
and only applies to same-company native agent callers. Conflicting
explicit runtime settings are rejected. The provider event field is
optional for existing producers.

## What Changed

- Added the native `hire_agent` protocol action, catalog entry, API
contract, and runner authority checks.
- Added `inheritRuntimeFrom: "caller"` validation and a closed native
runtime inheritance allowlist.
- Preserved managed AI binding normalization, default-environment
validation, approval snapshots, permissions, requalification, and fresh
child instructions.
- Added provider `inputUpdated` schema support and Rust forwarding.
- Added first-activity and server tool timing evidence with conservative
identity-join handling.
- Added route, authority, provider-event, sidecar, API, catalog, and
Rust-focused tests.
- Kept private Honeycomb links, raw traces, and local result paths out
of this description.

## Verification

Focused checks passed:

- 458 timing/session checks.
- 61 native hire inheritance checks.
- 20 hire authority checks.
- 1,741 API checks.
- 106 catalog checks.
- 54 provider sidecar checks.
- 12 Rust provider checks.

Live R2 and R3 each passed 45 checks across 6 runs (361,135 ms for R2).
R1 stopped at missing Docker image setup. The final trace is available
at
https://ui.honeycomb.io/paperclip/environments/test/datasets/paperclip/result/BiMypLNvmiB?tab=traces.
Latest-head CI passed all required build, typecheck, Rust, static,
Vitest, serialized-server, workspace, chat, and E2E jobs. The focused
local checks listed above passed; the broad local suite was not run
before the live evaluation, while CI provides the full repository
verification.

## Risks

- Timing fields describe separate observed windows. They do not prove a
provider or MCP owner without a valid trace join.
- The inheritance allowlist must stay synchronized with native runner
configuration fields.
- Approval snapshots include resolved safe inherited settings and should
be reviewed when native configuration fields change.
- The focused local suite is narrower than the full repository suite;
latest-head CI covers the broader repository checks.

> Roadmap review: `ROADMAP.md` places this work within Paperclip's
bring-your-own-agent direction. It extends existing native runner hiring
and observability behavior.

## Model Used

OpenAI GPT-6 (exact serving model ID is not exposed), with extended
reasoning and repository tool use; GPT-5.6 Luna assisted with focused
implementation and verification work.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 23:16:11 -05:00
Devin FoleyandPaperclip 92d4868e79 fix(server): isolate run errors and redact runtime capability headers (#13826)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Operators use Sentry to investigate failed runs and server errors.
> - Run reports attach a task ID, run ID, error code, and adapter
fingerprint.
> - The server skips Sentry's OpenTelemetry setup to preserve its
separate tracing and privacy settings.
> - Without an async context manager, a scope mutation can attach old
run data to later errors.
> - The HTTP logger also retains a runtime credential capability header.
> - This change isolates run metadata and redacts that header so
diagnostics identify failures without leaking credentials.

## Linked Issues or Issue Description

Refs #13446 and #13719.

**What happened?**

After a terminal run failure, an unrelated server exception can inherit
that run's tags, context, and fingerprint. Sentry then groups a database
error with an earlier adapter failure. The real SDK reproduces this with
the application's `skipOpenTelemetrySetup: true` setting. HTTP request
logs also retain the `x-paperclip-github-capability` header, which must
be treated as a credential.

**Expected behavior**

Run metadata belongs to the terminal run event. Later exceptions must
not inherit it. Every genuine error must still be captured. Runtime
capability headers must be redacted on success and failure logs.

**Steps to reproduce**

1. Initialize the optional Sentry SDK with the application's options and
an in-memory transport.
2. Capture a terminal run failure.
3. Capture an unrelated exception.
4. Inspect the second event. Before this fix, it contains the first
run's identity and fingerprint.
5. Send a request with a fixture runtime GitHub capability header.
Before this fix, HTTP logs retain the fixture value.

## What Changed

- Pass tags, context, and fingerprint directly to `captureException`
instead of mutating the ambient scope.
- Preserve the existing run fields, grouping keys, ordinary exception
capture, and privacy settings.
- Test two run identities interleaved with unrelated exceptions against
the real optional SDK.
- Update the capture contract tests and document event-local run
metadata.
- Redact the runtime GitHub capability header through the existing HTTP
logger policy. Test successful, denied, and failed requests.
- Add a dedicated GitHub-hosted CI check that installs the exact
optional SDK version declared in `server/package.json`. It fails if the
real-SDK regression would be skipped. The SDK stays outside the
workspace and production dependency graph.

## Verification

- The real-SDK regression failed before the fix because the unrelated
event contained `contexts.run_failure`.
- Five focused suites passed: 123 tests, including all optional SDK
tests. Suites: `run-failure-sentry-real-sdk.test.ts`,
`run-failure-sentry.test.ts`, `sentry.test.ts`,
`run-failure-report.test.ts`, and `http-log-redaction.test.ts`. A custom
in-memory transport prevented outbound Sentry delivery.
- All three new header-redaction cases failed before the policy fix and
passed afterward.
- The dedicated CI command passed locally with
`PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1` and the audited SDK available
through `NODE_PATH`.
- Server TypeScript check passed with a scratch configuration that
resolves this checkout's workspace packages. The existing dependency
links point to another checkout.
- `node scripts/check-module-boundaries.mjs` and `git diff --check`
passed.
- Gitleaks and a separate private-data scan passed before push.
- Full local workspace typecheck, test, and build were not run. The
machine has less than 2 GiB free and those commands include Rust builds.
Full PR CI must pass before merge.
- The dedicated real-SDK GitHub check passed with 1 test executed and no
skips: https://github.com/paperclipai/paperclip/actions/runs/35774449002
- Greptile reviewed c9db03bcab at 5/5. Its only thread is resolved. Full
PR CI passed on that same head:
https://github.com/paperclipai/paperclip/actions/runs/35774449020

## Risks

Small change to error attribution. Unrelated errors may now form their
correct Sentry groups instead of reopening a prior run group. No errors
are filtered or suppressed. No tracing is enabled and no new event
fields are added. No schema or runtime-execution changes. HTTP logs
retain their request and status diagnostics while masking the capability
value. The new SDK job has read-only permissions, no secrets, and an
in-memory Sentry transport.

## Model Used

OpenAI GPT-6 (Codex), with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 12:53:31 -07:00
Devin FoleyandPaperclip c496acb570 Return client errors for known OAuth reconnect states (#13794)
Map missing OAuth refresh credentials and terminal reauthorization to HTTP 422. Preserve reconnect instructions and reporting of unexpected provider failures.

All 363 focused tests and server typecheck pass. Required CI and review checks passed; Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-21 20:50:01 -07:00
Devin FoleyandPaperclip a7d3b17a97 Explain disabled Slack MCP app access during discovery
Recognize Slack's exact disabled-app response and return actionable setup
instructions from catalog and health routes. Bound response parsing and keep
unknown upstream errors reportable without exposing provider settings links.

Verified 361 focused tests, server typecheck, and authenticated discovery.
The three route regressions fail before this change and pass afterward.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-21 20:25:06 -07:00
Devin FoleyandPaperclip 878734a061 fix(tools): treat OAuth sign-in challenges as client errors (#13786)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connected apps can require OAuth sign-in before they list their
tools.
> - Remote discovery recognizes this condition as `oauth_challenge`.
> - Discovery and catalog refresh currently return it as HTTP 502.
> - The server error handler reports that response as a crash.
> - This pull request returns HTTP 422 for the explicit sign-in
challenge.
> - The operator keeps the sign-in instructions, while unexpected
upstream failures remain reportable.

## Linked Issues or Issue Description

**What happened?**

Connecting a remote MCP app that answers with a recognized OAuth
challenge returns HTTP 502. Reading an empty catalog or explicitly
refreshing it does the same. The server error handler then sends the
expected sign-in condition to error monitoring.

**Expected behavior**

A known sign-in requirement returns HTTP 422 with the existing
`oauth_challenge` code, message, and setup/reconnect links. An
unexplained upstream HTTP 400 or an unavailable upstream service still
returns 502 and reaches error monitoring.

**Steps to reproduce**

1. Configure a remote MCP app that returns HTTP 401 with a Bearer
challenge.
2. Connect the app, read its empty catalog, or request a catalog
refresh.
3. Observe HTTP 502 and a server error report before this change.

**Paperclip version or commit**

Reproduced on `6de50ba594b15efaa3eae6ed869cd39b3a436456`.

**Deployment mode**

Server with remote MCP connections. The regression coverage uses local
PostgreSQL and mocked upstream HTTP responses.

Related: #9750 addresses MCP initialization and session recovery. It
does not change the classification of this recognized sign-in condition.
Targeted searches found no duplicate classification PR.

## What Changed

- Return 422 for `oauth_challenge` from discovery and from catalog
health-error normalization.
- Preserve the existing structured error and remediation links.
- Test automatic empty-catalog reads and explicit refreshes. Verify that
OAuth challenges produce no Sentry capture and that upstream 400/503
failures still do.
- Update the direct-connect and blocked-redirect expectations and
document the monitoring behavior.

## Verification

- Before the fix, three sign-in route regressions fail with 502 instead
of 422; all four upstream-error controls pass.
- After the fix, all 339 tool-access and error-handler tests pass,
including authorization and redirect protections.
- These suites ran against disposable Homebrew PostgreSQL 16.14 through
the existing test-constructor seam. The temporary setup and config
remain outside the repository. CI uses the ordinary embedded PostgreSQL
setup.
- Direct server `tsc --noEmit` passes.
- Full build and recursive typecheck were attempted; the Runner Rust
step cannot run because `cargo` is absent on this machine.
- Full `pnpm test:run`: 8,211 passed, 14 failed, 4,759 skipped. The 36
failed files match the existing embedded PostgreSQL startup/cleanup and
macOS runtime-cache `EACCES` limitations. The changed database-backed
service suite passed separately with local PostgreSQL.
- Greptile: 5/5 with no unresolved comments. Seven CI workers received a
simultaneous shutdown signal; the failed jobs are being retried through
the normal workflow. Other completed checks passed.

## Risks

Low risk. Clients now receive 422 instead of 502 for the explicit
`oauth_challenge` condition. The code, message, and remediation links
remain available. No permissions, credential handling, OAuth discovery
rules, retry policy, or schema change. Other upstream failures retain
their existing behavior.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, repository inspection, code
editing, and test execution. The session does not expose an exact model
snapshot or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (339 service and
error-handler tests; full workspace limitations are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 00:34:45 +00:00
Devin FoleyandPaperclip 6de50ba594 fix(sentry): carry the deployment environment to the browser (#13784)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Operators can enable Sentry for the server and the signed-in
browser.
> - The server SDK reads `SENTRY_ENVIRONMENT` from the process
environment.
> - The browser receives its DSN through the session response, but
receives no environment.
> - A browser in staging therefore reports errors under the SDK's
production default.
> - This pull request passes the configured environment through the
existing session and monitoring gate.
> - Browser errors then identify the deployment environment while
preserving the existing privacy settings.

## Linked Issues or Issue Description

**What happened?**

With `SENTRY_ENVIRONMENT=staging`, browser exceptions are tagged
`production`. This can send errors to the wrong environment's alerts and
makes deployment follow-up unreliable.

**Expected behavior**

The browser uses the server's configured Sentry environment. A reused
image works in either staging or production. A signed-out browser still
sends no events.

**Steps to reproduce**

1. Configure a frontend Sentry DSN and `SENTRY_ENVIRONMENT=staging`.
2. Sign in and capture a browser exception.
3. Inspect the event environment. Before this change, it is
`production`.

**Paperclip version or commit**

Reproduced on `a3749aac4680a901fa0fe1cc898907887abc9908` with the real
browser SDK and a local test transport.

**Deployment mode**

Authenticated server and browser with optional Sentry monitoring
enabled.

No duplicate environment-attribution issue or pull request was found in
the targeted GitHub search.

## What Changed

- Add `sentryEnvironment` to the authenticated session response and
shared schema. The optional field supports a newer browser reading an
older server response.
- Pass the environment to the browser SDK. An environment change
restarts the client through its existing serialized lifecycle.
- Cover environment attribution with a real SDK event, session
authorization, unchanged-session refetches, environment changes, and
legacy responses.
- Document configuration and compatibility. Keep the loaded bundle's
release identity and existing privacy filters.

## Verification

- The regression test emits `production` for a requested staging
environment before the fix.
- Focused route, schema, browser lifecycle and real-SDK tests: 69 pass.
- UI and shared-package typechecks, direct server `tsc --noEmit`, and
token gates pass.
- Full `pnpm build` and `pnpm -r typecheck` were attempted. Both stop at
the Runner Rust step because `cargo` is absent on this machine.
- Complete UI suite: 6,540 tests pass in 626 files.
- Full `pnpm test:run`: 8,210 passed, 14 failed, 4,753 skipped; 36 files
fail due to embedded PostgreSQL startup/cleanup and macOS runtime-cache
`EACCES`. These match the existing local baseline; none touch the
changed behavior.
- Greptile: 5/5, no unresolved review threads. Linux CI has passed
Build, Typecheck + Release Registry, and the completed test jobs so far.
Remaining jobs are running or queued: the AWS runner provisioner is
retrying EC2 CreateFleet `InternalError` responses. Full results will be
recorded before merge.

## Risks

Low risk. This adds one optional session field and changes Sentry
attribution only. No migration or new monitoring opt-in is introduced.
Missing settings keep the browser SDK default. Agent and unauthenticated
requests still receive 401 without monitoring settings. Existing loaded
browser bundles keep their old behavior until refreshed.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, repository inspection, code
editing, and test execution. The session does not expose an exact model
snapshot or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass (focused and full UI suites
pass; full-root environment failures documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 23:40:46 +00:00
Devin FoleyandPaperclip 600e552d7b fix: attribute Sentry errors to the loaded source release (#13719)
Attribute optional server and browser Sentry events to their source build.
Use validated build commits for Docker and source/npm artifacts, preserve
explicit server release overrides, and keep cached browser bundles tied
to the commit they loaded.

Verify 127 focused tests, server/UI typechecks, Docker and source build
stamps, all 53 CI checks, and Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-20 08:08:33 -07:00
Nicky LeachandPaperclip c276d3fdc3 feat(observability): report terminal run failures to Sentry (#13446)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server records the state of each agent run.
> - Terminal run failures need clear error tracking for operators.
> - The server did not report terminal `failed` or `timed_out`
transitions to Sentry.
> - This pull request reports each genuine terminal failure transition
with safe diagnostic data.
> - The benefit is faster diagnosis without changing run control flow or
exposing credentials.

## Linked Issues or Issue Description

**What happened?**

The server wrote terminal run failures but did not report them to
Sentry. Operators could not see these failures in error tracking.

**Expected behavior**

The server should report each genuine transition to `failed` or
`timed_out` to Sentry.

**Steps to reproduce**

1. Run an agent task that reaches a terminal failure state.
2. Inspect the Sentry events for the server.
3. Observe that the terminal run failure has no matching Sentry event.

**Paperclip version or commit**

The change targets the current `master` branch.

No public GitHub issue or pull request covers this change.

## What Changed

- Add `captureRunFailure()` as a fail-open Sentry entry point.
- Add `reportRunFailure()` to filter status, resolve the adapter, redact
text, and report the failure.
- Call `reportRunFailure()` beside each of the eight terminal status
writers.
- Report six diagnostic values: the instance host, task identifier, run
identifier, error message, error code, and agent adapter.
- Group events by error code and agent adapter while keeping the
redacted message in the event.
- Report only genuine transitions and avoid duplicate finalization
events.
- Keep Sentry failures outside run control flow.

## Verification

- `pnpm vitest run
server/src/services/__tests__/run-failure-report.test.ts
server/src/__tests__/run-failure-sentry.test.ts
server/src/__tests__/native-session-resumption.test.ts`
- `pnpm vitest run
server/src/services/execution-control-reconciliation.test.ts`
- `pnpm --filter @paperclipai/server typecheck`
- The full continuous-integration suite must run on this pull request.

## Risks

- The report path can add diagnostic events when Sentry is configured.
- The report path returns without action when Sentry is not configured.
- Redaction runs before length limits and before the event leaves the
process.
- The change has no migration and no schema change.

## Model Used

OpenAI Codex, GPT-5, tool use and code review support. The exact context
window and reasoning configuration are not exposed by the runtime.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 23:32:15 -07:00
Nicky LeachandPaperclip 6019e2bd6e feat(adapter-utils): carry binary bodies and attachment routes over the HTTP/2 sandbox bridge (#12923)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip runs agents in local and remote sandboxes through adapter
utilities
> - The HTTP/2 sandbox bridge decoded every body as UTF-8 text and
rejected non-JSON content
> - This stopped agents from uploading or downloading issue attachments
through that bridge
> - This pull request carries raw bytes, permits the two attachment
routes, and enforces a shared body limit
> - The benefit is correct attachment transfer with a process-wide
memory guard

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The HTTP/2 sandbox bridge forwards request bodies between an agent
sandbox and the Paperclip host. It now supports binary bodies and the
issue attachment routes.

**Current behavior**

The bridge decodes each body as UTF-8 text. It returns HTTP 415 for
content types outside the JSON route list. An agent cannot upload or
download an issue attachment through this transport.

**Proposed behavior**

The bridge carries raw bytes through the forward path. It permits the
attachment upload and content routes. The queue transport and file
gateway keep their existing route behavior. A shared 10 MiB body limit
and process-wide byte reservation protect memory use.

**Reason and benefit**

Attachment clients need byte-preserving transfer. The shared limit keeps
the gateway and host aligned. The reservation prevents concurrent
streams from exceeding the accepted process memory ceiling.

**Breaking changes**

The HTTP/2 bridge accepts two attachment routes and permits binary
content. The queue transport and file gateway keep their previous route
lists and HTTP 415 behavior. No schema or external endpoint changes.

## What Changed

- Carry request and response bodies as raw bytes through the HTTP/2
bridge.
- Permit attachment upload and attachment content routes on the HTTP/2
bridge only.
- Raise the resolved per-body limit to 10 MiB and share it between the
gateway and host.
- Reserve body bytes before allocation and release each stream
reservation on every terminal path.
- Document the body limit, process ceiling, and reservation behavior.

## Verification

- Run `pnpm exec vitest run
packages/adapter-utils/src/http2-bridge-server.test.ts
packages/adapter-utils/src/execution-target-sandbox.test.ts
packages/adapter-utils/src/sandbox-callback-bridge.test.ts`; 226 tests
pass.
- Run `pnpm --filter @paperclipai/adapter-utils typecheck`; it passes.
- Run the direct server TypeScript check with `tsc --noEmit` in
`server/`; it passes with zero errors.
- Verify multipart upload and binary download round trips over HTTP/2
without corruption.
- Verify the queue transport and file gateway return HTTP 415 for the
same routes.
- Verify the host rejects bodies over the resolved limit.
- Verify a denied reservation returns HTTP 503 and allocates no copy.
- Verify stream cleanup releases reservations after completion, error,
abort, timeout, and close.

## Risks

The bridge now accepts larger bodies and binary content. The
process-wide reservation limits total live body bytes to 1 GiB. Route
behavior changes only for the HTTP/2 bridge. The security review found
no blocking issue for this commit range.

## Model Used

OpenAI Codex, GPT-5. The runtime used tool calls and code execution. The
runtime did not expose the context window size. No model-generated code
changes were made for this pull request.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-08 14:18:12 -07:00
DottaandPaperclip c723bb4dfc fix(skills): reuse validated runtime revisions during preparation (#13042)
## Thinking Path

> - Paperclip manages AI agents and prepares their runtime inputs before
each turn.
> - Shared company skills are part of those inputs for native and legacy
adapters.
> - Runtime materialization refreshed the full inventory again for every
declared file.
> - Remote skill directories were also downloaded and rebuilt on every
turn.
> - Measured preparation took 42–73 seconds while runner execution took
7–9 seconds.
> - This change reads the inventory once and reuses validated installed
revisions.
> - Agents retain their selected skills while repeated preparation
avoids upstream work.

## Linked Issues or Issue Description

**What happened?**

One 114-skill preparation performed 407 inventory refreshes, 48
directory rebuilds, and 388 GitHub file fetches. Reusing existing local
copies took 151 ms.

**Expected behavior**

Each listing refreshes inventory once. Unchanged installed remote
revisions reuse complete, validated local copies. Local edits remain
visible. Explicit updates select new revisions.

**Steps to reproduce**

1. Import GitHub skills with supporting files.
2. Run an agent turn, then run another with the same installed
revisions.
3. Observe repeated inventory scans, downloads, and runtime directory
replacement before execution.

Related prior attempts: #2330 and #9268 (still open; #9268 last updated
July 9). Those use a marker compared with `updatedAt`. This patch
follows the required content validation, immutable revision, company
isolation, atomic publication, and read-only semantics, and removes
refresh-per-file multiplication.

## What Changed

- Split public file reading from reading an already loaded skill.
Runtime listing refreshes inventory once.
- Add a company-scoped revision cache with file manifests outside the
delivered skill directory. Fingerprints omit cosmetic metadata.
- Validate exact file inventory, sizes, and hashes before warm reuse.
Reject traversal and symlinks. Stage complete builds and serialize
atomic publication across processes.
- Preserve local/catalog direct sources, stored Markdown fallback,
explicit version snapshots, and legacy mutable-ref compatibility. Report
missing supporting files and keep older valid revisions readable.
- Clean both runtime layouts on rename/removal and record
`skills.prepare` under preparation timing.
- Add service/cache regressions and an isolated 114-skill benchmark,
including a new-process warm run.

## Verification

- Final targeted skill-service/cache/trace validation: 86 tests pass (61
embedded-PostgreSQL service tests, 19 cache tests, 6 trace tests).
Database tests executed rather than skipped. Focused skill routes,
adapter selection, and native runtime context also pass.
- `pnpm -r typecheck` and `pnpm build` pass locally at `22caa1fe4`.
- The full `pnpm test:run` matrix passes on supported Linux CI at the
final head: [CI
run](https://github.com/paperclipai/paperclip/actions/runs/34236097762).
Local full-suite execution encountered PostgreSQL startup contention, a
random allocated-port boundary, and a socket hang-up; every affected
suite passed on an isolated rerun. The interrupted local serialized run
is not claimed as a complete local pass.
- Repeatable benchmark: `pnpm --filter @paperclipai/server exec tsx
../scripts/benchmark-skill-preparation.ts`. Mixed 114-skill inventory
with 429 remote files on Linux: cold 286 ms, warm median 96 ms / maximum
153 ms including a new process. Every warm sample performs one refresh,
zero upstream fetches/rebuilds, and reports no missing entries; content
assertions pass.
- Controlled deployment against the previously deployed revision
completed with zero lost runs. Real inventory: 114 skills, 670 declared
files; 402 cached files match the prior installed copies byte-for-byte.
Ten post-deployment warm preparations: median 129 ms / maximum 208 ms;
new-process warm 194 ms, zero downloads/rebuilds/missing entries.
- Five sequential real browser questions persisted in 10.6–20.7 s
(median 12.2 s), versus 50–83 s before. Skill preparation median 240 ms,
with one 2.37 s outlier. Total preparation median 3.337 s / maximum
8.728 s **does not fully meet** the <3 s / <5 s target. The excluded
historical-run redaction query takes about 1.36 s per scan at two
preparation call sites; wider application latency coincided with the
outlier, without a cache rebuild. These residuals are reported rather
than discarded.
- Disposable skill reimport verified through actual selected-skill runs:
the next run read the changed code. Fixture removed and agent
configuration verified unchanged.
- Greptile 5/5, zero unresolved review threads, all final-head CI checks
green.

## Risks

- Cold preparation still requires upstream availability for supporting
files. An unavailable revision is reported missing and never falls back
to an older revision.
- Valid older revisions and quarantined invalid entries consume additive
disk space until skill cleanup. An abruptly killed publisher can leave a
lock that requires operator cleanup after confirming its PID is dead.
- Warm validation reads all cached file bytes. Very large inventories
still have proportional local I/O cost.
- No HTTP API, schema, agent configuration, or first-party Telemetry
changes. OpenTelemetry retains its operator endpoint gate.

## Model Used

OpenAI GPT-6 in Codex, with reasoning, repository inspection, code
editing, and test execution. The exact serving snapshot and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted and isolated
reruns; full Linux CI matrix passes, local full-run caveats above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-08 09:35:32 -05:00
Nicky LeachandPaperclip ed3559dd21 feat(server): split the Sentry DSN into front-end and backend variables (#12678)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip reports server and browser errors through optional Sentry
monitoring
> - One environment variable sends both error types to one Sentry
project
> - Operators need separate control for browser and server error data
> - This pull request adds specific variables and keeps the existing
variable as a fallback
> - The benefit is separate monitoring without breaking current
deployments

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The Sentry configuration for server and browser monitoring uses one
environment variable.

**Subsystem affected**

Cross-cutting (multiple of the above)

**Current behavior**

`SENTRY_DSN` supplies the server and browser clients. Both clients
therefore report to the same Sentry project.

**Proposed behavior**

`SENTRY_DSN_FRONTEND` supplies the browser client. `SENTRY_DSN_BACKEND`
supplies the server process. `SENTRY_DSN` remains a fallback for either
component.

**Reason and benefit**

Operators can send browser and server errors to separate Sentry
projects. Operators can also activate only one component.

**Breaking changes**

None. Existing deployments can continue to use `SENTRY_DSN`.

## What Changed

- Add `resolveSentryDsns(env)` and use it in the server and browser
configuration paths.
- Add precedence, empty-string, fallback, and route tests.
- Update the README, observability guide, and stale code comments.
- Log one warning when the server uses the legacy fallback without
exposing a DSN value.

## Verification

- `pnpm vitest run --project server sentry-dsn` — 8 tests pass.
- `pnpm vitest run --project server auth-routes` — 21 tests pass.
- The earlier run of the three targeted suites passed 40 tests.
- `tsc --noEmit` passes for the files in this diff.
- All required GitHub Actions checks pass, including the full
continuous-integration suite.

## Risks

The main risk is an incorrect environment variable precedence rule. Unit
tests cover specific values, empty strings, and legacy fallback
behavior. The existing `SENTRY_DSN` path remains compatible.

## Model Used

OpenAI Codex — GPT-5, current runtime, tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-01 11:02:04 -07:00
DottaandDev Agent 25cf079ec5 feat(runner): add Codex-native application integration (#12591)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The runner package is useful only when the application can start,
observe, and recover a native Codex run safely.
> - Existing direct adapters must keep their current execution and
finalization paths.
> - The application boundary therefore needs additive persistence,
authorization, coordination, and recovery behind an explicit
experimental adapter.
> - This pull request adds that Codex-only boundary without activating
generalized providers, remote environments, or the later task/SDK
surfaces.

## Linked Issues or Issue Description

**Subsystem affected**

Shared contracts, database persistence, adapter utilities, server
native-runtime services, and the experimental Paperclip Runner adapter.

**Problem or motivation**

The already-landed runner package has a qualified Codex path, but the
application needs durable native-run state, guarded runtime selection,
authenticated coordination, tool security, finalization, and recovery
before the experimental adapter can be exercised safely.

**Proposed solution**

Add a Codex-only `paperclip_runner` application path behind the existing
default-off native-runner setting. Bind native state and coordination to
company/run identity, preserve persisted-run recovery, and leave every
direct adapter on its existing legacy execution path.

**Alternatives considered**

The earlier stack boundary introduced a generalized executor and
remote-environment lifecycle here. That made this PR depend on
implementations in higher PRs and changed reusable sandbox behavior
globally. Those pieces are now deferred together to #12592.

**Roadmap alignment**

ROADMAP.md does not list a conflicting native-runner integration
project. This change adds the application boundary for the existing
Runner architecture.

## What Changed

- Added native run/result/finalization/provider-trace persistence,
shared validators, and idempotent migration/replay coverage.
- Added guarded Codex-only runtime selection, authenticated PRP
coordination, recovery, finalization, and interaction services.
- Added run/company-bound tool-gateway authorization, credential
redaction, SSRF protections, and replay-safe behavior.
- Added the explicit `paperclip_runner` adapter behind the default-off
rollout setting.
- Preserved legacy answered-question wake projection and direct-adapter
execution/finalization paths.
- Hardened cancellation so only owned in-memory child processes are
signaled; persisted recycled PIDs/process groups are never trusted.
- Retained the narrow Claude ACPX isolated-context security follow-up
discovered after #12590.
- Deferred the generalized executor, provider ingress, remote lifecycle,
SDK/lab/eval work, release-process changes, and lockfile.

## Verification

- Changed-file delta against `master`: 133 files.
- GitHub Actions is the authoritative verification environment for this
PR.
- Full CI, security, and Greptile review will run on this lowest
unmerged stack PR.
- Local tests/build/typecheck were not run because this checkout is
resource constrained.
- Static diff/reference checks pass, and `pnpm-lock.yaml` is unchanged.

## Risks

- This touches central heartbeat and agent-route code, so legacy
compatibility is the primary risk.
- Runtime selection remains Codex-only and explicit; direct Codex,
Claude, OpenCode, process, HTTP, and plugin adapters remain on their
existing paths.
- Fresh native starts fail closed while the rollout flag is off;
persisted native records remain readable and recoverable.
- Cancellation, company/run binding, tool calls, status decisions, and
completion writes are guarded or replay-safe.

> For core feature work, check [ROADMAP.md](ROADMAP.md) first and
discuss it in #dev before opening the PR. Feature PRs that overlap with
planned core work may need to be redirected.

## Model Used

OpenAI Codex, GPT-5.6, with repository tools, code execution, and
parallel agent review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked existing issues or described the issue in-PR
following the relevant issue template
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass — GitHub Actions is
authoritative for this resource-constrained checkout
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [ ] All Paperclip CI and security gates are green
- [ ] Greptile is 5/5 with no open actionable findings
- [x] I will address all Greptile and reviewer comments before merge

## Stack

- Position: 3 of 5 overall; lowest of 3 currently unmerged
- Base: `master`
- Previous:
[#12590](https://github.com/paperclipai/paperclip/pull/12590), qualified
Claude ACPX runtime — merged
- Next: [#12592](https://github.com/paperclipai/paperclip/pull/12592),
generalized Codex executor, task experience, and developer SDKs

---------

Co-authored-by: Dev Agent <dev@paperclip.ing>
2026-08-31 14:38:38 -05:00
Nicky LeachandPaperclip 64b7dce0ad refactor(adapter-utils): replace the process-wide byte ledger with route-local byte bounds (#12465)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The adapter layer carries sandbox requests to host processes.
> - The HTTP/2 bridge used one process-wide byte ledger for all routes.
> - One busy route could exhaust that shared budget and move another
route to file transport.
> - This pull request gives each host retention site a fixed byte bound
and limits concurrent HTTP/2 streams.
> - The benefit is local protection: one route cannot consume the byte
budget of another route.

## Linked Issues or Issue Description

**What happened?**

The HTTP/2 bridge used one aggregate byte ledger for retained bytes
across all routes. A busy route could exhaust the shared budget and
force an unrelated route to use file transport.

**Expected behavior**

Each route should protect its own retained bytes. A reset on one HTTP/2
stream should cancel only that stream's host forward.

**Steps to reproduce**

1. Start the HTTP/2 bridge with multiple sandbox routes.
2. Send enough retained data through one route to reach the aggregate
byte limit.
3. Send a request through a sibling route.
4. Observe that the sibling route can fall back to file transport
because the first route used the shared ledger.

**Paperclip version or commit**

`47639e227e78e3c5e0dd1a3c0e2d792fe86895a3`

**Deployment mode**

Built from source with the adapter-utils and server test suites.

## What Changed

- Bound each host retention site with a fixed local byte limit.
- Limited concurrent live HTTP/2 streams with one built-in stream limit.
- Bound each host forward and response-body read to its own HTTP/2
stream lifetime.
- Removed the process-wide byte ledger, its environment override, its
metrics, and its file-transport fallbacks.
- Added tests for the stream limit, host body budget, and sibling-stream
cancellation.

## Verification

- Run `pnpm vitest run --project adapter-utils`.
- Confirm that 996 adapter-utils tests pass.
- Confirm that `test_live_forward_work_never_passes_the_stream_limit`
passes.
- Confirm that `test_the_host_body_budget_matches_the_stream_limit`
passes.
- Confirm that the sibling-stream cancellation test passes.
- Run `pnpm tsc --noEmit`.
- Confirm that all pull request checks pass.

## Risks

The bridge no longer uses a process-wide byte ledger. A local bound or
stream limit that is too low can reject or delay valid work. The tests
cover the new limits and stream cancellation behavior.

## Model Used

OpenAI GPT-5 Codex. Runtime model ID: GPT-5. The model used code
execution and repository tools. The runtime does not expose the context
window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-28 14:36:18 -07:00
Nicky LeachandPaperclip 7895f7f2b0 Install the declared Sentry server package into the hosted image (#12330)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip supports opt-in Sentry error monitoring for server and
browser errors.
> - The hosted image must include the server package when an operator
sets SENTRY_DSN.
> - The server package is an optional peer in the source tree, so the
image did not include it.
> - This pull request installs the declared server package in the hosted
image and checks the result.
> - The benefit is a hosted tenant can send server errors without a
manual package install.

## Linked Issues or Issue Description

No public issue exists for this change.

**What happened?**

The hosted image did not include the declared @sentry/node server
package. A hosted tenant could set SENTRY_DSN, but the server could not
load the package from the image.

**Expected behavior**

The hosted image must include the exact @sentry/node version from
server/package.json. The self-hosted image must remain without this
optional package.

**Steps to reproduce**

1. Build or pull the hosted image.
2. Resolve @sentry/node from the server package path.
3. Compare its version with server/package.json.
4. Confirm that the tsx loader path still resolves.

**Paperclip version or commit**

Commit b6ff556a33ebdbe764b7f495951cd59009776608.

**Deployment mode**

Docker hosted image.

## What Changed

- Add a cloud-server-deps Docker stage that installs the declared
@sentry/node version in isolation.
- Copy the isolated package into the cloud image without changing the
production image.
- Add a probe that checks the tsx loader and the resolved Sentry
version.
- Run the probe after the hosted image push in the Docker workflow.
- Add server tests and update the observability documentation.

## Verification

- Run `pnpm exec vitest run --project @paperclipai/server
server/src/__tests__/cloud-image-sentry.test.ts`.
- Confirm that the changed test passes in CI.
- Confirm that all pull request checks pass.
- Note that the Docker workflow does not run for pull requests. It runs
after a push to master, for configured tags, or after manual dispatch.

## Risks

- Low risk. The production image body stays unchanged.
- The cloud image adds the declared Sentry package and a small
dependency tree.
- The workflow probe fails if the image loses the tsx loader or resolves
a different Sentry version.

## Model Used

OpenAI GPT-5; exact model version supplied by the execution service;
tool use and code execution; context window not specified.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-27 19:10:49 -07:00
Nicky LeachandPaperclip 1de105c475 fix(observability): pin the Sentry browser SDK and gate the optional Sentry server peer on the exact version (#12270)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip uses separate server and browser packages for runtime
services and the board.
> - Sentry integrations need an exact SDK version and safe optional
loading.
> - A version range can select an SDK that the privacy tests did not
audit.
> - Missing peer metadata does not describe the optional server SDK
contract.
> - This pull request pins the browser SDK and gates the optional server
SDK on its exact version.
> - The benefit is a clear SDK contract with fail-open startup behavior.

## Linked Issues or Issue Description

**What happened?**

The browser package used the range ^10.71.0, so a lockfile refresh could
select a newer SDK. The server loaded @sentry/node dynamically but did
not declare its optional peer contract.

**Expected behavior**

The browser package must use the audited 10.71.0 version. The server
must load @sentry/node only when the installed peer matches 10.71.0. The
server must start when the optional peer is absent.

**Steps to reproduce**

1. Install the project dependencies.
2. Inspect the browser Sentry version and the server package metadata.
3. Start the server without installing @sentry/node.
4. Confirm that the server starts and that the dynamic Sentry bootstrap
does not load an unsupported peer version.

**Paperclip version or commit**

9c57c0f119

**Deployment mode**

Built from source with pnpm dev or pnpm build.

**Installation method**

Built from source.

**Agent adapter(s) involved**

Not adapter-specific (core change).

**Database mode**

Not database-related.

## What Changed

- Pin @sentry/browser to exactly 10.71.0 as a UI development dependency.
- Declare @sentry/node as an optional server peer dependency at 10.71.0.
- Gate the dynamic server bootstrap on the exact peer version.
- Add tests for the browser pin, peer metadata, version gate, and
fail-open loading.
- Document the supported server SDK version.
- Keep the lockfile unchanged because the pull request workflow
regenerates it for manifest changes.

## Verification

- Server tests pass with six expected skips when @sentry/node is absent.
- UI tests pass.
- The UI build emits the lazy Sentry browser chunk.
- git diff --check passes.
- GitHub pull request checks must pass after this pull request opens.
- Greptile must return a 5/5 score with no open findings.

## Risks

The exact version gate prevents Sentry startup when an unsupported SDK
version exists. The integration remains optional and fail-open. The
lockfile workflow must regenerate the lockfile before frozen downstream
jobs run. The label-gated Storybook visual job must not run until it can
restore the generated lockfile artifact.

## Model Used

OpenAI Codex, GPT-5, tool use and code review support, exact context
window details are managed by the execution platform.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with Fixes: # / Closes: #
/ Refs: # OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-27 07:20:03 -07:00
Nicky LeachandPaperclip 06cd21ed0f fix(observability): declare the optional OpenTelemetry peer dependencies (#12249)
## Thinking Path

> - Paperclip manages AI agents for work.
> - Paperclip includes an observability path that operators can enable
for tracing.
> - The server loads several OpenTelemetry packages only when tracing is
enabled.
> - The documentation calls these packages optional peer dependencies,
but the server manifest does not declare them.
> - This gap hides supported versions and stops Dependabot from
maintaining the packages.
> - This pull request aligns package metadata, runtime checks, and
documentation with the opt-in tracing design.
> - The change gives operators clear installation behavior and keeps the
no-op default.

## Linked Issues or Issue Description

This pull request fixes a package metadata and installation defect.
Related observability work appears in
[#8476](https://github.com/paperclipai/paperclip/pull/8476) and
[#9672](https://github.com/paperclipai/paperclip/pull/9672).

The server documentation described optional OpenTelemetry peer
dependencies, but `server/package.json` did not declare them. Package
managers and Dependabot could not see the supported version ranges. The
UI and Claude local adapter also relied on automatic peer installation
for `yjs` and `@anthropic-ai/sdk`.

The package manifests now declare the optional runtime packages. A
default install does not install optional tracing peers. The server
keeps its no-op behavior when tracing is disabled or a peer is absent.

## What Changed

- Add seven optional OpenTelemetry packages to `server/package.json` and
mark each package as optional.
- Keep `@opentelemetry/api` as a normal dependency for the no-op
interface.
- Disable automatic peer installation in `.npmrc`.
- Declare `yjs` for the UI package and `@anthropic-ai/sdk` for the
Claude local adapter.
- Check declared peer versions before the server loads a dynamic
OpenTelemetry import.
- Keep the endpoint gate, dynamic imports, and fail-open behavior
unchanged.
- Update the observability and README documentation.
- Tell Dependabot that its npm parser does not read `peerDependencies`.

## Verification

- Targeted server tests pass: 34 passed and 2 skipped.
- The skipped tests require the real OpenTelemetry SDK and remain
pre-existing.
- The pull request workflow regenerates the lockfile because manifest
files and `.npmrc` changed.
- The policy job confirms that the pull request does not include
`pnpm-lock.yaml`.
- GitHub checks pass except `security/snyk (cryppadotta)`, which remains
pending after its authorized wait cap.
- Greptile Review reports 5/5 with no open findings.
- Server typecheck passes.

## Risks

- Optional peers can produce a diagnostic when the installed version
does not match the declared range.
- A missing optional peer does not stop the server.
- Disabling automatic peer installation can expose undeclared package
use in other workspaces.
- This pull request declares the affected packages and adds tests for
the changed behavior.
- This pull request makes no database or API changes.

## Model Used

OpenAI Codex, GPT-5, with repository inspection and pull request
preparation.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with Fixes / Closes /
Refs OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-26 17:34:53 -07:00
Nicky LeachandPaperclip 8f1e3cfe24 feat(observability): add opt-in Sentry error monitoring for the server and the browser (#12190)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server and the browser need clear error reports when an operator
enables external monitoring.
> - Paperclip already uses an opt-in OpenTelemetry pattern for server
traces.
> - Sentry can provide error reports for both runtime paths when the
operator sets one data source name.
> - This pull request adds one opt-in Sentry gate for the server and the
browser.
> - The benefit is faster diagnosis while the default setup sends no
Sentry data.

## Linked Issues or Issue Description

**What is improved?**

Paperclip gains optional error monitoring for server and browser
failures.

**Subsystem affected**

Cross-cutting (server, UI, and shared authentication data).

**Current behavior**

Paperclip has no built-in Sentry error capture for server failures or
browser boundary failures. Operators must inspect local logs and browser
tools.

**Proposed behavior**

When the operator sets `SENTRY_DSN`, the server and authenticated
browser use the same Sentry project. When the variable is absent, both
paths stay inactive. The server loads Sentry dynamically and fails open
when the optional package is absent.

**Reason and benefit**

Operators can inspect runtime errors in one Sentry project. The default
setup remains local and sends no monitoring data.

**Breaking changes**

None when `SENTRY_DSN` remains unset. Authenticated session responses
add the optional `sentryDsn` field.

**Additional context**

The implementation uses built-in Sentry privacy options. It disables
default HTTP context and breadcrumb integrations and keeps
`sendDefaultPii` false.

## What Changed

- Add an opt-in server Sentry gate with dynamic package loading and
fail-open behavior.
- Add the Sentry data source name to the authenticated session response.
- Add an authenticated browser Sentry gate and React error boundary
capture.
- Add tests for server, browser, route, and application error paths.
- Document activation, installation, privacy settings, capture behavior,
and operator controls.

## Verification

- Run `npx vitest run server/src/__tests__/sentry.test.ts`.
- Run `npx vitest run ui/src/lib/sentry.test.ts`.
- Run `npx vitest run server/src/__tests__/auth-routes.test.ts
server/src/__tests__/shutdown.test.ts`.
- Confirm that the full continuous integration suite passes on this pull
request.
- Leave `SENTRY_DSN` unset and confirm that the server and browser gates
stay inactive.
- Set `SENTRY_DSN` and install the optional Sentry packages before a
manual capture check.

## Risks

The operator controls the Sentry project and accepts the data risk when
the operator enables the feature. Error objects can contain messages,
stacks, or cause chains with private values. The default configuration
sends no data because the feature stays off without `SENTRY_DSN`. A
missing optional server package does not stop server boot.

## Model Used

OpenAI Codex, GPT-5, with tool use, repository inspection, GitHub CLI
operations, and code review support.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-26 11:15:45 -07:00
Nicky LeachandPaperclip 445547c989 feat(duplex): run the Daytona sandbox callback bridge over Node HTTP/2 (#12120)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Sandbox providers carry agent work through controlled execution
channels
> - The Daytona callback bridge uses a bespoke line-framed protocol over
its duplex channel
> - The bespoke protocol adds framing work and does not use the Node
transport that already supports multiplexed streams
> - This pull request carries raw bytes across the channel, adds a Node
HTTP/2 bridge, and selects it for Daytona
> - The benefit is one authenticated, multiplexed callback session with
queue_v1 as the bounded fallback

## Linked Issues or Issue Description

**Subsystem affected**

The packages/plugins Daytona provider and the shared duplex execution
path.

**Problem or motivation**

The Daytona callback bridge uses a bespoke line-framed protocol over the
provider duplex channel. This adds protocol work and limits stream
handling.

**Proposed solution**

Carry raw bytes through the cross-layer channel. Add an authenticated
Node HTTP/2 host server and sandbox client gateway. Select http2_v1 for
Daytona and retain queue_v1 as the fallback.

**Alternatives considered**

Keep the current duplex_v1 protocol. This keeps the bespoke framing path
and does not provide one HTTP/2 session for callback streams.

**Roadmap alignment**

ROADMAP.md lists Daytona under cloud and sandbox agents. This change
improves the shipped Daytona provider path.

**Additional context**

The branch adds no dependency. Node 24 provides the http2 module. The
host token check and canonical path parser remain the single dispatch
path.

## What Changed

- Carry raw Uint8Array chunks through the adapter, plugin, worker,
runtime, and Daytona layers.
- Encode bytes as base64 only across the JSON-RPC hop, because JSON has
no binary type.
- Add the bounded host HTTP/2 server and the in-sandbox HTTP/2 client
gateway.
- Authenticate every stream with the per-run bridge token before route
work.
- Parse the path once and reuse the canonical result for route and
forwarding work.
- Select http2_v1 for Daytona and fall back once to queue_v1 when the
client preface is absent.
- Add transport, session, stream, and fallback telemetry.
- Mark HTTP/2 as the preferred transport and queue_v1 as the
soft-deprecated fallback.

## Verification

- `npx vitest run packages/adapter-utils/src` — 990 passed and 4
skipped.
- `npx vitest run
server/src/__tests__/plugin-worker-manager-duplex.test.ts` — 32 passed.
- `npx vitest run --config
packages/plugins/sandbox-providers/daytona/vitest.config.ts` — 220
passed and 6 skipped.
- `npx tsc --noEmit` in `packages/adapter-utils`, `packages/shared`,
`packages/plugins/sdk`, and `server` — clean.
- No `package.json` or `pnpm-lock.yaml` file changed.
- The live Daytona test skips when `DAYTONA_API_KEY` is absent.
- The root `npx tsc --noEmit` command has a pre-existing missing
`packages/adapters/droid-local` reference on this branch and on
`master`.

## Risks

- The transport change affects several duplex layers and could expose
byte-boundary errors.
- A missing HTTP/2 client preface falls back once to queue_v1 and
records `preface_missing`.
- The host token check and canonical path parser must remain on the
shared dispatch path.
- The live Daytona test needs `DAYTONA_API_KEY` and does not run in this
agent sandbox.

## Model Used

OpenAI GPT-5, tool-enabled coding agent with repository inspection,
GitHub CLI, and shell execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-25 07:35:39 -07:00
Nicky LeachandPaperclip d1573244b5 refactor: disambiguate the Telemetry and Observability data paths (#12128)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip records first-party events, OpenTelemetry data, and local
run-log events
> - The code and documents used one term for these three data paths
> - This naming made the required review level unclear
> - This pull request names each data path in the module names,
documents, and code comments
> - The benefit is a clear review rule without a runtime change

## Linked Issues or Issue Description

**Issue type**

Unclear or confusing.

**Where is the issue?**

`packages/shared/src/telemetry/README.md`, `doc/observability.md`,
`doc/run-log-events.md`, and the duplex instrumentation modules.

**What's wrong?**

The repository used Telemetry for first-party events, OpenTelemetry
data, and local run-log events. This usage made the data path and review
level unclear.

**Suggested fix**

Use Telemetry only for Paperclip first-party events. Use Observability
for OpenTelemetry data. Use the run log for rows in
`heartbeat_run_events`.

Related public pull requests: #8476 and #9672.

## What Changed

- Rename the duplex instrumentation modules and identifiers from
`Telemetry` to `Observability`.
- Move the Observability and run-log contracts out of the Telemetry
README.
- Add `doc/observability.md` and `doc/run-log-events.md` as the
canonical documents.
- Add a file-path review rule to `AGENTS.md`.
- Correct the remaining code comments that name the wrong data path.
- Keep all event names, payloads, database records, spans, configuration
keys, environment variables, and runtime paths unchanged.

## Verification

- `npx vitest run packages/shared/src/telemetry/readme-contract.test.ts`
passes.
- `npx vitest run packages/adapter-utils/src/published-exports.test.ts`
passes.
- `npx vitest run
packages/adapter-utils/src/acpx-engine/startup-timing.test.ts` passes
with 42 tests.
- `pnpm --filter @paperclipai/adapter-utils typecheck` passes.
- `pnpm --filter server typecheck` passes.
- The old module name does not remain in TypeScript or JSON files,
except for the intentional publication guard.
- CI and Greptile checks remain pending after PR creation.

## Risks

- The old duplex module subpath no longer has a compatibility shim. The
board accepted this intentional hard break.
- The new duplex module subpath stays blocked from package publication.
- The change has no runtime effect. The main risk is an incorrect
document or module reference.

## Model Used

OpenAI GPT-5 Codex, exact model ID `gpt-5`, with tool use and code
review support.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR with the documentation issue
fields
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-24 16:42:33 -07:00
Nicky LeachandPaperclip 5a1ce7aed8 fix(server): stamp built commit into service.version (#11748)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The server emits OpenTelemetry spans so operators can trace agent
work
> - Each span needs a service version that identifies the code that
produced it
> - The current service version comes from a static environment value
and can become stale after a rebuild
> - This pull request records the built commit and resolves the service
version from the build stamp, runtime Git, the environment, or an
unknown fallback
> - The benefit is trace data that identifies the correct built commit
during development and deployment

## Linked Issues or Issue Description

**What happened?**

The server used a static `OTEL_SERVICE_VERSION` value for every
OpenTelemetry span. Rebuilds could produce traces with an old commit
value.

**Expected behavior**

The server should report the built commit when a build stamp exists. It
should use runtime Git, the environment value, or `unknown` as fallback.

**Steps to reproduce**

1. Set `OTEL_SERVICE_VERSION` to an old commit value.
2. Build the server at a different commit.
3. Start the server and inspect the OpenTelemetry service version.
4. Confirm that the built commit takes precedence over the old
environment value.

## What Changed

- Add a build script that writes the short Git commit to
`dist/build-info.json`.
- Resolve `service.version` from the build stamp, runtime Git, the
environment, or `unknown`.
- Log the resolved service version once during server startup.
- Add tests for the resolution order and safe behavior without Git.
- Document the resolution order in `doc/observability.md`.

## Verification

- `pnpm --filter @paperclipai/server build`
- `npx vitest run server/src/__tests__/service-version.test.ts`
- `pnpm --filter @paperclipai/server typecheck`
- Confirm that the build stamp contains the short commit.
- Confirm that the stamp wins over the environment value.
- Confirm that a build without Git exits successfully without a stamp.

## Risks

The server now prefers the built commit over `OTEL_SERVICE_VERSION`. A
build without Git uses the existing environment value or `unknown`. The
change needs no schema migration and has a single-commit rollback path.

## Model Used

OpenAI Codex, GPT-5, tool use and code execution. The runtime does not
expose the context window size or reasoning mode.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-19 21:20:40 -07:00
Jannes StubbemannandClaude Opus 4.7 362c30ccdc feat(server): opt-in OpenTelemetry auto-instrumentation (#3735)
## Thinking Path

> - Paperclip orchestrates AI agents for zero-human companies
> - Production self-hosters increasingly expect telemetry out of the box
— Jaeger, Tempo, Honeycomb, Datadog, Grafana Cloud, Dynatrace all speak
OTLP
> - Today there is no OpenTelemetry bootstrap in the server, so
operators who want traces have to patch their fork or run a sidecar that
captures only HTTP-level info
> - An opt-in bootstrap that costs nothing when disabled is the
minimum-viable surface for this audience
> - The OpenTelemetry packages are heavyweight enough that we don't want
them in the default dependency graph — they should load only when the
operator configures an OTLP endpoint
> - This pull request adds a self-contained
`server/src/instrumentation.ts` that dynamically imports the OTel SDK
and starts it when `OTEL_EXPORTER_OTLP_ENDPOINT` is set, and is a
complete no-op otherwise

## Linked Issues or Issue Description

No existing issue covers this directly — feature-gap description
following the feature-request template:

**Problem or motivation**

Production self-hosters increasingly expect telemetry out of the box —
Jaeger, Tempo, Honeycomb, Datadog, Grafana Cloud, Dynatrace all speak
OTLP — but the server has no OpenTelemetry bootstrap. Operators who want
traces today must patch their fork or run a sidecar that captures only
HTTP-level information.

**Proposed solution**

An opt-in OTel bootstrap gated on `OTEL_EXPORTER_OTLP_ENDPOINT`, loaded
via dynamic `import()` only when configured, so the heavyweight OTel
packages stay out of the default dependency graph.

**Alternatives considered**

Related open PRs found during the duplicate-PR search approach
observability differently: #4894 adds OTLP instrumentation to Paperclip
core unconditionally, and #3752 proposes an observability plugin. Not
duplicates — different layering: this PR keeps the default install
dependency-free via opt-in dynamic import.

## What Changed

- New `server/src/instrumentation.ts` — opt-in OpenTelemetry
auto-instrumentation. Gated on `OTEL_EXPORTER_OTLP_ENDPOINT`. Respects
the standard OTel env vars (`OTEL_SERVICE_NAME`, `OTEL_SERVICE_VERSION`,
`OTEL_EXPORTER_OTLP_ENDPOINT`). Skips the fs/dns/net
auto-instrumentations (too chatty). `sdk.start()` is wrapped in
try/catch so a bad endpoint or missing native bindings doesn't crash the
server. `process.once("SIGTERM" / "SIGINT", …)` for clean shutdown on
the first signal only. OTel packages are loaded via dynamic `import()`
so they are true optional runtime dependencies — no entries in
`package.json`, no lockfile churn.
- `server/src/index.ts` — import `./instrumentation.js` as the very
first statement so auto-instrumentation can patch `http` / `express` /
`pg` before they are evaluated by downstream modules.

## Verification

- `OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317 pnpm start` after
`pnpm install
@opentelemetry/{sdk-node,auto-instrumentations-node,exporter-trace-otlp-grpc,resources,semantic-conventions}`
in `server/` — traces show up in the configured collector; HTTP,
Express, and Postgres spans are populated.
- `OTEL_EXPORTER_OTLP_ENDPOINT` unset — server starts with no
OTel-shaped output in logs, no behavior change.
- `OTEL_EXPORTER_OTLP_ENDPOINT=…` set but packages not installed —
single `console.warn` at startup telling the operator which packages to
install.

## Risks

Low. No behavior change unless the env var is set. The bootstrap never
throws into the caller; every failure path ends in `console.warn` /
`console.error` and falls through to non-traced operation.

## Model Used

Claude Opus 4.6 (1M context), extended thinking mode.

## Checklist

- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] Thinking path traces from project context to this change
- [x] Model used specified
- [x] Tests run locally and pass
- [x] CI green
- [x] Greptile review addressed

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-12 10:44:22 -07:00