Commit Graph
11 Commits
Author SHA1 Message Date
DottaandPaperclip 8cfedd7df8 fix(runner): restore task monitors and durable timed waits (#15446)
## Thinking Path

> - Paperclip manages AI agents and their task execution.
> - Agents need a durable way to return to work after a delayed check.
> - The issue monitor scheduler already provides a one-shot wake for an
assignee.
> - Native runners reject generic execution-policy writes and had no
bound monitor tool.
> - Scheduling alone is insufficient because native completion also
needs to accept a timed wait.
> - This pull request adds an authorized monitor tool and connects it to
completion and the existing scheduler.
> - An agent can now schedule its next check, end the run, and resume on
the same task.

## Linked Issues or Issue Description

**What happened?**

A native runner could not set its own task monitor. `call_api` correctly
rejected execution-policy writes, while `schedule_wake` had no
production binding. `paperclip_finish` also rejected monitor waits.

**Expected behavior**

A standard native run can set a one-shot monitor on its current task or
another accessible task assigned to the same agent. After a confirmed
schedule on the current task, it can yield. The scheduler later delivers
`issue_monitor_due`.

**Steps to reproduce**

1. Start a standard native task.
2. Ask the agent to check the task again later and end its current run.
3. Inspect available tools and try the generic issue execution-policy
update.
4. Observe the missing native tool and the lifecycle-write denial.

Related PRs: #14680 concerns monitor notes in the shared wake prompt.
#11919 changes attempt-limit scope. This PR adds native scheduling and
completion authority and retains the existing cumulative attempt bounds.
It does not depend on either PR.

## What Changed

- Add provider-neutral `set_task_monitor` with a default current-task
target, future timestamp, required notes, existing bounds, and explicit
clearing.
- Check company, task visibility, ownership, runtime permissions, work
mode, and active-run authority. Preserve review-only restrictions.
Reject the reserved server-owned quota-recovery name before saving or
accepting a native wait.
- Commit the monitor, audit event, and retry receipt together. Retry
receipts survive a successor run without re-arming cleared or consumed
timers.
- Permit `paperclip_finish` to yield to a persisted monitor. Recheck
ownership and the schedule when committing final disposition. Release
execution without an immediate continuation.
- Preserve due monitors during native execution. Fence wake admission
and consumption against replacement, clearing, reassignment, and
completion. Preserve unrelated review policy.
- Expose scheduled and consumed monitor instructions in task context.
Update provider schemas, Rust validation, generated contracts, and
execution documentation.
- Add an opt-in live Codex smoke script with isolated data and explicit
run/session/runner/process evidence.

## Verification

- Repository `pnpm -r typecheck` and `pnpm build` passed after rebase.
Server typecheck passed again after review fixes. All CI test shards
pass on `3def77b1b`, including runner TypeScript/Rust, server,
serialized server, workspace, and browser tests. All CI gates are green,
including the canary dry run. Greptile is 5/5 on the same commit with
zero unresolved threads.
- The local monolithic `pnpm test:run`, started before the rebase, was
interrupted after current-head CI test coverage passed. It is not
counted as a standalone full-suite pass; the focused local regression
suites passed.
- Targeted server tests cover scheduling, replacement, clearing, policy
preservation, cumulative bounds, cross-run retries, permissions,
provider-neutral discovery, review restrictions, completion authority,
and scheduler/finalizer races.
- Runner contract/catalog/semantic tests and Rust terminal-tool tests
cover the new operation and monitor completion.
- Live Codex test passed twice (latest live run on `e0bcd63e6`) in a
temporary database and workspace, with a 300,000 ms warm window. First
run `67bd7709-c089-4d4a-9d2b-0d6b618a34b0` yielded at
`2026-10-07T13:15:03.274Z`. Second run
`97c323f9-595a-4cc5-a007-db5a2fbb937c` started at `13:15:30.952Z`,
received `issue_monitor_due`, and completed the same task. Exactly one
monitor wake was recorded.
- Both live runs used native session
`7b1dd753-1c9b-4e7a-b22f-a125dbc3748c`, runner
`e90d9a1b-3502-4ee3-b15e-edc024c555d4`, provider session
`01a11680-6d01-70c0-9a55-db7246ed66c3`, and PID `64218` with the same
process start time. This proves warm reuse for that local Codex test,
not only successful scheduling.
- Reproduce the paid live test with `node --import
./server/node_modules/tsx/dist/loader.mjs
server/scripts/smoke-native-task-monitor.ts --run`, with the installed
Codex binary on `PATH` and a valid local login.

## Risks

- The scheduler now defers monitor dispatch while the task has an active
native run. A stuck run still depends on the existing recovery
lifecycle.
- Idempotency uses the existing run ledger; no table or migration is
added.
- Other providers share the tested tool and completion contracts. Only
Codex received a live model test.
- Existing `call_api` lifecycle restrictions remain enforced. Monitor
waits do not bypass task blockers, reviews, or approvals.

## Model Used

OpenAI Codex, GPT-6 family, with tool use, code execution, and
TypeScript/Rust editing. The session does not expose the exact deployed
model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 08:41:24 -05:00
DottaandPaperclip b508a05c43 feat: add internal agent complaints and suggestions (#15367)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use legacy skills or native runner tools to work on tasks.
> - Those agents can encounter friction that does not belong in the task
thread.
> - A complaint should preserve the raw reaction. A suggestion should
describe an improvement.
> - This pull request adds attributed local storage and both submission
paths.
> - Agents can submit feedback once and continue their primary work.

## Linked Issues or Issue Description

**Subsystem affected**

Server, database, shared contracts, runtime skills, and native runner
tools.

**Problem or motivation**

Agents have no default internal channel for incidental complaints and
suggestions. Sending this feedback through task comments adds noise and
can alter task workflows.

**Proposed solution**

Store free-form feedback in the current instance database. Derive agent,
run, company, and task attribution from active authority. Provide
default legacy skills and provider-neutral native actions. Keep the
instructions close to Warp's MIT-licensed originals.

**Alternatives considered**

Task comments and external Slack delivery add unwanted side effects.
Mandatory suggestion fields and short editorial limits would discard
useful feedback. This release has no listing API, UI, read tool,
automatic triage, or external forwarding.

**Roadmap alignment**

This is a maintainer-requested addition to the existing runtime skills
and runner tool paths. It does not duplicate a listed roadmap milestone.
Searches for complaint tooling, suggestion-box, and agent commentary
found no overlapping public PR or issue.

## What Changed

- Add the company-scoped `agent_commentary` table, shared validation,
and idempotent migration `0310`.
- Add one transactional service and the agent-only POST route. Validate
active authority before writes or replay. Redact known credentials.
Commit a content-free audit with each new record.
- Add `submit_complaint` and `submit_suggestion` to standard, ask, and
planning modes. Keep review, revocation, and completion restrictions.
Store replay identity on the commentary row.
- Mount `complain` and `suggestion-box` by default for legacy agents.
Bundle a dependency-free Node.js stdin helper in the operational skill
and allow its POST through the sandbox bridge.
- Preserve Warp's complaint voice and suggestion guidance, with
attribution and local transport adaptations. Keep source attribution and
MIT notices in each skill's LICENSE, outside runtime instructions.
- Document custom-runtime HTTP use and database inspection. Add
real-database tests and a repeatable live Codex smoke for local and
Daytona execution.
- Pin the lagging-source migration fixture before the identity-repair
migration so later migrations preserve its regression coverage.

## Verification

- Personally ran real Codex submissions in all four environments on
2026-10-06. Local runs passed at 20:35 UTC. Daytona native passed at
20:31 UTC; Daytona legacy passed at 20:33 UTC. Each stored exactly two
rows with company, agent, run, and task attribution, wrote the
continuation marker, exited zero, created no task comments, and left
task status unchanged. Each recorded two content-free activity entries.
- Daytona used production provider hooks, real remote execution and file
transfer, the legacy queue callback bridge, and native private WebSocket
ingress. The current Linux runner was built from `abf47b595`, staged,
and verified against controller contracts. Both sandboxes were confirmed
deleted. This is a focused feedback transport smoke; it does not claim
full Runner E2E catalog or browser qualification.
- The immutable base image and Linux binary digest are recorded in [the
verification
documentation](https://github.com/paperclipai/paperclip/blob/codex/agent-commentary/doc/agent-commentary.md#verification).
The smoke script can save content-free JSON evidence. No credentials or
feedback bodies are in these reports.

| Environment | Runner | Complaint row | Suggestion row |
| --- | --- | --- | --- |
| local | legacy Codex | `59413a00-1de2-4bb1-bcc6-9c4b54c64aa6` |
`3db2364d-3e15-4f47-846f-875d3902999d` |
| local | native Codex | `5da22b5f-41df-4de5-8ba0-d9345ab01267` |
`2d5abe17-dd41-403c-a5ee-4729f2d58921` |
| daytona | legacy Codex | `27c9d0aa-8477-409f-9da0-e8ffa48dee50` |
`209681c9-d1e9-4ce1-999e-48fa07692389` |
| daytona | native Codex | `6eb001bb-4bcf-43f7-8717-f662f53dc7c3` |
`77c383d8-a997-49e5-a33e-25c70e15c0b2` |

- Run the local check with `node cli/node_modules/tsx/dist/cli.mjs
server/scripts/verify-agent-commentary-live.ts`. The documentation gives
the Daytona invocation. Both use disposable instance databases and
normal Codex provider usage.
- Repository `pnpm -r typecheck` and `pnpm build` passed after the test
extension. The build includes runner generation, contracts, and replay
checks. The smoke scripts also passed a separate TypeScript check. The
lagging-source migration regression passed. All equivalent current-head
Vitest CI shards passed. The local monolithic `pnpm test:run` invocation
was stopped after CI supplied that coverage; it did not complete
locally.
- Focused tests cover company isolation, spoofing, revoked credentials,
stale ownership, post-finish rejection, concurrent replay, conflicting
keys, atomic rollback, and deletion through existing services. Boundary
tests cover empty text, Unicode, text beyond 8,000 characters, and the
524,288-character ceiling without truncation. Mounting tests cover
Codex, Claude, and sandbox staging. Helper tests cover standalone Node
execution, stdin, invalid UTF-8, redirects, HTTP failure, and its
deadline. Privacy and bridge tests cover successful and rejected
requests.
- Instructions were compared with Warp's originals. MIT notices and
source credits live only in LICENSE files. Native tools preserve
truthful disclosure when asked, without routine announcements.
- [Full
CI](https://github.com/paperclipai/paperclip/actions/runs/37508559190)
and Greptile 5/5 passed on the earlier feature commit `5209c3501`. The
later head found the migration-fixture assumption fixed in this update.
On `8a4965164`, all 55 check contexts passed after one browser shard
rerun. Its initial reviewer signoff failure also passed an isolated
local browser run (1 test). Greptile scored that head 5/5 and identified
one smoke cleanup gap. `6ecbafb0b` fixes failed-acquisition cleanup with
four passing tests and a passing smoke-script typecheck. Fresh CI is
pending for this final test-only fix. No commentary production code
changed during verification.

## Risks

- Feedback is internally attributed. It is not anonymous. Existing
redaction removes known credentials, but agents must still omit
sensitive content. Normal provider transcripts can include their
submitted arguments.
- Default skill availability changes for existing legacy agents. Runtime
policy filtering still applies. The helper uses the existing Node.js
runtime with no extra dependencies; custom runtimes can call the HTTP
endpoint.
- Feedback is removed with its run, agent, or company. Task deletion
clears only the issue pointer. Normal database backups include the
table.
- The migration is additive and has no backfill. Writes serialize on the
active run for replay consistency. No server suggestion quota is
imposed.

## Model Used

OpenAI `gpt-6-astra` through Codex, with `xhigh` reasoning effort and a
reported 258,400-token context window. Capabilities used: repository
inspection, code execution, and live runtime verification. No subagents
were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 16:14:06 -05:00
DottaandPaperclip eb049aebf2 feat(skills): let agents update company skills safely (#15049)
## Thinking Path

> - Paperclip is an open source control plane for AI-agent companies.
> - Company skills give agents reusable work instructions.
> - Skill Studio can edit skill files and save version history.
> - Agents can create a skill, but they do not have a first-class update
tool.
> - An agent update needs a version check and safe retry behavior to
prevent lost edits.
> - This pull request adds `update_skill` through the existing company
skill file API.
> - The change keeps company policy, version history, and audit records
in one path.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: the server API, shared validation, and runner tool
catalog.

**Problem or motivation**

An agent can create a company skill but cannot update its `SKILL.md`
through a first-class tool. An unguarded retry can also create duplicate
versions or overwrite a newer edit.

**Proposed solution**

Add `update_skill` with a required current version ID and a retry key.
Route it through the existing skill file API. Reject stale versions and
changed-input retries. Save the version and audit event together.

**Alternatives considered**

A separate write endpoint would duplicate the Skill Studio mutation path
and policy checks. This PR reuses that path instead.

**Roadmap alignment**

This work extends Skills Manager and Skill Studio, which are listed in
`ROADMAP.md`.

**Additional context**

The tool accepts a complete `SKILL.md`, not a partial patch. Callers
must read the current version before they edit it.

## What Changed

- Add optional version and retry fields to the existing skill file
update contract.
- Add a guarded API update with a stable retry receipt and attributed
audit event.
- Add `update_skill` to native and semantic runner tool catalogs, with
mode and policy gates.
- Add unit, integration, protocol, and semantic-tool regression
coverage.
- Document agent use and extend the OpenAPI request contract.

## Verification

- Focused tests and direct server and runner TypeScript checks passed
before this PR.
- `git diff --check` passed after the rebase onto `master`.
- CI passed on the latest PR head, including the full test matrix,
typecheck, and build. Local full typecheck and build stopped because
`cargo` is not installed. The local full test run ended without a
verdict.
- No dedicated end-to-end eval scenario was added or run. The protocol
coverage and semantic-tool test cover the new action deterministically.
- Reviewers can read a skill version, call `update_skill`, repeat the
same key, then try a stale version and a changed-input key. Only the
first edit must create a new version.

## Risks

- File writes and database transactions must stay in sync when a write
fails. The integration tests cover failed writes and retry behavior, but
CI must verify them on the PR head.
- Existing Skill Studio callers do not send the new optional
coordination fields. Their request shape remains valid.

## Model Used

- OpenAI Codex CLI assisted with this change. The runner did not expose
the exact model ID or context window. The agent used code execution and
repository tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details; exact model ID and context window were not exposed)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused tests; full suite
is pending CI)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (53 pass, 4 skip on the latest
head)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 12:37:55 -05:00
DottaandPaperclip 018993140f feat: let agents name prompt-only tasks (#14761)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users create tasks with a title and a description.
> - A required title adds work when the prompt already explains the
request.
> - An agent can name the task once it reads that request.
> - This pull request accepts prompt-only tasks and starts them with a
short prompt slice.
> - A scoped title tool lets the assigned agent replace that slice early
without changing execution state.
> - A live browser eval checks the real agent call, saved title, audit
entry, and preservation of user titles.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: task creation, shared contracts, database, server, runner
tools, and board UI.

**Problem or motivation**

Users must currently write a title before they can submit a detailed
task prompt. The agent has enough context to write a useful title
itself.

**Proposed solution**

Make the title optional when a description is present. Save the first
120 characters of the normalized prompt as a provisional title. Ask the
assigned agent to call `set_task_title` early. Use an atomic
provisional-title guard to preserve titles supplied or edited by users.
Keep explicit titles supported.

Related: #14543 and #14556 concern empty-title submission. This change
intentionally enables that submission when a prompt is present, instead
of requiring a title.

## What Changed

- Add the `titleNeedsGeneration` field with an idempotent migration.
Keep existing titles unchanged.
- Add `PUT /api/issues/:id/title` and the native and legacy
`set_task_title` tool. Enforce company access, active-run ownership,
shared, bounded retry receipts across native/HTTP calls, and
transactional audit logging. Refresh external-object links after commit,
with the same feature gate and plugin detectors as ordinary title edits.
- Add early naming guidance in Standard, Ask, and Plan task context.
Preserve the description, status, and assignment.
- Allow prompt-only root and child task creation, plus draft restoration
in the New Task dialog. Keep user titles supported.
- Add an opt-in Product E2E suite for prompt-only Standard and Ask
tasks, plus an explicit-title control. It checks actual provider calls
within the first five tools, persisted state, audit attribution, and the
reloaded UI.
- Preserve a closed vocabulary of API key maintenance phrases in
declared prose while rejecting opaque credential suffixes. Add one
bounded naming retry after wording is rejected, without treating the
rejected call as a saved title.
- Repair the native cleanup receipt check exposed during full
verification: accept matching input digests, retain legacy input checks,
and reject conflicting receipts.

## Verification

- Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3
passed** with native Codex `gpt-5.4-mini`, first attempts only,
automatic retries disabled. Standard and Ask each saved “Rotate expired
API key” on their first tool call, with matching persisted state and a
single same-run audit entry. The explicit-title control retained its
user title with zero title writes. All three verified the reloaded
browser UI.
- Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns
are retained separately; they exposed credential-prose handling and
prompted the naming recovery fix. No failed result was regraded or
deleted.
- Reproduce with `pnpm test:e2e:runner -- --id
task-titles.runner-codex-mini.local.prompt-title-standard --id
task-titles.runner-codex-mini.local.prompt-title-ask --id
task-titles.runner-codex-mini.local.preserve-explicit-title
--max-automatic-retries 0` and an authorized provider key.
- Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit.
The runner build used the configured external eval source tree.
- Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck
and UI token gates passed.
- Title API/native regressions cover prompt-only and explicit child
creation, user edits, ownership/company isolation, external reference
refresh, cross-surface retry replay, and the 64-key limit without
receipt eviction. All passed. Prompt-context coverage: **44 tests
passed**.
- Rust credential regressions: **35 tests passed**, including benign
maintenance qualifiers and opaque credential rejection in every declared
prose field. Catalog/report reconciliation: **28 tests passed**. Native
recovery: **560 tests passed**.
- Broad local `pnpm test:run`: **14,555 tests passed** in the general
server group; two suites failed to initialize embedded PostgreSQL and
the existing 40,000-file Git streaming stress test exceeded its
300-second macOS timeout. All three suites then passed in isolation (**5
tests passed**) without code or timeout changes. The original full local
command exited nonzero and is not being represented as a clean full run.
- Latest-head GitHub checks are green: **53 passed, 4 skipped, zero
failed or pending**, including all test shards and the canary packaging
dry run. Greptile reviewed the same commit at **5/5**, with zero
unresolved review threads.

## Risks

- The additive database field must reach the server and UI together. The
migration uses `IF NOT EXISTS` and defaults existing tasks to a final
title.
- Title generation depends on the assigned agent running. Tasks without
a run keep their provisional title.
- Live qualification covers the native Codex path in Standard and Ask
modes. API/legacy and Plan behavior have deterministic coverage.
- The credential-prose exception validates the entire suffix against a
closed maintenance vocabulary. Unknown suffixes, assignments, quoted
values, credential prefixes, and diagnostics retain strict checks.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, tool use, and code
execution. The exact deployment ID and context window are not exposed in
this session. The live eval uses the native Codex `gpt-5.4-mini`
profile.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 16:48:24 -05:00
DottaandFry 3ca196b0a6 feat(agents): persist agent files across tasks without revision history (#14420)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - An agent needs personal files across tasks and sessions.
> - AGENTS.md is one file in that directory. Supporting files need the
same persistence.
> - The Instructions Editor and agent runs must share one current
directory.
> - Concurrent runs should apply only the files they change. The last
sync of the same file wins.
> - This pull request uses existing file transport and removes temporary
copies after sync.
> - Old instruction-only sessions keep their restore contract. New saves
do not create revision history.

## Linked Issues or Issue Description

Refs #14325. This replaces its revision-oriented design with persistent
agent files. Keep #14325 unmerged.

Transport prerequisite #14416 merged first at
`d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master
and remains below 100 changed files.

Related work: #4513 and #8798 cover instruction tooling. This change
handles run synchronization, cross-task personal files, browser editing,
and old-session restoration.

## What Changed

- Keep one current directory per company and agent. Point AGENT_HOME at
a temporary working copy for each active run. Keep task files and
provider HOME separate.
- Restore text, binary files, and nested folders through workspace
transport. Exclude remote agent files from task Git snapshots with a
self-ignoring file inside the reserved runtime directory; never write
through repository-controlled Git metadata.
- Collect after the provider and child processes have stopped. Keep
resumable conversation state.
- Apply changed and deleted files under the agent lock. The last sync
wins for the same file. Unrelated concurrent changes survive.
- Remove temporary copies after successful sync, rejected sync, and
staging failure. Register ownership before copying so restart recovery
can remove interrupted preparation. Retry transient synchronization up
to three times. Preserve the original remote lease reference until
deletion succeeds; restart cleanup never acquires a replacement sandbox.
Do not create captured directories or a conflict-review queue for new
runs.
- Keep browser editing, stale-draft protection, and streaming binary
downloads. Keep the instruction entry and text editor limited to 1 MiB.
- Keep historical agent-folder sync failures on their affected runs
instead of repeating them above current saved instructions. Preserve
legacy candidate review and current browser-save errors. Avoid duplicate
quota warnings while retaining separate sync failures when they describe
a different problem.
- Require target-scoped caller grants for peer instruction access, while
preserving self edits, responsible-user checks, and protected-change
consent.
- Treat full storage as a nonblocking run warning, never an agent pause
or run-admission failure. Restore already-over-quota saved folders so
ordinary agent cleanup can recover; warn on each run until cleanup. The
run detail view shows the warning.
- Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash
large files as streams. Check editor-save quotas with metadata instead
of hashing unrelated files.
- Preserve old native inputs, instruction-only copies, paths, digests,
and pending legacy candidates. Adopt old revision heads once. New writes
do not append history rows.
- Add idempotent migration 0287 and verify upgrades from the preview
tables and receipts.
- Add nine interactive stories under **Agents / Persistent files**,
including automatic incoming edits, stale browser drafts, and
storage-limit diagnostics.

## Verification

- Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after
merging current master and the landed transport prerequisite.
Integration required no manual conflict resolution; the feature remains
99 changed files. Full workspace typecheck, production build, token
gates, and 715 focused tests passed on this merge candidate. Fresh
Greptile review is 5/5 with no unresolved findings. All 55 checks
passed, with four conditional skips, including the build, typecheck,
browser E2E, and canary dry run. A single retry recovered four jobs
interrupted by runner shutdowns; no source changes were required.
- Historical-warning UI fix: all 6,834 UI tests across 640 files passed,
including regression coverage for three old failures, legacy preserved
edits, and warnings scoped to the affected run. Full workspace
typecheck, production build, Storybook build, and token gates passed.
Browser-verified Storybook playtests passed for Historical Failures
After Successful Save, Storage Limit, and Full Storage Run Warning.
- Review follow-ups at `4e20c9fb2`: all 18 focused tests passed,
including external Git directories, linked worktrees, symlinks,
hardlinks, and distinct I/O failures alongside storage warnings. Server
and UI typechecks, token gates, and the production build passed.
- Storage warning regressions at `0724f3012`: all 33 directory tests and
all five heartbeat-list tests passed, with no skips in their successful
runs. They cover repeated runs while full, an already-over-quota saved
folder, cleanup, warnings retained after unrelated save failures, and
bounded warnings in large result JSON. Server typecheck passed after the
final warning fixes.
- Full workspace typecheck, production build, and token gates passed
during this follow-up. Product E2E harness: 631 tests passed across 52
files; harness typecheck passed. Earlier native session/context and
directory/legacy collection suites passed 537 tests; Runner
unit/transport suites passed 329 tests.
- **Real E2E at `0724f3012` (before this follow-up):** legacy local
Codex and native Daytona Codex each passed six tasks, one server
restart, seven independent assertions, and cleanup verification. Both
prove browser-to-agent edits, agent-to-browser edits, nested/binary
restoration, per-file last-sync-wins, a successful run after an
oversized save rejection, and cleanup clearing the warning.
- Native local Codex also passed the six-task quota flow before the
final warning-retention fixes. That pass began at `918d1ed02` while the
bounded-result warning fix was being edited, so it is not claimed as
exact-final-head evidence. Its final-head rerun failed during embedded
PostgreSQL bootstrap before any provider run: the macOS host had 87,365
of 87,381 SysV semaphores occupied. No unrelated services or kernel
limits were changed.
- The final-source report intentionally records **2/3 cells passed**,
preserving the blocked native-local attempt:
`tests/runner-e2e/results/agent-files-quota-final-20260928-report/`.
Earlier failed attempts and provenance notes remain under
`tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and
the original campaign directories.
- Daytona used immutable image
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf`
and its exact Linux runner binary. Controller source is `0724f3012`;
image source is recorded separately.
- Legacy-session compatibility and all three ACP Stop/resume browser
regressions passed on the prior validated feature head
`169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same
provider session is retained and interrupted writes are not replayed.
Migration upgrade tests also passed earlier.
- Nine interactive stories are under **Agents / Persistent files**,
including **Full Storage Run Warning**. Its playtest and visual browser
inspection passed; the warning states that runs continue and the editor
remains available.
- Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs
skipped, no failures or pending checks. All eight browser E2E shards and
their aggregate passed. Fresh Greptile review is 5/5 with no findings;
all review threads are resolved, the security scan passed, and GitHub
reports no merge conflicts.
- The broad local follow-up test run was interrupted after host
semaphore exhaustion affected isolated PostgreSQL instances. It also
encountered the existing macOS long-path fixture failure and two timeout
failures. This is not a claim that the full local suite passed. Logs are
retained; focused storage/warning tests passed.

## Risks

- A later sync can overwrite an earlier edit to the same file, including
a saved browser edit. There is no text merge or retained version. This
is the intended last-sync-wins policy.
- A save that exceeds a storage limit is rejected and its temporary copy
is discarded. The run itself continues normally, and later runs restore
the last saved files with a warning until cleanup. Transient sync
failures get bounded retries. An I/O failure partway through a sync can
leave some files updated; a failed receipt does not claim whole-folder
success.
- Larger folders increase copy time, network traffic, and temporary disk
usage. Active runs still need working copies. Terminal runs do not
accumulate archives. Operators must provision disk for agents and
configured concurrency; these limits are not company-wide quotas.
- A restored old native session remains instruction-only until a fresh
session starts. Its original conflict fence and existing pending
candidates remain compatible.
- Provider processes close at the collection boundary. Conversation
resume remains available, but warm process reuse is lost.
- Backups must include the instance filesystem and database. External
bundles keep their existing behavior until explicitly moved to managed
storage.

## Model Used

OpenAI Codex, GPT-6 family. The session does not expose a more specific
model ID or context-window size. Reasoning, code execution, and browser
tools assisted this change. Real provider E2E uses `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing>
2026-09-29 08:25:56 -05:00
DottaandPaperclip 7944ed3d97 fix(runner): preserve hire runtime safety and first-activity timing (#13852)
## Thinking Path

> - Paperclip is the open source control plane for companies of AI
agents.
> - Native runner agents need governed tools, durable runtime state, and
useful execution evidence.
> - A first activity trace waited 53.467 seconds even though tool
activity took 6.274 seconds; provider input arrived before the server
API call executed.
> - Native agents also need a safe way to hire teammates without asking
the model to rebuild runtime configuration.
> - This pull request separates the observed ACP input-stream window
from the actual server `tool.execute` span and adds a server-owned
native hire contract.
> - The benefit is clearer latency evidence and safer native teammates
with existing approval, auth, and company boundaries preserved.

## Linked Issues or Issue Description

Related Daytona provenance work is in
[#13814](https://github.com/paperclipai/paperclip/pull/13814). No
duplicate public PR was found for this combined timing and native-hire
change.

**What existing behavior does this improve?**

Native runner agents can use governed tools and request hires. The
server did not expose a safe native hire operation that reused the
caller's validated runtime settings. First-activity traces also mixed
provider input timing with server tool execution timing.

**Current behavior**

A native hire must construct a separate runner configuration. Full
configuration copying could expose paths, instructions, secrets, or
sessions. Timing evidence could make a provider or MCP identity join
appear proven when the trace did not contain that join.

**Proposed behavior**

The native `hire_agent` operation accepts identity and persona inputs.
The server sends `adapterType: "paperclip_runner"` with
`inheritRuntimeFrom: "caller"`, then copies only validated provider,
model, permission, lifecycle, and bounded execution settings. It
inherits and validates the default environment, derives the managed AI
binding through existing normalization, preserves approval and
permissions, and creates fresh child instructions. Caller secrets,
paths, prompts, and sessions are excluded.

Provider events now include the optional boolean `inputUpdated`, with
Rust forwarding support. Timing evidence separately records the ACP
input-stream window and the actual server `tool.execute` activity. It
does not claim a provider or MCP join without matching evidence.

**Reason and benefit**

Native agents can hire teammates that start with the caller's approved
execution policy. Operators retain company boundaries, auth rules,
approval gates, and requalification. Reviewers can distinguish provider
streaming time from server API execution time when diagnosing
first-activity delays.

**Breaking changes**

None for existing hires or tool calls. `inheritRuntimeFrom` is optional
and only applies to same-company native agent callers. Conflicting
explicit runtime settings are rejected. The provider event field is
optional for existing producers.

## What Changed

- Added the native `hire_agent` protocol action, catalog entry, API
contract, and runner authority checks.
- Added `inheritRuntimeFrom: "caller"` validation and a closed native
runtime inheritance allowlist.
- Preserved managed AI binding normalization, default-environment
validation, approval snapshots, permissions, requalification, and fresh
child instructions.
- Added provider `inputUpdated` schema support and Rust forwarding.
- Added first-activity and server tool timing evidence with conservative
identity-join handling.
- Added route, authority, provider-event, sidecar, API, catalog, and
Rust-focused tests.
- Kept private Honeycomb links, raw traces, and local result paths out
of this description.

## Verification

Focused checks passed:

- 458 timing/session checks.
- 61 native hire inheritance checks.
- 20 hire authority checks.
- 1,741 API checks.
- 106 catalog checks.
- 54 provider sidecar checks.
- 12 Rust provider checks.

Live R2 and R3 each passed 45 checks across 6 runs (361,135 ms for R2).
R1 stopped at missing Docker image setup. The final trace is available
at
https://ui.honeycomb.io/paperclip/environments/test/datasets/paperclip/result/BiMypLNvmiB?tab=traces.
Latest-head CI passed all required build, typecheck, Rust, static,
Vitest, serialized-server, workspace, chat, and E2E jobs. The focused
local checks listed above passed; the broad local suite was not run
before the live evaluation, while CI provides the full repository
verification.

## Risks

- Timing fields describe separate observed windows. They do not prove a
provider or MCP owner without a valid trace join.
- The inheritance allowlist must stay synchronized with native runner
configuration fields.
- Approval snapshots include resolved safe inherited settings and should
be reviewed when native configuration fields change.
- The focused local suite is narrower than the full repository suite;
latest-head CI covers the broader repository checks.

> Roadmap review: `ROADMAP.md` places this work within Paperclip's
bring-your-own-agent direction. It extends existing native runner hiring
and observability behavior.

## Model Used

OpenAI GPT-6 (exact serving model ID is not exposed), with extended
reasoning and repository tool use; GPT-5.6 Luna assisted with focused
implementation and verification work.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 23:16:11 -05:00
DottaandPaperclip 7bc03e0acd feat(runner): default harnesses to full auto and support task reassignment (#13686)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agent Chat uses native runners to save plans and coordinate tasks.
> - Provider defaults differed across harnesses and could stop
unattended work at a second permission gate.
> - Agents also lacked a dedicated tool to move existing work to another
agent safely.
> - This change defaults native providers to full automatic permission
for provider tools and connected tools.
> - A guarded reassignment tool preserves task identity, stops the
previous run, and schedules the new owner once.
> - Codex and Claude chat acceptance tests now use production permission
defaults.

## Linked Issues or Issue Description

**Subsystem affected**

Native runner, ACPX Claude permission policy, task authority, and Agent
Chat acceptance tests.

**Problem or motivation**

A user can authorize an agent to save a plan or create a task, but
Claude's default provider gate can still stop that action. Reassignment
needs a dedicated operation that preserves context and avoids concurrent
owners or unintended recovery runs.

**Proposed solution**

Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to
`never`. Apply the defaults at configuration, execution, fresh-session,
resume, driver, and proxy boundaries. Keep explicit permission settings
and server-side company, claim, task-mode, and approval checks. Add
`reassign_task` with version checks, durable idempotency, audited
cancellation, and guarded successor scheduling.

**Alternatives considered**

A Paperclip-only allowlist still blocks provider tools and other
connections during unattended work. Full automatic permission is the
requested product default. Recreating a task discards its identity and
history. Updating assignment without stopping the previous run can leave
two agents working on the same task.

**Roadmap alignment**

This extends the existing planning, delegated work, governed tool
access, and recovery features. It adds no new service or schema
migration. Recent related tasks and open PRs were checked for duplicate
work.

**Additional context**

Related: #13678 (Agent Chat tools and recovery), #13677 (remote runner
startup). The stacked legacy-adapter companion is #13693. This also
fixes the deployed-server artifact fallback needed to stage the current
runner binary.

## What Changed

- Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex
to `never`, including missing settings at direct driver and proxy entry
points. These defaults cover provider tools and connected tools.
Preserve explicitly configured restrictive modes.
- Include assigned approval reads using canonical side-effect
classifications, so verifying a recorded approval does not trigger
another provider gate. Paperclip approval decisions still enforce
controller authority.
- Carry the new permission mode through server configuration, execution
contracts, recovery identity, TypeScript, and Rust. Keep
`approve-paperclip` as an optional restricted mode, with exact SDK rules
and closed unknown requests. It is not a default.
- Add `reassign_task` to the semantic catalog, controller, mock
authority, and generated contracts.
- Guard reassignment with company authorization, expected owner and
version, protected-state checks, and durable retry receipts.
- Honor explicit backlog task creation atomically with the initial plan,
without scheduling a wake. Preserve backlog holds regardless of
dependency readiness.
- Stop active work before changing ownership. Restore the prior owner
through a guarded, idempotent wake if final handoff validation fails.
Keep intentional reassignment stops out of failure recovery. Preserve
backlog and blocked states without waking them early.
- Add authorization, concurrency, replay, stop, and permission boundary
regressions. Add Codex and Claude chat reassignment cases and run native
chat cases with production defaults.
- Clarify shared runner guidance: save plans and Paperclip documents
directly with `write_document`; create and register a local file only
when a downloadable file is requested.
- Document provider defaults and the operator choices for existing
agents.

## Verification

- Current head `d82fbb0f03546d27cecf072250e4172e0b1ee662`: **55 checks
passed**, with two intentional skips. [PR
checks](https://github.com/paperclipai/paperclip/pull/13686/checks).
- Greptile reviewed that exact head at **5/5**. The security reviewer
acknowledged the intended full-auto default, and the acknowledged
discussions are resolved.
- Full workspace `pnpm -r typecheck` and `pnpm build` passed locally
after rebasing onto current master. Targeted adapter/server, runner,
API, default/resume, and heartbeat configuration tests passed.
- **All six real-provider acceptance cases passed on their first
attempt, with cleanup passing:** plan handoff, task reassignment, and
backlog creation/status, each on native Claude and Codex. Evidence
records Claude's effective `approve-all` mode. [Campaign and
downloadable
evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548).
- The live campaign tested combined revision
`a37881c824dcd7170380fc4b788732fc743e5da7`. The final PR heads add only
a heartbeat test expectation correction; application code is unchanged
from that live-tested revision.
- The campaign's result-enforcement job passed. Its separate report
publisher failed because the trusted workflow's `patchedDependencies`
configuration differs from its frozen lockfile. All six results and
screenshots remain available as GitHub artifacts. The overall manual
workflow is red for this publishing failure.
- Full-suite coverage is supplied by the passing CI partitions. The
separate unsharded local run was stopped after the corresponding CI
partitions passed; it is not counted as a completed local run.
- Reassignment tests cover stale state, cross-company access, denied
authority, cancellation failure, compensating wake, and idempotent
retries. Backlog tests verify the original creation audit, saved plan,
exact task count, and absence of task-bound runs.

## Risks

- Agents with no explicit permission mode now receive full provider tool
permission, including connected tools. This is a deliberate broad
default. Existing explicit restrictive modes still apply. Controller
authorization, company isolation, workspace boundaries, and Paperclip
governance remain in force.
- Reassignment crosses run cancellation and task ownership transactions.
Durable stop intent, revalidation, audit receipts, and guarded queue
dispatch cover interruptions and retries.
- The new permission enum requires a current runner artifact. The remote
artifact fallback uses the same resolved controller binary for upload
and execution.
- Live provider behavior remains subject to the selected model. Targeted
live results do not qualify the full catalog.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The exact deployment model ID and context-window size are not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 16:49:18 -05:00
DottaandPaperclip 9fd2e50310 feat: create company skills from runner tasks (#13538)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Runner gives agents tools to change company resources.
> - Users need agents to save reusable skills during a task.
> - A saved skill needs a visible result that users can inspect and
edit.
> - This pull request adds `create_skill` and a task feed card linked to
Skill Studio.
> - Users can open the saved skill from the task and edit the same
resource.

## Linked Issues or Issue Description

**Subsystem affected**

Runner tools, company skill storage, task feed, and Skill Studio.

**Problem or motivation**

The Runner has no dedicated tool to create a company skill. A user
cannot follow a creation result from the task feed to the saved skill.

**Proposed solution**

Add a company-scoped `create_skill` tool. Save the skill with the
existing company policy. Add one creation card to the task. Open a named
sidebar tab from that card. Let the user open the same skill in Skill
Studio.

**Alternatives considered**

An agent can write a local file, but that file is not a company skill. A
second document copy in the task would become stale after a Studio edit.
The sidebar therefore reads the saved skill directly.

**Roadmap alignment**

This extends the shipped Skills Manager, Skill Studio, and Skills Store
milestone. The maintainer requested and approved this scope. Search
found no duplicate `create_skill` PR or issue. Related UI validation
work: #8715. This PR does not change that validation display.

## What Changed

- Add the real Runner tool, its contract, and its mock implementation.
- Validate the complete SKILL.md and derive company, task, agent, and
run identity from authentication.
- Apply the existing company skill policy. Do not assign the skill to an
agent.
- Make keyed retries return one skill and one creation event. Reject
conflicting retries.
- Make concurrent file creation safe. Never replace an existing
published skill during creation.
- Add a creation card, a named sidebar tab, and an Open in Skill Studio
action.
- Show saved Studio edits when the user returns to the task.
- Add storage, policy, mode, retry, UI, and Product E2E tests. Document
the tool.
- Fix deleted-name reuse, onboarding panel persistence, immediate feed
refresh, and mock validation parity from review.
- Serialize Studio file edits and renames with skill deletion and
recreation. Reject stale editor requests before they can change a
replacement skill.
- Generate the standalone mock parser and validator from the production
contract. Use portable UUIDs so the browser scenario bundle builds.

## Verification

- All latest-head PR checks pass on `145dd76a5`, including all server
shards, browser E2E, Runner verification, build, typecheck, and release
dry run. Greptile: 5/5 with no open findings. An interrupted CI runner
was retried successfully.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm check:token-gates`: passed.
- Review regressions: 73 storage tests, 6 real API tests, 63 UI tests,
and 61 semantic runtime tests passed. Parser synchronization passed.
- CI exposed existing fire-and-forget Sentry test races. Reproduced the
resumption race locally, then synchronized the related sweep and
finalizer assertions on the actual report; all 27 tests across the three
affected files pass.
- Runner scenario browser build and strict content-security-policy
check: passed.
- Runner suite: 2,012 tests passed; 10 skipped.
- `pnpm test:run`: the general-server batch had 12,416 passes and two
failures. The old tool-count assertion was fixed; all 16 authority tests
then passed. The chat webhook test had a socket error; it passed four
isolated reruns.
- Both workspace test groups passed. The isolated route suites
completed. Two socket failures in the initial route batches passed on
individual reruns; all remaining 61 files passed.
- Product E2E `create-skill-studio`: passed with local Codex and local
ACPX Claude.
- Manual browser test: submit a task, observe the real tool call and
creation card, open the sidebar, edit in Studio, save, and return. The
task reached Done. The saved second revision and sidebar tab survived a
server restart.
- The new companion headless Runner Eval passed. Companion coverage PR:
https://github.com/paperclipai/paperclip-evals/pull/23. Daytona was not
run because no immutable runner image was configured.

## Risks

- Database writes and local file writes cannot share one transaction.
Recovery accepts only an exact file-for-file retry after a database
rollback. Conflicting files remain untouched.
- The sidebar displays the current skill. The feed card remains the
historical creation receipt.
- No database migration, dependency, or workflow change is included.
- Remote Daytona behavior still needs a run with a configured immutable
image.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) handled design, integration, review, and
browser verification. OpenAI `gpt-5.6-luna` assisted with bounded
implementation and eval work. Both used code execution and tool access.
The host did not expose the context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 13:01:58 -05:00
DottaandPaperclip ab15aff390 feat: add experimental persistent agent chat (#13284)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Conversations must use the same tasks, controls, and execution
history.
> - Users need an ongoing chat with an agent without managing task
properties.
> - Agents should clarify and plan work, then hand execution to assigned
project tasks.
> - This pull request combines the reviewed Agent Chat stack for one
squash merge.
> - The benefit is persistent conversation with normal task governance
and shared UI.

## Linked Issues or Issue Description

**Subsystem affected**

Task lifecycle, agent runtime tools, shared task UI, and browser/paid
runner tests.

**Problem or motivation**

Users need one persistent conversation with each agent. A separate chat
store or renderer would duplicate task behavior and bypass existing
controls.

**Proposed solution**

Use a task-backed chat per company, user, and agent. Reuse the task
composer and transcript. Clarify and plan in chat, then create assigned
project tasks with the relevant plan. Keep Agent Chat behind its own
disabled-by-default experimental setting.

**Roadmap alignment**

This implements the task-backed direction in [CEO
Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat).
Related proposals: #2504 and #9693. Related request: #7981. The
maintainer requested one squash merge of the complete stack.

Consolidates the reviewed runtime
[#13281](https://github.com/paperclipai/paperclip/pull/13281), backend
[#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI
[#13283](https://github.com/paperclipai/paperclip/pull/13283) layers
with this PR's E2E coverage. All four layers passed CI and received
Greptile 5/5 before consolidation. This PR targets master and includes
the complete feature.

## What Changed

- Add personal canonical chat tasks with ordinary company visibility,
immutable identity, idempotent first sends, and an idle waiting state.
- Process `/new` in queue order. Preserve history, release a chat pause,
and fence old provider context and delayed writes.
- Keep chat lifecycle rules across recovery, finalization, assignment,
task lists, and rollups.
- Support research and plan revision in chat. Hand plans to ordinary
assigned project tasks before execution starts. Reject new chat
subtasks.
- Add repository-aware project creation and discovery tools, including
multiple repository IDs and GitHub URLs, authorization, idempotency, and
durable project-created cards.
- Reuse task UI components for chat, with starred/recent agent
navigation and a separate `enableAgentChat` experimental flag.
- Add deterministic browser tests and 24 paid chat cells across four
Codex/Claude profiles, with validated reports and screenshots.
- Integrate current master recovery, controller lease, queued-message,
and task UI changes. Gate chat interruption and deferred promotion on
ownership/feature policy. Guarantee lease renewal and active controls
are stopped even if teardown fails.
- Preserve master's migration 0273 and generate chat migration 0274 with
idempotent replay for development databases.

## Verification

- Prior exact heads of all four PRs passed Linux CI, including build,
typecheck, general/serialized tests, and browser E2E. Each had Greptile
5/5 and no unresolved findings.
- Integrated local verification passed: full repository typecheck and
production build, Storybook build, token gates, 340 focused UI tests,
all 20 deterministic chat browser tests, two migration replay tests, 88
focused chat/queue/native/controller tests, and provider/session
regressions including real lease expiry. These include the three
lifecycle regressions for the final admission/teardown fixes; server
typecheck also passes. Current head
`1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no
unresolved findings and passing security scans. All final-head CI gates
passed: build, full Runner verification, typecheck/release registry,
canary, all general/serialized test shards, and all browser E2E shards
([CI
run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)).
Local PostgreSQL startup contention required serialized retries; skipped
fixtures do not count as passing coverage.
- The earlier paid campaign passed all 24 chat cells and retained 32
screenshots:
[report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat).
It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior
evidence, not a paid run of this integrated head.
- Manual check: enable Agent Chat in Experimental settings, open an
agent, clarify and revise a plan, then hand off to an assigned project
task. Stop a reply, send `/new`, and verify fresh context with retained
history. Disable the setting and verify agent shortcuts/new chat turns
are blocked.

## Risks

- Queue/session integration can affect retries and delayed writes. Tests
cover ownership, cancellation, reset boundaries, idle recovery, and
ordinary task behavior.
- Migration 0274 adds conversation fields and constraints. Replay is
idempotent and preserves existing development chat history.
- This combines the previously reviewed stack at the maintainer's
request. Agent Chat remains off by default and is separate from
Conference Room.

## Model Used

OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code
execution, browser tools, and parallel review. The exact context-window
size is not exposed in this session. Codex and Claude also ran as test
subjects in the linked paid campaign.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 08:56:04 -05:00
DottaandPaperclip 5bddff0920 feat(runner): add guarded API search and call fallback (#13003)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - The new runner gives agents dedicated tools for common tasks.
> - Some API operations and parameters have no dedicated tool.
> - Agents need a controlled way to find and use those operations.
> - This pull request adds API search and calls through the real server
routes.
> - Existing tools remain the preferred path. The new tools are disabled
by default.
> - Paired tests measure correctness, tool choice, cost and time.

## Linked Issues or Issue Description

**Subsystem affected**

Paperclip Runner contracts, production tool authority and the server API
catalog.

**Problem or motivation**

The runner cannot use much of the API described by the old Paperclip
skill. A generic HTTP client would also let agents bypass runner control
rules.

**Proposed solution**

Add `search_api` and `call_api`. Resolve calls from the mounted API
catalog. Use server-held, run-bound credentials. Preserve route checks
and runner lifecycle rules. Keep the tools disabled until an operator
enables selected companies.

**Alternatives considered**

A dedicated tool for every endpoint would add a large initial prompt. An
unrestricted HTTP tool would weaken authorization and replay controls.

**Roadmap alignment**

This extends the native runner tooling. The repository owner requested
this design and implementation. The roadmap and related open PRs were
checked. No duplicate API escape-hatch PR was found.

## What Changed

- Register two compact fallback tools in canonical contracts and
provider projections.
- Build deterministic API discovery from OpenAPI, mounted experimental
routes and the old skill reference.
- Execute bounded JSON, text, file and download requests through
authenticated HTTP routes.
- Recheck active runs, company access and work modes. Block runner
lifecycle, scheduling, credential and approval bypasses. Keep routine
annotation collaboration available.
- Retain mutation receipts. Report uncertain outcomes without blindly
repeating writes.
- Add a company rollout gate and a durable eval worker with complete
cost accounting checks.
- Record child-task creation in the activity log with the agent and run.
- Add contract, authorization, file, replay and real runnerd/PRP/HTTP
tests.
- Document rollout gates and paid coverage limits. The companion eval
repository retains immutable attempts and reports.

## Verification

- Final app commit `da58370524c3626a744eec20164397c5fb6ba9ef`: all 32
checks passed; the unrelated Storybook visual check was skipped.
Greptile 5/5; no unresolved review threads.

- Full Linux build and recursive typecheck passed. Repository tests were
run by project and serialized shard; all 143 serialized server suites
passed.
- Runner TypeScript: 1,599 passed, two skipped. Rust release: 451
passing test reports. Conformance and replay parity passed. The required
API check passed 837 tests, including runnerd → PRP → authority → real
HTTP.
- Bindings cannot enable API tools without the explicit deployment flag.
Unit and real-authority tests prove the default-off boundary.
- The standalone API check builds and stages its own binary. It passed
after existing staged and debug binaries were removed from the test
container.
- UI and CLI tests passed. Initial environment failures (missing jq,
Docker overlay file identity, and parallel linker memory pressure) and
focused passing reruns are retained. The macOS full runner suite has
platform-specific failures; Linux is the qualified full-check platform.
- Eval harness: 27 tests passed; existing CI discovery ran 86 tests with
two unrelated skips. Credential export rejection is tested against the
actual report command.
- Luna and OpenRouter Sonnet each passed 60 common-workflow runs: ten
workflows, three repetitions per arm, zero unnecessary API fallback.
- Sonnet passed 11 selected capability/contract cases after fixes.
Gemini passed three smoke cases. DeepSeek exceeded the 120-second limit
and remains unqualified.
- Luna's two cost flags received focused follow-up. The original flags
and a later n=1 latency flag remain visible. Sonnet had no cost or
latency increase above 20%.
- The catalog contains 785 entries; 58 were exercised across all stages.
Most operation probes remain unrun and some need additional fixtures.
Authored probes do not establish successful coverage.
- Total conservative accounted cost: $9.875960. Active paid-campaign
time: 88.16/90 minutes. No missing accounting. Later security and
harness fixes have provider-free verification; no paid validation is
claimed for those revisions.
- Inspect the [qualification
report](https://github.com/paperclipai/paperclip-evals/blob/codex/seach-call-api-tools/evals/runner-api-tools/reports/2026-09-07-production/READINESS.md)
and [verification
record](https://github.com/paperclipai/paperclip-evals/blob/codex/seach-call-api-tools/evals/runner-api-tools/reports/2026-09-07-production/verification.json).

## Risks

- This is a broad authenticated API surface. Keep the default-off gate
until an operator selects initial rollout companies.
- Paid coverage is incomplete. Small regression samples do not prove all
workflows are unchanged.
- A timeout or server failure can follow a committed mutation. The
result reports an unknown outcome and requires state inspection.
- The new definitions add prompt tokens. The report retains cost flags
and cache variation.
- No database migration is required.
- Repository rules require code-owner approval before merge. Technical
CI and automated review are complete.

## Model Used

OpenAI Codex based on GPT-6 assisted with code, tests and review. The
exact serving model ID and context window are not exposed in this
session. It used reasoning, tool calls and code execution.

Eval models: `gpt-5.6-luna` with low reasoning,
`openrouter/anthropic/claude-sonnet-5`,
`openrouter/google/gemini-3.8-flash`, and
`openrouter/deepseek/deepseek-v4-flash-0731`. Attempts retain runtime
versions, model identity, usage and source provenance.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-07 14:14:43 -05:00
Dotta 560e7e48b5 feat(runner): add SDK and developer tooling (#12608)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The runner package already provides the production protocol and
execution spine.
> - Contributors still need stable SDK surfaces, deterministic test
tools, and local inspection tools.
> - Those surfaces share generated contracts and must change as one
package boundary.
> - This pull request adds the package-local SDK, labs, examples, and
drift checks.
> - The benefit is a reviewable developer platform that does not change
application execution selection.

## Linked Issues or Issue Description

**Subsystem affected**

`packages/paperclip-runner` — runner SDK, conformance tools, and
developer tooling.

**Problem or motivation**

The production runner spine is present, but package consumers cannot
build deterministic integrations, inspect sessions, or verify
provider-neutral behavior through supported surfaces.

**Proposed solution**

Add browser, React, standalone, live-session, scenario, conformance, and
evaluation surfaces. Add generated contract inventories and
package-local verification scripts. Keep production application routing
unchanged.

**Alternatives considered**

We considered splitting each generated catalog, SDK surface, and demo
into separate pull requests. Those changes share exports, fixtures, and
drift gates. Splitting them would create intermediate package states
that do not build.

**Roadmap alignment**

No overlapping item appears in `ROADMAP.md`. This work extends the
runner package that is already on `master`.

## What Changed

- Add browser, React, standalone, live-session, and issue-thread SDK
surfaces.
- Add deterministic mock control-plane, scenario, conformance, replay,
and evaluation tools.
- Add bounded Codex, OpenCode, and ACPX development transports and
fixtures.
- Keep deferred managed-provider execution fail-closed. Persisted
compatibility data remains readable.
- Add generated capability inventories with their source files and drift
checks.
- Add examples, package documentation, browser checks, and
clean-consumer checks.
- Preserve the reviewed protocol bounds, replay compatibility aliases,
process environment isolation, and semantic redaction limits.
- Update the ACPX package patch that the existing workspace patch
registry already tracks.
- Do not change `pnpm-lock.yaml`, repository workflows, server runtime
selection, or the application UI.

## Verification

GitHub Actions is the verification authority for this pull request. The
repository CI, package TypeScript and Rust checks, package tests,
generated-output drift checks, browser checks, security scans, and
Greptile review must pass on the exact head.

Local test suites were not run because this series uses parallel GitHub
Actions for verification.

## Risks

This is a large greenfield package change. The main risks are public
export drift, generated-output drift, and optional React consumer
compatibility. Package boundary checks, clean-consumer checks, and
browser tests cover those risks. Production adapter selection and server
execution are outside this pull request.

## Stack

1. **This PR:** runner SDK and developer tooling.
2. [Codex production server
integration](https://github.com/paperclipai/paperclip/pull/12616).
3. [Provider-neutral task-thread
UI](https://github.com/paperclipai/paperclip/pull/12617).

## Model Used

OpenAI Codex, GPT-5, high-reasoning mode, with tool use and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR following the feature request
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-31 21:33:11 -05:00