mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 05:41:56 +02:00
## Thinking Path > - Paperclip manages agents, tasks, permissions, and execution budgets. > - Native tasks can resume after a connection decision in the same provider conversation. > - Each turn still needs an accepted completion report. > - Compact continuation messages did not explain that reports from earlier turns cannot finish the new turn. > - The connection evaluator could also grade before the final reply was stored or reject valid unavailable-access wording. > - This PR clarifies the current-turn report requirement and fixes those observation boundaries. > - The original connection instructions and strict native completion gate stay in place. ## Linked Issues or Issue Description Refs #15489. The reduction remains draft while this separate repair is qualified. Refs #15471 for the earlier connection continuation work. ## What Changed - Add a current-turn completion reminder to compact continuation inputs. - Keep final prose insufficient for completion. Preserve permissions and retry policy. - Wait for the final successful task run's saved, attributed decline reply within the existing deadline. - Use one bounded explanation matcher for both decline checks. - Wait for a recorded tool-action rejection to dispatch its bound continuation, with strict company, task, agent and source-run checks. - Retain the exact grading input before later API refreshes. - Add failure and delay regressions and update the Runner and evaluator docs. ## Verification The fresh comparison has **15/15 original passes on each variant**: 15 unchanged pass pairs, zero new failures, and no pending pair. There are 30 case attempts and **65 actual agent runs** (baseline 33; candidate 32). All runs succeeded. All 30 cleanups passed. No model attempt was retried. | Profile | Baseline | Candidate | | --- | --- | --- | | Native Codex `gpt-5.6-sol` | 5/5 | 5/5 | | ACPX Claude `claude-sonnet-5` | 5/5 | 5/5 | | OpenCode `openrouter/deepseek/deepseek-v4-flash-0731` | 5/5 | 5/5 | Each profile covers service approval, service decline, connection decline, provider decline, and selection of the second provider. The saved replies, approved briefings, decisions, fixture observations, final task states, and native completion records were inspected. All 18 saved decline-grade snapshots match their original captured inputs and checks. Result, API snapshot, and final ledger run sets agree. - [Candidate campaign](https://github.com/paperclipai/paperclip/actions/runs/37711658378) · [public candidate report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37711658378-1/index.html) - [Baseline campaign](https://github.com/paperclipai/paperclip/actions/runs/37711675579) · [public baseline report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37711675579-1/index.html) - Candidate source: `79905343bba280d462765faad19a26e7f179259e`. Baseline source: `7c5e120158f1385a1fc5f20be41f66c58a605534`. Both use master context `fc6304dfe5f2e446e09bd052a7b45f51e930f250`. The trusted workflow source is separately frozen at `dd777f4b7343305c4e6f44c422f44a1d78e12e4f`. - Both variants use the same evaluator, fixtures, models, permissions, 720-second cell deadline, and 1,000-cent company and agent hard stops. The only production difference is the compact continuation reminder. - Suite hash: `aee30b74b4d38ada08777798db0932fbc64e368bb427d29caa8cfd86c7f59747`. Definition hash: `ea9e17f54fe0af1acbb2337ddfab8a7e92db8488bd0f59510a63095ca0229b60`. - Provider-free transport capture: startup/resume input stays at 55,726 bytes. Compact continuation input grows from 53,334 to 53,535 bytes. The 42-tool catalog stays unchanged. These are Paperclip input bytes, not complete vendor prompt tokens. - Focused evaluator tests: 31 pass. Native contract, transport delivery, and session tests: 194 pass. Evaluator support: 1,819 TypeScript tests and 128 Node tests pass, with one intentional skip. - Full build, workspace typecheck, and evaluator typecheck pass. Current-head CI passes all required gates. The current rollup has 51 successful check runs, two intentional Storybook skips, and a successful Snyk status. Review is 5/5 with zero unresolved threads. - CI attempt 1 had one initial runtime-fixture health timeout. The exact test and its full 164-test file pass locally. One CI shard retry passed. The original CI failure, its dependent verify failure, and the retry remain visible in [CI history](https://github.com/paperclipai/paperclip/actions/runs/37711235060). - The broad local `pnpm test:run` attempt was interrupted after about 49 minutes (exit 130). It recorded one failure in the unchanged Zep memory-connector disabled-setting test. That test and the full 388-test tool-access file pass in separate local checks; current-head CI also passes. The local cause is not established, and this broad local attempt is **not** claimed as passing. Three earlier local failures also pass in their isolated checks; their original logs remain retained. - The first two setup admissions were cancelled before provider jobs to include the review correction. They made no provider calls. The completed campaigns above are the first and only model attempts for these corrected variants. Cost evidence stays separate from behavior. Original result summaries report only OpenCode amounts: baseline $0.039192096 and candidate $0.063221620. Final run ledgers also retain estimates for Claude (baseline $1.232439000; candidate $1.333611200) and Codex (baseline $2.340324400; candidate $2.129335600). These estimates do not replace the original summaries. Local and GitHub runtime are unmetered here. Invoices are unknown. This is not a cheaper or faster claim. ## Risks - A single matched trial cannot prove general equivalence or causation. The reminder is an instruction change, not a new completion enforcement rule. - Candidate OpenCode service-decline finished within one continuous run; its baseline used two. That pair passed the task outcome, but it does not qualify the reminder on a resumed decline turn. No extra paid run was used to replace it. - The text matcher is bounded evidence of an explanation. It does not prove reasoning or consumption of feedback. Bounded stdout excerpts do not prove that every extra attempted tool call is absent. - Missing saved replies or continuations still fail at the original deadline. Failed native completion remains a failure even when final prose is correct. - The unresolved local-suite discrepancy above remains a validation limit. Full remote CI and both focused local reproductions pass. - The original connection instructions stay in place. These results do not qualify the reduction in #15489. Its original 11/15 versus 12/15 grades and two new failing pairs remain unchanged. ## Model Used OpenAI Codex, based on GPT-6. The exact deployment ID and context window are not exposed in this session. Capabilities used: reasoning, repository editing, code execution, test inspection and eval analysis. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the focused local tests listed above and they pass; the interrupted broad local run is disclosed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
263 lines
16 KiB
Markdown
263 lines
16 KiB
Markdown
# Everyday Paperclip workflow evals
|
|
|
|
This manual suite tests useful work through the production browser, public API,
|
|
native runner, and normal agent instructions. It complements the tightly
|
|
scripted runner contract fixtures. It does not add a scheduled or default paid
|
|
run: `--all` and generic profile selectors exclude it. Select the suite or an
|
|
exact execution ID explicitly.
|
|
|
|
## Stories and assertions
|
|
|
|
| Story | Cases | Required evidence |
|
|
|---|---|---|
|
|
| Build a small project and revise it | `build-revise` | Download both ZIPs through the UI; independently execute the delivered CLI and import its function; test the revision; retrieve the original bytes again. |
|
|
| Delegate and incorporate late feedback | `delegate-feedback` | One child assigned to Riley; send feedback while the child runs; find it in the child history and independently test `--max-length` in the delivered ZIP. The worker must not execute on the parent. |
|
|
| Hire a teammate and use them again | `hire-reuse` | One Morgan QA reporting to the lead, native runner and the same encrypted connection bindings, real child execution, then a second usable delivery from that same agent. |
|
|
| Decide on an installed service action | `service-approve`, `service-decline` | Assign an authenticated local MCP fixture with Ask first; match its tool action and connection ID; no provider call before approval; exactly one after approval and a verified document; none after decline. |
|
|
| Decline a new connection | `connection-decline` | Start without service connections; match a Notion connection intent; click Not now; verify the saved rejection, no new connection or repeated request, and an explanation followed by Done. |
|
|
| Request an email address | `agentmail-setup` | Enable Chat connectors, ask for an email address, and require a durable AgentMail card for the requesting agent and user. Reload and verify one password field, the direct API-key link, and no access selectors or modal. Decline and verify no connection or repeated request. |
|
|
| Continue work after a controller restart | `recover-controller` | Observe saved source, persist a user message, restart the isolated controller, and independently test the delivered result. |
|
|
| Stop work and change direction | `stop-redirect` | Click Stop, send one new request, reload, observe exactly one stored user message and the new answer, and reach Done. |
|
|
| Create and edit a company skill | `create-skill-studio` | Create one skill through the runner, verify its persisted library entry and activity-feed card, open Skill Studio, save an edit, and verify the edit after returning. Local Codex, local ACPX Claude, and warm Daytona cells are explicit. |
|
|
|
|
For normal completion, all story tasks must reach Done, with no active run,
|
|
pending completion confirmation, or scheduled recovery. Runs must prove native
|
|
identity and native terminal contracts. A workspace-contention cancellation is
|
|
not provider execution only when the persisted pre-dispatch record explicitly
|
|
says `providerWorkStarted: false` and no process/session/runner identity exists.
|
|
Other unexplained cancellations remain failures. Twelve total run records bound
|
|
each story, including contention and recovery.
|
|
|
|
Arbitrary runner-process termination is not part of the model scorecard. The
|
|
historical `recover-runner`, `recover-runner-safe`, and `recover-runner-uncertain`
|
|
attempts remain available as diagnostics, with their original grades and costs.
|
|
The first two did not establish a safe restart boundary, and the uncertainty
|
|
case measures a deterministic safety rule. None supports ranking models.
|
|
See [controlled recovery tests](../runner-recovery/README.md).
|
|
|
|
## Matrix and running
|
|
|
|
The local matrix has fourteen cases on native Codex `gpt-5.6-sol`, native ACPX Claude
|
|
`claude-sonnet-5`, and native Codex `gpt-5.4-mini`: 42 cells. The two core profiles
|
|
also declare build/revise, delegation, controller-restart, and skill-creation cases
|
|
on Daytona: eight cells. OpenCode adds only local hiring/reuse and delegation,
|
|
for 52 cells total. Hiring/reuse and delegation have one attempt and a
|
|
1,000-cent company and lead-agent hard stop. Remote runner-process killing is not supported. For remote controller
|
|
restart, a verified first download supplies the persistence checkpoint; the
|
|
controller is interrupted during a subsequent revision with another queued
|
|
requirement.
|
|
|
|
```sh
|
|
pnpm test:e2e:runner -- --list --suite everyday-workflows
|
|
pnpm test:e2e:runner -- --suite everyday-workflows --environment local --max-parallel 2
|
|
pnpm test:e2e:runner -- --id everyday-workflows.runner-codex-mini.local.build-revise
|
|
pnpm test:e2e:runner -- --suite everyday-workflows --environment daytona --max-parallel 2
|
|
```
|
|
|
|
Before project stories or the Python calibration tests, start Docker on the
|
|
harness host and fetch the pinned oracle image. CI prepares and verifies this
|
|
same pinned image before the paid project-story cells; artifact checks run on
|
|
the harness host. The workflow verifies the exact repository digest after the
|
|
pull. This is required for local and Daytona stories.
|
|
|
|
```sh
|
|
docker pull python@sha256:9d2e5553305c7c7b0097999bb17187c69b921ccd6bc9d40e4bb5ebe652c00285
|
|
python3 tests/runner-e2e/everyday-artifact.py --preflight
|
|
```
|
|
|
|
The harness checks this prerequisite before it creates the task. It does not
|
|
pull an image during a model attempt or fall back to host execution.
|
|
|
|
Use the credential and immutable Daytona image setup in [README.md](README.md).
|
|
Provider calls cost money. Each cell owns an isolated instance and project.
|
|
There are no real third-party mutations in the service fixture; it exercises
|
|
production connection, transport, tool approval, and document delivery paths.
|
|
|
|
## Deterministic checks and calibration
|
|
|
|
```sh
|
|
pnpm test:e2e:runner:typecheck
|
|
pnpm test:e2e:runner:unit
|
|
python3 -m unittest discover -s tests/runner-e2e -p test_everyday_artifact.py
|
|
```
|
|
|
|
The independent oracle rejects wrong output, ignored late feedback, trailing
|
|
separator bugs, invalid argument acceptance, duplicate source modules, archive
|
|
path traversal, and symlinks. Passing agent-authored tests cannot override it.
|
|
Lifecycle calibration rejects legacy execution, missing runner identity,
|
|
unexpected crashes, workers on the parent, and answers left in review.
|
|
|
|
ZIP evaluation runs delivered Python in a Docker container with a read-only
|
|
project mount and root filesystem, no network, a non-root user, no Linux
|
|
capabilities, and bounded CPU, memory, process count, output, and duration. Only
|
|
the extracted delivery enters the container. The container is removed after
|
|
grading. Calibration includes attempts to read a host file and reach a host
|
|
loopback service.
|
|
|
|
## Evalbook evidence and qualification
|
|
|
|
Each packaged attempt retains `snapshots/everyday-workflow.json`, downloaded
|
|
ZIPs, assertions, actual task comments and run records, timing, accounting,
|
|
source provenance, and screenshots. The story records a digest of its harness
|
|
sources. Infrastructure failures and failed attempts must remain inspectable.
|
|
|
|
Import packaged results with `paperclip-evals/evals/everyday-workflows/import_results.py`.
|
|
It uses the canonical Runner Evalbook generator and the built Runner Lab viewer.
|
|
It does not invent provider transcripts, tool counts, model observations, or
|
|
cost estimates. The selected model is checked against persisted native execution
|
|
inputs; that is distinct from provider-side model identity verification.
|
|
|
|
Initial live results are diagnostic. They are not a reliability estimate or a
|
|
model ranking. Before promotion, freeze both source revisions and harness
|
|
digest, run at least three independent local repetitions, qualify the eight
|
|
remote cells against a verified image, and review every failure. Keep model
|
|
quality, lifecycle correctness, infrastructure availability, and latency separate.
|
|
|
|
## Revised evaluation contract (14 September, second campaign)
|
|
|
|
That campaign used 32 cells: the original local stories plus two local Codex
|
|
text-only safe-replacement probes, and the unchanged six remote cells. The old
|
|
`recover-runner` results remain historical; `recover-runner-uncertain` is a new
|
|
case that expects a visible Blocked safety stop, preserved source and queued
|
|
input, and no unverified provider replay. Its Retry control is inspected, not
|
|
claimed to restore work. Successful manual recovery remains unqualified.
|
|
|
|
`recover-runner-safe` interrupts a text-only Codex turn and queues new direction.
|
|
A pass requires the server's durable `verified_safe_replacement` evidence and the
|
|
new answer. No safety proof is injected or fabricated. If that premise cannot be
|
|
verified in a live probe, report it as an unqualified recovery boundary, not an
|
|
established product defect. Claude has no catalog cell for this Codex-specific
|
|
replacement proof. Deterministic native-safe-replacement tests cover its proof
|
|
and admission gates independently of model behavior.
|
|
|
|
Delegation now submits feedback through the existing child task composer and
|
|
records the delivered comment ID. The child must consume the message and deliver
|
|
the revised program. This does not require a lead to relay a parent comment.
|
|
The separate issue-update-comment-wakeup route tests exercise exact supported
|
|
mention routing, including access, dependency, identity, and duplicate-wake gates.
|
|
|
|
The approval case provisions an authenticated local service through the public
|
|
API, with a random server-held credential that never enters the agent environment
|
|
or browser trace. Approval/decline interactions still use the browser. Provider
|
|
captures distinguish rejected unauthenticated requests from accepted calls. The
|
|
old public-endpoint attempts remain boundary evidence, not an isolation promise.
|
|
|
|
Stop now waits for the owned runner to exit, records project file hashes, and
|
|
checks them again after the new response. This proves stability over that interval,
|
|
not indefinite monitoring. Hiring and declined-access policy changes are deferred
|
|
by user decision; their old results must not be presented as new campaign runs.
|
|
|
|
## Decline correction (14 September, third campaign)
|
|
|
|
That campaign used 35 cells (29 local, six remote). `service-decline` tests
|
|
rejection of a protected action on an already installed service; its former
|
|
"connection request" title was misleading. `connection-decline` separately tests
|
|
Not now on new Notion setup. Both permit a brief explanation as the complete
|
|
fallback, so Done is expected after that explanation. Neither test requires
|
|
completion after refusing work that is still required.
|
|
|
|
The installed-service decline fixture now uses the same server-held credential
|
|
as approval. The harness requires one pending interaction, validates its kind
|
|
and connection/provider identity before clicking, and waits for the exact
|
|
interaction's saved decision. Wrong interactions fail `decision-request-matches-story`
|
|
with a screenshot; they are not evidence of an ignored decline. Both decline
|
|
stories check a new explanation after the decision and reject repeated requests.
|
|
|
|
Historical attempts remain unchanged. This campaign resumes the previously
|
|
deferred decline cases; hiring remains deferred. Notion setup is declined in
|
|
the UI, so this test neither authenticates to nor reads real Notion data.
|
|
|
|
Decision screenshots are included in the evidence package. Before capturing the
|
|
final screen, the harness waits for the thread and latest persisted agent comment
|
|
to render, then scrolls that comment into view. A Done header alone is not proof
|
|
that the final response was visible.
|
|
|
|
|
|
## Recovery scope correction (14 September)
|
|
|
|
This correction reduced the catalog to **30 cells: 24 local and six remote**. Forced runner
|
|
crash probes are retired from paid selection. Their original attempt IDs remain
|
|
in Evalbook's Diagnostics history and Latest pages; they are excluded from the
|
|
main matrix without changing grades or deleting evidence. Reported spend still
|
|
includes all attempts.
|
|
|
|
`recover-controller` and `stop-redirect` retain concrete supported journeys:
|
|
restart the controller while preserving the runner, or use Stop and submit a new
|
|
direction. Their assertions verify pending input, saved work, and the next
|
|
usable result. Neither claims recovery from an arbitrary provider-process crash.
|
|
|
|
A future user-facing crash-recovery case needs a reproducible recoverable fault,
|
|
an identified supported recovery action, and evidence through the final usable
|
|
result. A missing test premise must be reported as unexercised, not a model
|
|
failure. Do not introduce a new paid case just to replace a retired row.
|
|
|
|
## Skill creation (16 September)
|
|
|
|
`create-skill-studio` adds five cells: three local profiles and the two core
|
|
profiles on Daytona. The current catalog has **35 cells: 27 local and eight
|
|
remote**. The test opens the created skill from its task-feed card, checks the
|
|
canonical skill identity in Studio, saves an edit, and returns to the same skill
|
|
in the task sidebar. A model's authored document heading is not used as the
|
|
identity check.
|
|
|
|
## External-provider fallback
|
|
|
|
Three explicit local cases cover aggregator routing with the normal production
|
|
agent guidance. They add nine local cells. The AgentMail setup case adds three
|
|
local cells; the suite now has 50 cells total.
|
|
|
|
The AgentMail case uses real model discovery and production interaction/UI paths,
|
|
but declines before sending credentials to AgentMail. Its independent grader is
|
|
calibrated against missing, duplicate, misaddressed, hidden, and malformed cards.
|
|
The email integration suite separately proves credential/inbox creation, access
|
|
defaults, assignment checks, and completion using a fixture provider. Neither
|
|
test qualifies live AgentMail delivery. Run a bounded local model cell with:
|
|
|
|
```sh
|
|
pnpm test:e2e:runner -- --id everyday-workflows.runner-codex-mini.local.agentmail-setup --max-automatic-retries 0
|
|
```
|
|
|
|
- `provider-native`: Jira is supported natively and by aggregators. Require the
|
|
Jira connection card directly, decline it in the browser, and verify no provider
|
|
question, connection creation, or repeated request.
|
|
- `provider-decline`: HubSpot has no built-in connector in this fixture. Require
|
|
Composio, Arcade, Zapier, and None in that order with external-service disclosure.
|
|
Restart the controller, reload the question, choose None through the UI, and
|
|
verify one saved answer, no connection changes, and no fabricated result.
|
|
- `provider-second`: Install a deterministic Arcade gateway with a read-only
|
|
HubSpot action through public APIs. Choose Arcade in the browser after restart.
|
|
Require no calls before selection, exactly one call afterwards, the independently
|
|
generated contact marker in the agent response, and no duplicate connection.
|
|
|
|
These use the existing native profiles and 12-minute local attempt deadline;
|
|
expected provider runs are two per case. There are no real third-party mutations.
|
|
The fixture server is closed and the harness cleans its disposable instance.
|
|
Provider-choice screenshots, persisted interactions, gateway call counts, source
|
|
revision, harness digest, and existing usage/cost evidence accompany each attempt.
|
|
`connection-routing-evidence.test.ts` calibrates the grader against undisclosed
|
|
routing, incorrect ordering, early calls, duplicate questions, and fabricated reads.
|
|
Run a single `everyday-workflows.runner-codex-mini.local.provider-decline` cell
|
|
first; do not treat these fixtures as live provider compatibility tests.
|
|
|
|
### Connection-guidance observation boundaries
|
|
|
|
The explicit `native-connection-guidance` suite keeps one attempt per cell.
|
|
For decline cases, Done and a succeeded run are insufficient: settlement waits
|
|
within the existing cell deadline for a saved reply attributed by run ID to the
|
|
lead agent's final successful run on this task. Readiness checks storage and
|
|
identity, not favorable wording. A missing reply times out; an incorrect reply
|
|
settles and fails the independent explanation check. Both decline explanation
|
|
checks use the same bounded matcher, including unavailable/rejected access and
|
|
“wasn't able to pull” wording. This is textual evidence, not proof of cognition.
|
|
|
|
An executed approval or a recorded tool-action rejection with `wake_assignee`
|
|
may precede its continuation run. Only a response bound to the same task, company,
|
|
agent, and completed source run can defer the stranded-blocker check, and only
|
|
within the original deadline. A failed execution, unrelated card, or already
|
|
consumed response cannot extend that wait.
|
|
|
|
`connection-guidance-decline-grade.json` and the workflow's
|
|
`declineGradeEvidence` retain the exact cloned assertion input and original
|
|
checks. Later final-state or cleanup observations cannot overwrite that input.
|
|
Earlier campaign grades remain unchanged when the evaluator is corrected.
|