mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 16:11:46 +02:00
## Thinking Path > - Paperclip manages work across persistent agent sessions. > - The runner validates completion calls before accepting their results. > - Generic validation errors can leave the agent unable to repair a rejected call. > - Restart tests also exempted every later failure on an intentionally interrupted run. > - This change gives bounded schema feedback and limits the test exemption to expected interruption outcomes. > - Failures become easier to repair and diagnose without changing authorization or task prompts. ## Linked Issues or Issue Description Refs #13674 and #13676. Related environment and Agent Chat fixes landed in #13677 and #13678. Those changes do not cover these diagnostics. **What happened?** A malformed completion call received a general field list without the failed schema location. The everyday restart test hid later adapter errors on an intentionally interrupted run until its deadline. A clean pnpm install also broke the shutdown test because it resolved an undeclared Playwright package. **Expected behavior** Return enough schema information to repair completion calls without returning submitted values. Fail promptly on an unexpected recovery error. Resolve the declared test package's CLI. **Steps to reproduce** Run the new completion-validation and everyday lifecycle regressions against the parent commit. The new assertions fail there. Run the shutdown test in a clean workspace installation. **Paperclip version or commit** Based on masterc1f6c3310. **Deployment mode** Native runner and Runner E2E harness, local and Daytona. ## What Changed - Add up to three schema locations and missing field names to rejected completion-tool feedback. Limit the message to 512 bytes. Do not echo input values or unknown input keys. - Preserve generic errors for other or unauthorized tools. - Exempt only expected cancellation, shutdown, and process-loss outcomes in fault-injection tests. Apply the same rule while polling and grading. - Fail promptly when persisted review timestamps prove the required blocked-parent ordering was not exercised. Wait when separately fetched snapshots lack that evidence. Keep the boundary requirement. - Use the CLI export from the declared Playwright test dependency. - Add red/green regressions and document the interruption rule. ## Verification - 393 Runner E2E harness tests pass. - 291 runner-core library tests pass. - Harness typecheck and runner TypeScript build pass; Rust formatting and git diff checks pass. - The baseline completion regression and three lifecycle assertions failed before the fix and pass after it. - Fresh worker-setup rerun: https://github.com/paperclipai/paperclip/actions/runs/35449301275 (7/8 passed). Mini completed the child review and both tasks, but did not enter the blocked-parent-before-review ordering required by that case. The new diagnostic identifies this unexercised boundary promptly; it does not turn the case into a pass. - Candidate Mini rerun: https://github.com/paperclipai/paperclip/actions/runs/35449686308. The recording has one rejected completion followed by success, compared with 12 validation errors in the baseline. The case still fails because it does not hire the requested teammate, despite delivering a ZIP that passes artifact checks. This is not a behavioral pass claim. - Updated report: https://pages.paperclip.ing/runner-e2e-operational-35444497313/investigation.html. - All 55 PR checks are successful or skipped on4d2c91869, including repository typecheck, build, tests, native Runner checks, and Greptile 5/5. The reviewer withdrew the diagnostic finding after confirming the conservative behavior avoids false rejection from non-atomic snapshots. Local verification was limited to the affected harness and runner. ## Risks Schema feedback must remain bounded and exclude submitted values. Stricter fault grading can expose previously hidden failures. The Daytona restart identity bug and remaining behavior failures are not fixed here. Runtime authorization, completion validation, and production prompts remain unchanged. ## Model Used OpenAI GPT-6 via Codex, with repository inspection, code editing, and test execution. The exact deployment identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>