mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat needs reliable native execution before native runners become the onboarding default. > - Existing stories covered idle reassignment and controller restart, but not an executing worker handoff or worker process loss. > - Status answer tests also need to reject stale claims and invented facts. > - This pull request adds six opt-in full-stack cells with independent state assertions and retained evidence. > - The probes exposed a misleading Retry across server projection and recovery-banner paths; the fix reports the blocked recovery honestly. > - The tests preserve failures without changing recovery policy, production prompts, or onboarding defaults. ## Linked Issues or Issue Description Refs: #13762. Related: #13765 (Retry targets the latest failed attempt), #13753 (task context ownership), #13746 (native recovery work). ## What Changed - Add active reassignment with saved draft and plan preservation, old-worker cancellation, and successor completion checks. - Preserve recovery-needed projection when native cleanup fails before its coordinator exists, refuse a generic retry that would immediately fail again, and replace the recovery banner's misleading Retry with Inspect run. - Add verified local worker process loss with a required successful continuation; retain a failing qualification result when recovery is unavailable, while independently verifying the UI/API refuse doomed retries. - Add two-turn factual answer checks for current blockers, stale claims, inactive backlog work, and unknown facts. Retain prose for separate semantic review. - Add positive and negative oracle calibration and document fault isolation, cleanup, billing, and qualification limits. ## Verification - Eval TypeScript check passes. - All 442 eval support tests pass locally. The 89 focused server tests and server typecheck pass. Six recovery-banner UI tests and token gates pass. - Initial new-cell campaign: https://github.com/paperclipai/paperclip/actions/runs/35657128077. All six results are retained; four failed on fixture-contract issues and two exposed real worker cleanup quarantine. - All 26 existing native onboarding cells: https://github.com/paperclipai/paperclip/actions/runs/35656761484 (26/26 passed on master846336e5a, all cleanup passed). - Intermediate handoff/fault campaign: https://github.com/paperclipai/paperclip/actions/runs/35657945095 (four retained failures: two overly strict draft oracles, two real crash quarantines). - Final active handoff: https://github.com/paperclipai/paperclip/actions/runs/35659014397 (2/2 passed on cf6d4ae3a; both cleanup passed). - Clarified answer-quality fixtures: https://github.com/paperclipai/paperclip/actions/runs/35658262695 (2/2 passed on 4a26f10be; both cleanup passed; all four answers semantically reviewed). - Quarantine guard regression campaign: https://github.com/paperclipai/paperclip/actions/runs/35659580100 (both API requests correctly refused with 409/no second run, but exposed a separate misleading Retry in the recovery banner and a fixture wait on a non-admitted run). - Final quarantine guard verification: https://github.com/paperclipai/paperclip/actions/runs/35661067305 (147e42f7e: both providers verify Inspect run/no Retry, HTTP 409, one retained run, unchanged saved plan, and successful disposable cleanup. Both evals intentionally remain red with `worker_crash_recovery_unqualified`; no successful continuation exists). The preceding campaign 35658772755 never ran provider cases because GitHub artifact finalization returned HTTP 403. - Full repository CI passes on147e42f7e: typecheck, tests, build, and browser gates. One unchanged local-service-supervisor readiness test failed initially; its six-test file passed in isolation and the failed shard passed on its single rerun. Latest-head rollup: 54 successful, 2 intentionally skipped, no failed or pending checks. Greptile is 5/5 with zero unresolved findings. - See tests/runner-e2e/QUALIFICATION-2026-09-21.md for retained attempts and semantic review. Published reports: [onboarding](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35656761484-1/), [handoff](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35659014397-1/), [grounded answers](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35658262695-1/), [crash guards and unqualified recovery](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35661067305-1/). ## Risks - Paid cells are explicit-only and local-only. The fault fixture signals only the exact native run PID after checking its identity. - Live worker-loss probes currently fail on cleanup quarantine for both providers. The eval must remain red until there is a usable recovery, even when preservation and refusal checks pass. Verified cleanup with a fresh attempt versus exact-session resume remains a product decision. - Structured facts alone do not qualify prose quality; semantic review remains separate. - Onboarding uses the existing runtime switch after the real wizard and before provider execution. Native UI selection and public defaults remain unchanged. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact served snapshot and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
80 lines
3.3 KiB
Python
80 lines
3.3 KiB
Python
import contextlib
|
|
import io
|
|
import json
|
|
from pathlib import Path
|
|
import runpy
|
|
import signal
|
|
import subprocess
|
|
import sys
|
|
import unittest
|
|
from unittest.mock import patch
|
|
|
|
HELPER = str(Path(__file__).with_name("worker-fault.py"))
|
|
|
|
|
|
class FaultIdentityTests(unittest.TestCase):
|
|
def invoke(self, expected="1234", run="fixture-run", live_pid="43210", mode="kill"):
|
|
def fake_open(name, *args, **kwargs):
|
|
if name.endswith("/stat"):
|
|
return io.StringIO("43210 (worker) " + " ".join(["S"] + ["0"] * 18 + ["1234"]))
|
|
if name.endswith("/cmdline"):
|
|
return io.BytesIO(b"node\0runner\0--run-id\0fixture-run\0")
|
|
if "/fdinfo/" in name:
|
|
return io.StringIO("Pid:\t" + live_pid + "\n")
|
|
raise AssertionError(name)
|
|
with patch.object(sys, "platform", "linux"), patch.object(sys, "argv", [HELPER, mode, "43210", run, expected]), \
|
|
patch("os.pidfd_open", return_value=42, create=True) as opened, \
|
|
patch("signal.pidfd_send_signal", create=True) as sent, patch("os.close") as closed, \
|
|
patch("builtins.open", fake_open), contextlib.redirect_stdout(io.StringIO()):
|
|
try:
|
|
runpy.run_path(HELPER, run_name="__main__")
|
|
except RuntimeError:
|
|
sent.assert_not_called()
|
|
closed.assert_called_once_with(42)
|
|
raise
|
|
opened.assert_called_once_with(43210)
|
|
closed.assert_called_once_with(42)
|
|
if mode == "kill":
|
|
sent.assert_called_once_with(42, signal.SIGKILL, None, 0)
|
|
else:
|
|
sent.assert_not_called()
|
|
|
|
def test_signals_owned_handle_not_numeric_pid(self):
|
|
self.invoke()
|
|
|
|
def test_inspection_does_not_signal(self):
|
|
self.invoke(mode="inspect")
|
|
|
|
def test_refuses_changed_start_identity(self):
|
|
with self.assertRaisesRegex(RuntimeError, "start identity changed"):
|
|
self.invoke(expected="earlier-process")
|
|
|
|
def test_refuses_wrong_run(self):
|
|
with self.assertRaisesRegex(RuntimeError, "run ID"):
|
|
self.invoke(run="other-run")
|
|
|
|
def test_refuses_dead_original_handle_even_if_pid_was_reused(self):
|
|
with self.assertRaisesRegex(RuntimeError, "exited"):
|
|
self.invoke(live_pid="-1")
|
|
|
|
@unittest.skipUnless(sys.platform == "linux", "Real pidfd fault qualification runs on Linux CI")
|
|
def test_real_owned_child_with_wrong_then_correct_start_identity(self):
|
|
child = subprocess.Popen([sys.executable, "-c", "import time; time.sleep(60)", "--run-id", "fault-test"])
|
|
try:
|
|
args = [sys.executable, HELPER]
|
|
identity = json.loads(subprocess.check_output(args + ["inspect", str(child.pid), "fault-test"]))
|
|
wrong = subprocess.run(args + ["kill", str(child.pid), "fault-test", "wrong"], capture_output=True)
|
|
self.assertNotEqual(wrong.returncode, 0)
|
|
self.assertIsNone(child.poll())
|
|
result = json.loads(subprocess.check_output(args + ["kill", str(child.pid), "fault-test", identity["startTicks"]]))
|
|
self.assertTrue(result["signalled"])
|
|
self.assertEqual(child.wait(timeout=5), -signal.SIGKILL)
|
|
finally:
|
|
if child.poll() is None:
|
|
child.kill()
|
|
child.wait(timeout=5)
|
|
|
|
|
|
if __name__ == "__main__":
|
|
unittest.main()
|