Files
cceeb0aa66 test(runner): add everyday workflow evaluation harness (#13474)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner must support project work, delegation, hiring, and
service access.
> - Browser tests exposed lost connection access, rejected helper
events, and stalled recovery.
> - Some eval failures also came from incorrect fixtures and decision
controls.
> - This pull request fixes those paths and adds eight everyday workflow
stories.
> - The tests retain observed failures and verify delivered files
independently.
> - The benefit is repeatable evidence for common user tasks and their
remaining gaps.

## Linked Issues or Issue Description

Related work: #13404 contains earlier workflow fixes. #13300 and #13470
changed the CI contracts used by the harness security tests. Merged
companion:
[paperclip-evals#22](https://github.com/paperclipai/paperclip-evals/pull/22).

**What happened?**

Native ACPX sessions did not receive the assigned connection gateway.
Codex helper events could arrive before their spawn receipt and fail
thread validation. A parent continuation could take a shared workspace
before its child retried. A failed native continuation could leave the
task status without a clear recovery blocker. The eval harness also
confused tool approvals with new connection requests and could reject a
valid delegated download.

**Expected behavior**

Keep assigned gateway access and its approval checks. Verify helper
lineage before accepting helper progress. Let a waiting child proceed
before automatic parent recovery. Preserve a failed task's recovery
ownership. Grade the actual requested workflow and its delivered files.

**Steps to reproduce**

Run the everyday workflow suite with the native Codex and Claude
profiles. Exercise service approval, connection refusal, delegated
project work, and teammate reuse. The commands and case requirements are
in `tests/runner-e2e/EVERYDAY-WORKFLOWS.md`. Use `pnpm
test:runner-recovery` for controlled crash and replacement cases.

## What Changed

- Pass the scoped connection gateway binding through the native ACPX
host and sidecar.
- Recognize Codex helper lineage from parent metadata and spawn
receipts. Verify early helper events with `thread/read`. Keep helper
events separate from root completion authority.
- Guide agents to use persistent hiring, child tasks, dependency
records, and a blocked handoff while waiting for a child.
- Defer automatic parent recovery while a child has an active execution
path in the same shared workspace. Allow parent recovery when the child
needs review.
- Record Blocked status and recovery evidence when a failed native
continuation needs reconciliation, including existing active or
escalated incidents. Preserve their owner and retry budget.
- Add eight browser-driven workflow cases. Use real decision controls,
explicit child feedback delivery, managed hiring credentials, and
independent ZIP checks inside a bounded Docker sandbox. Verify sandbox
availability before task creation. Record screenshot SHA-256 at capture.
- Keep runner crash probes in controlled recovery tests. Preserve the
original failure when cleanup also fails.
- Display missing accounting and replay revisions as unavailable. Align
harness security assertions with the approved CI changes.
- Make the channel-rejection browser fixture bind its file after the
send captures its payload. This prevents live refresh from removing the
file before the simulated race.

## Verification

- Full workspace `pnpm -r typecheck` passed after merging current
master.
- Runner E2E typecheck passed. Harness unit tests passed: 216/216.
- Wake-queue database tests passed: 55/55. The two added
existing-incident tests failed before the fix and pass after it.
- Docker artifact calibration passed: 12/12. Host-file and host-loopback
isolation tests failed before the fix and pass after it. Read-only
delivery and output limits are also verified.
- Full `pnpm build` passed. Targeted recovery tests passed: 83/83.
- The channel-rejection browser test passed five consecutive runs after
fixing the fixture race found in CI.
- Local general-server (12,351 tests), UI (6,250), CLI (485), and
workspace package groups passed. The monolithic run stopped at an
unchanged lock-heartbeat fixture race; the isolated workspace group
passed on rerun (shared: 747/747). A separate local serialized run
passed 97 files before two socket errors in the unchanged issue-list
route suite; that suite passed 15/15 on isolated rerun. These local full
commands did not finish uninterrupted; the complete CI matrix below
covers the remaining suites.
- Final head `0fb293733fe307be7e6667ae8f1364077d0c6455`: **34 successful
checks, 2 expected skips**, including every server/workspace shard,
browser shard, native runner verification, build, and typecheck. [Final
CI
run](https://github.com/paperclipai/paperclip/actions/runs/34989136700).
- Greptile reviewed this exact head at **5/5**; all review threads are
resolved. Both Superagent security checks are successful.
- ACPX credential-boundary tests passed: 118/118. Superagent accepted
the runner/sidecar versus provider-environment trace and cleared its
finding.
- The latest paid local campaign on source
`f6a2fdf7ac2af859826a2ae627ff4125a5478529` passed 22/24 cases: Sol 8/8,
Claude 7/8, Mini 7/8. These results predate the merge with current
master.
- The two remaining failures are in `hire-reuse`: Claude exceeded the
attempt deadline during final review; Mini made invalid deliverable tool
calls and remained Blocked.
- Six Daytona cases were not run because the matching immutable runner
image was unavailable. This PR does not claim new remote model results.

## Risks

The changes affect connection admission, helper identity, and recovery
scheduling. Assigned gateway grants and user approval still govern
service calls. The workspace admission gate still exists; the broader
folder-sync design is separate work. Provider behavior can still cause
the two recorded hiring failures. No database migration is required.
Paid cases are opt-in and have bounded attempt deadlines. Project
stories now require Docker and the documented pinned Python image on the
harness host.

## Model Used

OpenAI `gpt-6-astra` performed implementation, diagnosis, and
substantive review. OpenAI `gpt-5.6-luna` assisted with verification, PR
preparation, and review tracking. Both used repository tools and code
execution. Context-window sizes were not recorded.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks and
isolated reruns; full-run limitations are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: OpenAI GPT-5.6 Luna <noreply@openai.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 11:04:16 -05:00

135 lines
7.7 KiB
Python

"""Independent acceptance oracle. Never runs the project's own test assertions."""
import argparse, json, os, pathlib, selectors, stat, subprocess, sys, tempfile, time, uuid, zipfile
BASE_CASES = [("Hello World", "hello-world"), (" Queue--Ready!! ", "queue-ready"),
("Already-Fine", "already-fine"), ("Café 東京", "caf"), ("!!!", ""),
("a__b c", "a-b-c"), ("123", "123"), ("", "")]
# Published multi-platform python:3.13-slim digest. Never pull implicitly during a model run.
SANDBOX_IMAGE = 'python@sha256:9d2e5553305c7c7b0097999bb17187c69b921ccd6bc9d40e4bb5ebe652c00285'
def preflight():
for args in [['docker', 'version', '--format', '{{.Server.Version}}'],
['docker', 'image', 'inspect', SANDBOX_IMAGE]]:
result = subprocess.run(args, capture_output=True, text=True, timeout=10)
if result.returncode:
raise RuntimeError(f'Artifact sandbox qualification failed. Start Docker and run: docker pull {SANDBOX_IMAGE}')
class ArtifactSandbox:
"""Only extracted delivery files enter the container. No host secrets or network."""
def __init__(self, root, source):
self.root = root
self.cwd = '/project/' + str(source.parent.relative_to(root))
self.source = '/project/' + str(source.relative_to(root))
self.name = 'paperclip-artifact-oracle-' + uuid.uuid4().hex
def __enter__(self):
preflight()
self.root.chmod(0o755)
args = ['docker', 'run', '--detach', '--rm', '--pull=never', '--name', self.name,
'--network=none', '--read-only', '--cap-drop=ALL',
'--security-opt=no-new-privileges', '--pids-limit=64', '--memory=128m',
'--memory-swap=128m', '--cpus=1', '--user=65534:65534',
'--log-driver=none', '--tmpfs=/tmp:rw,noexec,nosuid,size=16m',
'--mount', f'type=bind,source={self.root},target=/project,readonly',
'--workdir', self.cwd, SANDBOX_IMAGE, 'python', '-I', '-c',
'import time; time.sleep(180)']
try:
result = subprocess.run(args, capture_output=True, text=True, timeout=15)
if result.returncode:
raise RuntimeError('Artifact sandbox could not start: ' + result.stderr[:500])
except BaseException:
self.close()
raise
return self
def close(self):
subprocess.run(['docker', 'rm', '--force', self.name], capture_output=True, timeout=10)
def __exit__(self, *unused):
self.close()
def invoke(self, args, *, imported=False):
code = ['-c', "import sys; sys.path.insert(0, " + repr(self.cwd) + "); from slugify import slugify; assert slugify(' A B! ') == 'a-b'"] if imported else [self.source, *args]
command = ['docker', 'exec', '--workdir', self.cwd, self.name, 'python', '-I', '-B', *code]
# Bound output as it arrives. A generated program must not exhaust host memory or disk.
proc = subprocess.Popen(command, stdout=subprocess.PIPE, stderr=subprocess.PIPE)
output = {'stdout': bytearray(), 'stderr': bytearray()}
deadline = time.monotonic() + 10
try:
with selectors.DefaultSelector() as selector:
selector.register(proc.stdout, selectors.EVENT_READ, 'stdout')
selector.register(proc.stderr, selectors.EVENT_READ, 'stderr')
while selector.get_map():
remaining = deadline - time.monotonic()
if remaining <= 0:
raise ValueError('Artifact command exceeded the time limit')
for key, _ in selector.select(remaining):
chunk = os.read(key.fd, 4096)
if not chunk:
selector.unregister(key.fileobj)
continue
output[key.data].extend(chunk)
if sum(map(len, output.values())) > 65_536:
raise ValueError('Artifact command exceeded the output limit')
code = proc.wait(timeout=max(0.01, deadline-time.monotonic()))
return subprocess.CompletedProcess(command, code,
output['stdout'].decode('utf-8', errors='replace'),
output['stderr'].decode('utf-8', errors='replace'))
finally:
if proc.poll() is None:
proc.kill()
proc.wait()
proc.stdout.close()
proc.stderr.close()
def inspect(archive, mode):
checks = []
def check(name, passed, detail=""):
checks.append(dict(id=name, passed=bool(passed), detail=detail))
with tempfile.TemporaryDirectory(prefix="paperclip-artifact-oracle-") as tmp:
root=pathlib.Path(tmp)
with zipfile.ZipFile(archive) as z:
entries=z.infolist()
if len(entries)>250 or sum(e.file_size for e in entries)>10_000_000:
raise ValueError("Archive exceeds the bounded source project size")
for entry in entries:
p=pathlib.PurePosixPath(entry.filename)
if p.is_absolute() or ".." in p.parts or "\\" in entry.filename or stat.S_ISLNK(entry.external_attr >> 16):
raise ValueError("Unsafe archive member")
z.extractall(root)
sources=list(root.rglob("slugify.py"))
check("one-slugify-source",len(sources)==1)
check("readme-present",any(p.name.lower().startswith('readme') for p in root.rglob('*')))
check("project-tests-present",any(p.name.startswith('test') and p.suffix=='.py' for p in root.rglob('*')))
if len(sources)!=1:return checks
source=sources[0]
with ArtifactSandbox(root, source) as sandbox:
invoke = sandbox.invoke
for index,(text,expected) in enumerate(BASE_CASES):
result=invoke([text]);check(f"base-{index}",result.returncode==0 and result.stdout.rstrip('\r\n')==expected,
f"exit={result.returncode}; expected={expected!r}; observed={result.stdout[:160]!r}")
# Import from the delivered module, not an evaluator reimplementation.
imported=invoke([], imported=True)
check('importable-function',imported.returncode==0)
if mode=='separator':
for separator,expected in [('_','queue_ready'),('-','queue-ready')]:
result=invoke([' Queue--Ready!! ','--separator',separator]);check('separator-'+separator,result.returncode==0 and result.stdout.strip()==expected)
check('reject-invalid-separator',invoke(['hello','--separator','/']).returncode!=0)
if mode=='max-length':
for size,expected in [('7','queue-r'),('6','queue'),('1','q')]:
result=invoke([' Queue--Ready!! ','--max-length',size]);check('length-'+size,result.returncode==0 and result.stdout.strip()==expected)
check('reject-zero-length',invoke(['hello','--max-length','0']).returncode!=0)
return checks
if __name__=='__main__':
parser=argparse.ArgumentParser();parser.add_argument('archive',nargs='?');parser.add_argument('--preflight',action='store_true');parser.add_argument('--mode',choices=['base','separator','max-length'],default='base');args=parser.parse_args()
try:
if args.preflight:
preflight();checks=[dict(id='artifact-sandbox-ready',passed=True,detail='Pinned image available')]
elif args.archive:checks=inspect(args.archive,args.mode)
else:parser.error('archive is required unless --preflight is set')
except Exception as error:checks=[dict(id='artifact-readable',passed=False,detail=str(error))]
print(json.dumps(dict(passed=all(c['passed'] for c in checks),checks=checks)))
sys.exit(0 if all(c['passed'] for c in checks) else 1)