Files
PaperClipAI/tests/runner-e2e/test_everyday_artifact.py
T
cceeb0aa66 test(runner): add everyday workflow evaluation harness (#13474)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner must support project work, delegation, hiring, and
service access.
> - Browser tests exposed lost connection access, rejected helper
events, and stalled recovery.
> - Some eval failures also came from incorrect fixtures and decision
controls.
> - This pull request fixes those paths and adds eight everyday workflow
stories.
> - The tests retain observed failures and verify delivered files
independently.
> - The benefit is repeatable evidence for common user tasks and their
remaining gaps.

## Linked Issues or Issue Description

Related work: #13404 contains earlier workflow fixes. #13300 and #13470
changed the CI contracts used by the harness security tests. Merged
companion:
[paperclip-evals#22](https://github.com/paperclipai/paperclip-evals/pull/22).

**What happened?**

Native ACPX sessions did not receive the assigned connection gateway.
Codex helper events could arrive before their spawn receipt and fail
thread validation. A parent continuation could take a shared workspace
before its child retried. A failed native continuation could leave the
task status without a clear recovery blocker. The eval harness also
confused tool approvals with new connection requests and could reject a
valid delegated download.

**Expected behavior**

Keep assigned gateway access and its approval checks. Verify helper
lineage before accepting helper progress. Let a waiting child proceed
before automatic parent recovery. Preserve a failed task's recovery
ownership. Grade the actual requested workflow and its delivered files.

**Steps to reproduce**

Run the everyday workflow suite with the native Codex and Claude
profiles. Exercise service approval, connection refusal, delegated
project work, and teammate reuse. The commands and case requirements are
in `tests/runner-e2e/EVERYDAY-WORKFLOWS.md`. Use `pnpm
test:runner-recovery` for controlled crash and replacement cases.

## What Changed

- Pass the scoped connection gateway binding through the native ACPX
host and sidecar.
- Recognize Codex helper lineage from parent metadata and spawn
receipts. Verify early helper events with `thread/read`. Keep helper
events separate from root completion authority.
- Guide agents to use persistent hiring, child tasks, dependency
records, and a blocked handoff while waiting for a child.
- Defer automatic parent recovery while a child has an active execution
path in the same shared workspace. Allow parent recovery when the child
needs review.
- Record Blocked status and recovery evidence when a failed native
continuation needs reconciliation, including existing active or
escalated incidents. Preserve their owner and retry budget.
- Add eight browser-driven workflow cases. Use real decision controls,
explicit child feedback delivery, managed hiring credentials, and
independent ZIP checks inside a bounded Docker sandbox. Verify sandbox
availability before task creation. Record screenshot SHA-256 at capture.
- Keep runner crash probes in controlled recovery tests. Preserve the
original failure when cleanup also fails.
- Display missing accounting and replay revisions as unavailable. Align
harness security assertions with the approved CI changes.
- Make the channel-rejection browser fixture bind its file after the
send captures its payload. This prevents live refresh from removing the
file before the simulated race.

## Verification

- Full workspace `pnpm -r typecheck` passed after merging current
master.
- Runner E2E typecheck passed. Harness unit tests passed: 216/216.
- Wake-queue database tests passed: 55/55. The two added
existing-incident tests failed before the fix and pass after it.
- Docker artifact calibration passed: 12/12. Host-file and host-loopback
isolation tests failed before the fix and pass after it. Read-only
delivery and output limits are also verified.
- Full `pnpm build` passed. Targeted recovery tests passed: 83/83.
- The channel-rejection browser test passed five consecutive runs after
fixing the fixture race found in CI.
- Local general-server (12,351 tests), UI (6,250), CLI (485), and
workspace package groups passed. The monolithic run stopped at an
unchanged lock-heartbeat fixture race; the isolated workspace group
passed on rerun (shared: 747/747). A separate local serialized run
passed 97 files before two socket errors in the unchanged issue-list
route suite; that suite passed 15/15 on isolated rerun. These local full
commands did not finish uninterrupted; the complete CI matrix below
covers the remaining suites.
- Final head `0fb293733fe307be7e6667ae8f1364077d0c6455`: **34 successful
checks, 2 expected skips**, including every server/workspace shard,
browser shard, native runner verification, build, and typecheck. [Final
CI
run](https://github.com/paperclipai/paperclip/actions/runs/34989136700).
- Greptile reviewed this exact head at **5/5**; all review threads are
resolved. Both Superagent security checks are successful.
- ACPX credential-boundary tests passed: 118/118. Superagent accepted
the runner/sidecar versus provider-environment trace and cleared its
finding.
- The latest paid local campaign on source
`f6a2fdf7ac2af859826a2ae627ff4125a5478529` passed 22/24 cases: Sol 8/8,
Claude 7/8, Mini 7/8. These results predate the merge with current
master.
- The two remaining failures are in `hire-reuse`: Claude exceeded the
attempt deadline during final review; Mini made invalid deliverable tool
calls and remained Blocked.
- Six Daytona cases were not run because the matching immutable runner
image was unavailable. This PR does not claim new remote model results.

## Risks

The changes affect connection admission, helper identity, and recovery
scheduling. Assigned gateway grants and user approval still govern
service calls. The workspace admission gate still exists; the broader
folder-sync design is separate work. Provider behavior can still cause
the two recorded hiring failures. No database migration is required.
Paid cases are opt-in and have bounded attempt deadlines. Project
stories now require Docker and the documented pinned Python image on the
harness host.

## Model Used

OpenAI `gpt-6-astra` performed implementation, diagnosis, and
substantive review. OpenAI `gpt-5.6-luna` assisted with verification, PR
preparation, and review tracking. Both used repository tools and code
execution. Context-window sizes were not recorded.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks and
isolated reruns; full-run limitations are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: OpenAI GPT-5.6 Luna <noreply@openai.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 11:04:16 -05:00

92 lines
4.6 KiB
Python

"""Calibrate the independent oracle with correct and deliberately broken deliveries."""
import importlib.util
import pathlib
import stat
import socket
import tempfile
import unittest
import zipfile
spec = importlib.util.spec_from_file_location("oracle", pathlib.Path(__file__).with_name("everyday-artifact.py"))
oracle = importlib.util.module_from_spec(spec)
spec.loader.exec_module(oracle)
GOOD = '''import argparse,re
def slugify(text): return re.sub(r'[^a-z0-9]+','-',text.strip().lower()).strip('-')
if __name__ == '__main__':
p=argparse.ArgumentParser();p.add_argument('text');p.add_argument('--separator',choices=['-','_'],default='-');p.add_argument('--max-length',type=int)
a=p.parse_args()
if a.max_length is not None and a.max_length<=0:p.error('positive length required')
value=slugify(a.text).replace('-',a.separator)
if a.max_length is not None:value=value[:a.max_length].rstrip(a.separator)
print(value)
'''
class ArtifactOracleTests(unittest.TestCase):
def grade(self, source=GOOD, mode='base', extras=None):
with tempfile.TemporaryDirectory() as tmp:
archive=pathlib.Path(tmp)/'project.zip'
with zipfile.ZipFile(archive,'w') as z:
z.writestr('project/slugify.py',source)
z.writestr('project/README.md','Usage: python3 slugify.py "Hello World"')
# Deliberately passing but useless agent-authored tests must not determine our verdict.
z.writestr('project/test_slugify.py','assert True')
for name,content in (extras or {}).items():z.writestr(name,content)
return oracle.inspect(archive,mode)
def test_correct_delivery_passes_all_modes(self):
for mode in ['base','separator','max-length']:
with self.subTest(mode=mode):self.assertTrue(all(c['passed'] for c in self.grade(mode=mode)))
def test_wrong_output_fails_despite_agent_tests(self):
checks=self.grade(GOOD.replace("print(value)","print('hello-world')"))
self.assertFalse(all(c['passed'] for c in checks))
def test_missing_late_requirement_fails(self):
source=GOOD.replace("if a.max_length is not None:value=value[:a.max_length].rstrip(a.separator)","if False:pass")
self.assertFalse(all(c['passed'] for c in self.grade(source,'max-length')))
def test_trailing_separator_bug_fails(self):
checks=self.grade(GOOD.replace(".rstrip(a.separator)",""),'max-length')
self.assertFalse(next(c['passed'] for c in checks if c['id']=='length-6'))
def test_invalid_separator_bug_fails(self):
checks=self.grade(GOOD.replace("choices=['-','_'],", ""),'separator')
self.assertFalse(next(c['passed'] for c in checks if c['id']=='reject-invalid-separator'))
def test_artifact_cannot_read_host_files(self):
with tempfile.TemporaryDirectory() as tmp:
marker=pathlib.Path(tmp)/'host-private-marker';marker.write_text('private fixture')
prefix=f"import pathlib\nassert not pathlib.Path({str(marker)!r}).exists(), 'host filesystem is visible'\n"
self.assertTrue(all(c['passed'] for c in self.grade(prefix+GOOD)))
def test_artifact_cannot_reach_host_loopback(self):
with socket.socket() as server:
server.bind(('127.0.0.1',0));server.listen(64)
port=server.getsockname()[1]
prefix=f"import socket\ns=socket.socket();s.settimeout(0.2)\nassert s.connect_ex(('127.0.0.1',{port})) != 0, 'host network is visible'\ns.close()\n"
self.assertTrue(all(c['passed'] for c in self.grade(prefix+GOOD)))
def test_artifact_cannot_modify_the_read_only_delivery(self):
prefix="import pathlib\ntry: pathlib.Path(__file__).write_text('changed')\nexcept OSError: pass\nelse: raise AssertionError('project is writable')\n"
self.assertTrue(all(c['passed'] for c in self.grade(prefix+GOOD)))
def test_generated_output_is_bounded(self):
with self.assertRaisesRegex(ValueError, 'output limit'):
self.grade("print('x'*1000000)")
def test_duplicate_source_fails(self):
self.assertFalse(all(c['passed'] for c in self.grade(extras={'other/slugify.py':GOOD})))
def test_path_traversal_is_rejected_before_execution(self):
for name in ['../escape','/absolute','back\\slash']:
with self.subTest(name=name),self.assertRaisesRegex(ValueError,'Unsafe archive'):
self.grade(extras={name:'bad'})
def test_symlink_is_rejected(self):
member=zipfile.ZipInfo('link');member.external_attr=(stat.S_IFLNK|0o777)<<16
with self.assertRaisesRegex(ValueError,'Unsafe archive'):
self.grade(extras={member:'/tmp/target'})
if __name__=='__main__':unittest.main()