mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 00:54:38 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Runner tasks must continue after approval and stop when the user presses Stop. > - Live evals found races at approval delivery and provider startup. > - Browser readiness and CI setup errors also hid the actual task results. > - This pull request fixes those races and the related test infrastructure. > - Regression tests and saved live reports show which cases now pass. ## Linked Issues or Issue Description Companion eval definitions PR: https://github.com/paperclipai/paperclip-evals/pull/25 (AgentCore paused and provider/environment infrastructure). Related: #13741 now supplies the late-startup Stop fence and warm-attachment recovery; this PR retains that fence and extends startup tracking and regression coverage to both native backend paths. #13539 introduced queued approvals during active runs. #13738 fixes child assignment, task replies, and warm process continuity and is already in the base. #13291 concerns automatic continuation of interrupted legacy sandbox runs; this PR fixes native startup cancellation and does not change that recovery policy. **What happened?** An accepted service approval could wait after its source run stopped. Stop could return success before the provider handle existed. Work could then start after Stop, or a cancelled run could be recorded as failed. Some E2E tests also failed on unloaded browser content or irrelevant reply wording. Runner CI could fail before model work because of dependency or sandbox setup. **Expected behavior** Deliver each settled approval once after its source run stops. Do not start work after an acknowledged Stop. Preserve the audited cancellation. Test the intended product behavior with a ready browser and verified runtime dependencies. **Steps to reproduce** 1. Approve a service request while its source run is active. Let the run finish. Check that its result starts one continuation. 2. Delay provider startup. Press Stop before its handle is available. Check cancellation, then submit `/new`. 3. Run the browser, warm-workspace, and Stop-and-redirect cases from the linked report. **Paperclip version or commit** The branch includes master at `9d19f98b5`. The report records the original source for each focused attempt. **Deployment mode** Isolated local development instances and disposable Daytona sandboxes. ## What Changed - Deliver settled tool-action results for the exact company and source run during final cleanup. Keep the existing idempotent receipt and periodic recovery sweep. - Wait for startup to hand off its provider handle before acknowledging Stop. Reject first-turn admission after cancellation. Preserve a matching audited pending or acknowledged cancellation. - Wait for mounted task history and connector controls in browser tests. Record failure evidence. Grade workspace contents and process continuity separately from exact reply wording. Require each warm-turn marker once and in order, allowing surrounding prose. - Stop-and-redirect now checks that the source file exists and work is active before Stop. - Resolve target dependency locks in an uncredentialed CI job. Verify the lock artifact hash. Keep orchestration and publication on the trusted workflow revision. - Materialize the pinned OpenCode executable and configure the exact Codex executable's user-namespace profile before provider credentials are available. - Compress Daytona directory uploads with gzip. Preserve files, executable modes, symlinks, empty directories, and confinement checks. - Classify file-transfer RPC deadlines as infrastructure. Keep unrelated runner RPC failures visible. ## Verification - [Focused live report with screenshots and original attempts](https://pages.paperclip.ing/runner-reliability-20260921/): 14 of 15 selected Product E2E cases pass across the recorded revisions. Claude and Codex Stop → `/new`, Claude service approval, delegation, both hiring/reuse cases, and native Daytona warm continuity pass. - Two credentialed Runner smoke cases pass. These are not full protocol coverage. - E2E harness after the master merge: 429 tests pass. E2E and server TypeScript checks pass. - Daytona plugin: 239 tests pass, 6 skipped. Plugin TypeScript build passes. The compression test fails against the old code and passes with the change. - Runner backend/runtime regression group: 161 tests pass. Cancellation/startup selection: 26 tests pass. Approval delivery: 34 real-database tests pass. - Workflow security: 7 tests pass. Both edited workflows pass actionlint. Runner TypeScript and Rust builds pass. - After merging master, all 389 native executor tests pass, including both native backend paths and late startup after the Stop deadline. - Post-merge `pnpm -r typecheck` and `pnpm build` pass. The monolithic local `pnpm test:run` was interrupted to integrate master and is inconclusive. The [hosted CI test partitions](https://github.com/paperclipai/paperclip/actions/runs/35620461738) pass on `50a3e43822bcba1e0d07b1b45b0be91cbf9312da`. An unchanged sandbox callback schema test initially received HTTP 503. It passed five isolated local runs, its full local test file, and one failed-job CI retry. No assertion was weakened. ## Risks - Stop can wait for the bounded startup handoff. If it cannot settle, the existing pending-recovery state remains instead of a false acknowledgement. - Immediate approval delivery must remain idempotent across cleanup and recovery sweeps. Tests cover duplicate delivery and company/run boundaries. - The workflow changes still need hosted Linux verification. They retain the trusted workflow and credential boundaries. - Gzip reduces the observed provider upload from about 1.8 GB to 663 MB. It does not yet fix the remaining Claude Daytona transfer timeout. That recovery test never reached Claude, so recovery remains unverified. Use a matching image with the verified provider package preinstalled for the next recovery test; retain cold-upload coverage separately. - The report preserves diagnostic runs with missing source metadata and marks them as such. It does not claim a new full-suite pass. - This PR adds no new prompt policy or historical status reconciliation. ## Model Used OpenAI GPT-6 through Codex performed the primary implementation and review. The exact primary backend model ID is not exposed in this session. OpenAI `gpt-5.6-luna` assisted with bounded infrastructure work and verification. The agents used repository tools, code execution, and browser tests. The exact backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: OpenAI GPT-6 <noreply@openai.com>
848 lines
39 KiB
YAML
848 lines
39 KiB
YAML
name: Runner Direct Live Protocol Evals
|
|
|
|
on:
|
|
schedule:
|
|
- cron: "23 9 * * 0"
|
|
workflow_dispatch:
|
|
inputs:
|
|
target_branch:
|
|
description: "Branch in paperclipai/paperclip to evaluate; trusted orchestration still runs from master"
|
|
type: string
|
|
required: false
|
|
evals_sha:
|
|
description: "Exact 40-character paperclipai/paperclip-evals commit to execute"
|
|
type: string
|
|
required: false
|
|
rosters:
|
|
description: "Comma-separated live roster IDs/files, or all for the maintained enabled direct suite"
|
|
type: string
|
|
default: "all"
|
|
required: false
|
|
max_infrastructure_retries:
|
|
description: "Automatic retries only for explicitly retryable infrastructure failures (0-3)"
|
|
type: number
|
|
default: 1
|
|
|
|
permissions:
|
|
contents: read
|
|
|
|
concurrency:
|
|
group: runner-protocol-live-evals-${{ github.event_name == 'workflow_dispatch' && inputs.target_branch != '' && inputs.target_branch != github.event.repository.default_branch && format('development-{0}', inputs.target_branch) || format('protected-{0}', github.run_id) }}
|
|
cancel-in-progress: ${{ github.event_name == 'workflow_dispatch' && inputs.target_branch != '' && inputs.target_branch != github.event.repository.default_branch }}
|
|
|
|
jobs:
|
|
authorize:
|
|
name: Authorize paid direct eval campaign
|
|
if: github.event_name != 'schedule' || vars.RUNNER_PROTOCOL_EVAL_NIGHTLY_ENABLED == 'true'
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 5
|
|
permissions:
|
|
contents: read
|
|
outputs:
|
|
test_runner: ${{ steps.runner.outputs.runner }}
|
|
max_parallel_default: ${{ steps.runner.outputs.max_parallel_default }}
|
|
max_parallel_limit: ${{ steps.runner.outputs.max_parallel_limit }}
|
|
target_sha: ${{ steps.target.outputs.sha }}
|
|
target_ref: ${{ steps.target.outputs.ref }}
|
|
evals_sha: ${{ steps.evals.outputs.sha }}
|
|
steps:
|
|
- name: Require default branch and allowlisted numeric actor IDs
|
|
env:
|
|
GH_TOKEN: ${{ github.token }}
|
|
REPOSITORY: ${{ github.repository }}
|
|
REF: ${{ github.ref }}
|
|
DEFAULT_BRANCH: ${{ github.event.repository.default_branch }}
|
|
ACTOR: ${{ github.actor }}
|
|
ACTOR_ID: ${{ github.actor_id }}
|
|
TRIGGERING_ACTOR: ${{ github.triggering_actor }}
|
|
ALLOWED_ACTOR_IDS: ${{ vars.RUNNER_E2E_ALLOWED_ACTOR_IDS }}
|
|
run: |
|
|
set -euo pipefail
|
|
if [ "$REF" != "refs/heads/$DEFAULT_BRANCH" ]; then
|
|
echo "Paid direct Runner eval campaigns may run only from the default branch." >&2
|
|
exit 1
|
|
fi
|
|
if ! jq -e 'type == "array" and length > 0 and all(.[]; type == "number" and . > 0 and floor == .)' <<< "${ALLOWED_ACTOR_IDS:-}" >/dev/null; then
|
|
echo "RUNNER_E2E_ALLOWED_ACTOR_IDS must be a non-empty JSON array of numeric GitHub user IDs." >&2
|
|
exit 1
|
|
fi
|
|
triggering_actor_id="$(gh api "users/$TRIGGERING_ACTOR" --jq .id)"
|
|
if [ "$triggering_actor_id" != "$ACTOR_ID" ] && [ "$TRIGGERING_ACTOR" = "$ACTOR" ]; then
|
|
echo "GitHub actor identity contexts disagree; refusing the paid run." >&2
|
|
exit 1
|
|
fi
|
|
for candidate in "$triggering_actor_id" "$ACTOR_ID"; do
|
|
if ! jq -e --argjson candidate "$candidate" 'index($candidate) != null' <<< "$ALLOWED_ACTOR_IDS" >/dev/null; then
|
|
echo "The initiating GitHub account is not authorized to run paid Runner eval campaigns." >&2
|
|
exit 1
|
|
fi
|
|
done
|
|
|
|
- name: Resolve requested Paperclip branch to an immutable commit
|
|
id: target
|
|
env:
|
|
GH_TOKEN: ${{ github.token }}
|
|
REPOSITORY: ${{ github.repository }}
|
|
TARGET_BRANCH: ${{ inputs.target_branch || github.event.repository.default_branch }}
|
|
run: |
|
|
set -euo pipefail
|
|
if [ -z "$TARGET_BRANCH" ] || [[ "$TARGET_BRANCH" == refs/* ]]; then
|
|
echo "target_branch must name a branch in this repository without a refs/ prefix." >&2
|
|
exit 1
|
|
fi
|
|
encoded_branch="$(jq -rn --arg branch "$TARGET_BRANCH" '$branch | @uri')"
|
|
target_sha="$(gh api -X GET "repos/$REPOSITORY/branches/$encoded_branch" --jq .commit.sha)"
|
|
[[ "$target_sha" =~ ^[0-9a-f]{40}$ ]]
|
|
echo "sha=$target_sha" >> "$GITHUB_OUTPUT"
|
|
echo "ref=refs/heads/$TARGET_BRANCH" >> "$GITHUB_OUTPUT"
|
|
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
ref: ${{ github.sha }}
|
|
persist-credentials: false
|
|
|
|
- name: Generate private eval-repository token
|
|
id: evals_token
|
|
env:
|
|
COMMITPERCLIP_KEY: ${{ secrets.COMMITPERCLIP_KEY }}
|
|
GH_REPO: paperclipai/paperclip-evals
|
|
run: |
|
|
set -euo pipefail
|
|
token="$(node .github/scripts/get-bot-token.mjs)"
|
|
echo "::add-mask::$token"
|
|
echo "value=$token" >> "$GITHUB_OUTPUT"
|
|
|
|
- name: Verify the private eval program is pinned to an exact commit
|
|
id: evals
|
|
env:
|
|
GH_TOKEN: ${{ steps.evals_token.outputs.value }}
|
|
EVALS_SHA: ${{ inputs.evals_sha || vars.RUNNER_PROTOCOL_EVALS_SHA }}
|
|
run: |
|
|
set -euo pipefail
|
|
if ! [[ "$EVALS_SHA" =~ ^[0-9a-f]{40}$ ]]; then
|
|
echo "evals_sha (or RUNNER_PROTOCOL_EVALS_SHA for schedules) must be an exact 40-character commit." >&2
|
|
exit 1
|
|
fi
|
|
resolved="$(gh api -X GET "repos/paperclipai/paperclip-evals/commits/$EVALS_SHA" --jq .sha)"
|
|
test "$resolved" = "$EVALS_SHA"
|
|
echo "sha=$resolved" >> "$GITHUB_OUTPUT"
|
|
|
|
- name: Validate retry envelope
|
|
env:
|
|
RETRIES: ${{ github.event_name == 'schedule' && 1 || inputs.max_infrastructure_retries }}
|
|
run: |
|
|
set -euo pipefail
|
|
[[ "$RETRIES" =~ ^[0-3]$ ]]
|
|
|
|
- name: Select paid test runner
|
|
id: runner
|
|
env:
|
|
AWS_PAID_RUNNER_ENABLED: ${{ vars.RUNNER_E2E_AWS_ENABLED }}
|
|
run: |
|
|
set -euo pipefail
|
|
if [ "$AWS_PAID_RUNNER_ENABLED" = true ]; then
|
|
{
|
|
echo 'runner=runs-on/fleet=paperclip-public-pr-x64/env=public-ci'
|
|
echo 'max_parallel_default=100'
|
|
echo 'max_parallel_limit=100'
|
|
} >> "$GITHUB_OUTPUT"
|
|
else
|
|
{
|
|
echo 'runner=ubuntu-latest'
|
|
echo 'max_parallel_default=32'
|
|
echo 'max_parallel_limit=57'
|
|
} >> "$GITHUB_OUTPUT"
|
|
fi
|
|
|
|
target_lock:
|
|
name: Resolve target pnpm lockfile
|
|
needs: authorize
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 10
|
|
permissions:
|
|
contents: read
|
|
outputs:
|
|
artifact_id: ${{ steps.upload.outputs.artifact-id }}
|
|
lock_sha256: ${{ steps.lock.outputs.sha256 }}
|
|
steps:
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
ref: ${{ needs.authorize.outputs.target_sha }}
|
|
persist-credentials: false
|
|
|
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
|
|
with:
|
|
node-version: 24
|
|
|
|
- uses: pnpm/action-setup@0977fd99725f1db4007ccb2928dbb4e90d06cc86 # v6
|
|
env:
|
|
NPM_CONFIG_AUDIT: "false"
|
|
NPM_CONFIG_FUND: "false"
|
|
NPM_CONFIG_UPDATE_NOTIFIER: "false"
|
|
with:
|
|
version: 9.15.4
|
|
|
|
- name: Resolve target lockfile without lifecycle scripts
|
|
id: lock
|
|
run: |
|
|
set -euo pipefail
|
|
pnpm install --ignore-scripts --no-frozen-lockfile --lockfile-only
|
|
test -s pnpm-lock.yaml
|
|
unexpected="$(git status --short | awk '$2 != "pnpm-lock.yaml" { print }')"
|
|
if [ -n "$unexpected" ]; then
|
|
echo "Lockfile resolution changed files other than pnpm-lock.yaml:" >&2
|
|
echo "$unexpected" >&2
|
|
exit 1
|
|
fi
|
|
echo "sha256=$(sha256sum pnpm-lock.yaml | cut -d ' ' -f 1)" >> "$GITHUB_OUTPUT"
|
|
|
|
- name: Upload resolved target lockfile
|
|
id: upload
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
with:
|
|
name: runner-protocol-target-pnpm-lock-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: pnpm-lock.yaml
|
|
retention-days: 30
|
|
if-no-files-found: error
|
|
|
|
catalog:
|
|
name: Pin and fan out the direct Evalbook roster
|
|
needs: authorize
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 10
|
|
permissions:
|
|
contents: read
|
|
outputs:
|
|
matrix_0: ${{ steps.catalog.outputs.matrix_0 }}
|
|
matrix_1: ${{ steps.catalog.outputs.matrix_1 }}
|
|
matrix_1_present: ${{ steps.catalog.outputs.matrix_1_present }}
|
|
max_parallel_per_shard: ${{ steps.catalog.outputs.max_parallel_per_shard }}
|
|
selected: ${{ steps.catalog.outputs.selected }}
|
|
steps:
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
ref: ${{ github.sha }}
|
|
persist-credentials: false
|
|
|
|
- name: Generate private eval-repository token
|
|
id: evals_token
|
|
env:
|
|
COMMITPERCLIP_KEY: ${{ secrets.COMMITPERCLIP_KEY }}
|
|
GH_REPO: paperclipai/paperclip-evals
|
|
run: |
|
|
set -euo pipefail
|
|
token="$(node .github/scripts/get-bot-token.mjs)"
|
|
echo "::add-mask::$token"
|
|
echo "value=$token" >> "$GITHUB_OUTPUT"
|
|
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
repository: paperclipai/paperclip-evals
|
|
ref: ${{ needs.authorize.outputs.evals_sha }}
|
|
path: .paperclip-evals
|
|
token: ${{ steps.evals_token.outputs.value }}
|
|
persist-credentials: false
|
|
|
|
- name: Build the two bounded roster-plus-case matrices
|
|
id: catalog
|
|
env:
|
|
PAPERCLIP_PROTOCOL_EVAL_SOURCE_SHA: ${{ needs.authorize.outputs.target_sha }}
|
|
PAPERCLIP_PROTOCOL_EVALS_SHA: ${{ needs.authorize.outputs.evals_sha }}
|
|
MAX_PARALLEL: ${{ vars.RUNNER_E2E_MAX_PARALLEL || needs.authorize.outputs.max_parallel_default }}
|
|
MAX_PARALLEL_LIMIT: ${{ needs.authorize.outputs.max_parallel_limit }}
|
|
ROSTERS: ${{ inputs.rosters || 'all' }}
|
|
run: |
|
|
set -euo pipefail
|
|
if ! [[ "$MAX_PARALLEL" =~ ^[1-9][0-9]*$ ]] || [ "$MAX_PARALLEL" -lt 2 ] || [ "$MAX_PARALLEL" -gt "$MAX_PARALLEL_LIMIT" ]; then
|
|
echo "RUNNER_E2E_MAX_PARALLEL must be an integer from 2 through $MAX_PARALLEL_LIMIT for the two-shard direct suite." >&2
|
|
exit 1
|
|
fi
|
|
node packages/paperclip-runner/scripts/runner-protocol-eval-campaign.mjs catalog \
|
|
--evals-root .paperclip-evals \
|
|
--rosters "$ROSTERS" \
|
|
--campaign-id "gha-${GITHUB_RUN_ID}-${GITHUB_RUN_ATTEMPT}" \
|
|
--max-parallel "$MAX_PARALLEL" \
|
|
--output runner-protocol-eval-catalog.json
|
|
|
|
- name: Require the chat-report renderer before paid execution
|
|
run: |
|
|
set -euo pipefail
|
|
python3 .paperclip-evals/evals/paperclip-runner/tools/eval_program.py report --help | grep -q -- --public-viewer
|
|
|
|
- name: Upload immutable campaign catalog
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
with:
|
|
name: runner-protocol-eval-catalog-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: runner-protocol-eval-catalog.json
|
|
retention-days: 30
|
|
if-no-files-found: error
|
|
|
|
build_runner:
|
|
name: Build portable direct-eval runner once
|
|
needs: [authorize, target_lock, catalog]
|
|
runs-on: ${{ needs.authorize.outputs.test_runner }}
|
|
timeout-minutes: 30
|
|
permissions:
|
|
contents: read
|
|
steps:
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
ref: ${{ needs.authorize.outputs.target_sha }}
|
|
persist-credentials: false
|
|
|
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
|
|
with:
|
|
node-version: 24
|
|
|
|
- uses: pnpm/action-setup@0977fd99725f1db4007ccb2928dbb4e90d06cc86 # v6
|
|
env:
|
|
NPM_CONFIG_AUDIT: "false"
|
|
NPM_CONFIG_FUND: "false"
|
|
NPM_CONFIG_UPDATE_NOTIFIER: "false"
|
|
with:
|
|
version: 9.15.4
|
|
|
|
- name: Download resolved target lockfile
|
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
|
|
with:
|
|
artifact-ids: ${{ needs.target_lock.outputs.artifact_id }}
|
|
path: ${{ runner.temp }}/runner-protocol-target-lock
|
|
|
|
- name: Restore resolved target lockfile
|
|
env:
|
|
TARGET_SHA: ${{ needs.authorize.outputs.target_sha }}
|
|
EXPECTED_LOCK_SHA256: ${{ needs.target_lock.outputs.lock_sha256 }}
|
|
run: |
|
|
set -euo pipefail
|
|
test "$(git rev-parse HEAD)" = "$TARGET_SHA"
|
|
lock="$RUNNER_TEMP/runner-protocol-target-lock/pnpm-lock.yaml"
|
|
test -f "$lock"
|
|
test "$(find "$(dirname "$lock")" -type f | wc -l | tr -d ' ')" = 1
|
|
test "$(sha256sum "$lock" | cut -d ' ' -f 1)" = "$EXPECTED_LOCK_SHA256"
|
|
cp "$lock" pnpm-lock.yaml
|
|
test "$(sha256sum pnpm-lock.yaml | cut -d ' ' -f 1)" = "$EXPECTED_LOCK_SHA256"
|
|
|
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
|
|
with:
|
|
node-version: 24
|
|
cache: pnpm
|
|
|
|
- run: pnpm install --frozen-lockfile --ignore-scripts
|
|
|
|
- name: Materialize the pinned OpenCode executable before packaging
|
|
run: node packages/paperclip-runner/scripts/materialize-opencode-binary.mjs
|
|
|
|
- name: Build runner CLI, daemon, and canonical attempt viewer
|
|
run: |
|
|
set -euo pipefail
|
|
pnpm --filter @paperclipai/paperclip-runner build:typescript
|
|
pnpm --filter @paperclipai/paperclip-runner build:runner-binaries
|
|
pnpm --filter @paperclipai/paperclip-runner build:issue-thread
|
|
# Older target refs must fail before paid cells, not publish an empty viewer.
|
|
grep -q 'paperclip-eval-report' packages/paperclip-runner/dist-issue-thread/assets/*.js
|
|
grep -q 'evalbook-site' packages/paperclip-runner/dist-issue-thread/assets/*.css
|
|
|
|
- name: Package a portable provider runtime
|
|
run: |
|
|
set -euo pipefail
|
|
mkdir -p "$RUNNER_TEMP/runner-protocol-build/package" "$RUNNER_TEMP/runner-protocol-build/portable"
|
|
pnpm --dir packages/paperclip-runner pack \
|
|
--pack-destination "$RUNNER_TEMP/runner-protocol-build/package"
|
|
package="$(find "$RUNNER_TEMP/runner-protocol-build/package" -maxdepth 1 -type f -name '*.tgz' -print -quit)"
|
|
test -f "$package"
|
|
pnpm --filter @paperclipai/paperclip-runner deploy --prod \
|
|
"$RUNNER_TEMP/runner-protocol-build/portable"
|
|
cp "$package" "$RUNNER_TEMP/runner-protocol-build/paperclip-runner.tgz"
|
|
cp packages/paperclip-runner/runner/target/debug/paperclip-runnerd "$RUNNER_TEMP/runner-protocol-build/paperclip-runnerd"
|
|
cp -R packages/paperclip-runner/dist-issue-thread "$RUNNER_TEMP/runner-protocol-build/dist-issue-thread"
|
|
test -f "$RUNNER_TEMP/runner-protocol-build/portable/dist/cli/eval-session.js"
|
|
test -d "$RUNNER_TEMP/runner-protocol-build/portable/node_modules/.pnpm"
|
|
test -x "$RUNNER_TEMP/runner-protocol-build/paperclip-runnerd"
|
|
tar --create --gzip --file runner-protocol-build.tar.gz -C "$RUNNER_TEMP/runner-protocol-build" .
|
|
sha256sum runner-protocol-build.tar.gz > runner-protocol-build.tar.gz.sha256
|
|
|
|
- name: Upload immutable portable runner
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
with:
|
|
name: runner-protocol-build-${{ needs.authorize.outputs.target_sha }}-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: |
|
|
runner-protocol-build.tar.gz
|
|
runner-protocol-build.tar.gz.sha256
|
|
retention-days: 1
|
|
compression-level: 0
|
|
if-no-files-found: error
|
|
|
|
- name: Upload canonical viewer for publisher byte verification
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
with:
|
|
name: runner-protocol-viewer-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: packages/paperclip-runner/dist-issue-thread/
|
|
retention-days: 30
|
|
if-no-files-found: error
|
|
|
|
eval_shard_0:
|
|
name: Direct eval ${{ matrix.rosterId }} / ${{ matrix.caseId }}
|
|
needs: [authorize, catalog, build_runner]
|
|
runs-on: ${{ needs.authorize.outputs.test_runner }}
|
|
timeout-minutes: 18
|
|
permissions:
|
|
contents: read
|
|
id-token: write
|
|
environment:
|
|
name: runner-e2e-paid
|
|
strategy:
|
|
fail-fast: false
|
|
max-parallel: ${{ fromJSON(needs.catalog.outputs.max_parallel_per_shard) }}
|
|
matrix: ${{ fromJSON(needs.catalog.outputs.matrix_0) }}
|
|
steps: &direct_eval_steps
|
|
- name: Reauthorize paid execution before provider access
|
|
env:
|
|
GH_TOKEN: ${{ github.token }}
|
|
REF: ${{ github.ref }}
|
|
DEFAULT_BRANCH: ${{ github.event.repository.default_branch }}
|
|
ACTOR: ${{ github.actor }}
|
|
ACTOR_ID: ${{ github.actor_id }}
|
|
TRIGGERING_ACTOR: ${{ github.triggering_actor }}
|
|
ALLOWED_ACTOR_IDS: ${{ vars.RUNNER_E2E_ALLOWED_ACTOR_IDS }}
|
|
run: |
|
|
set -euo pipefail
|
|
test "$REF" = "refs/heads/$DEFAULT_BRANCH"
|
|
triggering_actor_id="$(gh api "users/$TRIGGERING_ACTOR" --jq .id)"
|
|
test "$triggering_actor_id" = "$ACTOR_ID" || test "$TRIGGERING_ACTOR" != "$ACTOR"
|
|
for candidate in "$triggering_actor_id" "$ACTOR_ID"; do
|
|
jq -e --argjson candidate "$candidate" 'type == "array" and index($candidate) != null' <<< "$ALLOWED_ACTOR_IDS" >/dev/null
|
|
done
|
|
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
ref: ${{ github.sha }}
|
|
persist-credentials: false
|
|
|
|
- name: Generate private eval-repository token
|
|
id: evals_token
|
|
env:
|
|
COMMITPERCLIP_KEY: ${{ secrets.COMMITPERCLIP_KEY }}
|
|
GH_REPO: paperclipai/paperclip-evals
|
|
run: |
|
|
set -euo pipefail
|
|
token="$(node .github/scripts/get-bot-token.mjs)"
|
|
echo "::add-mask::$token"
|
|
echo "value=$token" >> "$GITHUB_OUTPUT"
|
|
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
repository: paperclipai/paperclip-evals
|
|
ref: ${{ needs.authorize.outputs.evals_sha }}
|
|
path: .paperclip-evals
|
|
token: ${{ steps.evals_token.outputs.value }}
|
|
persist-credentials: false
|
|
|
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
|
|
with:
|
|
node-version: 24
|
|
|
|
- name: Download portable runner
|
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
|
|
with:
|
|
name: runner-protocol-build-${{ needs.authorize.outputs.target_sha }}-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: runner-protocol-build
|
|
|
|
- name: Verify and extract portable runner
|
|
run: |
|
|
set -euo pipefail
|
|
cd runner-protocol-build
|
|
sha256sum --check runner-protocol-build.tar.gz.sha256
|
|
mkdir extracted
|
|
tar --extract --gzip --file runner-protocol-build.tar.gz --directory extracted
|
|
test -x extracted/paperclip-runnerd
|
|
|
|
# Keep the host policy aligned with runner-full-stack-e2e.yml. This is
|
|
# deliberately before any provider credential or web-identity step.
|
|
- name: Provision Codex sandbox on the disposable trusted runner
|
|
if: matrix.provider == 'codex' || matrix.rosterId == 'protocol-live-acpx-codex-control'
|
|
run: |
|
|
node --input-type=module <<'NODE'
|
|
import { execFileSync } from "node:child_process";
|
|
import { createHash } from "node:crypto";
|
|
import { readFileSync, realpathSync, writeFileSync } from "node:fs";
|
|
import { createRequire } from "node:module";
|
|
import path from "node:path";
|
|
if (process.platform !== "linux") process.exit(0);
|
|
let restricted = "0";
|
|
try { restricted = readFileSync("/proc/sys/kernel/apparmor_restrict_unprivileged_userns", "utf8").trim(); } catch {}
|
|
if (restricted !== "1") process.exit(0);
|
|
const root = realpathSync(path.join(process.env.GITHUB_WORKSPACE, "runner-protocol-build/extracted/portable"));
|
|
const runnerRequire = createRequire(path.join(root, "package.json"));
|
|
const acpRequire = createRequire(runnerRequire.resolve("@agentclientprotocol/codex-acp/package.json"));
|
|
const codexRequire = createRequire(acpRequire.resolve("@openai/codex/package.json"));
|
|
const arch = process.arch === "x64" ? "x64" : process.arch === "arm64" ? "arm64" : null;
|
|
if (!arch) throw new Error("Unsupported Codex CI architecture");
|
|
const platformPackage = codexRequire.resolve(`@openai/codex-linux-${arch}/package.json`);
|
|
const triple = arch === "x64" ? "x86_64-unknown-linux-musl" : "aarch64-unknown-linux-musl";
|
|
const suffix = `/vendor/${triple}/bin/codex`;
|
|
const binary = realpathSync(path.join(path.dirname(platformPackage), suffix));
|
|
if (!binary.startsWith(root + "/node_modules/.pnpm/") || !binary.endsWith(suffix) || !/^[/A-Za-z0-9_.@+\-]+$/.test(binary)) {
|
|
throw new Error("Codex executable is outside the resolved dependency tree");
|
|
}
|
|
const name = `paperclip-e2e-codex-${createHash("sha256").update(binary).digest("hex").slice(0,16)}`;
|
|
const profilePath = path.join(process.env.RUNNER_TEMP, "paperclip-codex-userns.apparmor");
|
|
writeFileSync(profilePath, `abi <abi/4.0>,\ninclude <tunables/global>\nprofile ${name} "${binary}" flags=(unconfined) {\n userns,\n}\n`, {mode:0o600, flag:"wx"});
|
|
execFileSync("sudo", ["-n", "apparmor_parser", "-r", profilePath], {timeout:15000, stdio:"pipe"});
|
|
NODE
|
|
|
|
- name: Prepare short-lived AgentCore web identity
|
|
if: matrix.credentialName == 'AWS_AGENTCORE_OIDC'
|
|
env:
|
|
AGENTCORE_ROLE_ARN: ${{ vars.PAPERCLIP_AWS_AGENTCORE_EXECUTION_ROLE_ARN }}
|
|
run: |
|
|
set -euo pipefail
|
|
test -n "$AGENTCORE_ROLE_ARN"
|
|
token="$(curl --fail --silent --show-error \
|
|
-H "Authorization: Bearer $ACTIONS_ID_TOKEN_REQUEST_TOKEN" \
|
|
"${ACTIONS_ID_TOKEN_REQUEST_URL}&audience=sts.amazonaws.com" | jq -r .value)"
|
|
test -n "$token"
|
|
echo "::add-mask::$token"
|
|
token_file="$RUNNER_TEMP/runner-protocol-agentcore-token"
|
|
printf '%s' "$token" > "$token_file"
|
|
chmod 600 "$token_file"
|
|
{
|
|
echo "AWS_WEB_IDENTITY_TOKEN_FILE=$token_file"
|
|
echo "AWS_ROLE_ARN=$AGENTCORE_ROLE_ARN"
|
|
echo "AWS_ROLE_SESSION_NAME=runner-protocol-${GITHUB_RUN_ID}-${GITHUB_RUN_ATTEMPT}"
|
|
} >> "$GITHUB_ENV"
|
|
|
|
- name: Run one immutable direct protocol cell
|
|
id: direct_eval
|
|
env:
|
|
CELL_ID: ${{ matrix.cellId }}
|
|
ROSTER_FILE: ${{ matrix.rosterFile }}
|
|
CASE_ID: ${{ matrix.caseId }}
|
|
CREDENTIAL_NAME: ${{ matrix.credentialName }}
|
|
PROVIDER: ${{ matrix.provider }}
|
|
MAX_INFRASTRUCTURE_RETRIES: ${{ github.event_name == 'schedule' && 1 || inputs.max_infrastructure_retries }}
|
|
OPENAI_API_KEY: ${{ matrix.credentialName == 'OPENAI_API_KEY' && secrets.OPENAI_API_KEY || '' }}
|
|
ANTHROPIC_API_KEY: ${{ matrix.credentialName == 'ANTHROPIC_API_KEY' && secrets.ANTHROPIC_API_KEY || '' }}
|
|
OPENROUTER_API_KEY: ${{ matrix.credentialName == 'OPENROUTER_API_KEY' && secrets.OPENROUTER_API_KEY || '' }}
|
|
PAPERCLIP_CLAUDE_MANAGED_PROFILE_ID: ${{ vars.PAPERCLIP_CLAUDE_MANAGED_PROFILE_ID }}
|
|
PAPERCLIP_CLAUDE_MANAGED_AGENT_ID: ${{ vars.PAPERCLIP_CLAUDE_MANAGED_AGENT_ID }}
|
|
PAPERCLIP_CLAUDE_MANAGED_AGENT_VERSION: ${{ vars.PAPERCLIP_CLAUDE_MANAGED_AGENT_VERSION }}
|
|
PAPERCLIP_CLAUDE_MANAGED_ENVIRONMENT_ID: ${{ vars.PAPERCLIP_CLAUDE_MANAGED_ENVIRONMENT_ID }}
|
|
AWS_REGION: ${{ vars.PAPERCLIP_AWS_AGENTCORE_REGION }}
|
|
AWS_DEFAULT_REGION: ${{ vars.PAPERCLIP_AWS_AGENTCORE_REGION }}
|
|
PAPERCLIP_AWS_AGENTCORE_PROFILE_ID: ${{ vars.PAPERCLIP_AWS_AGENTCORE_PROFILE_ID }}
|
|
PAPERCLIP_AWS_AGENTCORE_ACCOUNT_ID: ${{ vars.PAPERCLIP_AWS_AGENTCORE_ACCOUNT_ID }}
|
|
PAPERCLIP_AWS_AGENTCORE_HARNESS_ARN: ${{ vars.PAPERCLIP_AWS_AGENTCORE_HARNESS_ARN }}
|
|
PAPERCLIP_AWS_AGENTCORE_HARNESS_VERSION: ${{ vars.PAPERCLIP_AWS_AGENTCORE_HARNESS_VERSION }}
|
|
PAPERCLIP_AWS_AGENTCORE_ENDPOINT_ARN: ${{ vars.PAPERCLIP_AWS_AGENTCORE_ENDPOINT_ARN }}
|
|
PAPERCLIP_AWS_AGENTCORE_ENDPOINT_QUALIFIER: ${{ vars.PAPERCLIP_AWS_AGENTCORE_ENDPOINT_QUALIFIER }}
|
|
PAPERCLIP_AWS_AGENTCORE_RUNTIME_ARN: ${{ vars.PAPERCLIP_AWS_AGENTCORE_RUNTIME_ARN }}
|
|
PAPERCLIP_AWS_AGENTCORE_MEMORY_ARN: ${{ vars.PAPERCLIP_AWS_AGENTCORE_MEMORY_ARN }}
|
|
PAPERCLIP_AWS_AGENTCORE_MEMORY_ID: ${{ vars.PAPERCLIP_AWS_AGENTCORE_MEMORY_ID }}
|
|
PAPERCLIP_AWS_AGENTCORE_INVOCATION_ROLE_ARN: ${{ vars.PAPERCLIP_AWS_AGENTCORE_INVOCATION_ROLE_ARN }}
|
|
PAPERCLIP_AWS_AGENTCORE_CONTEXT_BUCKET: ${{ vars.PAPERCLIP_AWS_AGENTCORE_CONTEXT_BUCKET }}
|
|
PAPERCLIP_AWS_AGENTCORE_CONTEXT_PREFIX: ${{ vars.PAPERCLIP_AWS_AGENTCORE_CONTEXT_PREFIX }}
|
|
PAPERCLIP_AWS_AGENTCORE_CONTEXT_KMS_KEY_ARN: ${{ vars.PAPERCLIP_AWS_AGENTCORE_CONTEXT_KMS_KEY_ARN }}
|
|
PAPERCLIP_AWS_AGENTCORE_QUALIFICATION_REVISION: ${{ vars.PAPERCLIP_AWS_AGENTCORE_QUALIFICATION_REVISION }}
|
|
run: |
|
|
set -euo pipefail
|
|
mkdir -p cell-output/runs
|
|
if [ "$CREDENTIAL_NAME" != "AWS_AGENTCORE_OIDC" ]; then
|
|
test -n "${!CREDENTIAL_NAME:-}"
|
|
fi
|
|
if [ "$PROVIDER" = "claude_managed" ]; then
|
|
test -n "$PAPERCLIP_CLAUDE_MANAGED_PROFILE_ID"
|
|
test -n "$PAPERCLIP_CLAUDE_MANAGED_AGENT_ID"
|
|
test -n "$PAPERCLIP_CLAUDE_MANAGED_AGENT_VERSION"
|
|
test -n "$PAPERCLIP_CLAUDE_MANAGED_ENVIRONMENT_ID"
|
|
fi
|
|
set +e
|
|
python3 .paperclip-evals/evals/paperclip-runner/tools/run_live_roster.py run \
|
|
--roster ".paperclip-evals/evals/paperclip-runner/rosters/$ROSTER_FILE" \
|
|
--case "$CASE_ID" \
|
|
--runner-cli runner-protocol-build/extracted/portable/dist/cli/eval-session.js \
|
|
--runner-package runner-protocol-build/extracted/paperclip-runner.tgz \
|
|
--runnerd runner-protocol-build/extracted/paperclip-runnerd \
|
|
--runs-root cell-output/runs \
|
|
--max-infrastructure-retries "$MAX_INFRASTRUCTURE_RETRIES" \
|
|
--run-id "gha-${GITHUB_RUN_ID}-${GITHUB_RUN_ATTEMPT}-${CELL_ID}"
|
|
status=$?
|
|
set -e
|
|
CELL_EXIT_CODE="$status" node --input-type=module <<'NODE'
|
|
import { writeFileSync } from "node:fs";
|
|
writeFileSync("cell-output/cell.json", `${JSON.stringify({
|
|
schema: "paperclip.runner-protocol-eval.cell/v1",
|
|
cellId: process.env.CELL_ID,
|
|
rosterFile: process.env.ROSTER_FILE,
|
|
caseId: process.env.CASE_ID,
|
|
exitCode: Number(process.env.CELL_EXIT_CODE),
|
|
}, null, 2)}\n`, { mode: 0o600 });
|
|
NODE
|
|
exit "$status"
|
|
|
|
- name: Upload access-controlled cell attempt
|
|
if: always()
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
with:
|
|
name: runner-protocol-eval-${{ github.run_id }}-${{ github.run_attempt }}-${{ matrix.cellId }}
|
|
path: cell-output/
|
|
retention-days: 30
|
|
if-no-files-found: error
|
|
|
|
eval_shard_1:
|
|
name: Direct eval ${{ matrix.rosterId }} / ${{ matrix.caseId }}
|
|
if: needs.catalog.outputs.matrix_1_present == 'true'
|
|
needs: [authorize, catalog, build_runner]
|
|
runs-on: ${{ needs.authorize.outputs.test_runner }}
|
|
timeout-minutes: 18
|
|
permissions:
|
|
contents: read
|
|
id-token: write
|
|
environment:
|
|
name: runner-e2e-paid
|
|
strategy:
|
|
fail-fast: false
|
|
max-parallel: ${{ fromJSON(needs.catalog.outputs.max_parallel_per_shard) }}
|
|
matrix: ${{ fromJSON(needs.catalog.outputs.matrix_1) }}
|
|
steps: *direct_eval_steps
|
|
|
|
report:
|
|
name: Merge attempts and render canonical Evalbook
|
|
if: always() && !cancelled() && needs.catalog.result == 'success' && needs.build_runner.result == 'success'
|
|
needs: [authorize, catalog, build_runner, eval_shard_0, eval_shard_1]
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 30
|
|
permissions:
|
|
actions: read
|
|
contents: read
|
|
outputs:
|
|
public_report_ready: ${{ steps.public_report.outputs.ready }}
|
|
steps:
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
ref: ${{ github.sha }}
|
|
persist-credentials: false
|
|
|
|
- name: Generate private eval-repository token
|
|
id: evals_token
|
|
env:
|
|
COMMITPERCLIP_KEY: ${{ secrets.COMMITPERCLIP_KEY }}
|
|
GH_REPO: paperclipai/paperclip-evals
|
|
run: |
|
|
set -euo pipefail
|
|
token="$(node .github/scripts/get-bot-token.mjs)"
|
|
echo "::add-mask::$token"
|
|
echo "value=$token" >> "$GITHUB_OUTPUT"
|
|
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
repository: paperclipai/paperclip-evals
|
|
ref: ${{ needs.authorize.outputs.evals_sha }}
|
|
path: .paperclip-evals
|
|
token: ${{ steps.evals_token.outputs.value }}
|
|
persist-credentials: false
|
|
|
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
|
|
with:
|
|
node-version: 24
|
|
|
|
- uses: pnpm/action-setup@0977fd99725f1db4007ccb2928dbb4e90d06cc86 # v6
|
|
env:
|
|
NPM_CONFIG_AUDIT: "false"
|
|
NPM_CONFIG_FUND: "false"
|
|
NPM_CONFIG_UPDATE_NOTIFIER: "false"
|
|
with:
|
|
version: 9.15.4
|
|
|
|
- name: Resolve trusted report lockfile without lifecycle scripts
|
|
run: |
|
|
set -euo pipefail
|
|
pnpm install --ignore-scripts --no-frozen-lockfile --lockfile-only
|
|
test -s pnpm-lock.yaml
|
|
|
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
|
|
with:
|
|
node-version: 24
|
|
cache: pnpm
|
|
|
|
- name: Download immutable campaign catalog
|
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
|
|
with:
|
|
name: runner-protocol-eval-catalog-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: runner-protocol-catalog
|
|
|
|
- name: Download portable runner and viewer
|
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
|
|
with:
|
|
name: runner-protocol-build-${{ needs.authorize.outputs.target_sha }}-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: runner-protocol-build
|
|
|
|
- name: Download every access-controlled cell
|
|
id: download_cells
|
|
continue-on-error: true
|
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
|
|
with:
|
|
pattern: runner-protocol-eval-${{ github.run_id }}-${{ github.run_attempt }}-*
|
|
path: downloaded-runner-protocol-evals
|
|
merge-multiple: false
|
|
|
|
- name: Retry cell download after artifact transport failure
|
|
if: steps.download_cells.outcome == 'failure'
|
|
continue-on-error: true
|
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
|
|
with:
|
|
pattern: runner-protocol-eval-${{ github.run_id }}-${{ github.run_attempt }}-*
|
|
path: downloaded-runner-protocol-evals
|
|
merge-multiple: false
|
|
|
|
- name: Materialize an empty download root when every cell failed early
|
|
run: mkdir -p downloaded-runner-protocol-evals
|
|
|
|
- name: Verify portable viewer
|
|
run: |
|
|
set -euo pipefail
|
|
cd runner-protocol-build
|
|
sha256sum --check runner-protocol-build.tar.gz.sha256
|
|
mkdir extracted
|
|
tar --extract --gzip --file runner-protocol-build.tar.gz --directory extracted
|
|
test -f extracted/dist-issue-thread/index.html
|
|
|
|
- name: Aggregate every expected cell, including missing infrastructure cells
|
|
env:
|
|
PAPERCLIP_PROTOCOL_EVAL_SOURCE_SHA: ${{ needs.authorize.outputs.target_sha }}
|
|
PAPERCLIP_PROTOCOL_EVAL_SOURCE_REF: ${{ needs.authorize.outputs.target_ref }}
|
|
PAPERCLIP_PROTOCOL_EVALS_SHA: ${{ needs.authorize.outputs.evals_sha }}
|
|
PAPERCLIP_PROTOCOL_EVAL_WORKFLOW_URL: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}
|
|
run: |
|
|
node packages/paperclip-runner/scripts/runner-protocol-eval-campaign.mjs aggregate \
|
|
--catalog runner-protocol-catalog/runner-protocol-eval-catalog.json \
|
|
--downloads downloaded-runner-protocol-evals \
|
|
--evals-root .paperclip-evals \
|
|
--runs-out runner-protocol-merged/runs \
|
|
--campaign-out runner-protocol-merged/campaign.json
|
|
|
|
- name: Render the access-controlled canonical Evalbook report
|
|
run: |
|
|
python3 .paperclip-evals/evals/paperclip-runner/tools/eval_program.py report \
|
|
--runs-root runner-protocol-merged/runs \
|
|
--output runner-protocol-merged/report \
|
|
--viewer-root runner-protocol-build/extracted/dist-issue-thread \
|
|
--inventory .paperclip-evals/evals/paperclip-runner/inventory.json \
|
|
--coverage-matrix .paperclip-evals/evals/paperclip-runner/coverage-matrix.json
|
|
cp runner-protocol-merged/campaign.json runner-protocol-merged/report/campaign.json
|
|
|
|
- name: Render the same canonical grid from a public-safe evidence projection
|
|
run: |
|
|
node packages/paperclip-runner/scripts/runner-protocol-eval-campaign.mjs sanitize \
|
|
--runs-root runner-protocol-merged/runs \
|
|
--output runner-protocol-merged/public-runs
|
|
python3 .paperclip-evals/evals/paperclip-runner/tools/eval_program.py report \
|
|
--runs-root runner-protocol-merged/public-runs \
|
|
--output runner-protocol-merged/public-report \
|
|
--viewer-root runner-protocol-build/extracted/dist-issue-thread \
|
|
--public-viewer \
|
|
--inventory .paperclip-evals/evals/paperclip-runner/inventory.json \
|
|
--coverage-matrix .paperclip-evals/evals/paperclip-runner/coverage-matrix.json
|
|
cp runner-protocol-merged/campaign.json runner-protocol-merged/public-report/campaign.json
|
|
|
|
- name: Set up report browser verification
|
|
uses: pnpm/action-setup@0977fd99725f1db4007ccb2928dbb4e90d06cc86 # v6
|
|
with:
|
|
version: 9.15.4
|
|
|
|
- name: Verify the actual chat viewer before publication
|
|
run: |
|
|
pnpm install --frozen-lockfile --ignore-scripts
|
|
pnpm --filter @paperclipai/paperclip-runner exec playwright install --with-deps chromium
|
|
node packages/paperclip-runner/scripts/verify-runner-evalbook-viewer.mjs --report-root runner-protocol-merged/public-report --screenshots runner-protocol-merged/viewer-proof
|
|
node packages/paperclip-runner/scripts/verify-runner-evalbook-viewer.mjs --report-root runner-protocol-merged/report
|
|
|
|
- name: Enforce the static public allowlist
|
|
id: public_report
|
|
run: |
|
|
node --input-type=module -e 'import { validatePublicProtocolEvalReport } from "./packages/paperclip-runner/scripts/publish-runner-protocol-eval-history.mjs"; await validatePublicProtocolEvalReport("runner-protocol-merged/public-report", { viewerRoot: "runner-protocol-build/extracted/dist-issue-thread" });'
|
|
echo "ready=true" >> "$GITHUB_OUTPUT"
|
|
|
|
- name: Add campaign result to the workflow summary
|
|
run: |
|
|
{
|
|
echo '## Runner direct live protocol evals'
|
|
echo
|
|
jq -r '"- Cells: \(.totals.passed)/\(.totals.selected) passed\n- Behavior failures: \(.totals.behaviorFailures)\n- Infrastructure failures: \(.totals.infrastructureFailures)\n- Paperclip: `\(.source.paperclip.sha)`\n- Evals: `\(.source.evals.sha)`"' runner-protocol-merged/campaign.json
|
|
} >> "$GITHUB_STEP_SUMMARY"
|
|
|
|
- name: Upload access-controlled canonical Evalbook and raw attempts
|
|
if: always()
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
with:
|
|
name: runner-protocol-eval-report-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: runner-protocol-merged/
|
|
retention-days: 30
|
|
if-no-files-found: error
|
|
|
|
- name: Upload publisher-only sanitized Evalbook
|
|
if: steps.public_report.outputs.ready == 'true'
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
with:
|
|
name: runner-protocol-eval-public-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: runner-protocol-merged/public-report/
|
|
retention-days: 1
|
|
if-no-files-found: error
|
|
|
|
- name: Enforce complete green campaign
|
|
if: always()
|
|
run: jq -e '.complete == true and .allPassed == true' runner-protocol-merged/campaign.json >/dev/null
|
|
|
|
publish_history:
|
|
name: Publish immutable Evalbook and mutable campaign index
|
|
needs: [authorize, catalog, report]
|
|
if: always() && needs.report.outputs.public_report_ready == 'true'
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 15
|
|
concurrency:
|
|
group: runner-protocol-eval-history-publish
|
|
cancel-in-progress: false
|
|
permissions:
|
|
contents: read
|
|
id-token: write
|
|
environment:
|
|
name: runner-e2e-history
|
|
url: ${{ steps.publish.outputs.report_url }}
|
|
steps:
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
# AWS credentials can execute only the publisher from the trusted workflow revision.
|
|
ref: ${{ github.sha }}
|
|
persist-credentials: false
|
|
|
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
|
|
with:
|
|
node-version: 24
|
|
|
|
- name: Download only the sanitized canonical Evalbook
|
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
|
|
with:
|
|
name: runner-protocol-eval-public-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: runner-protocol-public-report
|
|
|
|
- name: Download the same-run canonical viewer for byte verification
|
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
|
|
with:
|
|
name: runner-protocol-viewer-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: runner-protocol-trusted-viewer
|
|
|
|
- name: Exchange GitHub OIDC identity for scoped AWS credentials
|
|
uses: aws-actions/configure-aws-credentials@e6de054238d6b7531b4efff3b6587d9aade6a06c # v6
|
|
with:
|
|
role-to-assume: ${{ vars.RUNNER_PROTOCOL_EVAL_HISTORY_AWS_ROLE_ARN || vars.RUNNER_E2E_HISTORY_AWS_ROLE_ARN }}
|
|
aws-region: ${{ vars.RUNNER_PROTOCOL_EVAL_HISTORY_AWS_REGION || vars.RUNNER_E2E_HISTORY_AWS_REGION }}
|
|
|
|
- name: Publish versioned report and refresh the root index
|
|
id: publish
|
|
env:
|
|
PAPERCLIP_RUNNER_PROTOCOL_EVAL_PUBLIC_REPORT_DIR: ${{ github.workspace }}/runner-protocol-public-report
|
|
PAPERCLIP_RUNNER_PROTOCOL_EVAL_VIEWER_DIR: ${{ github.workspace }}/runner-protocol-trusted-viewer
|
|
RUNNER_PROTOCOL_EVAL_HISTORY_S3_BUCKET: ${{ vars.RUNNER_PROTOCOL_EVAL_HISTORY_S3_BUCKET || vars.RUNNER_E2E_HISTORY_S3_BUCKET }}
|
|
RUNNER_PROTOCOL_EVAL_HISTORY_PREFIX: ${{ vars.RUNNER_PROTOCOL_EVAL_HISTORY_PREFIX || 'runner-protocol-evals' }}
|
|
RUNNER_PROTOCOL_EVAL_HISTORY_PUBLIC_BASE_URL: ${{ vars.RUNNER_PROTOCOL_EVAL_HISTORY_PUBLIC_BASE_URL || vars.RUNNER_E2E_HISTORY_PUBLIC_BASE_URL }}
|
|
run: node packages/paperclip-runner/scripts/publish-runner-protocol-eval-history.mjs
|