mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 21:03:51 +02:00
## Thinking Path > - Paperclip manages agents through a shared native runner. > - Built-in harness support should ship with Paperclip's public distribution. > - Grok already speaks ACP; it does not require a new public bridge package. > - Sandbox provisioning owns the native executable and its pinned version. > - The runner must verify that prerequisite without downloading it during npm installation. > - This change separates built-in launcher identity from external runtime identity. > - Clean npm installation and live staging checks verify the distribution boundary. ## Linked Issues or Issue Description Refs #13882, #13973, #13977, #13979. This follow-up now targets master after #13882 was squash-merged. It replaces the private `@paperclipai/grok-acp` workspace package with runner-owned assets. Current master is included so the branch also contains the merged scheduler, complete-event capture, and durable cleanup fixes. ## What Changed - Ship Grok launcher and qualification metadata inside the runner's compiled output and the public server's vendored runner tree. - Remove the separate Grok npm package and all package-manager install hooks for this runtime. - Require the checksum-verified Grok Build 1.0.13 binary at `/opt/paperclip/providers/grok/1.0.13/grok` in the selected execution environment. Provision it explicitly in the Daytona image and CI setup. - Keep native binaries outside the provider pack. Bind the built-in launcher into the pack manifest. - Preserve executable leases, descriptor-backed startup, credential fences, permissions, and exact ACP model admission. - Use `builtin:grok-acp` and `native:grok` as profile identities. Historical package-profile sessions fail closed on resume rather than being silently reinterpreted. - Resolve built-in assets from the authenticated sidecar location, including public server npm layouts. Keep the controller path out of provider environments. - Add clean npm tarball installation verification to the existing trusted canary CI job and the admitted manual EC2 verification path. It stages a unified release version and runs npm lifecycle scripts, then verifies missing-prerequisite rejection and admission after separate provisioning without credentials or inference. - Include the controller-owned provider pack in stamped Cloud images. Unstamped local images omit the pack and remain usable; remote ACPX requires full source provenance. - Correct CLI approval-page metadata for an already authenticated Cloud board user; approval authorization remains unchanged. - Honor explicit native-runner enablement in the Cloud agent picker and direct setup page, keeping the flag disabled by default. - Allow selecting the execution environment before connecting credentials. Include Grok in the existing authenticated hello-probe flow, targeting its pinned native prerequisite for runner setup. - Recover an existing subscription sign-in conflict through an explicit cancel-and-retry action, serialized after cancellation succeeds. - Preserve the selected ACPX harness before normalizing config fields, so new Grok agents use the Grok default model. - Keep the credential-free Cloud provider pack root-owned and readable after runtime UID remapping; verify manifest and referenced asset access under an unrelated unprivileged UID during image builds. - Archive prior failover backups alongside explicitly replaced harness state, preserving evidence while preventing stale backups from blocking a fresh replacement. - Update Daytona image content inputs and contract tests for the built-in assets and explicit provisioner. - Document and regression-test the shared `approve-all` default for Grok setup, saved configuration, and native execution. Explicitly saved restrictions remain unchanged. ## Verification Current merge-repair head `df09eb3e1a619430ad8419a0ee9aedd486689b05` incorporates master `f1a394bd30cb56fb9e479f98b9f50176fe921858` after the base PR was squash-merged. All 12 conflicts came from incoming files identical to the tested pre-squash base. The final tree exactly matches a three-way merge using that original base, preserving built-in Grok distribution and removal of the obsolete private package. All 252 focused runner/UI tests, six npm-isolation tests, and token gates pass. Fresh exact-head Greptile review is 5/5 with no outstanding findings; security scans and EC2 native compilation pass. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([run 36468768035](https://github.com/paperclipai/paperclip/actions/runs/36468768035)). The repository owner explicitly authorized bypassing code-owner approval after all checks passed; no CI checks or repository protection settings are bypassed or changed. The only remaining PR was removed from the completed stack metadata to permit native auto-merge. Earlier integration head `78cb306ecc41b5c96577c26c1d89153b0ef865a1` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36447124691)). The initial attempt lost two EC2 runners to shutdown signals and stalled a third shard during dependency preparation; all three passed the same-commit failed-job-only retry. Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. [Final public npm verification](https://github.com/paperclipai/paperclip/actions/runs/36445542764) passed on `76ea70cd4d13786a042af9df82f0fd7a8c85ae30`: 17 public packages, an executed offline lifecycle sentinel, unchanged consumer lock, built-in launcher, missing-prerequisite rejection, and verified separately provisioned binary/command lease. Provisioning and cleanup require no host privilege elevation; only the positive probe mounts the temporary native binary read-only. The verifier is unchanged by the final master merge. All six isolation tests and an offline npm smoke test pass. The prior head had 56 green CI checks and a 5/5 review after two unchanged tests timed out and passed a failed-job-only retry ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36444597313)). All 56 recovery-display/lineage tests pass; re-review cleared the already-covered missed-retry concern. Earlier EC2 failures remain retained: [npm lockfile rejection](https://github.com/paperclipai/paperclip/actions/runs/36436311203), [missing compiler in the slim image](https://github.com/paperclipai/paperclip/actions/runs/36440210984), and the aggregate 15-minute test timeouts in those broad runs. Both broad attempts passed typecheck, token gates, Product E2E type/unit checks and build. The focused EC2 lane preserves the existing trusted-actor and immutable-source gates. Earlier documentation/test checkpoint `ff244c4fd78a7ede5a3e00efe09f475f133ef33e` leaves runtime behavior unchanged. 154 focused tests pass across configuration building, native provider resolution, permission policy, credentials, UI configuration, and new-agent setup (including both Grok auth modes); token gates pass. All fresh CI is green for this head: 56 successful checks/statuses and two intentional skips ([run 36367065119](https://github.com/paperclipai/paperclip/actions/runs/36367065119)). Greptile is 5/5 with no new findings. Grok already inherits the shared `approve-all` default, so unattended setup requires no manual permission change. Runtime head `bb5a9307991f1ac567b781970ef11b39d518e19b` fixes a final staging continuation failure before provider startup: explicit replacement archived the old harness but left its failover backups active, which caused `runner_harness_state_mismatch`. The regression fails before the fix and passes after it; all eight adjacent recovery-safety cases also pass. Old backups remain inspectable inside the continuity archive. All fresh CI is green at this head ([run 36360839248](https://github.com/paperclipai/paperclip/actions/runs/36360839248)), with a 5/5 review. One unrelated Cursor test timed out in the initial server shard; the same-commit failed-job rerun passed, and both attempts are retained. Staging deployment is confirmed healthy on this revision. The controller image is `ghcr.io/paperclipai/paperclip@sha256:6ad91c487910ccd2596ff7aed0a3a3ea5233d12b51b83cd6e1402237749b9673`. The final browser-created staging task passed on this exact revision with API authentication: context read → structured human question → controller restart → answer submission → same native provider session resumed → document saved → task Done. The two turns took approximately 119s and 77s. The actual write receipt was applied, and the saved document has exactly one revision containing the selected answer and requested marker. Usage and cost were not reported. [Controller image build](https://github.com/paperclipai/paperclip/actions/runs/36360889243). - Previous integration head `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`: all CI green (53 successful checks/statuses, two intentional skips), including repository typecheck/build/tests, native Runner tests, browser shards, and canary installation checks. [CI run 36358672529](https://github.com/paperclipai/paperclip/actions/runs/36358672529). Greptile is 5/5 with no unresolved findings. - Focused checks cover Grok credentials, executable admission, launcher assets, provider-pack paths/permissions, workflow contracts, setup defaults, CLI authorization, and subscription conflict recovery. All 39 protocol definitions validate. Final integration checks pass 124 catalog/evidence/cache tests and nine project-form tests; token gates pass. Some local dependency checks could not load the stale installed dependency tree; the corresponding fresh EC2 checks pass. - Clean public npm installation passed on EC2 at `8b172ebcf8e02e30662d830c00f3961e3bd459ec` ([run 36164964900](https://github.com/paperclipai/paperclip/actions/runs/36164964900)): 17 unified-version packages, lifecycle scripts enabled, built-in launcher present, no separate Grok package or npm-downloaded binary, missing prerequisite rejected, separately provisioned native executable and command lease verified. No credentials or inference were used. Subsequent changes preserve this npm asset layout. - The immutable Daytona prerequisite image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:98957d5be0ac774d086b6402b5849e8e6356fec70fb8c09fca6eb4ed6de918e0`, built from `5a2db471f3ddabe77f9f80e76ed27f996cb97fba`. The previous Cloud controller image was `ghcr.io/paperclipai/paperclip@sha256:fd914e1ab1e45f741e8e078ff452d16f082d7ac05f9b4b3506d3a3c64150d204`, built from `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`; it is superseded by the latest image above. Its EC2 build verified provider-pack access under an unrelated unprivileged UID. - Browser staging at `40f898bc4cba73c1dff4e6344a3983ba0fb247ef` passed full Grok onboarding with the correct `grok-4.7` model, saved credential delivery, and pinned Daytona execution. A browser-created task read context and asked the structured human question. After a controller restart, answering the persisted question resumed the same native provider session, saved the requested document, and completed the task. Actual tool outcomes and durable state agree: one question and one document revision. The two successful turns took 42.7s and 63.1s; usage and cost were not reported. - Restricted policy returned the expected `approval_required` outcome. Functional staging tests explicitly selected `approve-all`; controller authorization and governed approvals remain enforced. Temporary board CLI access was revoked and verified rejected (HTTP 401), and the disposable onboarding agent was paused. Failures remain retained: the pre-fix continuation failure (its task remains blocked; the passing final task is fresh), the original Cloud provider-pack permission failure, the expected restricted-policy denial, the superseded npm staging failure, and an earlier monolithic CI infrastructure timeout. Browser CI exposed a project alias/form race; the final stack uses master's stronger draft-preservation fix and all browser shards pass. Historical full subscription/API protocol and Product rosters retain their original source revisions and do not qualify this packaging revision. No local Docker or Rust build was used. ## Risks The branch includes master’s draft-preservation fix for project URL aliases. It keeps the same project’s edit form mounted and clears prior data when the project or company changes. Custom sandboxes and local execution hosts must provision the pinned binary before Grok starts. Missing, changed, unsupported-platform, and symlinked executables fail admission. The new builtin profile cannot resume sessions created with the former private-package profile. Existing Claude/Codex npm bridge profiles retain their package pins. Grok restricted modes preserve the selected policy but cannot automatically admit Paperclip calls: ACP permission metadata does not independently bind tool authority, so those calls stop with `approval_required`. New Grok configurations default to `approve-all`, including API configurations that omit the mode. Existing explicitly restricted configurations remain restricted; controller authorization and governed approvals remain enforced. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
895 lines
42 KiB
YAML
895 lines
42 KiB
YAML
name: Runner Direct Live Protocol Evals
|
|
|
|
on:
|
|
schedule:
|
|
- cron: "23 9 * * 0"
|
|
workflow_dispatch:
|
|
inputs:
|
|
target_branch:
|
|
description: "Branch in paperclipai/paperclip to evaluate; trusted orchestration still runs from master"
|
|
type: string
|
|
required: false
|
|
evals_sha:
|
|
description: "Exact 40-character paperclipai/paperclip-evals commit to execute"
|
|
type: string
|
|
required: false
|
|
rosters:
|
|
description: "Comma-separated live roster IDs/files, or all for the maintained enabled direct suite"
|
|
type: string
|
|
default: "all"
|
|
required: false
|
|
max_parallel:
|
|
description: "Optional lower concurrency (at least 2, no higher than the configured campaign limit)"
|
|
type: string
|
|
required: false
|
|
grok_authentication:
|
|
description: "Explicit Grok credential source; subscription uses protected GROK_AUTH_JSON"
|
|
type: choice
|
|
options:
|
|
- api_key
|
|
- subscription
|
|
default: api_key
|
|
required: false
|
|
max_infrastructure_retries:
|
|
description: "Automatic retries only for explicitly retryable infrastructure failures (0-3)"
|
|
type: number
|
|
default: 1
|
|
|
|
permissions:
|
|
contents: read
|
|
|
|
concurrency:
|
|
group: runner-protocol-live-evals-${{ github.event_name == 'workflow_dispatch' && inputs.target_branch != '' && inputs.target_branch != github.event.repository.default_branch && format('development-{0}', inputs.target_branch) || format('protected-{0}', github.run_id) }}
|
|
cancel-in-progress: ${{ github.event_name == 'workflow_dispatch' && inputs.target_branch != '' && inputs.target_branch != github.event.repository.default_branch }}
|
|
|
|
jobs:
|
|
authorize:
|
|
name: Authorize paid direct eval campaign
|
|
if: github.event_name != 'schedule' || vars.RUNNER_PROTOCOL_EVAL_NIGHTLY_ENABLED == 'true'
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 5
|
|
permissions:
|
|
contents: read
|
|
outputs:
|
|
test_runner: ${{ steps.runner.outputs.runner }}
|
|
max_parallel_default: ${{ steps.runner.outputs.max_parallel_default }}
|
|
max_parallel_limit: ${{ steps.runner.outputs.max_parallel_limit }}
|
|
target_sha: ${{ steps.target.outputs.sha }}
|
|
target_ref: ${{ steps.target.outputs.ref }}
|
|
evals_sha: ${{ steps.evals.outputs.sha }}
|
|
steps:
|
|
- name: Require default branch and allowlisted numeric actor IDs
|
|
env:
|
|
GH_TOKEN: ${{ github.token }}
|
|
REPOSITORY: ${{ github.repository }}
|
|
REF: ${{ github.ref }}
|
|
DEFAULT_BRANCH: ${{ github.event.repository.default_branch }}
|
|
ACTOR: ${{ github.actor }}
|
|
ACTOR_ID: ${{ github.actor_id }}
|
|
TRIGGERING_ACTOR: ${{ github.triggering_actor }}
|
|
ALLOWED_ACTOR_IDS: ${{ vars.RUNNER_E2E_ALLOWED_ACTOR_IDS }}
|
|
run: |
|
|
set -euo pipefail
|
|
if [ "$REF" != "refs/heads/$DEFAULT_BRANCH" ]; then
|
|
echo "Paid direct Runner eval campaigns may run only from the default branch." >&2
|
|
exit 1
|
|
fi
|
|
if ! jq -e 'type == "array" and length > 0 and all(.[]; type == "number" and . > 0 and floor == .)' <<< "${ALLOWED_ACTOR_IDS:-}" >/dev/null; then
|
|
echo "RUNNER_E2E_ALLOWED_ACTOR_IDS must be a non-empty JSON array of numeric GitHub user IDs." >&2
|
|
exit 1
|
|
fi
|
|
triggering_actor_id="$(gh api "users/$TRIGGERING_ACTOR" --jq .id)"
|
|
if [ "$triggering_actor_id" != "$ACTOR_ID" ] && [ "$TRIGGERING_ACTOR" = "$ACTOR" ]; then
|
|
echo "GitHub actor identity contexts disagree; refusing the paid run." >&2
|
|
exit 1
|
|
fi
|
|
for candidate in "$triggering_actor_id" "$ACTOR_ID"; do
|
|
if ! jq -e --argjson candidate "$candidate" 'index($candidate) != null' <<< "$ALLOWED_ACTOR_IDS" >/dev/null; then
|
|
echo "The initiating GitHub account is not authorized to run paid Runner eval campaigns." >&2
|
|
exit 1
|
|
fi
|
|
done
|
|
|
|
- name: Resolve requested Paperclip branch to an immutable commit
|
|
id: target
|
|
env:
|
|
GH_TOKEN: ${{ github.token }}
|
|
REPOSITORY: ${{ github.repository }}
|
|
TARGET_BRANCH: ${{ inputs.target_branch || github.event.repository.default_branch }}
|
|
run: |
|
|
set -euo pipefail
|
|
if [ -z "$TARGET_BRANCH" ] || [[ "$TARGET_BRANCH" == refs/* ]]; then
|
|
echo "target_branch must name a branch in this repository without a refs/ prefix." >&2
|
|
exit 1
|
|
fi
|
|
encoded_branch="$(jq -rn --arg branch "$TARGET_BRANCH" '$branch | @uri')"
|
|
target_sha="$(gh api -X GET "repos/$REPOSITORY/branches/$encoded_branch" --jq .commit.sha)"
|
|
[[ "$target_sha" =~ ^[0-9a-f]{40}$ ]]
|
|
echo "sha=$target_sha" >> "$GITHUB_OUTPUT"
|
|
echo "ref=refs/heads/$TARGET_BRANCH" >> "$GITHUB_OUTPUT"
|
|
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
ref: ${{ github.sha }}
|
|
persist-credentials: false
|
|
|
|
- name: Generate private eval-repository token
|
|
id: evals_token
|
|
env:
|
|
COMMITPERCLIP_KEY: ${{ secrets.COMMITPERCLIP_KEY }}
|
|
GH_REPO: paperclipai/paperclip-evals
|
|
run: |
|
|
set -euo pipefail
|
|
token="$(node .github/scripts/get-bot-token.mjs)"
|
|
echo "::add-mask::$token"
|
|
echo "value=$token" >> "$GITHUB_OUTPUT"
|
|
|
|
- name: Verify the private eval program is pinned to an exact commit
|
|
id: evals
|
|
env:
|
|
GH_TOKEN: ${{ steps.evals_token.outputs.value }}
|
|
EVALS_SHA: ${{ inputs.evals_sha || vars.RUNNER_PROTOCOL_EVALS_SHA }}
|
|
run: |
|
|
set -euo pipefail
|
|
if ! [[ "$EVALS_SHA" =~ ^[0-9a-f]{40}$ ]]; then
|
|
echo "evals_sha (or RUNNER_PROTOCOL_EVALS_SHA for schedules) must be an exact 40-character commit." >&2
|
|
exit 1
|
|
fi
|
|
resolved="$(gh api -X GET "repos/paperclipai/paperclip-evals/commits/$EVALS_SHA" --jq .sha)"
|
|
test "$resolved" = "$EVALS_SHA"
|
|
echo "sha=$resolved" >> "$GITHUB_OUTPUT"
|
|
|
|
- name: Validate retry envelope
|
|
env:
|
|
RETRIES: ${{ github.event_name == 'schedule' && 1 || inputs.max_infrastructure_retries }}
|
|
run: |
|
|
set -euo pipefail
|
|
[[ "$RETRIES" =~ ^[0-3]$ ]]
|
|
|
|
- name: Select paid test runner
|
|
id: runner
|
|
env:
|
|
AWS_PAID_RUNNER_ENABLED: ${{ vars.RUNNER_E2E_AWS_ENABLED }}
|
|
run: |
|
|
set -euo pipefail
|
|
if [ "$AWS_PAID_RUNNER_ENABLED" = true ]; then
|
|
{
|
|
echo 'runner=runs-on/fleet=paperclip-public-pr-x64/env=public-ci'
|
|
echo 'max_parallel_default=100'
|
|
echo 'max_parallel_limit=100'
|
|
} >> "$GITHUB_OUTPUT"
|
|
else
|
|
{
|
|
echo 'runner=ubuntu-latest'
|
|
echo 'max_parallel_default=32'
|
|
echo 'max_parallel_limit=57'
|
|
} >> "$GITHUB_OUTPUT"
|
|
fi
|
|
|
|
target_lock:
|
|
name: Resolve target pnpm lockfile
|
|
needs: authorize
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 10
|
|
permissions:
|
|
contents: read
|
|
outputs:
|
|
artifact_id: ${{ steps.upload.outputs.artifact-id }}
|
|
lock_sha256: ${{ steps.lock.outputs.sha256 }}
|
|
steps:
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
ref: ${{ needs.authorize.outputs.target_sha }}
|
|
persist-credentials: false
|
|
|
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
|
|
with:
|
|
node-version: 24
|
|
|
|
- uses: pnpm/action-setup@0977fd99725f1db4007ccb2928dbb4e90d06cc86 # v6
|
|
env:
|
|
NPM_CONFIG_AUDIT: "false"
|
|
NPM_CONFIG_FUND: "false"
|
|
NPM_CONFIG_UPDATE_NOTIFIER: "false"
|
|
with:
|
|
version: 9.15.4
|
|
|
|
- name: Resolve target lockfile without lifecycle scripts
|
|
id: lock
|
|
run: |
|
|
set -euo pipefail
|
|
pnpm install --ignore-scripts --no-frozen-lockfile --lockfile-only
|
|
test -s pnpm-lock.yaml
|
|
unexpected="$(git status --short | awk '$2 != "pnpm-lock.yaml" { print }')"
|
|
if [ -n "$unexpected" ]; then
|
|
echo "Lockfile resolution changed files other than pnpm-lock.yaml:" >&2
|
|
echo "$unexpected" >&2
|
|
exit 1
|
|
fi
|
|
echo "sha256=$(sha256sum pnpm-lock.yaml | cut -d ' ' -f 1)" >> "$GITHUB_OUTPUT"
|
|
|
|
- name: Upload resolved target lockfile
|
|
id: upload
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
with:
|
|
name: runner-protocol-target-pnpm-lock-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: pnpm-lock.yaml
|
|
retention-days: 30
|
|
if-no-files-found: error
|
|
|
|
catalog:
|
|
name: Pin and fan out the direct Evalbook roster
|
|
needs: authorize
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 10
|
|
permissions:
|
|
contents: read
|
|
outputs:
|
|
matrix_0: ${{ steps.catalog.outputs.matrix_0 }}
|
|
matrix_1: ${{ steps.catalog.outputs.matrix_1 }}
|
|
matrix_1_present: ${{ steps.catalog.outputs.matrix_1_present }}
|
|
max_parallel_per_shard: ${{ steps.catalog.outputs.max_parallel_per_shard }}
|
|
selected: ${{ steps.catalog.outputs.selected }}
|
|
steps:
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
ref: ${{ github.sha }}
|
|
persist-credentials: false
|
|
|
|
- name: Generate private eval-repository token
|
|
id: evals_token
|
|
env:
|
|
COMMITPERCLIP_KEY: ${{ secrets.COMMITPERCLIP_KEY }}
|
|
GH_REPO: paperclipai/paperclip-evals
|
|
run: |
|
|
set -euo pipefail
|
|
token="$(node .github/scripts/get-bot-token.mjs)"
|
|
echo "::add-mask::$token"
|
|
echo "value=$token" >> "$GITHUB_OUTPUT"
|
|
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
repository: paperclipai/paperclip-evals
|
|
ref: ${{ needs.authorize.outputs.evals_sha }}
|
|
path: .paperclip-evals
|
|
token: ${{ steps.evals_token.outputs.value }}
|
|
persist-credentials: false
|
|
|
|
- name: Build the two bounded roster-plus-case matrices
|
|
id: catalog
|
|
env:
|
|
PAPERCLIP_PROTOCOL_EVAL_SOURCE_SHA: ${{ needs.authorize.outputs.target_sha }}
|
|
PAPERCLIP_PROTOCOL_EVALS_SHA: ${{ needs.authorize.outputs.evals_sha }}
|
|
MAX_PARALLEL: ${{ vars.RUNNER_E2E_MAX_PARALLEL || needs.authorize.outputs.max_parallel_default }}
|
|
MAX_PARALLEL_LIMIT: ${{ needs.authorize.outputs.max_parallel_limit }}
|
|
REQUESTED_MAX_PARALLEL: ${{ inputs.max_parallel }}
|
|
ROSTERS: ${{ inputs.rosters || 'all' }}
|
|
GROK_AUTHENTICATION: ${{ inputs.grok_authentication || 'api_key' }}
|
|
run: |
|
|
set -euo pipefail
|
|
if ! [[ "$MAX_PARALLEL" =~ ^[1-9][0-9]*$ ]] || [ "$MAX_PARALLEL" -lt 2 ] || [ "$MAX_PARALLEL" -gt "$MAX_PARALLEL_LIMIT" ]; then
|
|
echo "RUNNER_E2E_MAX_PARALLEL must be an integer from 2 through $MAX_PARALLEL_LIMIT for the two-shard direct suite." >&2
|
|
exit 1
|
|
fi
|
|
if [ -n "${REQUESTED_MAX_PARALLEL:-}" ]; then
|
|
if ! [[ "$REQUESTED_MAX_PARALLEL" =~ ^[1-9][0-9]{0,2}$ ]] || [ "$REQUESTED_MAX_PARALLEL" -lt 2 ] || [ "$REQUESTED_MAX_PARALLEL" -gt "$MAX_PARALLEL" ]; then
|
|
echo "max_parallel must be an integer from 2 through the configured campaign limit." >&2
|
|
exit 1
|
|
fi
|
|
MAX_PARALLEL="$REQUESTED_MAX_PARALLEL"
|
|
fi
|
|
node packages/paperclip-runner/scripts/runner-protocol-eval-campaign.mjs catalog \
|
|
--evals-root .paperclip-evals \
|
|
--rosters "$ROSTERS" \
|
|
--grok-authentication "$GROK_AUTHENTICATION" \
|
|
--campaign-id "gha-${GITHUB_RUN_ID}-${GITHUB_RUN_ATTEMPT}" \
|
|
--max-parallel "$MAX_PARALLEL" \
|
|
--output runner-protocol-eval-catalog.json
|
|
|
|
- name: Require the chat-report renderer before paid execution
|
|
run: |
|
|
set -euo pipefail
|
|
python3 .paperclip-evals/evals/paperclip-runner/tools/eval_program.py report --help | grep -q -- --public-viewer
|
|
|
|
- name: Upload immutable campaign catalog
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
with:
|
|
name: runner-protocol-eval-catalog-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: runner-protocol-eval-catalog.json
|
|
retention-days: 30
|
|
if-no-files-found: error
|
|
|
|
build_runner:
|
|
name: Build portable direct-eval runner once
|
|
needs: [authorize, target_lock, catalog]
|
|
runs-on: ${{ needs.authorize.outputs.test_runner }}
|
|
timeout-minutes: 30
|
|
permissions:
|
|
contents: read
|
|
steps:
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
ref: ${{ needs.authorize.outputs.target_sha }}
|
|
persist-credentials: false
|
|
|
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
|
|
with:
|
|
node-version: 24
|
|
|
|
- uses: pnpm/action-setup@0977fd99725f1db4007ccb2928dbb4e90d06cc86 # v6
|
|
env:
|
|
NPM_CONFIG_AUDIT: "false"
|
|
NPM_CONFIG_FUND: "false"
|
|
NPM_CONFIG_UPDATE_NOTIFIER: "false"
|
|
with:
|
|
version: 9.15.4
|
|
|
|
- name: Download resolved target lockfile
|
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
|
|
with:
|
|
artifact-ids: ${{ needs.target_lock.outputs.artifact_id }}
|
|
path: ${{ runner.temp }}/runner-protocol-target-lock
|
|
|
|
- name: Restore resolved target lockfile
|
|
env:
|
|
TARGET_SHA: ${{ needs.authorize.outputs.target_sha }}
|
|
EXPECTED_LOCK_SHA256: ${{ needs.target_lock.outputs.lock_sha256 }}
|
|
run: |
|
|
set -euo pipefail
|
|
test "$(git rev-parse HEAD)" = "$TARGET_SHA"
|
|
lock="$RUNNER_TEMP/runner-protocol-target-lock/pnpm-lock.yaml"
|
|
test -f "$lock"
|
|
test "$(find "$(dirname "$lock")" -type f | wc -l | tr -d ' ')" = 1
|
|
test "$(sha256sum "$lock" | cut -d ' ' -f 1)" = "$EXPECTED_LOCK_SHA256"
|
|
cp "$lock" pnpm-lock.yaml
|
|
test "$(sha256sum pnpm-lock.yaml | cut -d ' ' -f 1)" = "$EXPECTED_LOCK_SHA256"
|
|
|
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
|
|
with:
|
|
node-version: 24
|
|
cache: pnpm
|
|
|
|
- run: pnpm install --frozen-lockfile --ignore-scripts
|
|
|
|
- name: Materialize the pinned OpenCode executable before packaging
|
|
run: node packages/paperclip-runner/scripts/materialize-opencode-binary.mjs
|
|
|
|
- name: Materialize the pinned Grok executable when the target includes it
|
|
run: |
|
|
if [ -f packages/paperclip-runner/scripts/provision-grok.mjs ]; then
|
|
sudo node packages/paperclip-runner/scripts/provision-grok.mjs /opt/paperclip/providers/grok/1.0.13/grok
|
|
fi
|
|
|
|
- name: Build runner CLI, daemon, and canonical attempt viewer
|
|
run: |
|
|
set -euo pipefail
|
|
pnpm --filter @paperclipai/paperclip-runner build:typescript
|
|
pnpm --filter @paperclipai/paperclip-runner build:runner-binaries
|
|
pnpm --filter @paperclipai/paperclip-runner build:issue-thread
|
|
# Older target refs must fail before paid cells, not publish an empty viewer.
|
|
grep -q 'paperclip-eval-report' packages/paperclip-runner/dist-issue-thread/assets/*.js
|
|
grep -q 'evalbook-site' packages/paperclip-runner/dist-issue-thread/assets/*.css
|
|
|
|
- name: Package a portable provider runtime
|
|
run: |
|
|
set -euo pipefail
|
|
mkdir -p "$RUNNER_TEMP/runner-protocol-build/package" "$RUNNER_TEMP/runner-protocol-build/portable"
|
|
pnpm --dir packages/paperclip-runner pack \
|
|
--pack-destination "$RUNNER_TEMP/runner-protocol-build/package"
|
|
package="$(find "$RUNNER_TEMP/runner-protocol-build/package" -maxdepth 1 -type f -name '*.tgz' -print -quit)"
|
|
test -f "$package"
|
|
pnpm --filter @paperclipai/paperclip-runner deploy --prod \
|
|
"$RUNNER_TEMP/runner-protocol-build/portable"
|
|
cp "$package" "$RUNNER_TEMP/runner-protocol-build/paperclip-runner.tgz"
|
|
cp packages/paperclip-runner/runner/target/debug/paperclip-runnerd "$RUNNER_TEMP/runner-protocol-build/paperclip-runnerd"
|
|
cp -R packages/paperclip-runner/dist-issue-thread "$RUNNER_TEMP/runner-protocol-build/dist-issue-thread"
|
|
test -f "$RUNNER_TEMP/runner-protocol-build/portable/dist/cli/eval-session.js"
|
|
test -d "$RUNNER_TEMP/runner-protocol-build/portable/node_modules/.pnpm"
|
|
test -x "$RUNNER_TEMP/runner-protocol-build/paperclip-runnerd"
|
|
tar --create --gzip --file runner-protocol-build.tar.gz -C "$RUNNER_TEMP/runner-protocol-build" .
|
|
sha256sum runner-protocol-build.tar.gz > runner-protocol-build.tar.gz.sha256
|
|
|
|
- name: Upload immutable portable runner
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
with:
|
|
name: runner-protocol-build-${{ needs.authorize.outputs.target_sha }}-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: |
|
|
runner-protocol-build.tar.gz
|
|
runner-protocol-build.tar.gz.sha256
|
|
retention-days: 1
|
|
compression-level: 0
|
|
if-no-files-found: error
|
|
|
|
- name: Upload canonical viewer for publisher byte verification
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
with:
|
|
name: runner-protocol-viewer-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: packages/paperclip-runner/dist-issue-thread/
|
|
retention-days: 30
|
|
if-no-files-found: error
|
|
|
|
eval_shard_0:
|
|
name: Direct eval ${{ matrix.rosterId }} / ${{ matrix.caseId }}
|
|
needs: [authorize, catalog, build_runner]
|
|
runs-on: ${{ needs.authorize.outputs.test_runner }}
|
|
timeout-minutes: 18
|
|
permissions:
|
|
contents: read
|
|
id-token: write
|
|
environment:
|
|
name: runner-e2e-paid
|
|
strategy:
|
|
fail-fast: false
|
|
max-parallel: ${{ fromJSON(needs.catalog.outputs.max_parallel_per_shard) }}
|
|
matrix: ${{ fromJSON(needs.catalog.outputs.matrix_0) }}
|
|
steps: &direct_eval_steps
|
|
- name: Reauthorize paid execution before provider access
|
|
env:
|
|
GH_TOKEN: ${{ github.token }}
|
|
REF: ${{ github.ref }}
|
|
DEFAULT_BRANCH: ${{ github.event.repository.default_branch }}
|
|
ACTOR: ${{ github.actor }}
|
|
ACTOR_ID: ${{ github.actor_id }}
|
|
TRIGGERING_ACTOR: ${{ github.triggering_actor }}
|
|
ALLOWED_ACTOR_IDS: ${{ vars.RUNNER_E2E_ALLOWED_ACTOR_IDS }}
|
|
run: |
|
|
set -euo pipefail
|
|
test "$REF" = "refs/heads/$DEFAULT_BRANCH"
|
|
triggering_actor_id="$(gh api "users/$TRIGGERING_ACTOR" --jq .id)"
|
|
test "$triggering_actor_id" = "$ACTOR_ID" || test "$TRIGGERING_ACTOR" != "$ACTOR"
|
|
for candidate in "$triggering_actor_id" "$ACTOR_ID"; do
|
|
jq -e --argjson candidate "$candidate" 'type == "array" and index($candidate) != null' <<< "$ALLOWED_ACTOR_IDS" >/dev/null
|
|
done
|
|
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
ref: ${{ github.sha }}
|
|
persist-credentials: false
|
|
|
|
- name: Generate private eval-repository token
|
|
id: evals_token
|
|
env:
|
|
COMMITPERCLIP_KEY: ${{ secrets.COMMITPERCLIP_KEY }}
|
|
GH_REPO: paperclipai/paperclip-evals
|
|
run: |
|
|
set -euo pipefail
|
|
token="$(node .github/scripts/get-bot-token.mjs)"
|
|
echo "::add-mask::$token"
|
|
echo "value=$token" >> "$GITHUB_OUTPUT"
|
|
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
repository: paperclipai/paperclip-evals
|
|
ref: ${{ needs.authorize.outputs.evals_sha }}
|
|
path: .paperclip-evals
|
|
token: ${{ steps.evals_token.outputs.value }}
|
|
persist-credentials: false
|
|
|
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
|
|
with:
|
|
node-version: 24
|
|
|
|
- name: Download portable runner
|
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
|
|
with:
|
|
name: runner-protocol-build-${{ needs.authorize.outputs.target_sha }}-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: runner-protocol-build
|
|
|
|
- name: Verify and extract portable runner
|
|
run: |
|
|
set -euo pipefail
|
|
cd runner-protocol-build
|
|
sha256sum --check runner-protocol-build.tar.gz.sha256
|
|
mkdir extracted
|
|
tar --extract --gzip --file runner-protocol-build.tar.gz --directory extracted
|
|
test -x extracted/paperclip-runnerd
|
|
|
|
# Keep the host policy aligned with runner-full-stack-e2e.yml. This is
|
|
# deliberately before any provider credential or web-identity step.
|
|
- name: Provision Codex sandbox on the disposable trusted runner
|
|
if: matrix.provider == 'codex' || matrix.rosterId == 'protocol-live-acpx-codex-control'
|
|
run: |
|
|
node --input-type=module <<'NODE'
|
|
import { execFileSync } from "node:child_process";
|
|
import { createHash } from "node:crypto";
|
|
import { readFileSync, realpathSync, writeFileSync } from "node:fs";
|
|
import { createRequire } from "node:module";
|
|
import path from "node:path";
|
|
if (process.platform !== "linux") process.exit(0);
|
|
let restricted = "0";
|
|
try { restricted = readFileSync("/proc/sys/kernel/apparmor_restrict_unprivileged_userns", "utf8").trim(); } catch {}
|
|
if (restricted !== "1") process.exit(0);
|
|
const root = realpathSync(path.join(process.env.GITHUB_WORKSPACE, "runner-protocol-build/extracted/portable"));
|
|
const runnerRequire = createRequire(path.join(root, "package.json"));
|
|
const acpRequire = createRequire(runnerRequire.resolve("@agentclientprotocol/codex-acp/package.json"));
|
|
const codexRequire = createRequire(acpRequire.resolve("@openai/codex/package.json"));
|
|
const arch = process.arch === "x64" ? "x64" : process.arch === "arm64" ? "arm64" : null;
|
|
if (!arch) throw new Error("Unsupported Codex CI architecture");
|
|
const platformPackage = codexRequire.resolve(`@openai/codex-linux-${arch}/package.json`);
|
|
const triple = arch === "x64" ? "x86_64-unknown-linux-musl" : "aarch64-unknown-linux-musl";
|
|
const suffix = `/vendor/${triple}/bin/codex`;
|
|
const binary = realpathSync(path.join(path.dirname(platformPackage), suffix));
|
|
if (!binary.startsWith(root + "/node_modules/.pnpm/") || !binary.endsWith(suffix) || !/^[/A-Za-z0-9_.@+\-]+$/.test(binary)) {
|
|
throw new Error("Codex executable is outside the resolved dependency tree");
|
|
}
|
|
const name = `paperclip-e2e-codex-${createHash("sha256").update(binary).digest("hex").slice(0,16)}`;
|
|
const profilePath = path.join(process.env.RUNNER_TEMP, "paperclip-codex-userns.apparmor");
|
|
writeFileSync(profilePath, `abi <abi/4.0>,\ninclude <tunables/global>\nprofile ${name} "${binary}" flags=(unconfined) {\n userns,\n}\n`, {mode:0o600, flag:"wx"});
|
|
execFileSync("sudo", ["-n", "apparmor_parser", "-r", profilePath], {timeout:15000, stdio:"pipe"});
|
|
NODE
|
|
|
|
- name: Prepare short-lived AgentCore web identity
|
|
if: matrix.credentialName == 'AWS_AGENTCORE_OIDC'
|
|
env:
|
|
AGENTCORE_ROLE_ARN: ${{ vars.PAPERCLIP_AWS_AGENTCORE_EXECUTION_ROLE_ARN }}
|
|
run: |
|
|
set -euo pipefail
|
|
test -n "$AGENTCORE_ROLE_ARN"
|
|
token="$(curl --fail --silent --show-error \
|
|
-H "Authorization: Bearer $ACTIONS_ID_TOKEN_REQUEST_TOKEN" \
|
|
"${ACTIONS_ID_TOKEN_REQUEST_URL}&audience=sts.amazonaws.com" | jq -r .value)"
|
|
test -n "$token"
|
|
echo "::add-mask::$token"
|
|
token_file="$RUNNER_TEMP/runner-protocol-agentcore-token"
|
|
printf '%s' "$token" > "$token_file"
|
|
chmod 600 "$token_file"
|
|
{
|
|
echo "AWS_WEB_IDENTITY_TOKEN_FILE=$token_file"
|
|
echo "AWS_ROLE_ARN=$AGENTCORE_ROLE_ARN"
|
|
echo "AWS_ROLE_SESSION_NAME=runner-protocol-${GITHUB_RUN_ID}-${GITHUB_RUN_ATTEMPT}"
|
|
} >> "$GITHUB_ENV"
|
|
|
|
- name: Run one immutable direct protocol cell
|
|
id: direct_eval
|
|
env:
|
|
CELL_ID: ${{ matrix.cellId }}
|
|
ROSTER_FILE: ${{ matrix.rosterFile }}
|
|
CASE_ID: ${{ matrix.caseId }}
|
|
CREDENTIAL_NAME: ${{ matrix.credentialName }}
|
|
PROVIDER: ${{ matrix.provider }}
|
|
MAX_INFRASTRUCTURE_RETRIES: ${{ github.event_name == 'schedule' && 1 || inputs.max_infrastructure_retries }}
|
|
OPENAI_API_KEY: ${{ matrix.credentialName == 'OPENAI_API_KEY' && secrets.OPENAI_API_KEY || '' }}
|
|
ANTHROPIC_API_KEY: ${{ matrix.credentialName == 'ANTHROPIC_API_KEY' && secrets.ANTHROPIC_API_KEY || '' }}
|
|
OPENROUTER_API_KEY: ${{ matrix.credentialName == 'OPENROUTER_API_KEY' && secrets.OPENROUTER_API_KEY || '' }}
|
|
XAI_API_KEY: ${{ matrix.credentialName == 'XAI_API_KEY' && secrets.XAI_API_KEY || '' }}
|
|
PAPERCLIP_ACPX_GROK_AUTH_JSON_SECRET: ${{ matrix.credentialName == 'PAPERCLIP_ACPX_GROK_AUTH_JSON_SECRET' && secrets.GROK_AUTH_JSON || '' }}
|
|
PAPERCLIP_CLAUDE_MANAGED_PROFILE_ID: ${{ vars.PAPERCLIP_CLAUDE_MANAGED_PROFILE_ID }}
|
|
PAPERCLIP_CLAUDE_MANAGED_AGENT_ID: ${{ vars.PAPERCLIP_CLAUDE_MANAGED_AGENT_ID }}
|
|
PAPERCLIP_CLAUDE_MANAGED_AGENT_VERSION: ${{ vars.PAPERCLIP_CLAUDE_MANAGED_AGENT_VERSION }}
|
|
PAPERCLIP_CLAUDE_MANAGED_ENVIRONMENT_ID: ${{ vars.PAPERCLIP_CLAUDE_MANAGED_ENVIRONMENT_ID }}
|
|
AWS_REGION: ${{ vars.PAPERCLIP_AWS_AGENTCORE_REGION }}
|
|
AWS_DEFAULT_REGION: ${{ vars.PAPERCLIP_AWS_AGENTCORE_REGION }}
|
|
PAPERCLIP_AWS_AGENTCORE_PROFILE_ID: ${{ vars.PAPERCLIP_AWS_AGENTCORE_PROFILE_ID }}
|
|
PAPERCLIP_AWS_AGENTCORE_ACCOUNT_ID: ${{ vars.PAPERCLIP_AWS_AGENTCORE_ACCOUNT_ID }}
|
|
PAPERCLIP_AWS_AGENTCORE_HARNESS_ARN: ${{ vars.PAPERCLIP_AWS_AGENTCORE_HARNESS_ARN }}
|
|
PAPERCLIP_AWS_AGENTCORE_HARNESS_VERSION: ${{ vars.PAPERCLIP_AWS_AGENTCORE_HARNESS_VERSION }}
|
|
PAPERCLIP_AWS_AGENTCORE_ENDPOINT_ARN: ${{ vars.PAPERCLIP_AWS_AGENTCORE_ENDPOINT_ARN }}
|
|
PAPERCLIP_AWS_AGENTCORE_ENDPOINT_QUALIFIER: ${{ vars.PAPERCLIP_AWS_AGENTCORE_ENDPOINT_QUALIFIER }}
|
|
PAPERCLIP_AWS_AGENTCORE_RUNTIME_ARN: ${{ vars.PAPERCLIP_AWS_AGENTCORE_RUNTIME_ARN }}
|
|
PAPERCLIP_AWS_AGENTCORE_MEMORY_ARN: ${{ vars.PAPERCLIP_AWS_AGENTCORE_MEMORY_ARN }}
|
|
PAPERCLIP_AWS_AGENTCORE_MEMORY_ID: ${{ vars.PAPERCLIP_AWS_AGENTCORE_MEMORY_ID }}
|
|
PAPERCLIP_AWS_AGENTCORE_INVOCATION_ROLE_ARN: ${{ vars.PAPERCLIP_AWS_AGENTCORE_INVOCATION_ROLE_ARN }}
|
|
PAPERCLIP_AWS_AGENTCORE_CONTEXT_BUCKET: ${{ vars.PAPERCLIP_AWS_AGENTCORE_CONTEXT_BUCKET }}
|
|
PAPERCLIP_AWS_AGENTCORE_CONTEXT_PREFIX: ${{ vars.PAPERCLIP_AWS_AGENTCORE_CONTEXT_PREFIX }}
|
|
PAPERCLIP_AWS_AGENTCORE_CONTEXT_KMS_KEY_ARN: ${{ vars.PAPERCLIP_AWS_AGENTCORE_CONTEXT_KMS_KEY_ARN }}
|
|
PAPERCLIP_AWS_AGENTCORE_QUALIFICATION_REVISION: ${{ vars.PAPERCLIP_AWS_AGENTCORE_QUALIFICATION_REVISION }}
|
|
run: |
|
|
set -euo pipefail
|
|
mkdir -p cell-output/runs
|
|
if [ "$CREDENTIAL_NAME" != "AWS_AGENTCORE_OIDC" ]; then
|
|
test -n "${!CREDENTIAL_NAME:-}"
|
|
fi
|
|
if [ "$PROVIDER" = "claude_managed" ]; then
|
|
test -n "$PAPERCLIP_CLAUDE_MANAGED_PROFILE_ID"
|
|
test -n "$PAPERCLIP_CLAUDE_MANAGED_AGENT_ID"
|
|
test -n "$PAPERCLIP_CLAUDE_MANAGED_AGENT_VERSION"
|
|
test -n "$PAPERCLIP_CLAUDE_MANAGED_ENVIRONMENT_ID"
|
|
fi
|
|
set +e
|
|
python3 .paperclip-evals/evals/paperclip-runner/tools/run_live_roster.py run \
|
|
--roster ".paperclip-evals/evals/paperclip-runner/rosters/$ROSTER_FILE" \
|
|
--case "$CASE_ID" \
|
|
--runner-cli runner-protocol-build/extracted/portable/dist/cli/eval-session.js \
|
|
--runner-package runner-protocol-build/extracted/paperclip-runner.tgz \
|
|
--runnerd runner-protocol-build/extracted/paperclip-runnerd \
|
|
--runs-root cell-output/runs \
|
|
--summary-path cell-output/roster-summary.json \
|
|
--max-infrastructure-retries "$MAX_INFRASTRUCTURE_RETRIES" \
|
|
--run-id "gha-${GITHUB_RUN_ID}-${GITHUB_RUN_ATTEMPT}-${CELL_ID}"
|
|
status=$?
|
|
set -e
|
|
CELL_EXIT_CODE="$status" node --input-type=module <<'NODE'
|
|
import { readFileSync, writeFileSync } from "node:fs";
|
|
let authenticationMode;
|
|
let authenticationEvidenceFailure;
|
|
if (["XAI_API_KEY", "PAPERCLIP_ACPX_GROK_AUTH_JSON_SECRET"].includes(process.env.CREDENTIAL_NAME)) {
|
|
const expected = process.env.CREDENTIAL_NAME === "XAI_API_KEY" ? "api_key" : "subscription";
|
|
try {
|
|
const observed = JSON.parse(readFileSync("cell-output/roster-summary.json", "utf8")).authenticationMode;
|
|
// Never copy arbitrary provider output into cell metadata or errors.
|
|
if (["api_key", "subscription"].includes(observed)) authenticationMode = observed;
|
|
if (authenticationMode !== expected) authenticationEvidenceFailure = "grok_authentication_evidence_mismatch";
|
|
} catch {
|
|
authenticationEvidenceFailure = "grok_authentication_evidence_unreadable";
|
|
}
|
|
}
|
|
writeFileSync("cell-output/cell.json", `${JSON.stringify({
|
|
schema: "paperclip.runner-protocol-eval.cell/v1",
|
|
cellId: process.env.CELL_ID,
|
|
rosterFile: process.env.ROSTER_FILE,
|
|
caseId: process.env.CASE_ID,
|
|
...(authenticationMode ? { authenticationMode } : {}),
|
|
...(authenticationEvidenceFailure ? { authenticationEvidenceFailure } : {}),
|
|
exitCode: Number(process.env.CELL_EXIT_CODE),
|
|
}, null, 2)}\n`, { mode: 0o600 });
|
|
if (authenticationEvidenceFailure) throw new Error(authenticationEvidenceFailure);
|
|
NODE
|
|
exit "$status"
|
|
|
|
- name: Upload access-controlled cell attempt
|
|
if: always()
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
with:
|
|
name: runner-protocol-eval-${{ github.run_id }}-${{ github.run_attempt }}-${{ matrix.cellId }}
|
|
path: cell-output/
|
|
retention-days: 30
|
|
if-no-files-found: error
|
|
|
|
eval_shard_1:
|
|
name: Direct eval ${{ matrix.rosterId }} / ${{ matrix.caseId }}
|
|
if: needs.catalog.outputs.matrix_1_present == 'true'
|
|
needs: [authorize, catalog, build_runner]
|
|
runs-on: ${{ needs.authorize.outputs.test_runner }}
|
|
timeout-minutes: 18
|
|
permissions:
|
|
contents: read
|
|
id-token: write
|
|
environment:
|
|
name: runner-e2e-paid
|
|
strategy:
|
|
fail-fast: false
|
|
max-parallel: ${{ fromJSON(needs.catalog.outputs.max_parallel_per_shard) }}
|
|
matrix: ${{ fromJSON(needs.catalog.outputs.matrix_1) }}
|
|
steps: *direct_eval_steps
|
|
|
|
report:
|
|
name: Merge attempts and render canonical Evalbook
|
|
if: always() && !cancelled() && needs.catalog.result == 'success' && needs.build_runner.result == 'success'
|
|
needs: [authorize, catalog, build_runner, eval_shard_0, eval_shard_1]
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 30
|
|
permissions:
|
|
actions: read
|
|
contents: read
|
|
outputs:
|
|
public_report_ready: ${{ steps.public_report.outputs.ready }}
|
|
steps:
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
ref: ${{ github.sha }}
|
|
persist-credentials: false
|
|
|
|
- name: Generate private eval-repository token
|
|
id: evals_token
|
|
env:
|
|
COMMITPERCLIP_KEY: ${{ secrets.COMMITPERCLIP_KEY }}
|
|
GH_REPO: paperclipai/paperclip-evals
|
|
run: |
|
|
set -euo pipefail
|
|
token="$(node .github/scripts/get-bot-token.mjs)"
|
|
echo "::add-mask::$token"
|
|
echo "value=$token" >> "$GITHUB_OUTPUT"
|
|
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
repository: paperclipai/paperclip-evals
|
|
ref: ${{ needs.authorize.outputs.evals_sha }}
|
|
path: .paperclip-evals
|
|
token: ${{ steps.evals_token.outputs.value }}
|
|
persist-credentials: false
|
|
|
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
|
|
with:
|
|
node-version: 24
|
|
|
|
- uses: pnpm/action-setup@0977fd99725f1db4007ccb2928dbb4e90d06cc86 # v6
|
|
env:
|
|
NPM_CONFIG_AUDIT: "false"
|
|
NPM_CONFIG_FUND: "false"
|
|
NPM_CONFIG_UPDATE_NOTIFIER: "false"
|
|
with:
|
|
version: 9.15.4
|
|
|
|
- name: Resolve trusted report lockfile without lifecycle scripts
|
|
run: |
|
|
set -euo pipefail
|
|
pnpm install --ignore-scripts --no-frozen-lockfile --lockfile-only
|
|
test -s pnpm-lock.yaml
|
|
|
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
|
|
with:
|
|
node-version: 24
|
|
cache: pnpm
|
|
|
|
- name: Download immutable campaign catalog
|
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
|
|
with:
|
|
name: runner-protocol-eval-catalog-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: runner-protocol-catalog
|
|
|
|
- name: Download portable runner and viewer
|
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
|
|
with:
|
|
name: runner-protocol-build-${{ needs.authorize.outputs.target_sha }}-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: runner-protocol-build
|
|
|
|
- name: Download every access-controlled cell
|
|
id: download_cells
|
|
continue-on-error: true
|
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
|
|
with:
|
|
pattern: runner-protocol-eval-${{ github.run_id }}-${{ github.run_attempt }}-*
|
|
path: downloaded-runner-protocol-evals
|
|
merge-multiple: false
|
|
|
|
- name: Retry cell download after artifact transport failure
|
|
if: steps.download_cells.outcome == 'failure'
|
|
continue-on-error: true
|
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
|
|
with:
|
|
pattern: runner-protocol-eval-${{ github.run_id }}-${{ github.run_attempt }}-*
|
|
path: downloaded-runner-protocol-evals
|
|
merge-multiple: false
|
|
|
|
- name: Materialize an empty download root when every cell failed early
|
|
run: mkdir -p downloaded-runner-protocol-evals
|
|
|
|
- name: Verify portable viewer
|
|
run: |
|
|
set -euo pipefail
|
|
cd runner-protocol-build
|
|
sha256sum --check runner-protocol-build.tar.gz.sha256
|
|
mkdir extracted
|
|
tar --extract --gzip --file runner-protocol-build.tar.gz --directory extracted
|
|
test -f extracted/dist-issue-thread/index.html
|
|
|
|
- name: Aggregate every expected cell, including missing infrastructure cells
|
|
env:
|
|
PAPERCLIP_PROTOCOL_EVAL_SOURCE_SHA: ${{ needs.authorize.outputs.target_sha }}
|
|
PAPERCLIP_PROTOCOL_EVAL_SOURCE_REF: ${{ needs.authorize.outputs.target_ref }}
|
|
PAPERCLIP_PROTOCOL_EVALS_SHA: ${{ needs.authorize.outputs.evals_sha }}
|
|
PAPERCLIP_PROTOCOL_EVAL_WORKFLOW_URL: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}
|
|
run: |
|
|
node packages/paperclip-runner/scripts/runner-protocol-eval-campaign.mjs aggregate \
|
|
--catalog runner-protocol-catalog/runner-protocol-eval-catalog.json \
|
|
--downloads downloaded-runner-protocol-evals \
|
|
--evals-root .paperclip-evals \
|
|
--runs-out runner-protocol-merged/runs \
|
|
--campaign-out runner-protocol-merged/campaign.json
|
|
|
|
- name: Render the access-controlled canonical Evalbook report
|
|
run: |
|
|
python3 .paperclip-evals/evals/paperclip-runner/tools/eval_program.py report \
|
|
--runs-root runner-protocol-merged/runs \
|
|
--output runner-protocol-merged/report \
|
|
--viewer-root runner-protocol-build/extracted/dist-issue-thread \
|
|
--inventory .paperclip-evals/evals/paperclip-runner/inventory.json \
|
|
--coverage-matrix .paperclip-evals/evals/paperclip-runner/coverage-matrix.json
|
|
cp runner-protocol-merged/campaign.json runner-protocol-merged/report/campaign.json
|
|
|
|
- name: Render the same canonical grid from a public-safe evidence projection
|
|
run: |
|
|
node packages/paperclip-runner/scripts/runner-protocol-eval-campaign.mjs sanitize \
|
|
--runs-root runner-protocol-merged/runs \
|
|
--output runner-protocol-merged/public-runs
|
|
python3 .paperclip-evals/evals/paperclip-runner/tools/eval_program.py report \
|
|
--runs-root runner-protocol-merged/public-runs \
|
|
--output runner-protocol-merged/public-report \
|
|
--viewer-root runner-protocol-build/extracted/dist-issue-thread \
|
|
--public-viewer \
|
|
--inventory .paperclip-evals/evals/paperclip-runner/inventory.json \
|
|
--coverage-matrix .paperclip-evals/evals/paperclip-runner/coverage-matrix.json
|
|
cp runner-protocol-merged/campaign.json runner-protocol-merged/public-report/campaign.json
|
|
|
|
- name: Set up report browser verification
|
|
uses: pnpm/action-setup@0977fd99725f1db4007ccb2928dbb4e90d06cc86 # v6
|
|
with:
|
|
version: 9.15.4
|
|
|
|
- name: Verify the actual chat viewer before publication
|
|
run: |
|
|
pnpm install --frozen-lockfile --ignore-scripts
|
|
pnpm --filter @paperclipai/paperclip-runner exec playwright install --with-deps chromium
|
|
node packages/paperclip-runner/scripts/verify-runner-evalbook-viewer.mjs --report-root runner-protocol-merged/public-report --screenshots runner-protocol-merged/viewer-proof
|
|
node packages/paperclip-runner/scripts/verify-runner-evalbook-viewer.mjs --report-root runner-protocol-merged/report
|
|
|
|
- name: Enforce the static public allowlist
|
|
id: public_report
|
|
run: |
|
|
node --input-type=module -e 'import { validatePublicProtocolEvalReport } from "./packages/paperclip-runner/scripts/publish-runner-protocol-eval-history.mjs"; await validatePublicProtocolEvalReport("runner-protocol-merged/public-report", { viewerRoot: "runner-protocol-build/extracted/dist-issue-thread" });'
|
|
echo "ready=true" >> "$GITHUB_OUTPUT"
|
|
|
|
- name: Add campaign result to the workflow summary
|
|
run: |
|
|
{
|
|
echo '## Runner direct live protocol evals'
|
|
echo
|
|
jq -r '"- Cells: \(.totals.passed)/\(.totals.selected) passed\n- Behavior failures: \(.totals.behaviorFailures)\n- Infrastructure failures: \(.totals.infrastructureFailures)\n- Paperclip: `\(.source.paperclip.sha)`\n- Evals: `\(.source.evals.sha)`"' runner-protocol-merged/campaign.json
|
|
} >> "$GITHUB_STEP_SUMMARY"
|
|
|
|
- name: Upload access-controlled canonical Evalbook and raw attempts
|
|
if: always()
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
with:
|
|
name: runner-protocol-eval-report-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: runner-protocol-merged/
|
|
retention-days: 30
|
|
if-no-files-found: error
|
|
|
|
- name: Upload publisher-only sanitized Evalbook
|
|
if: steps.public_report.outputs.ready == 'true'
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
with:
|
|
name: runner-protocol-eval-public-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: runner-protocol-merged/public-report/
|
|
retention-days: 1
|
|
if-no-files-found: error
|
|
|
|
- name: Enforce complete green campaign
|
|
if: always()
|
|
run: jq -e '.complete == true and .allPassed == true' runner-protocol-merged/campaign.json >/dev/null
|
|
|
|
publish_history:
|
|
name: Publish immutable Evalbook and mutable campaign index
|
|
needs: [authorize, catalog, report]
|
|
if: always() && needs.report.outputs.public_report_ready == 'true'
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 15
|
|
concurrency:
|
|
group: runner-protocol-eval-history-publish
|
|
cancel-in-progress: false
|
|
permissions:
|
|
contents: read
|
|
id-token: write
|
|
environment:
|
|
name: runner-e2e-history
|
|
url: ${{ steps.publish.outputs.report_url }}
|
|
steps:
|
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
with:
|
|
# AWS credentials can execute only the publisher from the trusted workflow revision.
|
|
ref: ${{ github.sha }}
|
|
persist-credentials: false
|
|
|
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
|
|
with:
|
|
node-version: 24
|
|
|
|
- name: Download only the sanitized canonical Evalbook
|
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
|
|
with:
|
|
name: runner-protocol-eval-public-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: runner-protocol-public-report
|
|
|
|
- name: Download the same-run canonical viewer for byte verification
|
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
|
|
with:
|
|
name: runner-protocol-viewer-${{ github.run_id }}-${{ github.run_attempt }}
|
|
path: runner-protocol-trusted-viewer
|
|
|
|
- name: Exchange GitHub OIDC identity for scoped AWS credentials
|
|
uses: aws-actions/configure-aws-credentials@e6de054238d6b7531b4efff3b6587d9aade6a06c # v6
|
|
with:
|
|
role-to-assume: ${{ vars.RUNNER_PROTOCOL_EVAL_HISTORY_AWS_ROLE_ARN || vars.RUNNER_E2E_HISTORY_AWS_ROLE_ARN }}
|
|
aws-region: ${{ vars.RUNNER_PROTOCOL_EVAL_HISTORY_AWS_REGION || vars.RUNNER_E2E_HISTORY_AWS_REGION }}
|
|
|
|
- name: Publish versioned report and refresh the root index
|
|
id: publish
|
|
env:
|
|
PAPERCLIP_RUNNER_PROTOCOL_EVAL_PUBLIC_REPORT_DIR: ${{ github.workspace }}/runner-protocol-public-report
|
|
PAPERCLIP_RUNNER_PROTOCOL_EVAL_VIEWER_DIR: ${{ github.workspace }}/runner-protocol-trusted-viewer
|
|
RUNNER_PROTOCOL_EVAL_HISTORY_S3_BUCKET: ${{ vars.RUNNER_PROTOCOL_EVAL_HISTORY_S3_BUCKET || vars.RUNNER_E2E_HISTORY_S3_BUCKET }}
|
|
RUNNER_PROTOCOL_EVAL_HISTORY_PREFIX: ${{ vars.RUNNER_PROTOCOL_EVAL_HISTORY_PREFIX || 'runner-protocol-evals' }}
|
|
RUNNER_PROTOCOL_EVAL_HISTORY_PUBLIC_BASE_URL: ${{ vars.RUNNER_PROTOCOL_EVAL_HISTORY_PUBLIC_BASE_URL || vars.RUNNER_E2E_HISTORY_PUBLIC_BASE_URL }}
|
|
run: node packages/paperclip-runner/scripts/publish-runner-protocol-eval-history.mjs
|