fix(evals): account for hiring completion notifications (#15007)

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Product E2E evals check real hiring and delegated task completion.
> - The hiring fixture requires three requested CEO turns and two coder
executions.
> - The server can also wake the CEO when each delegated task completes.
> - Two exact-five-run guards rejected these valid completion turns in
all four retained cells.
> - This pull request validates bounded completion turns in both guards.
> - The benefit is accurate workflow grading while all actual runs and
coverage failures remain visible.

## Linked Issues or Issue Description

Refs: #14985, #14948, #14961.

**What happened?**

The original hiring comparison reports Codex Fail → Fail and Claude Fail
→ Fail. Each cell has seven successful runs. The five requested work
turns are accompanied by two server task-completion notifications. All
six other delivery checks pass.

**Expected behavior**

Require exactly three distinct user-requested CEO turns and one coder
execution for each of two known tasks. Admit at most two strictly
attributed server completion turns, including one turn that batches both
tasks. Reject unknown, duplicate, failed, retried or extra-work runs.

**Steps to reproduce**

Inspect the retained four-cell report linked below. Each original result
fails `five-successful-turns`. The same exact count was also enforced by
the final chat-flow guard.

## What Changed

- Add one typed lifecycle helper shared by the hiring scorer and the
hiring-only final chat guard.
- Validate public run ledgers, company/user/account identity, request
attribution, task origins, completion deliveries, timing and replies.
- Keep exactly five required work turns; declare seven maximum total
turns for cost and timeout planning.
- Count all actual runs, including notification runs and unexpected
resets. Keep other chat count guards unchanged.
- Version the hiring grader as v3 (turn accounting v2) and include the
helper and chat guard in its definition digest.
- Keep source-read, exact coder-body and all six other delivery checks
unchanged.
- Add 144 focused helper/scorer/settlement calibrations and separately
versioned exact retained-input replay reports.
- Retry complete bracketed observations, await both owed callbacks and
attributed replies, and refresh the final guard consistently.
- Reject unrelated completion writes and failed mutation attempts using
exact canonical/native action IDs. Missing identity mapping is
uncomparable action coverage.

## Verification

- All 977 credential-free E2E support tests pass across 64 files,
including 144 focused lifecycle/action/scorer/settlement calibrations.
- E2E typecheck, ordinary plugin SDK and Runner TypeScript dependency
builds, capability contract/inventory checks and the existing two-cell
hiring discovery pass.
- [Executable replay
report](https://github.com/paperclipai/paperclip/blob/fed1729018cc100f5f4bbfb692777e49009c423b/doc/plans/2026-10-02-hiring-executable-accounting-replay.md)
pins current code revision `e4077ade1818d98b9862ae79ee1d49a007dcf9c1`,
v3 definition digest, exact source/input hashes and each original/new
check.
- The stricter replay verifies both Codex variants through both
executable guards. ACPX Claude action attribution remains
unresolved/uncomparable because provider execution IDs cannot be exactly
joined to native request IDs; guards fail closed. No notification writes
are observed. All six other outcomes and every original source/template
coverage check stay unchanged. Original files and Fail → Fail machine
verdicts remain preserved; zero providers are called.
- Full attempts remain uncomparable in both profiles. Historical Claude
also keeps its six-backtick exact-template mismatch. This grading repair
does not prove model-performance equivalence.
- The limited sidecar-v1 and initial executable-v2 passes checked
notification-created tasks but could miss unrelated document writes.
Those assessments remain preserved and do not prove harmless
notifications. The stricter v3 replay is separate.
- [Original measurement and separate
sidecar](https://github.com/paperclipai/paperclip/blob/8eb517ca1497687237163bdef4dfc4d3332ea916/doc/plans/2026-10-02-hiring-template-live-comparison.md)
retain 28 actual runs, eight automatic notifications, four successful
cleanups and unknown actual model charges. No models are rerun.
- The branch is replayed on master `59c07ede7`. Intervening master
changes are UI-only; eval source bytes and replay verdicts match. The
four-cell provider-free replay was repeated against the reachable code
revision.
- Initial-head normal CI retained browser failures in agent-run denial
feedback and touch-picker scroll position. Those browser paths and
imports were unchanged, but their cause was not established. The
necessary review-fix head passes both browser checks; no blind rerun was
requested.
- Local full repository typecheck/test/build were not repeated.
Exact-head normal CI passes the required repository gates, including
typecheck, tests, build and browser shards. Fresh Greptile review
completed on `fed1729018cc100f5f4bbfb692777e49009c423b` with 5/5 and
zero unresolved threads. An independent rerun of the 144 focused
helper/scorer/settlement tests passes on the unchanged head.


**Merge readiness:** This PR repairs the evaluator. Its positive and
negative calibrations pass, both guards reject missing action
attribution, current-head CI and review pass, and there are no merge
conflicts. The retained ACPX cells remain uncomparable because their
action IDs cannot be joined. That coverage limit remains a separate
follow-up; it does not require relaxing this grader or changing the old
results. No model calls, production instructions, carrier changes, or
historical regrades are part of this readiness update.

## Risks

- Missing or inconsistent public lifecycle evidence fails the bounded
helper. The focused calibrations reject plausible false positives and
malformed observations. Unmatched action IDs fail closed and are
reported as uncomparable rather than a model task regression.
- Source-read evidence remains incomplete. This PR does not change
provider event carriers or relax the coverage oracle.
- The versioned count check differs from original v1 results. Reports
retain both versions and exact input hashes.

## Model Used

OpenAI Codex, GPT-6 family as identified by this session. The exact
deployment ID and context-window size are not exposed. The assistant
used reasoning, repository tools, code execution and delegated
calibration work.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip normal CI gates are green (exact head
`fed1729018cc100f5f4bbfb692777e49009c423b`; fresh review tracked
separately below)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(completed exact-head review; zero unresolved threads)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
This commit is contained in:
DottaandPaperclip authored and GitHub committed 2026-10-03 06:42:51 -05:00
1 parent dd868ed125
commit 78e0034498
13 files changed
+2631 -44

No files matched your search

+5 -1
View File
@@ -81,7 +81,9 @@ The explicit-only [production hiring templates suite](../tests/runner-e2e/README
adds two local native Codex/Claude cells. It exercises API-created production
CEO defaults, an explicitly requested hiring skill/reference read, a permanent
coder hire, independently computed saved JSON fixtures and worker reuse.
Each cell expects five turns. Source/read coverage and workflow outcome are
Each cell requires five work turns and admits at most two strictly attributed
server task-completion turns. Every actual run remains counted; unknown or
extra-work turns fail. Source/read coverage and workflow outcome are
separate: missing read provenance leaves the candidate/baseline pair
uncomparable even if work succeeds. Baseline bundles and coder examples derive
from their own source revision, without requiring candidate wording or length.
@@ -386,3 +388,5 @@ The explicit-only Product E2E `confirmation-replies` suite tests conversational
approval and rejection, persisted message provenance, approval before execution,
ambiguous proposals, and the existing card-click path with native Claude/Codex.
See the [suite contract](../tests/runner-e2e/README.md#conversational-confirmation-replies-explicit-only).
Hiring notification accounting now also requires exact completed action attribution. Missing native/provider ID mapping is uncomparable evidence; it must not be reported as a model task regression or waived through name/order matching. The fixture waits for both known completion callbacks and settled bracketed observations, including the gap before pending outbox work becomes a wake. Strict action replay and original machine verdicts are retained separately.
@@ -0,0 +1,740 @@
{
"schema": "paperclip.hiring-executable-accounting-replay.v1",
"assessmentType": "Provider-free replay through both new executable guards; original model measurements and verdicts unchanged",
"evaluatedCodeRevision": "eef64009dc144b91b245a731ea9a1c3c07406a19",
"evaluatedCodeHashes": {
"tests/runner-e2e/hiring-template-cases.ts": "06915e33cc9a5241cd28a7184fa33d098c30aed4114accea4f93838aaba3738f",
"tests/runner-e2e/hiring-template-scoring.ts": "a9e1efa7e11bad3aa8dc873a28325b3153799251740ae87c4623a08f7d45a5cb",
"tests/runner-e2e/hiring-template-flow.ts": "f3c6d29dc5019955dda4372ead7a9fdfbde447b7b92ef103527b9393ec60bd30",
"tests/runner-e2e/hiring-template-turn-accounting.ts": "8f2ad8b09e419796b372929b59dc7f86d41dcbf3a3144c45e329a1aa988b383b",
"tests/runner-e2e/chat-flow.ts": "09dd87ef7aef299ddbd1fad65b72817c042595f9e4a4e699c430497ac053a2d4"
},
"graderVersion": "paperclip.hiring-templates.v2",
"definitionDigest": "4d527906a967ca71c4c0e62cfbdc3405419931bcd6741b0630c4223bed2d4b6a",
"providerCalls": 0,
"newProviderRuns": 0,
"automaticRetries": 0,
"summary": {
"originalPairs": "Codex Fail→Fail; Claude Fail→Fail",
"correctedWorkflowOutcomePairs": "Codex Pass→Pass; Claude Pass→Pass",
"fullAttemptComparison": "Uncomparable in both profiles and variants; source-read coverage unchanged",
"originalMachineFailuresPreserved": 4,
"correctedScorerGuardsPassed": 4,
"correctedFinalChatGuardsPassed": 4,
"unchangedCoverageAssessments": 4,
"unchangedOtherOutcomeAssessments": 4,
"actualMeasuredRuns": 28,
"actualModelChargesKnown": false,
"broadEquivalenceEstablished": false
},
"assessments": [
{
"variant": "candidate",
"profile": "runner-codex",
"original": {
"sourceRevision": "9f5404ad3aacbe76777952759414d34fd381e674",
"graderVersion": "paperclip.hiring-templates.v1",
"definitionDigest": "4d7f18420889325af89aa48eaaf021e8178aa6d26d504b92041635a15741f004",
"machineStatus": "failed",
"outcomePassed": false,
"countCheckPassed": false,
"comparisonStatus": "uncomparable"
},
"corrected": {
"graderVersion": "paperclip.hiring-templates.v2",
"definitionDigest": "4d527906a967ca71c4c0e62cfbdc3405419931bcd6741b0630c4223bed2d4b6a",
"outcomePassed": true,
"comparisonStatus": "uncomparable",
"scorerGuardPassed": true,
"finalChatGuardPassed": true,
"turnAccounting": {
"version": "paperclip.hiring-template-turn-accounting.v1",
"counts": {
"requiredWorkTurns": 5,
"maximumCompletionTurns": 2,
"maximumTotalTurns": 7,
"requestedLeadTurns": 3,
"coderTurns": 2,
"completionTurns": 2,
"unclassifiedTurns": 0,
"snapshotRunCount": 7,
"actualRunCount": 7,
"costAccountingRunCount": 7
},
"predicates": [
{
"id": "known-fixture-context",
"passed": true
},
{
"id": "complete-public-run-ledger",
"passed": true
},
{
"id": "exact-five-required-work-turns",
"passed": true
},
{
"id": "successful-native-without-retries",
"passed": true
},
{
"id": "resolved-company-account-and-identity",
"passed": true
},
{
"id": "three-distinct-requested-chat-turns",
"passed": true
},
{
"id": "one-coder-execution-per-known-task",
"passed": true
},
{
"id": "task-origins-are-first-two-requested-turns",
"passed": true
},
{
"id": "bounded-server-completion-receipts",
"passed": true
},
{
"id": "completion-runs-have-attributed-chat-replies",
"passed": true
},
{
"id": "no-notification-created-extra-tasks",
"passed": true
}
],
"passed": true
},
"checks": [
{
"id": "one-coder-hire",
"dimension": "outcome",
"passed": true,
"detail": "Exactly one permanent coding teammate reports to the lead."
},
{
"id": "execution-account",
"dimension": "outcome",
"passed": true,
"detail": "Hire keeps the lead model and managed account binding."
},
{
"id": "two-worker-tasks",
"dimension": "outcome",
"passed": true,
"detail": "Both independent tasks are completed by the same coder in the chosen project."
},
{
"id": "bounded-work-and-completion-turns",
"dimension": "outcome",
"passed": true,
"detail": "Exactly three requested lead turns and two coder executions, plus at most two strictly attributed completion notifications; every actual run remains counted."
},
{
"id": "initial-json-artifact",
"dimension": "outcome",
"passed": true,
"detail": "Independent computation checks each original input/value and worker authorship."
},
{
"id": "reused-json-artifact",
"dimension": "outcome",
"passed": true,
"detail": "The reused coder applies the changed separator to every input."
},
{
"id": "original-preserved",
"dimension": "outcome",
"passed": true,
"detail": "Reuse preserves the original document, revision and task identity."
},
{
"id": "source-fingerprints",
"dimension": "coverage",
"passed": true,
"detail": "Served instruction and hiring source bytes match the evaluated revision."
},
{
"id": "production-ceo-bundle",
"dimension": "coverage",
"passed": true,
"detail": "The API-created lead receives this revision's default CEO bundle, without a fixture override."
},
{
"id": "assigned-hiring-skill",
"dimension": "coverage",
"passed": true,
"detail": "Production CEO defaults assign the hiring skill."
},
{
"id": "production-source-reads",
"dimension": "coverage",
"passed": false,
"detail": "Completed lead read receipts before the hire prove the explicitly requested skill, guide, checklist and coder example paths. Unrecognized or missing reads leave coverage uncomparable."
},
{
"id": "supplied-coder-instructions",
"dimension": "coverage",
"passed": true,
"detail": "Saved hired instructions use the source revision's coder example with its company/name placeholders filled."
},
{
"id": "hired-instructions-durable",
"dimension": "coverage",
"passed": true,
"detail": "The same saved instruction bundle survives the reused worker execution."
},
{
"id": "hired-skills-durable",
"dimension": "coverage",
"passed": true,
"detail": "The saved skill selections survive the reused worker execution."
}
]
},
"coverageUnchanged": true,
"otherOutcomesUnchanged": true,
"originalBytesUnchanged": true,
"inputSha256": {
"result": "44cbf64cfabd0257f31fb4d6909018d346f5667b0a7b3d7e855a2ece7e29c292",
"hiringTemplate": "9609ddcb81c6bee3c5db97312138b4ea3e294da13ea540e41f363c826a514067",
"apiState": "ef02af29bcf61ff3eace0549ca164a28c7950d1a3e14fc58c89bf9cb0f944b4d"
}
},
{
"variant": "candidate",
"profile": "runner-acpx-claude",
"original": {
"sourceRevision": "9f5404ad3aacbe76777952759414d34fd381e674",
"graderVersion": "paperclip.hiring-templates.v1",
"definitionDigest": "4d7f18420889325af89aa48eaaf021e8178aa6d26d504b92041635a15741f004",
"machineStatus": "failed",
"outcomePassed": false,
"countCheckPassed": false,
"comparisonStatus": "uncomparable"
},
"corrected": {
"graderVersion": "paperclip.hiring-templates.v2",
"definitionDigest": "4d527906a967ca71c4c0e62cfbdc3405419931bcd6741b0630c4223bed2d4b6a",
"outcomePassed": true,
"comparisonStatus": "uncomparable",
"scorerGuardPassed": true,
"finalChatGuardPassed": true,
"turnAccounting": {
"version": "paperclip.hiring-template-turn-accounting.v1",
"counts": {
"requiredWorkTurns": 5,
"maximumCompletionTurns": 2,
"maximumTotalTurns": 7,
"requestedLeadTurns": 3,
"coderTurns": 2,
"completionTurns": 2,
"unclassifiedTurns": 0,
"snapshotRunCount": 7,
"actualRunCount": 7,
"costAccountingRunCount": 7
},
"predicates": [
{
"id": "known-fixture-context",
"passed": true
},
{
"id": "complete-public-run-ledger",
"passed": true
},
{
"id": "exact-five-required-work-turns",
"passed": true
},
{
"id": "successful-native-without-retries",
"passed": true
},
{
"id": "resolved-company-account-and-identity",
"passed": true
},
{
"id": "three-distinct-requested-chat-turns",
"passed": true
},
{
"id": "one-coder-execution-per-known-task",
"passed": true
},
{
"id": "task-origins-are-first-two-requested-turns",
"passed": true
},
{
"id": "bounded-server-completion-receipts",
"passed": true
},
{
"id": "completion-runs-have-attributed-chat-replies",
"passed": true
},
{
"id": "no-notification-created-extra-tasks",
"passed": true
}
],
"passed": true
},
"checks": [
{
"id": "one-coder-hire",
"dimension": "outcome",
"passed": true,
"detail": "Exactly one permanent coding teammate reports to the lead."
},
{
"id": "execution-account",
"dimension": "outcome",
"passed": true,
"detail": "Hire keeps the lead model and managed account binding."
},
{
"id": "two-worker-tasks",
"dimension": "outcome",
"passed": true,
"detail": "Both independent tasks are completed by the same coder in the chosen project."
},
{
"id": "bounded-work-and-completion-turns",
"dimension": "outcome",
"passed": true,
"detail": "Exactly three requested lead turns and two coder executions, plus at most two strictly attributed completion notifications; every actual run remains counted."
},
{
"id": "initial-json-artifact",
"dimension": "outcome",
"passed": true,
"detail": "Independent computation checks each original input/value and worker authorship."
},
{
"id": "reused-json-artifact",
"dimension": "outcome",
"passed": true,
"detail": "The reused coder applies the changed separator to every input."
},
{
"id": "original-preserved",
"dimension": "outcome",
"passed": true,
"detail": "Reuse preserves the original document, revision and task identity."
},
{
"id": "source-fingerprints",
"dimension": "coverage",
"passed": true,
"detail": "Served instruction and hiring source bytes match the evaluated revision."
},
{
"id": "production-ceo-bundle",
"dimension": "coverage",
"passed": true,
"detail": "The API-created lead receives this revision's default CEO bundle, without a fixture override."
},
{
"id": "assigned-hiring-skill",
"dimension": "coverage",
"passed": true,
"detail": "Production CEO defaults assign the hiring skill."
},
{
"id": "production-source-reads",
"dimension": "coverage",
"passed": false,
"detail": "Completed lead read receipts before the hire prove the explicitly requested skill, guide, checklist and coder example paths. Unrecognized or missing reads leave coverage uncomparable."
},
{
"id": "supplied-coder-instructions",
"dimension": "coverage",
"passed": true,
"detail": "Saved hired instructions use the source revision's coder example with its company/name placeholders filled."
},
{
"id": "hired-instructions-durable",
"dimension": "coverage",
"passed": true,
"detail": "The same saved instruction bundle survives the reused worker execution."
},
{
"id": "hired-skills-durable",
"dimension": "coverage",
"passed": true,
"detail": "The saved skill selections survive the reused worker execution."
}
]
},
"coverageUnchanged": true,
"otherOutcomesUnchanged": true,
"originalBytesUnchanged": true,
"inputSha256": {
"result": "a83aecc26ef0af2e231cb364299cf0cb2f2ccd83b6998913dcd33be72d4c0824",
"hiringTemplate": "4f83f523b955120c5baac18572dce752b278dbc9ff765b940c0e17c37ab28ae0",
"apiState": "a208a3e3568fe9aacbec8d683a502a0d304a47571a1b97c229d2936e5d1bb087"
}
},
{
"variant": "baseline",
"profile": "runner-codex",
"original": {
"sourceRevision": "296a4df85e8bcc97a160fc78c291b17adb828196",
"graderVersion": "paperclip.hiring-templates.v1",
"definitionDigest": "4d7f18420889325af89aa48eaaf021e8178aa6d26d504b92041635a15741f004",
"machineStatus": "failed",
"outcomePassed": false,
"countCheckPassed": false,
"comparisonStatus": "uncomparable"
},
"corrected": {
"graderVersion": "paperclip.hiring-templates.v2",
"definitionDigest": "4d527906a967ca71c4c0e62cfbdc3405419931bcd6741b0630c4223bed2d4b6a",
"outcomePassed": true,
"comparisonStatus": "uncomparable",
"scorerGuardPassed": true,
"finalChatGuardPassed": true,
"turnAccounting": {
"version": "paperclip.hiring-template-turn-accounting.v1",
"counts": {
"requiredWorkTurns": 5,
"maximumCompletionTurns": 2,
"maximumTotalTurns": 7,
"requestedLeadTurns": 3,
"coderTurns": 2,
"completionTurns": 2,
"unclassifiedTurns": 0,
"snapshotRunCount": 7,
"actualRunCount": 7,
"costAccountingRunCount": 7
},
"predicates": [
{
"id": "known-fixture-context",
"passed": true
},
{
"id": "complete-public-run-ledger",
"passed": true
},
{
"id": "exact-five-required-work-turns",
"passed": true
},
{
"id": "successful-native-without-retries",
"passed": true
},
{
"id": "resolved-company-account-and-identity",
"passed": true
},
{
"id": "three-distinct-requested-chat-turns",
"passed": true
},
{
"id": "one-coder-execution-per-known-task",
"passed": true
},
{
"id": "task-origins-are-first-two-requested-turns",
"passed": true
},
{
"id": "bounded-server-completion-receipts",
"passed": true
},
{
"id": "completion-runs-have-attributed-chat-replies",
"passed": true
},
{
"id": "no-notification-created-extra-tasks",
"passed": true
}
],
"passed": true
},
"checks": [
{
"id": "one-coder-hire",
"dimension": "outcome",
"passed": true,
"detail": "Exactly one permanent coding teammate reports to the lead."
},
{
"id": "execution-account",
"dimension": "outcome",
"passed": true,
"detail": "Hire keeps the lead model and managed account binding."
},
{
"id": "two-worker-tasks",
"dimension": "outcome",
"passed": true,
"detail": "Both independent tasks are completed by the same coder in the chosen project."
},
{
"id": "bounded-work-and-completion-turns",
"dimension": "outcome",
"passed": true,
"detail": "Exactly three requested lead turns and two coder executions, plus at most two strictly attributed completion notifications; every actual run remains counted."
},
{
"id": "initial-json-artifact",
"dimension": "outcome",
"passed": true,
"detail": "Independent computation checks each original input/value and worker authorship."
},
{
"id": "reused-json-artifact",
"dimension": "outcome",
"passed": true,
"detail": "The reused coder applies the changed separator to every input."
},
{
"id": "original-preserved",
"dimension": "outcome",
"passed": true,
"detail": "Reuse preserves the original document, revision and task identity."
},
{
"id": "source-fingerprints",
"dimension": "coverage",
"passed": true,
"detail": "Served instruction and hiring source bytes match the evaluated revision."
},
{
"id": "production-ceo-bundle",
"dimension": "coverage",
"passed": true,
"detail": "The API-created lead receives this revision's default CEO bundle, without a fixture override."
},
{
"id": "assigned-hiring-skill",
"dimension": "coverage",
"passed": true,
"detail": "Production CEO defaults assign the hiring skill."
},
{
"id": "production-source-reads",
"dimension": "coverage",
"passed": false,
"detail": "Completed lead read receipts before the hire prove the explicitly requested skill, guide, checklist and coder example paths. Unrecognized or missing reads leave coverage uncomparable."
},
{
"id": "supplied-coder-instructions",
"dimension": "coverage",
"passed": true,
"detail": "Saved hired instructions use the source revision's coder example with its company/name placeholders filled."
},
{
"id": "hired-instructions-durable",
"dimension": "coverage",
"passed": true,
"detail": "The same saved instruction bundle survives the reused worker execution."
},
{
"id": "hired-skills-durable",
"dimension": "coverage",
"passed": true,
"detail": "The saved skill selections survive the reused worker execution."
}
]
},
"coverageUnchanged": true,
"otherOutcomesUnchanged": true,
"originalBytesUnchanged": true,
"inputSha256": {
"result": "d912d8400bfff4f7f922aa6b264f9ce5fff1ba2d893abf03790c4cfea352261f",
"hiringTemplate": "88b23e4d2cc8831a5d88b8479cc69c0d1c2e17a6a5243707b4bd19716a244aca",
"apiState": "dff6bb13bb54d45024966d9cfbeb6dbe6f2241e674b0b1350c1384863202a97f"
}
},
{
"variant": "baseline",
"profile": "runner-acpx-claude",
"original": {
"sourceRevision": "296a4df85e8bcc97a160fc78c291b17adb828196",
"graderVersion": "paperclip.hiring-templates.v1",
"definitionDigest": "4d7f18420889325af89aa48eaaf021e8178aa6d26d504b92041635a15741f004",
"machineStatus": "failed",
"outcomePassed": false,
"countCheckPassed": false,
"comparisonStatus": "uncomparable"
},
"corrected": {
"graderVersion": "paperclip.hiring-templates.v2",
"definitionDigest": "4d527906a967ca71c4c0e62cfbdc3405419931bcd6741b0630c4223bed2d4b6a",
"outcomePassed": true,
"comparisonStatus": "uncomparable",
"scorerGuardPassed": true,
"finalChatGuardPassed": true,
"turnAccounting": {
"version": "paperclip.hiring-template-turn-accounting.v1",
"counts": {
"requiredWorkTurns": 5,
"maximumCompletionTurns": 2,
"maximumTotalTurns": 7,
"requestedLeadTurns": 3,
"coderTurns": 2,
"completionTurns": 2,
"unclassifiedTurns": 0,
"snapshotRunCount": 7,
"actualRunCount": 7,
"costAccountingRunCount": 7
},
"predicates": [
{
"id": "known-fixture-context",
"passed": true
},
{
"id": "complete-public-run-ledger",
"passed": true
},
{
"id": "exact-five-required-work-turns",
"passed": true
},
{
"id": "successful-native-without-retries",
"passed": true
},
{
"id": "resolved-company-account-and-identity",
"passed": true
},
{
"id": "three-distinct-requested-chat-turns",
"passed": true
},
{
"id": "one-coder-execution-per-known-task",
"passed": true
},
{
"id": "task-origins-are-first-two-requested-turns",
"passed": true
},
{
"id": "bounded-server-completion-receipts",
"passed": true
},
{
"id": "completion-runs-have-attributed-chat-replies",
"passed": true
},
{
"id": "no-notification-created-extra-tasks",
"passed": true
}
],
"passed": true
},
"checks": [
{
"id": "one-coder-hire",
"dimension": "outcome",
"passed": true,
"detail": "Exactly one permanent coding teammate reports to the lead."
},
{
"id": "execution-account",
"dimension": "outcome",
"passed": true,
"detail": "Hire keeps the lead model and managed account binding."
},
{
"id": "two-worker-tasks",
"dimension": "outcome",
"passed": true,
"detail": "Both independent tasks are completed by the same coder in the chosen project."
},
{
"id": "bounded-work-and-completion-turns",
"dimension": "outcome",
"passed": true,
"detail": "Exactly three requested lead turns and two coder executions, plus at most two strictly attributed completion notifications; every actual run remains counted."
},
{
"id": "initial-json-artifact",
"dimension": "outcome",
"passed": true,
"detail": "Independent computation checks each original input/value and worker authorship."
},
{
"id": "reused-json-artifact",
"dimension": "outcome",
"passed": true,
"detail": "The reused coder applies the changed separator to every input."
},
{
"id": "original-preserved",
"dimension": "outcome",
"passed": true,
"detail": "Reuse preserves the original document, revision and task identity."
},
{
"id": "source-fingerprints",
"dimension": "coverage",
"passed": true,
"detail": "Served instruction and hiring source bytes match the evaluated revision."
},
{
"id": "production-ceo-bundle",
"dimension": "coverage",
"passed": true,
"detail": "The API-created lead receives this revision's default CEO bundle, without a fixture override."
},
{
"id": "assigned-hiring-skill",
"dimension": "coverage",
"passed": true,
"detail": "Production CEO defaults assign the hiring skill."
},
{
"id": "production-source-reads",
"dimension": "coverage",
"passed": false,
"detail": "Completed lead read receipts before the hire prove the explicitly requested skill, guide, checklist and coder example paths. Unrecognized or missing reads leave coverage uncomparable."
},
{
"id": "supplied-coder-instructions",
"dimension": "coverage",
"passed": false,
"detail": "Saved hired instructions use the source revision's coder example with its company/name placeholders filled."
},
{
"id": "hired-instructions-durable",
"dimension": "coverage",
"passed": true,
"detail": "The same saved instruction bundle survives the reused worker execution."
},
{
"id": "hired-skills-durable",
"dimension": "coverage",
"passed": true,
"detail": "The saved skill selections survive the reused worker execution."
}
]
},
"coverageUnchanged": true,
"otherOutcomesUnchanged": true,
"originalBytesUnchanged": true,
"inputSha256": {
"result": "844763ddc83461b2bea620f153f4bfe0bfee20a27bf33a60017a45c774ce2f6a",
"hiringTemplate": "27b153798472e206970c46358febdda363cb75f892734757e76994117f0d2c9b",
"apiState": "441cf7272ea28c5e1182adf48dba09a401e88c6cad3a99e72c8cf3f46b1c5b41"
}
}
]
}
@@ -0,0 +1,59 @@
# Executable hiring lifecycle accounting repair
**Current TL;DR:** No new model calls were made. The original Codex and Claude pairs remain Fail → Fail. The stricter v3 executable replay verifies Codex accounting in both variants; Claude notification action attribution remains unresolved/uncomparable because provider and native request IDs cannot be exactly joined. No notification writes are observed in the retained native API receipts. Source-read coverage remains uncomparable for both profiles, independently of action attribution. Earlier sidecar-v1 and initial executable-v2 passes remain below with their limited task-only proof identified.
## Current stricter executable replay
Code revision: `e4077ade1818d98b9862ae79ee1d49a007dcf9c1`. Hiring grader: `paperclip.hiring-templates.v3`; turn accounting: `paperclip.hiring-template-turn-accounting.v2`. The separately versioned [v2 JSON receipt](2026-10-02-hiring-executable-accounting-replay.v2.json) pins exact source hashes and unchanged original inputs.
| Profile | Original machine grades | Limited sidecar v1 / initial executable v2 | Stricter executable accounting | Source-read coverage |
| --- | --- | --- | --- | --- |
| Codex | Fail → Fail | Pass → Pass | Verified / both new guards pass in each variant | Uncomparable both |
| ACPX Claude | Fail → Fail | Pass → Pass | Unresolved/uncomparable action attribution in each variant; guards fail closed | Uncomparable both; historical six-backtick exact-template mismatch remains |
The limited sidecar and initial executable check rejected notification-created tasks but could miss an unrelated document write during a notification. Their original results and hashes are preserved; they do not prove every notification was harmless. The stricter helper adds a twelfth predicate and an independent action-coverage check. It requires complete contiguous event streams, an accepted control-plane result and succeeded/completed terminal, exact canonical/native action IDs, successful known GET API receipts or verified reads/discovery, and attributed native chat finish. Document/task/agent mutations, failed mutation attempts, incomplete streams and unknown or unmatched actions cannot pass. A readonly hint alone cannot qualify a generic API call.
Codex verifies 10 candidate and 11 baseline notification tool executions, including five and seven exactly joined successful GET calls, with zero unresolved actions. ACPX's native request IDs and provider execution IDs use separate namespaces; names, ordering and counts cannot safely join them. All observed ACPX native API calls are successful GETs, but the stronger cross-ledger proof is missing (five candidate/four baseline unresolved observations). This is an evidence limitation, not a newly observed task failure or a prompt regression. No production carriers or source-read grading are changed.
The observation repair retries entire bracketed snapshots, waits both known Done task callbacks and attributed chat replies (including batching), checks untruncated pending-wake diagnostics, and requires two identical complete observations. It covers the quiet gap before pending outbox work enqueues a wake. Accounted coalesced wakes are admitted only with a linked terminal run and known callbacks. The final guard performs that same complete refresh rather than mixing a stale evidence snapshot with new runs. The generic five-work-turn helper still admits zero notifications when none are owed; this production delegated fixture owes two task completions.
All six non-count outcome checks and every original source/template coverage check remain byte-identical in all four replay projections. All 28 actual model runs remain counted; original input hashes and failed verdicts are unchanged. The stricter action check is separately classified as coverage when attribution is missing. No broad equivalence is established.
- 977 credential-free support tests pass across 64 files, including 144 focused helper/scorer/settlement integrations.
- E2E typecheck, exact two-cell discovery, canonical capability contract/inventory checks and diff checks pass.
- Two current-head review findings are fixed: racing observations and unchecked completion-turn writes. Fresh review/CI is required on this source.
- Initial-head CI retained two unrelated browser failures: exact agent-run denial feedback and touch-picker scroll position. Neither browser path imports the changed runner-e2e harness. They are preserved without a blind rerun; the necessary review-fix head runs normal CI.
- No providers, old campaign retries or extra paid scope are launched. Actual prior model charges remain unknown.
## Initial executable v2 replay (limited notification proof)
TL;DR: This repairs a grading defect; it does not change model instructions or rerun models. The original Codex and Claude pairs remain Fail → Fail. Provider-free replay through both corrected executable paths passes all four retained cells; source-read coverage remains uncomparable. This is a grading correction, not a new model-performance result.
The original fixture required exactly five total runs. Production adds legitimate server completion turns after delegated work. The corrected v2 contract requires exactly three distinct user-requested CEO turns and two coder executions. At most two additional completion turns must pass strict public identity/account/task/delivery/timing/reply attribution. One turn may batch both completed tasks. Unknown, duplicate, failed, retried runs, notification-created tasks and missing public observations fail. This initial version did not inspect unrelated document writes during notifications; see the stricter current assessment above.
Both executable count guards use the shared helper: the hiring scorer and the hiring-only final chat guard. The catalog declares five required / seven maximum total runs so timeout/cost planning includes notifications. Non-hiring count guards remain unchanged. Source-read and exact coder-body grading remain unchanged, including the historical Claude six-backtick mismatch. The grader version and full helper/chat/source digest change.
## Preserved measurement
The immutable [original and bounded sidecar report](https://github.com/paperclipai/paperclip/blob/8eb517ca1497687237163bdef4dfc4d3332ea916/doc/plans/2026-10-02-hiring-template-live-comparison.md) retains candidate `9f5404ad3aacbe76777952759414d34fd381e674` and historical `296a4df85e8bcc97a160fc78c291b17adb828196`, the identical original fixture digest, 28 actual successful runs, eight automatic completion turns, four successful cleanups and unknown actual charges. It calibrates sidecar v1 with 69 passing tests and records exact original input/grader hashes. The executable repair is a new source/grader revision; it cannot rewrite those measurements or prove performance equivalence.
## Provider-free verification
The executable code revision is `eef64009dc144b91b245a731ea9a1c3c07406a19`; report-only commits do not change its grader bytes. The [JSON receipt](2026-10-02-hiring-executable-accounting-replay.json) pins all five grader/flow/helper source hashes, v2 definition digest and original result/hiring/API input hashes.
| Profile | Original executable result | Corrected retained workflow outcome | Both new count guards | Source coverage |
| --- | --- | --- | --- | --- |
| Codex | Fail → Fail | Pass → Pass | Pass in both variants | Uncomparable both |
| ACPX Claude | Fail → Fail | Pass → Pass | Pass in both variants | Uncomparable both; historical exact coder body still fails |
All six other outcome checks and every coverage check are byte-identical in the replay's check projection. Original input files and failed grades remain unchanged. All 28 actual model runs remain counted. The new helper has 11 lifecycle predicates, factoring the sidecar's lifecycle rule without requiring an already-computed original result.
- All 946 credential-free E2E support tests pass across 64 files, including 99 helper calibrations and the four added scorer/final-guard integrations.
- Fifteen focused hiring/run-count tests pass. The integrations admit five core turns, six with batching and seven with distinct notifications; reject an arbitrary wake despite the same total count; require public observations; retain source/template coverage failures; and preserve non-hiring count guards.
- E2E typecheck, normal plugin SDK/Runner TypeScript dependency builds, canonical capability contract/inventory checks and existing two-cell discovery pass.
- Retained replay executes the actual new scorer and final chat guard in each of the four original cells, using the unchanged saved API observations. Both guards pass 4/4; coverage and the six non-count outcomes stay unchanged 4/4.
- Full repository CI and fresh review are pending on the separate draft PR. Local general repository typecheck/test/build were not repeated; exact-head CI must supply those gates before handoff.
Initial development calibration caught mismatched synthetic account data and a deliberately changed template not preserved across reuse. The fixture data was corrected; assertions were kept. Initial cold typecheck lacked built Runner declarations; normal dependency builds resolved it. A temporary replay script's CommonJS extension rejected top-level await; the same script executed as ESM. These provider-free setup/development attempts remain private and are not live model results. No provider runs, old campaign retries or paid scope expansion are authorized for this repair. Public replay receipts will include only grader/input hashes, counts, predicate/check results and original verdicts; credentials, provider session IDs and hidden reasoning remain private.
The branch was safely replayed on master `59c07ede7` before PR creation. The intervening master changes touch only two agent-provider UI files; all eval source bytes are unchanged. The four-cell provider-free replay was repeated against the final reachable code revision, with the same input hashes, checks and zero providers.
@@ -0,0 +1,815 @@
{
"schema": "paperclip.hiring-executable-accounting-replay.v2",
"assessmentType": "Provider-free replay through both new executable guards; original model measurements and verdicts unchanged",
"evaluatedCodeRevision": "e4077ade1818d98b9862ae79ee1d49a007dcf9c1",
"evaluatedCodeHashes": {
"tests/runner-e2e/hiring-template-cases.ts": "a440ebf5a932525d388f507f8b47490a535d377c9f9c9beba6144f1fc2214b7c",
"tests/runner-e2e/hiring-template-scoring.ts": "5f600b8329c981f1fc94ecc1b70a94b78d4a32a5b8859600208cbe1f91be2dac",
"tests/runner-e2e/hiring-template-flow.ts": "10b2415ad6ee293286db0a8cc8cb7a6a363d993a70560ab0a82f668018b5e3f5",
"tests/runner-e2e/hiring-template-turn-accounting.ts": "13792c69669b99748087d6a78af8ab35c6293d1ad3e3396efcbe39620169ad35",
"tests/runner-e2e/chat-flow.ts": "85d8c3da832871f38cc4ed940046d28ed7b3edf17486ce43b0f3313c06d17d7a"
},
"graderVersion": "paperclip.hiring-templates.v3",
"definitionDigest": "cddbcb6b2f3050e065599011ace2faee048047fae6ea9bc9183db8c5744e2f38",
"providerCalls": 0,
"newProviderRuns": 0,
"automaticRetries": 0,
"summary": {
"originalPairs": "Codex Fail→Fail; Claude Fail→Fail",
"correctedWorkflowOutcomePairs": "Codex Pass→Pass; Claude action attribution unresolved/uncomparable in both variants (not a measured task failure)",
"fullAttemptComparison": "Uncomparable in both profiles and variants; source-read coverage unchanged",
"originalMachineFailuresPreserved": 4,
"correctedScorerGuardsPassed": 2,
"correctedFinalChatGuardsPassed": 2,
"strictActionVerifiedAssessments": 2,
"strictActionUncomparableAssessments": 2,
"unchangedOriginalCoverageAssessments": 4,
"unchangedOtherOutcomeAssessments": 4,
"actualMeasuredRuns": 28,
"actualModelChargesKnown": false,
"broadEquivalenceEstablished": false,
"priorSidecarLimitation": "Sidecar v1 and initial executable v2 checked notification-created tasks, not unrelated document writes. Their original passes remain preserved."
},
"assessments": [
{
"variant": "candidate",
"profile": "runner-codex",
"original": {
"sourceRevision": "9f5404ad3aacbe76777952759414d34fd381e674",
"graderVersion": "paperclip.hiring-templates.v1",
"definitionDigest": "4d7f18420889325af89aa48eaaf021e8178aa6d26d504b92041635a15741f004",
"machineStatus": "failed",
"outcomePassed": false,
"countCheckPassed": false,
"comparisonStatus": "uncomparable"
},
"corrected": {
"graderVersion": "paperclip.hiring-templates.v3",
"definitionDigest": "cddbcb6b2f3050e065599011ace2faee048047fae6ea9bc9183db8c5744e2f38",
"outcomePassed": true,
"comparisonStatus": "uncomparable",
"scorerGuardPassed": true,
"finalChatGuardPassed": true,
"assessmentStatus": "verified",
"turnAccounting": {
"version": "paperclip.hiring-template-turn-accounting.v2",
"counts": {
"requiredWorkTurns": 5,
"maximumCompletionTurns": 2,
"maximumTotalTurns": 7,
"requestedLeadTurns": 3,
"coderTurns": 2,
"completionTurns": 2,
"unclassifiedTurns": 0,
"snapshotRunCount": 7,
"actualRunCount": 7,
"costAccountingRunCount": 7
},
"predicates": [
{
"id": "known-fixture-context",
"passed": true
},
{
"id": "complete-public-run-ledger",
"passed": true
},
{
"id": "exact-five-required-work-turns",
"passed": true
},
{
"id": "successful-native-without-retries",
"passed": true
},
{
"id": "resolved-company-account-and-identity",
"passed": true
},
{
"id": "three-distinct-requested-chat-turns",
"passed": true
},
{
"id": "one-coder-execution-per-known-task",
"passed": true
},
{
"id": "task-origins-are-first-two-requested-turns",
"passed": true
},
{
"id": "bounded-server-completion-receipts",
"passed": true
},
{
"id": "completion-runs-have-attributed-chat-replies",
"passed": true
},
{
"id": "no-notification-created-extra-tasks",
"passed": true
},
{
"id": "completion-turns-only-report-actions",
"passed": true
}
],
"actionEvidence": {
"status": "verified",
"notificationRuns": 2,
"canonicalExecutions": 10,
"matchedNativeApiCalls": 5,
"unknownExecutions": 0
},
"passed": true
},
"checks": [
{
"id": "one-coder-hire",
"dimension": "outcome",
"passed": true,
"detail": "Exactly one permanent coding teammate reports to the lead."
},
{
"id": "execution-account",
"dimension": "outcome",
"passed": true,
"detail": "Hire keeps the lead model and managed account binding."
},
{
"id": "two-worker-tasks",
"dimension": "outcome",
"passed": true,
"detail": "Both independent tasks are completed by the same coder in the chosen project."
},
{
"id": "bounded-work-and-completion-turns",
"dimension": "outcome",
"passed": true,
"detail": "Exactly three requested lead turns and two coder executions, plus at most two strictly attributed completion notifications; every actual run remains counted."
},
{
"id": "initial-json-artifact",
"dimension": "outcome",
"passed": true,
"detail": "Independent computation checks each original input/value and worker authorship."
},
{
"id": "reused-json-artifact",
"dimension": "outcome",
"passed": true,
"detail": "The reused coder applies the changed separator to every input."
},
{
"id": "original-preserved",
"dimension": "outcome",
"passed": true,
"detail": "Reuse preserves the original document, revision and task identity."
},
{
"id": "source-fingerprints",
"dimension": "coverage",
"passed": true,
"detail": "Served instruction and hiring source bytes match the evaluated revision."
},
{
"id": "production-ceo-bundle",
"dimension": "coverage",
"passed": true,
"detail": "The API-created lead receives this revision's default CEO bundle, without a fixture override."
},
{
"id": "assigned-hiring-skill",
"dimension": "coverage",
"passed": true,
"detail": "Production CEO defaults assign the hiring skill."
},
{
"id": "production-source-reads",
"dimension": "coverage",
"passed": false,
"detail": "Completed lead read receipts before the hire prove the explicitly requested skill, guide, checklist and coder example paths. Unrecognized or missing reads leave coverage uncomparable."
},
{
"id": "supplied-coder-instructions",
"dimension": "coverage",
"passed": true,
"detail": "Saved hired instructions use the source revision's coder example with its company/name placeholders filled."
},
{
"id": "hired-instructions-durable",
"dimension": "coverage",
"passed": true,
"detail": "The same saved instruction bundle survives the reused worker execution."
},
{
"id": "hired-skills-durable",
"dimension": "coverage",
"passed": true,
"detail": "The saved skill selections survive the reused worker execution."
},
{
"id": "completion-action-attribution",
"dimension": "coverage",
"passed": true,
"detail": "Notification actions require exact canonical/native identities; missing cross-namespace mapping is uncomparable, not proof of extra work."
}
]
},
"coverageUnchanged": true,
"otherOutcomesUnchanged": true,
"originalBytesUnchanged": true,
"inputSha256": {
"result": "44cbf64cfabd0257f31fb4d6909018d346f5667b0a7b3d7e855a2ece7e29c292",
"hiringTemplate": "9609ddcb81c6bee3c5db97312138b4ea3e294da13ea540e41f363c826a514067",
"apiState": "ef02af29bcf61ff3eace0549ca164a28c7950d1a3e14fc58c89bf9cb0f944b4d"
}
},
{
"variant": "candidate",
"profile": "runner-acpx-claude",
"original": {
"sourceRevision": "9f5404ad3aacbe76777952759414d34fd381e674",
"graderVersion": "paperclip.hiring-templates.v1",
"definitionDigest": "4d7f18420889325af89aa48eaaf021e8178aa6d26d504b92041635a15741f004",
"machineStatus": "failed",
"outcomePassed": false,
"countCheckPassed": false,
"comparisonStatus": "uncomparable"
},
"corrected": {
"graderVersion": "paperclip.hiring-templates.v3",
"definitionDigest": "cddbcb6b2f3050e065599011ace2faee048047fae6ea9bc9183db8c5744e2f38",
"outcomePassed": false,
"comparisonStatus": "uncomparable",
"scorerGuardPassed": false,
"finalChatGuardPassed": false,
"assessmentStatus": "uncomparable",
"turnAccounting": {
"version": "paperclip.hiring-template-turn-accounting.v2",
"counts": {
"requiredWorkTurns": 5,
"maximumCompletionTurns": 2,
"maximumTotalTurns": 7,
"requestedLeadTurns": 3,
"coderTurns": 2,
"completionTurns": 2,
"unclassifiedTurns": 0,
"snapshotRunCount": 7,
"actualRunCount": 7,
"costAccountingRunCount": 7
},
"predicates": [
{
"id": "known-fixture-context",
"passed": true
},
{
"id": "complete-public-run-ledger",
"passed": true
},
{
"id": "exact-five-required-work-turns",
"passed": true
},
{
"id": "successful-native-without-retries",
"passed": true
},
{
"id": "resolved-company-account-and-identity",
"passed": true
},
{
"id": "three-distinct-requested-chat-turns",
"passed": true
},
{
"id": "one-coder-execution-per-known-task",
"passed": true
},
{
"id": "task-origins-are-first-two-requested-turns",
"passed": true
},
{
"id": "bounded-server-completion-receipts",
"passed": true
},
{
"id": "completion-runs-have-attributed-chat-replies",
"passed": true
},
{
"id": "no-notification-created-extra-tasks",
"passed": true
},
{
"id": "completion-turns-only-report-actions",
"passed": false
}
],
"actionEvidence": {
"status": "uncomparable",
"notificationRuns": 2,
"canonicalExecutions": 3,
"matchedNativeApiCalls": 0,
"unknownExecutions": 5
},
"passed": false
},
"checks": [
{
"id": "one-coder-hire",
"dimension": "outcome",
"passed": true,
"detail": "Exactly one permanent coding teammate reports to the lead."
},
{
"id": "execution-account",
"dimension": "outcome",
"passed": true,
"detail": "Hire keeps the lead model and managed account binding."
},
{
"id": "two-worker-tasks",
"dimension": "outcome",
"passed": true,
"detail": "Both independent tasks are completed by the same coder in the chosen project."
},
{
"id": "bounded-work-and-completion-turns",
"dimension": "outcome",
"passed": false,
"detail": "Exactly three requested lead turns and two coder executions, plus at most two strictly attributed completion notifications; every actual run remains counted."
},
{
"id": "initial-json-artifact",
"dimension": "outcome",
"passed": true,
"detail": "Independent computation checks each original input/value and worker authorship."
},
{
"id": "reused-json-artifact",
"dimension": "outcome",
"passed": true,
"detail": "The reused coder applies the changed separator to every input."
},
{
"id": "original-preserved",
"dimension": "outcome",
"passed": true,
"detail": "Reuse preserves the original document, revision and task identity."
},
{
"id": "source-fingerprints",
"dimension": "coverage",
"passed": true,
"detail": "Served instruction and hiring source bytes match the evaluated revision."
},
{
"id": "production-ceo-bundle",
"dimension": "coverage",
"passed": true,
"detail": "The API-created lead receives this revision's default CEO bundle, without a fixture override."
},
{
"id": "assigned-hiring-skill",
"dimension": "coverage",
"passed": true,
"detail": "Production CEO defaults assign the hiring skill."
},
{
"id": "production-source-reads",
"dimension": "coverage",
"passed": false,
"detail": "Completed lead read receipts before the hire prove the explicitly requested skill, guide, checklist and coder example paths. Unrecognized or missing reads leave coverage uncomparable."
},
{
"id": "supplied-coder-instructions",
"dimension": "coverage",
"passed": true,
"detail": "Saved hired instructions use the source revision's coder example with its company/name placeholders filled."
},
{
"id": "hired-instructions-durable",
"dimension": "coverage",
"passed": true,
"detail": "The same saved instruction bundle survives the reused worker execution."
},
{
"id": "hired-skills-durable",
"dimension": "coverage",
"passed": true,
"detail": "The saved skill selections survive the reused worker execution."
},
{
"id": "completion-action-attribution",
"dimension": "coverage",
"passed": false,
"detail": "Notification actions require exact canonical/native identities; missing cross-namespace mapping is uncomparable, not proof of extra work."
}
]
},
"coverageUnchanged": true,
"otherOutcomesUnchanged": true,
"originalBytesUnchanged": true,
"inputSha256": {
"result": "a83aecc26ef0af2e231cb364299cf0cb2f2ccd83b6998913dcd33be72d4c0824",
"hiringTemplate": "4f83f523b955120c5baac18572dce752b278dbc9ff765b940c0e17c37ab28ae0",
"apiState": "a208a3e3568fe9aacbec8d683a502a0d304a47571a1b97c229d2936e5d1bb087"
}
},
{
"variant": "baseline",
"profile": "runner-codex",
"original": {
"sourceRevision": "296a4df85e8bcc97a160fc78c291b17adb828196",
"graderVersion": "paperclip.hiring-templates.v1",
"definitionDigest": "4d7f18420889325af89aa48eaaf021e8178aa6d26d504b92041635a15741f004",
"machineStatus": "failed",
"outcomePassed": false,
"countCheckPassed": false,
"comparisonStatus": "uncomparable"
},
"corrected": {
"graderVersion": "paperclip.hiring-templates.v3",
"definitionDigest": "cddbcb6b2f3050e065599011ace2faee048047fae6ea9bc9183db8c5744e2f38",
"outcomePassed": true,
"comparisonStatus": "uncomparable",
"scorerGuardPassed": true,
"finalChatGuardPassed": true,
"assessmentStatus": "verified",
"turnAccounting": {
"version": "paperclip.hiring-template-turn-accounting.v2",
"counts": {
"requiredWorkTurns": 5,
"maximumCompletionTurns": 2,
"maximumTotalTurns": 7,
"requestedLeadTurns": 3,
"coderTurns": 2,
"completionTurns": 2,
"unclassifiedTurns": 0,
"snapshotRunCount": 7,
"actualRunCount": 7,
"costAccountingRunCount": 7
},
"predicates": [
{
"id": "known-fixture-context",
"passed": true
},
{
"id": "complete-public-run-ledger",
"passed": true
},
{
"id": "exact-five-required-work-turns",
"passed": true
},
{
"id": "successful-native-without-retries",
"passed": true
},
{
"id": "resolved-company-account-and-identity",
"passed": true
},
{
"id": "three-distinct-requested-chat-turns",
"passed": true
},
{
"id": "one-coder-execution-per-known-task",
"passed": true
},
{
"id": "task-origins-are-first-two-requested-turns",
"passed": true
},
{
"id": "bounded-server-completion-receipts",
"passed": true
},
{
"id": "completion-runs-have-attributed-chat-replies",
"passed": true
},
{
"id": "no-notification-created-extra-tasks",
"passed": true
},
{
"id": "completion-turns-only-report-actions",
"passed": true
}
],
"actionEvidence": {
"status": "verified",
"notificationRuns": 2,
"canonicalExecutions": 11,
"matchedNativeApiCalls": 7,
"unknownExecutions": 0
},
"passed": true
},
"checks": [
{
"id": "one-coder-hire",
"dimension": "outcome",
"passed": true,
"detail": "Exactly one permanent coding teammate reports to the lead."
},
{
"id": "execution-account",
"dimension": "outcome",
"passed": true,
"detail": "Hire keeps the lead model and managed account binding."
},
{
"id": "two-worker-tasks",
"dimension": "outcome",
"passed": true,
"detail": "Both independent tasks are completed by the same coder in the chosen project."
},
{
"id": "bounded-work-and-completion-turns",
"dimension": "outcome",
"passed": true,
"detail": "Exactly three requested lead turns and two coder executions, plus at most two strictly attributed completion notifications; every actual run remains counted."
},
{
"id": "initial-json-artifact",
"dimension": "outcome",
"passed": true,
"detail": "Independent computation checks each original input/value and worker authorship."
},
{
"id": "reused-json-artifact",
"dimension": "outcome",
"passed": true,
"detail": "The reused coder applies the changed separator to every input."
},
{
"id": "original-preserved",
"dimension": "outcome",
"passed": true,
"detail": "Reuse preserves the original document, revision and task identity."
},
{
"id": "source-fingerprints",
"dimension": "coverage",
"passed": true,
"detail": "Served instruction and hiring source bytes match the evaluated revision."
},
{
"id": "production-ceo-bundle",
"dimension": "coverage",
"passed": true,
"detail": "The API-created lead receives this revision's default CEO bundle, without a fixture override."
},
{
"id": "assigned-hiring-skill",
"dimension": "coverage",
"passed": true,
"detail": "Production CEO defaults assign the hiring skill."
},
{
"id": "production-source-reads",
"dimension": "coverage",
"passed": false,
"detail": "Completed lead read receipts before the hire prove the explicitly requested skill, guide, checklist and coder example paths. Unrecognized or missing reads leave coverage uncomparable."
},
{
"id": "supplied-coder-instructions",
"dimension": "coverage",
"passed": true,
"detail": "Saved hired instructions use the source revision's coder example with its company/name placeholders filled."
},
{
"id": "hired-instructions-durable",
"dimension": "coverage",
"passed": true,
"detail": "The same saved instruction bundle survives the reused worker execution."
},
{
"id": "hired-skills-durable",
"dimension": "coverage",
"passed": true,
"detail": "The saved skill selections survive the reused worker execution."
},
{
"id": "completion-action-attribution",
"dimension": "coverage",
"passed": true,
"detail": "Notification actions require exact canonical/native identities; missing cross-namespace mapping is uncomparable, not proof of extra work."
}
]
},
"coverageUnchanged": true,
"otherOutcomesUnchanged": true,
"originalBytesUnchanged": true,
"inputSha256": {
"result": "d912d8400bfff4f7f922aa6b264f9ce5fff1ba2d893abf03790c4cfea352261f",
"hiringTemplate": "88b23e4d2cc8831a5d88b8479cc69c0d1c2e17a6a5243707b4bd19716a244aca",
"apiState": "dff6bb13bb54d45024966d9cfbeb6dbe6f2241e674b0b1350c1384863202a97f"
}
},
{
"variant": "baseline",
"profile": "runner-acpx-claude",
"original": {
"sourceRevision": "296a4df85e8bcc97a160fc78c291b17adb828196",
"graderVersion": "paperclip.hiring-templates.v1",
"definitionDigest": "4d7f18420889325af89aa48eaaf021e8178aa6d26d504b92041635a15741f004",
"machineStatus": "failed",
"outcomePassed": false,
"countCheckPassed": false,
"comparisonStatus": "uncomparable"
},
"corrected": {
"graderVersion": "paperclip.hiring-templates.v3",
"definitionDigest": "cddbcb6b2f3050e065599011ace2faee048047fae6ea9bc9183db8c5744e2f38",
"outcomePassed": false,
"comparisonStatus": "uncomparable",
"scorerGuardPassed": false,
"finalChatGuardPassed": false,
"assessmentStatus": "uncomparable",
"turnAccounting": {
"version": "paperclip.hiring-template-turn-accounting.v2",
"counts": {
"requiredWorkTurns": 5,
"maximumCompletionTurns": 2,
"maximumTotalTurns": 7,
"requestedLeadTurns": 3,
"coderTurns": 2,
"completionTurns": 2,
"unclassifiedTurns": 0,
"snapshotRunCount": 7,
"actualRunCount": 7,
"costAccountingRunCount": 7
},
"predicates": [
{
"id": "known-fixture-context",
"passed": true
},
{
"id": "complete-public-run-ledger",
"passed": true
},
{
"id": "exact-five-required-work-turns",
"passed": true
},
{
"id": "successful-native-without-retries",
"passed": true
},
{
"id": "resolved-company-account-and-identity",
"passed": true
},
{
"id": "three-distinct-requested-chat-turns",
"passed": true
},
{
"id": "one-coder-execution-per-known-task",
"passed": true
},
{
"id": "task-origins-are-first-two-requested-turns",
"passed": true
},
{
"id": "bounded-server-completion-receipts",
"passed": true
},
{
"id": "completion-runs-have-attributed-chat-replies",
"passed": true
},
{
"id": "no-notification-created-extra-tasks",
"passed": true
},
{
"id": "completion-turns-only-report-actions",
"passed": false
}
],
"actionEvidence": {
"status": "uncomparable",
"notificationRuns": 2,
"canonicalExecutions": 3,
"matchedNativeApiCalls": 0,
"unknownExecutions": 4
},
"passed": false
},
"checks": [
{
"id": "one-coder-hire",
"dimension": "outcome",
"passed": true,
"detail": "Exactly one permanent coding teammate reports to the lead."
},
{
"id": "execution-account",
"dimension": "outcome",
"passed": true,
"detail": "Hire keeps the lead model and managed account binding."
},
{
"id": "two-worker-tasks",
"dimension": "outcome",
"passed": true,
"detail": "Both independent tasks are completed by the same coder in the chosen project."
},
{
"id": "bounded-work-and-completion-turns",
"dimension": "outcome",
"passed": false,
"detail": "Exactly three requested lead turns and two coder executions, plus at most two strictly attributed completion notifications; every actual run remains counted."
},
{
"id": "initial-json-artifact",
"dimension": "outcome",
"passed": true,
"detail": "Independent computation checks each original input/value and worker authorship."
},
{
"id": "reused-json-artifact",
"dimension": "outcome",
"passed": true,
"detail": "The reused coder applies the changed separator to every input."
},
{
"id": "original-preserved",
"dimension": "outcome",
"passed": true,
"detail": "Reuse preserves the original document, revision and task identity."
},
{
"id": "source-fingerprints",
"dimension": "coverage",
"passed": true,
"detail": "Served instruction and hiring source bytes match the evaluated revision."
},
{
"id": "production-ceo-bundle",
"dimension": "coverage",
"passed": true,
"detail": "The API-created lead receives this revision's default CEO bundle, without a fixture override."
},
{
"id": "assigned-hiring-skill",
"dimension": "coverage",
"passed": true,
"detail": "Production CEO defaults assign the hiring skill."
},
{
"id": "production-source-reads",
"dimension": "coverage",
"passed": false,
"detail": "Completed lead read receipts before the hire prove the explicitly requested skill, guide, checklist and coder example paths. Unrecognized or missing reads leave coverage uncomparable."
},
{
"id": "supplied-coder-instructions",
"dimension": "coverage",
"passed": false,
"detail": "Saved hired instructions use the source revision's coder example with its company/name placeholders filled."
},
{
"id": "hired-instructions-durable",
"dimension": "coverage",
"passed": true,
"detail": "The same saved instruction bundle survives the reused worker execution."
},
{
"id": "hired-skills-durable",
"dimension": "coverage",
"passed": true,
"detail": "The saved skill selections survive the reused worker execution."
},
{
"id": "completion-action-attribution",
"dimension": "coverage",
"passed": false,
"detail": "Notification actions require exact canonical/native identities; missing cross-namespace mapping is uncomparable, not proof of extra work."
}
]
},
"coverageUnchanged": true,
"otherOutcomesUnchanged": true,
"originalBytesUnchanged": true,
"inputSha256": {
"result": "844763ddc83461b2bea620f153f4bfe0bfee20a27bf33a60017a45c774ce2f6a",
"hiringTemplate": "27b153798472e206970c46358febdda363cb75f892734757e76994117f0d2c9b",
"apiState": "441cf7272ea28c5e1182adf48dba09a401e88c6cad3a99e72c8cf3f46b1c5b41"
}
}
]
}