mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
fix(evals): account for hiring completion notifications (#15007)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Product E2E evals check real hiring and delegated task completion. > - The hiring fixture requires three requested CEO turns and two coder executions. > - The server can also wake the CEO when each delegated task completes. > - Two exact-five-run guards rejected these valid completion turns in all four retained cells. > - This pull request validates bounded completion turns in both guards. > - The benefit is accurate workflow grading while all actual runs and coverage failures remain visible. ## Linked Issues or Issue Description Refs: #14985, #14948, #14961. **What happened?** The original hiring comparison reports Codex Fail → Fail and Claude Fail → Fail. Each cell has seven successful runs. The five requested work turns are accompanied by two server task-completion notifications. All six other delivery checks pass. **Expected behavior** Require exactly three distinct user-requested CEO turns and one coder execution for each of two known tasks. Admit at most two strictly attributed server completion turns, including one turn that batches both tasks. Reject unknown, duplicate, failed, retried or extra-work runs. **Steps to reproduce** Inspect the retained four-cell report linked below. Each original result fails `five-successful-turns`. The same exact count was also enforced by the final chat-flow guard. ## What Changed - Add one typed lifecycle helper shared by the hiring scorer and the hiring-only final chat guard. - Validate public run ledgers, company/user/account identity, request attribution, task origins, completion deliveries, timing and replies. - Keep exactly five required work turns; declare seven maximum total turns for cost and timeout planning. - Count all actual runs, including notification runs and unexpected resets. Keep other chat count guards unchanged. - Version the hiring grader as v3 (turn accounting v2) and include the helper and chat guard in its definition digest. - Keep source-read, exact coder-body and all six other delivery checks unchanged. - Add 144 focused helper/scorer/settlement calibrations and separately versioned exact retained-input replay reports. - Retry complete bracketed observations, await both owed callbacks and attributed replies, and refresh the final guard consistently. - Reject unrelated completion writes and failed mutation attempts using exact canonical/native action IDs. Missing identity mapping is uncomparable action coverage. ## Verification - All 977 credential-free E2E support tests pass across 64 files, including 144 focused lifecycle/action/scorer/settlement calibrations. - E2E typecheck, ordinary plugin SDK and Runner TypeScript dependency builds, capability contract/inventory checks and the existing two-cell hiring discovery pass. - [Executable replay report](https://github.com/paperclipai/paperclip/blob/fed1729018cc100f5f4bbfb692777e49009c423b/doc/plans/2026-10-02-hiring-executable-accounting-replay.md) pins current code revision `e4077ade1818d98b9862ae79ee1d49a007dcf9c1`, v3 definition digest, exact source/input hashes and each original/new check. - The stricter replay verifies both Codex variants through both executable guards. ACPX Claude action attribution remains unresolved/uncomparable because provider execution IDs cannot be exactly joined to native request IDs; guards fail closed. No notification writes are observed. All six other outcomes and every original source/template coverage check stay unchanged. Original files and Fail → Fail machine verdicts remain preserved; zero providers are called. - Full attempts remain uncomparable in both profiles. Historical Claude also keeps its six-backtick exact-template mismatch. This grading repair does not prove model-performance equivalence. - The limited sidecar-v1 and initial executable-v2 passes checked notification-created tasks but could miss unrelated document writes. Those assessments remain preserved and do not prove harmless notifications. The stricter v3 replay is separate. - [Original measurement and separate sidecar](https://github.com/paperclipai/paperclip/blob/8eb517ca1497687237163bdef4dfc4d3332ea916/doc/plans/2026-10-02-hiring-template-live-comparison.md) retain 28 actual runs, eight automatic notifications, four successful cleanups and unknown actual model charges. No models are rerun. - The branch is replayed on master `59c07ede7`. Intervening master changes are UI-only; eval source bytes and replay verdicts match. The four-cell provider-free replay was repeated against the reachable code revision. - Initial-head normal CI retained browser failures in agent-run denial feedback and touch-picker scroll position. Those browser paths and imports were unchanged, but their cause was not established. The necessary review-fix head passes both browser checks; no blind rerun was requested. - Local full repository typecheck/test/build were not repeated. Exact-head normal CI passes the required repository gates, including typecheck, tests, build and browser shards. Fresh Greptile review completed on `fed1729018cc100f5f4bbfb692777e49009c423b` with 5/5 and zero unresolved threads. An independent rerun of the 144 focused helper/scorer/settlement tests passes on the unchanged head. **Merge readiness:** This PR repairs the evaluator. Its positive and negative calibrations pass, both guards reject missing action attribution, current-head CI and review pass, and there are no merge conflicts. The retained ACPX cells remain uncomparable because their action IDs cannot be joined. That coverage limit remains a separate follow-up; it does not require relaxing this grader or changing the old results. No model calls, production instructions, carrier changes, or historical regrades are part of this readiness update. ## Risks - Missing or inconsistent public lifecycle evidence fails the bounded helper. The focused calibrations reject plausible false positives and malformed observations. Unmatched action IDs fail closed and are reported as uncomparable rather than a model task regression. - Source-read evidence remains incomplete. This PR does not change provider event carriers or relax the coverage oracle. - The versioned count check differs from original v1 results. Reports retain both versions and exact input hashes. ## Model Used OpenAI Codex, GPT-6 family as identified by this session. The exact deployment ID and context-window size are not exposed. The assistant used reasoning, repository tools, code execution and delegated calibration work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip normal CI gates are green (exact head `fed1729018cc100f5f4bbfb692777e49009c423b`; fresh review tracked separately below) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (completed exact-head review; zero unresolved threads) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
This commit is contained in:
1 parent
dd868ed125
commit
78e0034498
13 files changed
+2631
-44
No files matched your search
+5
-1
@@ -81,7 +81,9 @@ The explicit-only [production hiring templates suite](../tests/runner-e2e/README
|
||||
adds two local native Codex/Claude cells. It exercises API-created production
|
||||
CEO defaults, an explicitly requested hiring skill/reference read, a permanent
|
||||
coder hire, independently computed saved JSON fixtures and worker reuse.
|
||||
Each cell expects five turns. Source/read coverage and workflow outcome are
|
||||
Each cell requires five work turns and admits at most two strictly attributed
|
||||
server task-completion turns. Every actual run remains counted; unknown or
|
||||
extra-work turns fail. Source/read coverage and workflow outcome are
|
||||
separate: missing read provenance leaves the candidate/baseline pair
|
||||
uncomparable even if work succeeds. Baseline bundles and coder examples derive
|
||||
from their own source revision, without requiring candidate wording or length.
|
||||
@@ -386,3 +388,5 @@ The explicit-only Product E2E `confirmation-replies` suite tests conversational
|
||||
approval and rejection, persisted message provenance, approval before execution,
|
||||
ambiguous proposals, and the existing card-click path with native Claude/Codex.
|
||||
See the [suite contract](../tests/runner-e2e/README.md#conversational-confirmation-replies-explicit-only).
|
||||
|
||||
Hiring notification accounting now also requires exact completed action attribution. Missing native/provider ID mapping is uncomparable evidence; it must not be reported as a model task regression or waived through name/order matching. The fixture waits for both known completion callbacks and settled bracketed observations, including the gap before pending outbox work becomes a wake. Strict action replay and original machine verdicts are retained separately.
|
||||
@@ -0,0 +1,740 @@
|
||||
{
|
||||
"schema": "paperclip.hiring-executable-accounting-replay.v1",
|
||||
"assessmentType": "Provider-free replay through both new executable guards; original model measurements and verdicts unchanged",
|
||||
"evaluatedCodeRevision": "eef64009dc144b91b245a731ea9a1c3c07406a19",
|
||||
"evaluatedCodeHashes": {
|
||||
"tests/runner-e2e/hiring-template-cases.ts": "06915e33cc9a5241cd28a7184fa33d098c30aed4114accea4f93838aaba3738f",
|
||||
"tests/runner-e2e/hiring-template-scoring.ts": "a9e1efa7e11bad3aa8dc873a28325b3153799251740ae87c4623a08f7d45a5cb",
|
||||
"tests/runner-e2e/hiring-template-flow.ts": "f3c6d29dc5019955dda4372ead7a9fdfbde447b7b92ef103527b9393ec60bd30",
|
||||
"tests/runner-e2e/hiring-template-turn-accounting.ts": "8f2ad8b09e419796b372929b59dc7f86d41dcbf3a3144c45e329a1aa988b383b",
|
||||
"tests/runner-e2e/chat-flow.ts": "09dd87ef7aef299ddbd1fad65b72817c042595f9e4a4e699c430497ac053a2d4"
|
||||
},
|
||||
"graderVersion": "paperclip.hiring-templates.v2",
|
||||
"definitionDigest": "4d527906a967ca71c4c0e62cfbdc3405419931bcd6741b0630c4223bed2d4b6a",
|
||||
"providerCalls": 0,
|
||||
"newProviderRuns": 0,
|
||||
"automaticRetries": 0,
|
||||
"summary": {
|
||||
"originalPairs": "Codex Fail→Fail; Claude Fail→Fail",
|
||||
"correctedWorkflowOutcomePairs": "Codex Pass→Pass; Claude Pass→Pass",
|
||||
"fullAttemptComparison": "Uncomparable in both profiles and variants; source-read coverage unchanged",
|
||||
"originalMachineFailuresPreserved": 4,
|
||||
"correctedScorerGuardsPassed": 4,
|
||||
"correctedFinalChatGuardsPassed": 4,
|
||||
"unchangedCoverageAssessments": 4,
|
||||
"unchangedOtherOutcomeAssessments": 4,
|
||||
"actualMeasuredRuns": 28,
|
||||
"actualModelChargesKnown": false,
|
||||
"broadEquivalenceEstablished": false
|
||||
},
|
||||
"assessments": [
|
||||
{
|
||||
"variant": "candidate",
|
||||
"profile": "runner-codex",
|
||||
"original": {
|
||||
"sourceRevision": "9f5404ad3aacbe76777952759414d34fd381e674",
|
||||
"graderVersion": "paperclip.hiring-templates.v1",
|
||||
"definitionDigest": "4d7f18420889325af89aa48eaaf021e8178aa6d26d504b92041635a15741f004",
|
||||
"machineStatus": "failed",
|
||||
"outcomePassed": false,
|
||||
"countCheckPassed": false,
|
||||
"comparisonStatus": "uncomparable"
|
||||
},
|
||||
"corrected": {
|
||||
"graderVersion": "paperclip.hiring-templates.v2",
|
||||
"definitionDigest": "4d527906a967ca71c4c0e62cfbdc3405419931bcd6741b0630c4223bed2d4b6a",
|
||||
"outcomePassed": true,
|
||||
"comparisonStatus": "uncomparable",
|
||||
"scorerGuardPassed": true,
|
||||
"finalChatGuardPassed": true,
|
||||
"turnAccounting": {
|
||||
"version": "paperclip.hiring-template-turn-accounting.v1",
|
||||
"counts": {
|
||||
"requiredWorkTurns": 5,
|
||||
"maximumCompletionTurns": 2,
|
||||
"maximumTotalTurns": 7,
|
||||
"requestedLeadTurns": 3,
|
||||
"coderTurns": 2,
|
||||
"completionTurns": 2,
|
||||
"unclassifiedTurns": 0,
|
||||
"snapshotRunCount": 7,
|
||||
"actualRunCount": 7,
|
||||
"costAccountingRunCount": 7
|
||||
},
|
||||
"predicates": [
|
||||
{
|
||||
"id": "known-fixture-context",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "complete-public-run-ledger",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "exact-five-required-work-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "successful-native-without-retries",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "resolved-company-account-and-identity",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "three-distinct-requested-chat-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "one-coder-execution-per-known-task",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "task-origins-are-first-two-requested-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "bounded-server-completion-receipts",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "completion-runs-have-attributed-chat-replies",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "no-notification-created-extra-tasks",
|
||||
"passed": true
|
||||
}
|
||||
],
|
||||
"passed": true
|
||||
},
|
||||
"checks": [
|
||||
{
|
||||
"id": "one-coder-hire",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Exactly one permanent coding teammate reports to the lead."
|
||||
},
|
||||
{
|
||||
"id": "execution-account",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Hire keeps the lead model and managed account binding."
|
||||
},
|
||||
{
|
||||
"id": "two-worker-tasks",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Both independent tasks are completed by the same coder in the chosen project."
|
||||
},
|
||||
{
|
||||
"id": "bounded-work-and-completion-turns",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Exactly three requested lead turns and two coder executions, plus at most two strictly attributed completion notifications; every actual run remains counted."
|
||||
},
|
||||
{
|
||||
"id": "initial-json-artifact",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Independent computation checks each original input/value and worker authorship."
|
||||
},
|
||||
{
|
||||
"id": "reused-json-artifact",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "The reused coder applies the changed separator to every input."
|
||||
},
|
||||
{
|
||||
"id": "original-preserved",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Reuse preserves the original document, revision and task identity."
|
||||
},
|
||||
{
|
||||
"id": "source-fingerprints",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Served instruction and hiring source bytes match the evaluated revision."
|
||||
},
|
||||
{
|
||||
"id": "production-ceo-bundle",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The API-created lead receives this revision's default CEO bundle, without a fixture override."
|
||||
},
|
||||
{
|
||||
"id": "assigned-hiring-skill",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Production CEO defaults assign the hiring skill."
|
||||
},
|
||||
{
|
||||
"id": "production-source-reads",
|
||||
"dimension": "coverage",
|
||||
"passed": false,
|
||||
"detail": "Completed lead read receipts before the hire prove the explicitly requested skill, guide, checklist and coder example paths. Unrecognized or missing reads leave coverage uncomparable."
|
||||
},
|
||||
{
|
||||
"id": "supplied-coder-instructions",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Saved hired instructions use the source revision's coder example with its company/name placeholders filled."
|
||||
},
|
||||
{
|
||||
"id": "hired-instructions-durable",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The same saved instruction bundle survives the reused worker execution."
|
||||
},
|
||||
{
|
||||
"id": "hired-skills-durable",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The saved skill selections survive the reused worker execution."
|
||||
}
|
||||
]
|
||||
},
|
||||
"coverageUnchanged": true,
|
||||
"otherOutcomesUnchanged": true,
|
||||
"originalBytesUnchanged": true,
|
||||
"inputSha256": {
|
||||
"result": "44cbf64cfabd0257f31fb4d6909018d346f5667b0a7b3d7e855a2ece7e29c292",
|
||||
"hiringTemplate": "9609ddcb81c6bee3c5db97312138b4ea3e294da13ea540e41f363c826a514067",
|
||||
"apiState": "ef02af29bcf61ff3eace0549ca164a28c7950d1a3e14fc58c89bf9cb0f944b4d"
|
||||
}
|
||||
},
|
||||
{
|
||||
"variant": "candidate",
|
||||
"profile": "runner-acpx-claude",
|
||||
"original": {
|
||||
"sourceRevision": "9f5404ad3aacbe76777952759414d34fd381e674",
|
||||
"graderVersion": "paperclip.hiring-templates.v1",
|
||||
"definitionDigest": "4d7f18420889325af89aa48eaaf021e8178aa6d26d504b92041635a15741f004",
|
||||
"machineStatus": "failed",
|
||||
"outcomePassed": false,
|
||||
"countCheckPassed": false,
|
||||
"comparisonStatus": "uncomparable"
|
||||
},
|
||||
"corrected": {
|
||||
"graderVersion": "paperclip.hiring-templates.v2",
|
||||
"definitionDigest": "4d527906a967ca71c4c0e62cfbdc3405419931bcd6741b0630c4223bed2d4b6a",
|
||||
"outcomePassed": true,
|
||||
"comparisonStatus": "uncomparable",
|
||||
"scorerGuardPassed": true,
|
||||
"finalChatGuardPassed": true,
|
||||
"turnAccounting": {
|
||||
"version": "paperclip.hiring-template-turn-accounting.v1",
|
||||
"counts": {
|
||||
"requiredWorkTurns": 5,
|
||||
"maximumCompletionTurns": 2,
|
||||
"maximumTotalTurns": 7,
|
||||
"requestedLeadTurns": 3,
|
||||
"coderTurns": 2,
|
||||
"completionTurns": 2,
|
||||
"unclassifiedTurns": 0,
|
||||
"snapshotRunCount": 7,
|
||||
"actualRunCount": 7,
|
||||
"costAccountingRunCount": 7
|
||||
},
|
||||
"predicates": [
|
||||
{
|
||||
"id": "known-fixture-context",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "complete-public-run-ledger",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "exact-five-required-work-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "successful-native-without-retries",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "resolved-company-account-and-identity",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "three-distinct-requested-chat-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "one-coder-execution-per-known-task",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "task-origins-are-first-two-requested-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "bounded-server-completion-receipts",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "completion-runs-have-attributed-chat-replies",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "no-notification-created-extra-tasks",
|
||||
"passed": true
|
||||
}
|
||||
],
|
||||
"passed": true
|
||||
},
|
||||
"checks": [
|
||||
{
|
||||
"id": "one-coder-hire",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Exactly one permanent coding teammate reports to the lead."
|
||||
},
|
||||
{
|
||||
"id": "execution-account",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Hire keeps the lead model and managed account binding."
|
||||
},
|
||||
{
|
||||
"id": "two-worker-tasks",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Both independent tasks are completed by the same coder in the chosen project."
|
||||
},
|
||||
{
|
||||
"id": "bounded-work-and-completion-turns",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Exactly three requested lead turns and two coder executions, plus at most two strictly attributed completion notifications; every actual run remains counted."
|
||||
},
|
||||
{
|
||||
"id": "initial-json-artifact",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Independent computation checks each original input/value and worker authorship."
|
||||
},
|
||||
{
|
||||
"id": "reused-json-artifact",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "The reused coder applies the changed separator to every input."
|
||||
},
|
||||
{
|
||||
"id": "original-preserved",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Reuse preserves the original document, revision and task identity."
|
||||
},
|
||||
{
|
||||
"id": "source-fingerprints",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Served instruction and hiring source bytes match the evaluated revision."
|
||||
},
|
||||
{
|
||||
"id": "production-ceo-bundle",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The API-created lead receives this revision's default CEO bundle, without a fixture override."
|
||||
},
|
||||
{
|
||||
"id": "assigned-hiring-skill",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Production CEO defaults assign the hiring skill."
|
||||
},
|
||||
{
|
||||
"id": "production-source-reads",
|
||||
"dimension": "coverage",
|
||||
"passed": false,
|
||||
"detail": "Completed lead read receipts before the hire prove the explicitly requested skill, guide, checklist and coder example paths. Unrecognized or missing reads leave coverage uncomparable."
|
||||
},
|
||||
{
|
||||
"id": "supplied-coder-instructions",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Saved hired instructions use the source revision's coder example with its company/name placeholders filled."
|
||||
},
|
||||
{
|
||||
"id": "hired-instructions-durable",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The same saved instruction bundle survives the reused worker execution."
|
||||
},
|
||||
{
|
||||
"id": "hired-skills-durable",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The saved skill selections survive the reused worker execution."
|
||||
}
|
||||
]
|
||||
},
|
||||
"coverageUnchanged": true,
|
||||
"otherOutcomesUnchanged": true,
|
||||
"originalBytesUnchanged": true,
|
||||
"inputSha256": {
|
||||
"result": "a83aecc26ef0af2e231cb364299cf0cb2f2ccd83b6998913dcd33be72d4c0824",
|
||||
"hiringTemplate": "4f83f523b955120c5baac18572dce752b278dbc9ff765b940c0e17c37ab28ae0",
|
||||
"apiState": "a208a3e3568fe9aacbec8d683a502a0d304a47571a1b97c229d2936e5d1bb087"
|
||||
}
|
||||
},
|
||||
{
|
||||
"variant": "baseline",
|
||||
"profile": "runner-codex",
|
||||
"original": {
|
||||
"sourceRevision": "296a4df85e8bcc97a160fc78c291b17adb828196",
|
||||
"graderVersion": "paperclip.hiring-templates.v1",
|
||||
"definitionDigest": "4d7f18420889325af89aa48eaaf021e8178aa6d26d504b92041635a15741f004",
|
||||
"machineStatus": "failed",
|
||||
"outcomePassed": false,
|
||||
"countCheckPassed": false,
|
||||
"comparisonStatus": "uncomparable"
|
||||
},
|
||||
"corrected": {
|
||||
"graderVersion": "paperclip.hiring-templates.v2",
|
||||
"definitionDigest": "4d527906a967ca71c4c0e62cfbdc3405419931bcd6741b0630c4223bed2d4b6a",
|
||||
"outcomePassed": true,
|
||||
"comparisonStatus": "uncomparable",
|
||||
"scorerGuardPassed": true,
|
||||
"finalChatGuardPassed": true,
|
||||
"turnAccounting": {
|
||||
"version": "paperclip.hiring-template-turn-accounting.v1",
|
||||
"counts": {
|
||||
"requiredWorkTurns": 5,
|
||||
"maximumCompletionTurns": 2,
|
||||
"maximumTotalTurns": 7,
|
||||
"requestedLeadTurns": 3,
|
||||
"coderTurns": 2,
|
||||
"completionTurns": 2,
|
||||
"unclassifiedTurns": 0,
|
||||
"snapshotRunCount": 7,
|
||||
"actualRunCount": 7,
|
||||
"costAccountingRunCount": 7
|
||||
},
|
||||
"predicates": [
|
||||
{
|
||||
"id": "known-fixture-context",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "complete-public-run-ledger",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "exact-five-required-work-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "successful-native-without-retries",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "resolved-company-account-and-identity",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "three-distinct-requested-chat-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "one-coder-execution-per-known-task",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "task-origins-are-first-two-requested-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "bounded-server-completion-receipts",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "completion-runs-have-attributed-chat-replies",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "no-notification-created-extra-tasks",
|
||||
"passed": true
|
||||
}
|
||||
],
|
||||
"passed": true
|
||||
},
|
||||
"checks": [
|
||||
{
|
||||
"id": "one-coder-hire",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Exactly one permanent coding teammate reports to the lead."
|
||||
},
|
||||
{
|
||||
"id": "execution-account",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Hire keeps the lead model and managed account binding."
|
||||
},
|
||||
{
|
||||
"id": "two-worker-tasks",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Both independent tasks are completed by the same coder in the chosen project."
|
||||
},
|
||||
{
|
||||
"id": "bounded-work-and-completion-turns",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Exactly three requested lead turns and two coder executions, plus at most two strictly attributed completion notifications; every actual run remains counted."
|
||||
},
|
||||
{
|
||||
"id": "initial-json-artifact",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Independent computation checks each original input/value and worker authorship."
|
||||
},
|
||||
{
|
||||
"id": "reused-json-artifact",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "The reused coder applies the changed separator to every input."
|
||||
},
|
||||
{
|
||||
"id": "original-preserved",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Reuse preserves the original document, revision and task identity."
|
||||
},
|
||||
{
|
||||
"id": "source-fingerprints",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Served instruction and hiring source bytes match the evaluated revision."
|
||||
},
|
||||
{
|
||||
"id": "production-ceo-bundle",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The API-created lead receives this revision's default CEO bundle, without a fixture override."
|
||||
},
|
||||
{
|
||||
"id": "assigned-hiring-skill",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Production CEO defaults assign the hiring skill."
|
||||
},
|
||||
{
|
||||
"id": "production-source-reads",
|
||||
"dimension": "coverage",
|
||||
"passed": false,
|
||||
"detail": "Completed lead read receipts before the hire prove the explicitly requested skill, guide, checklist and coder example paths. Unrecognized or missing reads leave coverage uncomparable."
|
||||
},
|
||||
{
|
||||
"id": "supplied-coder-instructions",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Saved hired instructions use the source revision's coder example with its company/name placeholders filled."
|
||||
},
|
||||
{
|
||||
"id": "hired-instructions-durable",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The same saved instruction bundle survives the reused worker execution."
|
||||
},
|
||||
{
|
||||
"id": "hired-skills-durable",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The saved skill selections survive the reused worker execution."
|
||||
}
|
||||
]
|
||||
},
|
||||
"coverageUnchanged": true,
|
||||
"otherOutcomesUnchanged": true,
|
||||
"originalBytesUnchanged": true,
|
||||
"inputSha256": {
|
||||
"result": "d912d8400bfff4f7f922aa6b264f9ce5fff1ba2d893abf03790c4cfea352261f",
|
||||
"hiringTemplate": "88b23e4d2cc8831a5d88b8479cc69c0d1c2e17a6a5243707b4bd19716a244aca",
|
||||
"apiState": "dff6bb13bb54d45024966d9cfbeb6dbe6f2241e674b0b1350c1384863202a97f"
|
||||
}
|
||||
},
|
||||
{
|
||||
"variant": "baseline",
|
||||
"profile": "runner-acpx-claude",
|
||||
"original": {
|
||||
"sourceRevision": "296a4df85e8bcc97a160fc78c291b17adb828196",
|
||||
"graderVersion": "paperclip.hiring-templates.v1",
|
||||
"definitionDigest": "4d7f18420889325af89aa48eaaf021e8178aa6d26d504b92041635a15741f004",
|
||||
"machineStatus": "failed",
|
||||
"outcomePassed": false,
|
||||
"countCheckPassed": false,
|
||||
"comparisonStatus": "uncomparable"
|
||||
},
|
||||
"corrected": {
|
||||
"graderVersion": "paperclip.hiring-templates.v2",
|
||||
"definitionDigest": "4d527906a967ca71c4c0e62cfbdc3405419931bcd6741b0630c4223bed2d4b6a",
|
||||
"outcomePassed": true,
|
||||
"comparisonStatus": "uncomparable",
|
||||
"scorerGuardPassed": true,
|
||||
"finalChatGuardPassed": true,
|
||||
"turnAccounting": {
|
||||
"version": "paperclip.hiring-template-turn-accounting.v1",
|
||||
"counts": {
|
||||
"requiredWorkTurns": 5,
|
||||
"maximumCompletionTurns": 2,
|
||||
"maximumTotalTurns": 7,
|
||||
"requestedLeadTurns": 3,
|
||||
"coderTurns": 2,
|
||||
"completionTurns": 2,
|
||||
"unclassifiedTurns": 0,
|
||||
"snapshotRunCount": 7,
|
||||
"actualRunCount": 7,
|
||||
"costAccountingRunCount": 7
|
||||
},
|
||||
"predicates": [
|
||||
{
|
||||
"id": "known-fixture-context",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "complete-public-run-ledger",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "exact-five-required-work-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "successful-native-without-retries",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "resolved-company-account-and-identity",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "three-distinct-requested-chat-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "one-coder-execution-per-known-task",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "task-origins-are-first-two-requested-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "bounded-server-completion-receipts",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "completion-runs-have-attributed-chat-replies",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "no-notification-created-extra-tasks",
|
||||
"passed": true
|
||||
}
|
||||
],
|
||||
"passed": true
|
||||
},
|
||||
"checks": [
|
||||
{
|
||||
"id": "one-coder-hire",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Exactly one permanent coding teammate reports to the lead."
|
||||
},
|
||||
{
|
||||
"id": "execution-account",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Hire keeps the lead model and managed account binding."
|
||||
},
|
||||
{
|
||||
"id": "two-worker-tasks",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Both independent tasks are completed by the same coder in the chosen project."
|
||||
},
|
||||
{
|
||||
"id": "bounded-work-and-completion-turns",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Exactly three requested lead turns and two coder executions, plus at most two strictly attributed completion notifications; every actual run remains counted."
|
||||
},
|
||||
{
|
||||
"id": "initial-json-artifact",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Independent computation checks each original input/value and worker authorship."
|
||||
},
|
||||
{
|
||||
"id": "reused-json-artifact",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "The reused coder applies the changed separator to every input."
|
||||
},
|
||||
{
|
||||
"id": "original-preserved",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Reuse preserves the original document, revision and task identity."
|
||||
},
|
||||
{
|
||||
"id": "source-fingerprints",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Served instruction and hiring source bytes match the evaluated revision."
|
||||
},
|
||||
{
|
||||
"id": "production-ceo-bundle",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The API-created lead receives this revision's default CEO bundle, without a fixture override."
|
||||
},
|
||||
{
|
||||
"id": "assigned-hiring-skill",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Production CEO defaults assign the hiring skill."
|
||||
},
|
||||
{
|
||||
"id": "production-source-reads",
|
||||
"dimension": "coverage",
|
||||
"passed": false,
|
||||
"detail": "Completed lead read receipts before the hire prove the explicitly requested skill, guide, checklist and coder example paths. Unrecognized or missing reads leave coverage uncomparable."
|
||||
},
|
||||
{
|
||||
"id": "supplied-coder-instructions",
|
||||
"dimension": "coverage",
|
||||
"passed": false,
|
||||
"detail": "Saved hired instructions use the source revision's coder example with its company/name placeholders filled."
|
||||
},
|
||||
{
|
||||
"id": "hired-instructions-durable",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The same saved instruction bundle survives the reused worker execution."
|
||||
},
|
||||
{
|
||||
"id": "hired-skills-durable",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The saved skill selections survive the reused worker execution."
|
||||
}
|
||||
]
|
||||
},
|
||||
"coverageUnchanged": true,
|
||||
"otherOutcomesUnchanged": true,
|
||||
"originalBytesUnchanged": true,
|
||||
"inputSha256": {
|
||||
"result": "844763ddc83461b2bea620f153f4bfe0bfee20a27bf33a60017a45c774ce2f6a",
|
||||
"hiringTemplate": "27b153798472e206970c46358febdda363cb75f892734757e76994117f0d2c9b",
|
||||
"apiState": "441cf7272ea28c5e1182adf48dba09a401e88c6cad3a99e72c8cf3f46b1c5b41"
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,59 @@
|
||||
# Executable hiring lifecycle accounting repair
|
||||
|
||||
**Current TL;DR:** No new model calls were made. The original Codex and Claude pairs remain Fail → Fail. The stricter v3 executable replay verifies Codex accounting in both variants; Claude notification action attribution remains unresolved/uncomparable because provider and native request IDs cannot be exactly joined. No notification writes are observed in the retained native API receipts. Source-read coverage remains uncomparable for both profiles, independently of action attribution. Earlier sidecar-v1 and initial executable-v2 passes remain below with their limited task-only proof identified.
|
||||
|
||||
## Current stricter executable replay
|
||||
|
||||
Code revision: `e4077ade1818d98b9862ae79ee1d49a007dcf9c1`. Hiring grader: `paperclip.hiring-templates.v3`; turn accounting: `paperclip.hiring-template-turn-accounting.v2`. The separately versioned [v2 JSON receipt](2026-10-02-hiring-executable-accounting-replay.v2.json) pins exact source hashes and unchanged original inputs.
|
||||
|
||||
| Profile | Original machine grades | Limited sidecar v1 / initial executable v2 | Stricter executable accounting | Source-read coverage |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| Codex | Fail → Fail | Pass → Pass | Verified / both new guards pass in each variant | Uncomparable both |
|
||||
| ACPX Claude | Fail → Fail | Pass → Pass | Unresolved/uncomparable action attribution in each variant; guards fail closed | Uncomparable both; historical six-backtick exact-template mismatch remains |
|
||||
|
||||
The limited sidecar and initial executable check rejected notification-created tasks but could miss an unrelated document write during a notification. Their original results and hashes are preserved; they do not prove every notification was harmless. The stricter helper adds a twelfth predicate and an independent action-coverage check. It requires complete contiguous event streams, an accepted control-plane result and succeeded/completed terminal, exact canonical/native action IDs, successful known GET API receipts or verified reads/discovery, and attributed native chat finish. Document/task/agent mutations, failed mutation attempts, incomplete streams and unknown or unmatched actions cannot pass. A readonly hint alone cannot qualify a generic API call.
|
||||
|
||||
Codex verifies 10 candidate and 11 baseline notification tool executions, including five and seven exactly joined successful GET calls, with zero unresolved actions. ACPX's native request IDs and provider execution IDs use separate namespaces; names, ordering and counts cannot safely join them. All observed ACPX native API calls are successful GETs, but the stronger cross-ledger proof is missing (five candidate/four baseline unresolved observations). This is an evidence limitation, not a newly observed task failure or a prompt regression. No production carriers or source-read grading are changed.
|
||||
|
||||
The observation repair retries entire bracketed snapshots, waits both known Done task callbacks and attributed chat replies (including batching), checks untruncated pending-wake diagnostics, and requires two identical complete observations. It covers the quiet gap before pending outbox work enqueues a wake. Accounted coalesced wakes are admitted only with a linked terminal run and known callbacks. The final guard performs that same complete refresh rather than mixing a stale evidence snapshot with new runs. The generic five-work-turn helper still admits zero notifications when none are owed; this production delegated fixture owes two task completions.
|
||||
|
||||
All six non-count outcome checks and every original source/template coverage check remain byte-identical in all four replay projections. All 28 actual model runs remain counted; original input hashes and failed verdicts are unchanged. The stricter action check is separately classified as coverage when attribution is missing. No broad equivalence is established.
|
||||
|
||||
- 977 credential-free support tests pass across 64 files, including 144 focused helper/scorer/settlement integrations.
|
||||
- E2E typecheck, exact two-cell discovery, canonical capability contract/inventory checks and diff checks pass.
|
||||
- Two current-head review findings are fixed: racing observations and unchecked completion-turn writes. Fresh review/CI is required on this source.
|
||||
- Initial-head CI retained two unrelated browser failures: exact agent-run denial feedback and touch-picker scroll position. Neither browser path imports the changed runner-e2e harness. They are preserved without a blind rerun; the necessary review-fix head runs normal CI.
|
||||
- No providers, old campaign retries or extra paid scope are launched. Actual prior model charges remain unknown.
|
||||
|
||||
## Initial executable v2 replay (limited notification proof)
|
||||
|
||||
TL;DR: This repairs a grading defect; it does not change model instructions or rerun models. The original Codex and Claude pairs remain Fail → Fail. Provider-free replay through both corrected executable paths passes all four retained cells; source-read coverage remains uncomparable. This is a grading correction, not a new model-performance result.
|
||||
|
||||
The original fixture required exactly five total runs. Production adds legitimate server completion turns after delegated work. The corrected v2 contract requires exactly three distinct user-requested CEO turns and two coder executions. At most two additional completion turns must pass strict public identity/account/task/delivery/timing/reply attribution. One turn may batch both completed tasks. Unknown, duplicate, failed, retried runs, notification-created tasks and missing public observations fail. This initial version did not inspect unrelated document writes during notifications; see the stricter current assessment above.
|
||||
|
||||
Both executable count guards use the shared helper: the hiring scorer and the hiring-only final chat guard. The catalog declares five required / seven maximum total runs so timeout/cost planning includes notifications. Non-hiring count guards remain unchanged. Source-read and exact coder-body grading remain unchanged, including the historical Claude six-backtick mismatch. The grader version and full helper/chat/source digest change.
|
||||
|
||||
## Preserved measurement
|
||||
|
||||
The immutable [original and bounded sidecar report](https://github.com/paperclipai/paperclip/blob/8eb517ca1497687237163bdef4dfc4d3332ea916/doc/plans/2026-10-02-hiring-template-live-comparison.md) retains candidate `9f5404ad3aacbe76777952759414d34fd381e674` and historical `296a4df85e8bcc97a160fc78c291b17adb828196`, the identical original fixture digest, 28 actual successful runs, eight automatic completion turns, four successful cleanups and unknown actual charges. It calibrates sidecar v1 with 69 passing tests and records exact original input/grader hashes. The executable repair is a new source/grader revision; it cannot rewrite those measurements or prove performance equivalence.
|
||||
|
||||
## Provider-free verification
|
||||
|
||||
The executable code revision is `eef64009dc144b91b245a731ea9a1c3c07406a19`; report-only commits do not change its grader bytes. The [JSON receipt](2026-10-02-hiring-executable-accounting-replay.json) pins all five grader/flow/helper source hashes, v2 definition digest and original result/hiring/API input hashes.
|
||||
|
||||
| Profile | Original executable result | Corrected retained workflow outcome | Both new count guards | Source coverage |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| Codex | Fail → Fail | Pass → Pass | Pass in both variants | Uncomparable both |
|
||||
| ACPX Claude | Fail → Fail | Pass → Pass | Pass in both variants | Uncomparable both; historical exact coder body still fails |
|
||||
|
||||
All six other outcome checks and every coverage check are byte-identical in the replay's check projection. Original input files and failed grades remain unchanged. All 28 actual model runs remain counted. The new helper has 11 lifecycle predicates, factoring the sidecar's lifecycle rule without requiring an already-computed original result.
|
||||
|
||||
- All 946 credential-free E2E support tests pass across 64 files, including 99 helper calibrations and the four added scorer/final-guard integrations.
|
||||
- Fifteen focused hiring/run-count tests pass. The integrations admit five core turns, six with batching and seven with distinct notifications; reject an arbitrary wake despite the same total count; require public observations; retain source/template coverage failures; and preserve non-hiring count guards.
|
||||
- E2E typecheck, normal plugin SDK/Runner TypeScript dependency builds, canonical capability contract/inventory checks and existing two-cell discovery pass.
|
||||
- Retained replay executes the actual new scorer and final chat guard in each of the four original cells, using the unchanged saved API observations. Both guards pass 4/4; coverage and the six non-count outcomes stay unchanged 4/4.
|
||||
- Full repository CI and fresh review are pending on the separate draft PR. Local general repository typecheck/test/build were not repeated; exact-head CI must supply those gates before handoff.
|
||||
|
||||
Initial development calibration caught mismatched synthetic account data and a deliberately changed template not preserved across reuse. The fixture data was corrected; assertions were kept. Initial cold typecheck lacked built Runner declarations; normal dependency builds resolved it. A temporary replay script's CommonJS extension rejected top-level await; the same script executed as ESM. These provider-free setup/development attempts remain private and are not live model results. No provider runs, old campaign retries or paid scope expansion are authorized for this repair. Public replay receipts will include only grader/input hashes, counts, predicate/check results and original verdicts; credentials, provider session IDs and hidden reasoning remain private.
|
||||
|
||||
The branch was safely replayed on master `59c07ede7` before PR creation. The intervening master changes touch only two agent-provider UI files; all eval source bytes are unchanged. The four-cell provider-free replay was repeated against the final reachable code revision, with the same input hashes, checks and zero providers.
|
||||
@@ -0,0 +1,815 @@
|
||||
{
|
||||
"schema": "paperclip.hiring-executable-accounting-replay.v2",
|
||||
"assessmentType": "Provider-free replay through both new executable guards; original model measurements and verdicts unchanged",
|
||||
"evaluatedCodeRevision": "e4077ade1818d98b9862ae79ee1d49a007dcf9c1",
|
||||
"evaluatedCodeHashes": {
|
||||
"tests/runner-e2e/hiring-template-cases.ts": "a440ebf5a932525d388f507f8b47490a535d377c9f9c9beba6144f1fc2214b7c",
|
||||
"tests/runner-e2e/hiring-template-scoring.ts": "5f600b8329c981f1fc94ecc1b70a94b78d4a32a5b8859600208cbe1f91be2dac",
|
||||
"tests/runner-e2e/hiring-template-flow.ts": "10b2415ad6ee293286db0a8cc8cb7a6a363d993a70560ab0a82f668018b5e3f5",
|
||||
"tests/runner-e2e/hiring-template-turn-accounting.ts": "13792c69669b99748087d6a78af8ab35c6293d1ad3e3396efcbe39620169ad35",
|
||||
"tests/runner-e2e/chat-flow.ts": "85d8c3da832871f38cc4ed940046d28ed7b3edf17486ce43b0f3313c06d17d7a"
|
||||
},
|
||||
"graderVersion": "paperclip.hiring-templates.v3",
|
||||
"definitionDigest": "cddbcb6b2f3050e065599011ace2faee048047fae6ea9bc9183db8c5744e2f38",
|
||||
"providerCalls": 0,
|
||||
"newProviderRuns": 0,
|
||||
"automaticRetries": 0,
|
||||
"summary": {
|
||||
"originalPairs": "Codex Fail→Fail; Claude Fail→Fail",
|
||||
"correctedWorkflowOutcomePairs": "Codex Pass→Pass; Claude action attribution unresolved/uncomparable in both variants (not a measured task failure)",
|
||||
"fullAttemptComparison": "Uncomparable in both profiles and variants; source-read coverage unchanged",
|
||||
"originalMachineFailuresPreserved": 4,
|
||||
"correctedScorerGuardsPassed": 2,
|
||||
"correctedFinalChatGuardsPassed": 2,
|
||||
"strictActionVerifiedAssessments": 2,
|
||||
"strictActionUncomparableAssessments": 2,
|
||||
"unchangedOriginalCoverageAssessments": 4,
|
||||
"unchangedOtherOutcomeAssessments": 4,
|
||||
"actualMeasuredRuns": 28,
|
||||
"actualModelChargesKnown": false,
|
||||
"broadEquivalenceEstablished": false,
|
||||
"priorSidecarLimitation": "Sidecar v1 and initial executable v2 checked notification-created tasks, not unrelated document writes. Their original passes remain preserved."
|
||||
},
|
||||
"assessments": [
|
||||
{
|
||||
"variant": "candidate",
|
||||
"profile": "runner-codex",
|
||||
"original": {
|
||||
"sourceRevision": "9f5404ad3aacbe76777952759414d34fd381e674",
|
||||
"graderVersion": "paperclip.hiring-templates.v1",
|
||||
"definitionDigest": "4d7f18420889325af89aa48eaaf021e8178aa6d26d504b92041635a15741f004",
|
||||
"machineStatus": "failed",
|
||||
"outcomePassed": false,
|
||||
"countCheckPassed": false,
|
||||
"comparisonStatus": "uncomparable"
|
||||
},
|
||||
"corrected": {
|
||||
"graderVersion": "paperclip.hiring-templates.v3",
|
||||
"definitionDigest": "cddbcb6b2f3050e065599011ace2faee048047fae6ea9bc9183db8c5744e2f38",
|
||||
"outcomePassed": true,
|
||||
"comparisonStatus": "uncomparable",
|
||||
"scorerGuardPassed": true,
|
||||
"finalChatGuardPassed": true,
|
||||
"assessmentStatus": "verified",
|
||||
"turnAccounting": {
|
||||
"version": "paperclip.hiring-template-turn-accounting.v2",
|
||||
"counts": {
|
||||
"requiredWorkTurns": 5,
|
||||
"maximumCompletionTurns": 2,
|
||||
"maximumTotalTurns": 7,
|
||||
"requestedLeadTurns": 3,
|
||||
"coderTurns": 2,
|
||||
"completionTurns": 2,
|
||||
"unclassifiedTurns": 0,
|
||||
"snapshotRunCount": 7,
|
||||
"actualRunCount": 7,
|
||||
"costAccountingRunCount": 7
|
||||
},
|
||||
"predicates": [
|
||||
{
|
||||
"id": "known-fixture-context",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "complete-public-run-ledger",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "exact-five-required-work-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "successful-native-without-retries",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "resolved-company-account-and-identity",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "three-distinct-requested-chat-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "one-coder-execution-per-known-task",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "task-origins-are-first-two-requested-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "bounded-server-completion-receipts",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "completion-runs-have-attributed-chat-replies",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "no-notification-created-extra-tasks",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "completion-turns-only-report-actions",
|
||||
"passed": true
|
||||
}
|
||||
],
|
||||
"actionEvidence": {
|
||||
"status": "verified",
|
||||
"notificationRuns": 2,
|
||||
"canonicalExecutions": 10,
|
||||
"matchedNativeApiCalls": 5,
|
||||
"unknownExecutions": 0
|
||||
},
|
||||
"passed": true
|
||||
},
|
||||
"checks": [
|
||||
{
|
||||
"id": "one-coder-hire",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Exactly one permanent coding teammate reports to the lead."
|
||||
},
|
||||
{
|
||||
"id": "execution-account",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Hire keeps the lead model and managed account binding."
|
||||
},
|
||||
{
|
||||
"id": "two-worker-tasks",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Both independent tasks are completed by the same coder in the chosen project."
|
||||
},
|
||||
{
|
||||
"id": "bounded-work-and-completion-turns",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Exactly three requested lead turns and two coder executions, plus at most two strictly attributed completion notifications; every actual run remains counted."
|
||||
},
|
||||
{
|
||||
"id": "initial-json-artifact",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Independent computation checks each original input/value and worker authorship."
|
||||
},
|
||||
{
|
||||
"id": "reused-json-artifact",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "The reused coder applies the changed separator to every input."
|
||||
},
|
||||
{
|
||||
"id": "original-preserved",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Reuse preserves the original document, revision and task identity."
|
||||
},
|
||||
{
|
||||
"id": "source-fingerprints",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Served instruction and hiring source bytes match the evaluated revision."
|
||||
},
|
||||
{
|
||||
"id": "production-ceo-bundle",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The API-created lead receives this revision's default CEO bundle, without a fixture override."
|
||||
},
|
||||
{
|
||||
"id": "assigned-hiring-skill",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Production CEO defaults assign the hiring skill."
|
||||
},
|
||||
{
|
||||
"id": "production-source-reads",
|
||||
"dimension": "coverage",
|
||||
"passed": false,
|
||||
"detail": "Completed lead read receipts before the hire prove the explicitly requested skill, guide, checklist and coder example paths. Unrecognized or missing reads leave coverage uncomparable."
|
||||
},
|
||||
{
|
||||
"id": "supplied-coder-instructions",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Saved hired instructions use the source revision's coder example with its company/name placeholders filled."
|
||||
},
|
||||
{
|
||||
"id": "hired-instructions-durable",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The same saved instruction bundle survives the reused worker execution."
|
||||
},
|
||||
{
|
||||
"id": "hired-skills-durable",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The saved skill selections survive the reused worker execution."
|
||||
},
|
||||
{
|
||||
"id": "completion-action-attribution",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Notification actions require exact canonical/native identities; missing cross-namespace mapping is uncomparable, not proof of extra work."
|
||||
}
|
||||
]
|
||||
},
|
||||
"coverageUnchanged": true,
|
||||
"otherOutcomesUnchanged": true,
|
||||
"originalBytesUnchanged": true,
|
||||
"inputSha256": {
|
||||
"result": "44cbf64cfabd0257f31fb4d6909018d346f5667b0a7b3d7e855a2ece7e29c292",
|
||||
"hiringTemplate": "9609ddcb81c6bee3c5db97312138b4ea3e294da13ea540e41f363c826a514067",
|
||||
"apiState": "ef02af29bcf61ff3eace0549ca164a28c7950d1a3e14fc58c89bf9cb0f944b4d"
|
||||
}
|
||||
},
|
||||
{
|
||||
"variant": "candidate",
|
||||
"profile": "runner-acpx-claude",
|
||||
"original": {
|
||||
"sourceRevision": "9f5404ad3aacbe76777952759414d34fd381e674",
|
||||
"graderVersion": "paperclip.hiring-templates.v1",
|
||||
"definitionDigest": "4d7f18420889325af89aa48eaaf021e8178aa6d26d504b92041635a15741f004",
|
||||
"machineStatus": "failed",
|
||||
"outcomePassed": false,
|
||||
"countCheckPassed": false,
|
||||
"comparisonStatus": "uncomparable"
|
||||
},
|
||||
"corrected": {
|
||||
"graderVersion": "paperclip.hiring-templates.v3",
|
||||
"definitionDigest": "cddbcb6b2f3050e065599011ace2faee048047fae6ea9bc9183db8c5744e2f38",
|
||||
"outcomePassed": false,
|
||||
"comparisonStatus": "uncomparable",
|
||||
"scorerGuardPassed": false,
|
||||
"finalChatGuardPassed": false,
|
||||
"assessmentStatus": "uncomparable",
|
||||
"turnAccounting": {
|
||||
"version": "paperclip.hiring-template-turn-accounting.v2",
|
||||
"counts": {
|
||||
"requiredWorkTurns": 5,
|
||||
"maximumCompletionTurns": 2,
|
||||
"maximumTotalTurns": 7,
|
||||
"requestedLeadTurns": 3,
|
||||
"coderTurns": 2,
|
||||
"completionTurns": 2,
|
||||
"unclassifiedTurns": 0,
|
||||
"snapshotRunCount": 7,
|
||||
"actualRunCount": 7,
|
||||
"costAccountingRunCount": 7
|
||||
},
|
||||
"predicates": [
|
||||
{
|
||||
"id": "known-fixture-context",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "complete-public-run-ledger",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "exact-five-required-work-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "successful-native-without-retries",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "resolved-company-account-and-identity",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "three-distinct-requested-chat-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "one-coder-execution-per-known-task",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "task-origins-are-first-two-requested-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "bounded-server-completion-receipts",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "completion-runs-have-attributed-chat-replies",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "no-notification-created-extra-tasks",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "completion-turns-only-report-actions",
|
||||
"passed": false
|
||||
}
|
||||
],
|
||||
"actionEvidence": {
|
||||
"status": "uncomparable",
|
||||
"notificationRuns": 2,
|
||||
"canonicalExecutions": 3,
|
||||
"matchedNativeApiCalls": 0,
|
||||
"unknownExecutions": 5
|
||||
},
|
||||
"passed": false
|
||||
},
|
||||
"checks": [
|
||||
{
|
||||
"id": "one-coder-hire",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Exactly one permanent coding teammate reports to the lead."
|
||||
},
|
||||
{
|
||||
"id": "execution-account",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Hire keeps the lead model and managed account binding."
|
||||
},
|
||||
{
|
||||
"id": "two-worker-tasks",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Both independent tasks are completed by the same coder in the chosen project."
|
||||
},
|
||||
{
|
||||
"id": "bounded-work-and-completion-turns",
|
||||
"dimension": "outcome",
|
||||
"passed": false,
|
||||
"detail": "Exactly three requested lead turns and two coder executions, plus at most two strictly attributed completion notifications; every actual run remains counted."
|
||||
},
|
||||
{
|
||||
"id": "initial-json-artifact",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Independent computation checks each original input/value and worker authorship."
|
||||
},
|
||||
{
|
||||
"id": "reused-json-artifact",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "The reused coder applies the changed separator to every input."
|
||||
},
|
||||
{
|
||||
"id": "original-preserved",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Reuse preserves the original document, revision and task identity."
|
||||
},
|
||||
{
|
||||
"id": "source-fingerprints",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Served instruction and hiring source bytes match the evaluated revision."
|
||||
},
|
||||
{
|
||||
"id": "production-ceo-bundle",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The API-created lead receives this revision's default CEO bundle, without a fixture override."
|
||||
},
|
||||
{
|
||||
"id": "assigned-hiring-skill",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Production CEO defaults assign the hiring skill."
|
||||
},
|
||||
{
|
||||
"id": "production-source-reads",
|
||||
"dimension": "coverage",
|
||||
"passed": false,
|
||||
"detail": "Completed lead read receipts before the hire prove the explicitly requested skill, guide, checklist and coder example paths. Unrecognized or missing reads leave coverage uncomparable."
|
||||
},
|
||||
{
|
||||
"id": "supplied-coder-instructions",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Saved hired instructions use the source revision's coder example with its company/name placeholders filled."
|
||||
},
|
||||
{
|
||||
"id": "hired-instructions-durable",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The same saved instruction bundle survives the reused worker execution."
|
||||
},
|
||||
{
|
||||
"id": "hired-skills-durable",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The saved skill selections survive the reused worker execution."
|
||||
},
|
||||
{
|
||||
"id": "completion-action-attribution",
|
||||
"dimension": "coverage",
|
||||
"passed": false,
|
||||
"detail": "Notification actions require exact canonical/native identities; missing cross-namespace mapping is uncomparable, not proof of extra work."
|
||||
}
|
||||
]
|
||||
},
|
||||
"coverageUnchanged": true,
|
||||
"otherOutcomesUnchanged": true,
|
||||
"originalBytesUnchanged": true,
|
||||
"inputSha256": {
|
||||
"result": "a83aecc26ef0af2e231cb364299cf0cb2f2ccd83b6998913dcd33be72d4c0824",
|
||||
"hiringTemplate": "4f83f523b955120c5baac18572dce752b278dbc9ff765b940c0e17c37ab28ae0",
|
||||
"apiState": "a208a3e3568fe9aacbec8d683a502a0d304a47571a1b97c229d2936e5d1bb087"
|
||||
}
|
||||
},
|
||||
{
|
||||
"variant": "baseline",
|
||||
"profile": "runner-codex",
|
||||
"original": {
|
||||
"sourceRevision": "296a4df85e8bcc97a160fc78c291b17adb828196",
|
||||
"graderVersion": "paperclip.hiring-templates.v1",
|
||||
"definitionDigest": "4d7f18420889325af89aa48eaaf021e8178aa6d26d504b92041635a15741f004",
|
||||
"machineStatus": "failed",
|
||||
"outcomePassed": false,
|
||||
"countCheckPassed": false,
|
||||
"comparisonStatus": "uncomparable"
|
||||
},
|
||||
"corrected": {
|
||||
"graderVersion": "paperclip.hiring-templates.v3",
|
||||
"definitionDigest": "cddbcb6b2f3050e065599011ace2faee048047fae6ea9bc9183db8c5744e2f38",
|
||||
"outcomePassed": true,
|
||||
"comparisonStatus": "uncomparable",
|
||||
"scorerGuardPassed": true,
|
||||
"finalChatGuardPassed": true,
|
||||
"assessmentStatus": "verified",
|
||||
"turnAccounting": {
|
||||
"version": "paperclip.hiring-template-turn-accounting.v2",
|
||||
"counts": {
|
||||
"requiredWorkTurns": 5,
|
||||
"maximumCompletionTurns": 2,
|
||||
"maximumTotalTurns": 7,
|
||||
"requestedLeadTurns": 3,
|
||||
"coderTurns": 2,
|
||||
"completionTurns": 2,
|
||||
"unclassifiedTurns": 0,
|
||||
"snapshotRunCount": 7,
|
||||
"actualRunCount": 7,
|
||||
"costAccountingRunCount": 7
|
||||
},
|
||||
"predicates": [
|
||||
{
|
||||
"id": "known-fixture-context",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "complete-public-run-ledger",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "exact-five-required-work-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "successful-native-without-retries",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "resolved-company-account-and-identity",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "three-distinct-requested-chat-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "one-coder-execution-per-known-task",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "task-origins-are-first-two-requested-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "bounded-server-completion-receipts",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "completion-runs-have-attributed-chat-replies",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "no-notification-created-extra-tasks",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "completion-turns-only-report-actions",
|
||||
"passed": true
|
||||
}
|
||||
],
|
||||
"actionEvidence": {
|
||||
"status": "verified",
|
||||
"notificationRuns": 2,
|
||||
"canonicalExecutions": 11,
|
||||
"matchedNativeApiCalls": 7,
|
||||
"unknownExecutions": 0
|
||||
},
|
||||
"passed": true
|
||||
},
|
||||
"checks": [
|
||||
{
|
||||
"id": "one-coder-hire",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Exactly one permanent coding teammate reports to the lead."
|
||||
},
|
||||
{
|
||||
"id": "execution-account",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Hire keeps the lead model and managed account binding."
|
||||
},
|
||||
{
|
||||
"id": "two-worker-tasks",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Both independent tasks are completed by the same coder in the chosen project."
|
||||
},
|
||||
{
|
||||
"id": "bounded-work-and-completion-turns",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Exactly three requested lead turns and two coder executions, plus at most two strictly attributed completion notifications; every actual run remains counted."
|
||||
},
|
||||
{
|
||||
"id": "initial-json-artifact",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Independent computation checks each original input/value and worker authorship."
|
||||
},
|
||||
{
|
||||
"id": "reused-json-artifact",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "The reused coder applies the changed separator to every input."
|
||||
},
|
||||
{
|
||||
"id": "original-preserved",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Reuse preserves the original document, revision and task identity."
|
||||
},
|
||||
{
|
||||
"id": "source-fingerprints",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Served instruction and hiring source bytes match the evaluated revision."
|
||||
},
|
||||
{
|
||||
"id": "production-ceo-bundle",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The API-created lead receives this revision's default CEO bundle, without a fixture override."
|
||||
},
|
||||
{
|
||||
"id": "assigned-hiring-skill",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Production CEO defaults assign the hiring skill."
|
||||
},
|
||||
{
|
||||
"id": "production-source-reads",
|
||||
"dimension": "coverage",
|
||||
"passed": false,
|
||||
"detail": "Completed lead read receipts before the hire prove the explicitly requested skill, guide, checklist and coder example paths. Unrecognized or missing reads leave coverage uncomparable."
|
||||
},
|
||||
{
|
||||
"id": "supplied-coder-instructions",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Saved hired instructions use the source revision's coder example with its company/name placeholders filled."
|
||||
},
|
||||
{
|
||||
"id": "hired-instructions-durable",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The same saved instruction bundle survives the reused worker execution."
|
||||
},
|
||||
{
|
||||
"id": "hired-skills-durable",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The saved skill selections survive the reused worker execution."
|
||||
},
|
||||
{
|
||||
"id": "completion-action-attribution",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Notification actions require exact canonical/native identities; missing cross-namespace mapping is uncomparable, not proof of extra work."
|
||||
}
|
||||
]
|
||||
},
|
||||
"coverageUnchanged": true,
|
||||
"otherOutcomesUnchanged": true,
|
||||
"originalBytesUnchanged": true,
|
||||
"inputSha256": {
|
||||
"result": "d912d8400bfff4f7f922aa6b264f9ce5fff1ba2d893abf03790c4cfea352261f",
|
||||
"hiringTemplate": "88b23e4d2cc8831a5d88b8479cc69c0d1c2e17a6a5243707b4bd19716a244aca",
|
||||
"apiState": "dff6bb13bb54d45024966d9cfbeb6dbe6f2241e674b0b1350c1384863202a97f"
|
||||
}
|
||||
},
|
||||
{
|
||||
"variant": "baseline",
|
||||
"profile": "runner-acpx-claude",
|
||||
"original": {
|
||||
"sourceRevision": "296a4df85e8bcc97a160fc78c291b17adb828196",
|
||||
"graderVersion": "paperclip.hiring-templates.v1",
|
||||
"definitionDigest": "4d7f18420889325af89aa48eaaf021e8178aa6d26d504b92041635a15741f004",
|
||||
"machineStatus": "failed",
|
||||
"outcomePassed": false,
|
||||
"countCheckPassed": false,
|
||||
"comparisonStatus": "uncomparable"
|
||||
},
|
||||
"corrected": {
|
||||
"graderVersion": "paperclip.hiring-templates.v3",
|
||||
"definitionDigest": "cddbcb6b2f3050e065599011ace2faee048047fae6ea9bc9183db8c5744e2f38",
|
||||
"outcomePassed": false,
|
||||
"comparisonStatus": "uncomparable",
|
||||
"scorerGuardPassed": false,
|
||||
"finalChatGuardPassed": false,
|
||||
"assessmentStatus": "uncomparable",
|
||||
"turnAccounting": {
|
||||
"version": "paperclip.hiring-template-turn-accounting.v2",
|
||||
"counts": {
|
||||
"requiredWorkTurns": 5,
|
||||
"maximumCompletionTurns": 2,
|
||||
"maximumTotalTurns": 7,
|
||||
"requestedLeadTurns": 3,
|
||||
"coderTurns": 2,
|
||||
"completionTurns": 2,
|
||||
"unclassifiedTurns": 0,
|
||||
"snapshotRunCount": 7,
|
||||
"actualRunCount": 7,
|
||||
"costAccountingRunCount": 7
|
||||
},
|
||||
"predicates": [
|
||||
{
|
||||
"id": "known-fixture-context",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "complete-public-run-ledger",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "exact-five-required-work-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "successful-native-without-retries",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "resolved-company-account-and-identity",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "three-distinct-requested-chat-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "one-coder-execution-per-known-task",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "task-origins-are-first-two-requested-turns",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "bounded-server-completion-receipts",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "completion-runs-have-attributed-chat-replies",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "no-notification-created-extra-tasks",
|
||||
"passed": true
|
||||
},
|
||||
{
|
||||
"id": "completion-turns-only-report-actions",
|
||||
"passed": false
|
||||
}
|
||||
],
|
||||
"actionEvidence": {
|
||||
"status": "uncomparable",
|
||||
"notificationRuns": 2,
|
||||
"canonicalExecutions": 3,
|
||||
"matchedNativeApiCalls": 0,
|
||||
"unknownExecutions": 4
|
||||
},
|
||||
"passed": false
|
||||
},
|
||||
"checks": [
|
||||
{
|
||||
"id": "one-coder-hire",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Exactly one permanent coding teammate reports to the lead."
|
||||
},
|
||||
{
|
||||
"id": "execution-account",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Hire keeps the lead model and managed account binding."
|
||||
},
|
||||
{
|
||||
"id": "two-worker-tasks",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Both independent tasks are completed by the same coder in the chosen project."
|
||||
},
|
||||
{
|
||||
"id": "bounded-work-and-completion-turns",
|
||||
"dimension": "outcome",
|
||||
"passed": false,
|
||||
"detail": "Exactly three requested lead turns and two coder executions, plus at most two strictly attributed completion notifications; every actual run remains counted."
|
||||
},
|
||||
{
|
||||
"id": "initial-json-artifact",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Independent computation checks each original input/value and worker authorship."
|
||||
},
|
||||
{
|
||||
"id": "reused-json-artifact",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "The reused coder applies the changed separator to every input."
|
||||
},
|
||||
{
|
||||
"id": "original-preserved",
|
||||
"dimension": "outcome",
|
||||
"passed": true,
|
||||
"detail": "Reuse preserves the original document, revision and task identity."
|
||||
},
|
||||
{
|
||||
"id": "source-fingerprints",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Served instruction and hiring source bytes match the evaluated revision."
|
||||
},
|
||||
{
|
||||
"id": "production-ceo-bundle",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The API-created lead receives this revision's default CEO bundle, without a fixture override."
|
||||
},
|
||||
{
|
||||
"id": "assigned-hiring-skill",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "Production CEO defaults assign the hiring skill."
|
||||
},
|
||||
{
|
||||
"id": "production-source-reads",
|
||||
"dimension": "coverage",
|
||||
"passed": false,
|
||||
"detail": "Completed lead read receipts before the hire prove the explicitly requested skill, guide, checklist and coder example paths. Unrecognized or missing reads leave coverage uncomparable."
|
||||
},
|
||||
{
|
||||
"id": "supplied-coder-instructions",
|
||||
"dimension": "coverage",
|
||||
"passed": false,
|
||||
"detail": "Saved hired instructions use the source revision's coder example with its company/name placeholders filled."
|
||||
},
|
||||
{
|
||||
"id": "hired-instructions-durable",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The same saved instruction bundle survives the reused worker execution."
|
||||
},
|
||||
{
|
||||
"id": "hired-skills-durable",
|
||||
"dimension": "coverage",
|
||||
"passed": true,
|
||||
"detail": "The saved skill selections survive the reused worker execution."
|
||||
},
|
||||
{
|
||||
"id": "completion-action-attribution",
|
||||
"dimension": "coverage",
|
||||
"passed": false,
|
||||
"detail": "Notification actions require exact canonical/native identities; missing cross-namespace mapping is uncomparable, not proof of extra work."
|
||||
}
|
||||
]
|
||||
},
|
||||
"coverageUnchanged": true,
|
||||
"otherOutcomesUnchanged": true,
|
||||
"originalBytesUnchanged": true,
|
||||
"inputSha256": {
|
||||
"result": "844763ddc83461b2bea620f153f4bfe0bfee20a27bf33a60017a45c774ce2f6a",
|
||||
"hiringTemplate": "27b153798472e206970c46358febdda363cb75f892734757e76994117f0d2c9b",
|
||||
"apiState": "441cf7272ea28c5e1182adf48dba09a401e88c6cad3a99e72c8cf3f46b1c5b41"
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -1508,15 +1508,26 @@ explicit fixture user request. It is not an additional production requirement.
|
||||
A follow-up delegates a second fixture to the same coder with underscore
|
||||
separators while preserving the original. A final read-only chat turn requests
|
||||
recorded task status. Three CEO turns and two actual worker executions make
|
||||
**five expected turns per cell**, with a **15-minute deadline** and 1,000-cent
|
||||
**five required work turns per cell**, plus at most **two strictly attributed
|
||||
automatic task-completion turns** (seven total maximum), with a **15-minute
|
||||
deadline** and 1,000-cent
|
||||
company/CEO budget hard stops. Normal managed-account fixture cleanup and
|
||||
company-wide cancellation apply. Both cells opt into the existing native API
|
||||
tools. No model-authored code is executed by the grading host.
|
||||
|
||||
The independent oracle checks every JSON input and computed value, authorship,
|
||||
two distinct completed tasks, project/reporting identity, exact five-run count,
|
||||
two distinct completed tasks, project/reporting identity, exactly three user-requested
|
||||
CEO turns and one coder execution per task,
|
||||
managed execution-account attribution, original document preservation, and
|
||||
worker reuse. It separately checks the production CEO bundle, assigned hiring
|
||||
worker reuse. Bounded completion turns must have the same company, managed
|
||||
account, responsible user, chat generation and known completed tasks; unique
|
||||
server delivery/update receipts; valid completion timing; and a run-attributed
|
||||
chat reply. One completion turn may batch both tasks. Unknown, duplicate,
|
||||
failed, retried or extra work runs, and notification-created tasks fail. Every
|
||||
actual run remains in usage/cost accounting. The hiring scorer and final chat
|
||||
count guard use the same rule. All other chat count guards stay unchanged.
|
||||
|
||||
The versioned `paperclip.hiring-templates.v3` oracle separately checks the production CEO bundle, assigned hiring
|
||||
skill, source hashes, completed pre-hire read receipts, the saved source-derived
|
||||
coder example, and durable instruction/skill selections.
|
||||
|
||||
@@ -1524,7 +1535,8 @@ coder example, and durable instruction/skill selections.
|
||||
bytes on each evaluated revision. A historical four-file CEO bundle and long
|
||||
coder example are admissible; the candidate is not imposed on the baseline.
|
||||
Instruction bytes and word counts are measurements, without a size pass/fail
|
||||
threshold. The definition digest fingerprints the cases, flow and grader;
|
||||
threshold. The definition digest fingerprints the cases, flow, grader, shared
|
||||
turn-accounting helper and final chat guard;
|
||||
source evidence also fingerprints the loader, generic execution contract,
|
||||
selected CEO files and production hiring references. Use the same fixture
|
||||
revision, scenario nonce, profile/model, managed account method and local
|
||||
@@ -1566,3 +1578,9 @@ The existing `first-task` suite uses the actual onboarding wizard and captures
|
||||
the changed chief-of-staff persona and skill selections; it needs no fixture
|
||||
change for that default selection. Hiring from that wizard-created chief of
|
||||
staff remains a separate follow-up qualification.
|
||||
|
||||
### Hiring completion accounting evidence
|
||||
|
||||
The v3 hiring grader uses turn-accounting v2 in both executable guards. It requires complete per-run public event streams, exact native tool-use/result pairing and canonical execution IDs for completion actions. Only successful known GET issue/document/comment operations, verified reads/discovery, and attributed native chat finish are admitted. Writes, failed mutation attempts, incomplete streams and unknown actions cannot pass. Separate ACPX host request IDs and provider execution IDs are not joined by name/order/count; missing mapping is uncomparable action coverage, not a measured task failure. The original source-read and exact template checks remain unchanged.
|
||||
|
||||
The live fixture retries entire bracketed observations, waits for both known task callbacks and attributed replies (including batching), checks untruncated pending-wake diagnostics, and requires two equal settled observations. Silence before outbox enqueue is not delivery. Five-turn generic accounting remains calibrated for no owed notifications; this delegated fixture owes two completions. All actual runs remain counted for usage and cost. Retained original, limited sidecar-v1, initial executable, and stricter v3 assessments remain separately versioned; no models are rerun by the repair.
|
||||
@@ -1,4 +1,6 @@
|
||||
import { runHiringTemplateFlow } from "./hiring-template-flow.js";
|
||||
import { gradeHiringTemplateTurns } from "./hiring-template-turn-accounting.js";
|
||||
import type { HiringTemplateEvidence } from "./hiring-template-scoring.js";
|
||||
import { runAmbiguousConfirmationReply, runUnansweredQuestionReturn } from "./confirmation-replies.js";
|
||||
import { expect, type Page } from "@playwright/test";
|
||||
import { pollUntil, type RunnerApi } from "./api.js";
|
||||
@@ -17,6 +19,21 @@ import { runActiveReassignment, runWorkerCrash, runAnswerQuality } from "./chat-
|
||||
import { matchesRunCount, minimumRunCount } from "./run-count.js";
|
||||
import { runChatCompletionUpdate } from "./completion-update-flow.js";
|
||||
|
||||
/** Hiring alone admits bounded, verified lifecycle notifications. Other suites keep their count contract. */
|
||||
export function assertChatFlowRunCount(input: {
|
||||
suiteId: string; task: Parameters<typeof matchesRunCount>[0]; runs: ChatRun[];
|
||||
hiringEvidence?: HiringTemplateEvidence; hiringApiState?: unknown;
|
||||
}) {
|
||||
expect(matchesRunCount(input.task, input.runs.length), "Declared total provider-run bounds").toBe(true);
|
||||
if (input.suiteId !== "hiring-templates") return;
|
||||
const accounting = gradeHiringTemplateTurns({
|
||||
evidence: input.hiringEvidence ? { ...input.hiringEvidence, runs: input.runs } : undefined,
|
||||
apiState: input.hiringApiState,
|
||||
});
|
||||
expect(accounting.passed, `Hiring lifecycle: ${accounting.predicates.filter(p => !p.passed).map(p => p.id).join(", ")}`).toBe(true);
|
||||
return accounting;
|
||||
}
|
||||
|
||||
// Public API observations only: this driver never fabricates provider results or writes DB state.
|
||||
export interface ChatIssue {
|
||||
id: string;
|
||||
@@ -419,6 +436,8 @@ export async function runChatFlow(input: ChatFlowInput) {
|
||||
await idle(count);
|
||||
};
|
||||
const noTasks = async () => expect(await tasks()).toHaveLength(0);
|
||||
let hiringEvidence: HiringTemplateEvidence | undefined;
|
||||
let refreshHiringEvidence: (() => Promise<HiringTemplateEvidence>) | undefined;
|
||||
try {
|
||||
if (caseId === "enable-disable-resume") await enableChatThroughSettings(input);
|
||||
else await api.patch("/api/instance/settings/experimental", {
|
||||
@@ -441,7 +460,9 @@ export async function runChatFlow(input: ChatFlowInput) {
|
||||
await runChatCompletionUpdate({ input, marker, allRuns, issue: () => issue!,
|
||||
refreshIssue: async () => { issue = await api.get<ChatIssue>(chatPath); if (issue) input.observe(issue, await allRuns()); } });
|
||||
} else if (execution.suite.id === "hiring-templates") {
|
||||
await runHiringTemplateFlow({ input, issue: () => issue!, turn, tasks, allRuns });
|
||||
const hiring = await runHiringTemplateFlow({ input, issue: () => issue!, turn, tasks, allRuns });
|
||||
hiringEvidence = hiring.evidence;
|
||||
refreshHiringEvidence = hiring.refresh;
|
||||
} else if (execution.suite.id === "agent-chat-qualification") {
|
||||
const context = { input, marker, issue: () => issue!, idle, allRuns, comments, expectedStops,
|
||||
refreshIssue: async () => { issue = await api.get<ChatIssue>(chatPath); input.observe(issue, await allRuns()); } };
|
||||
@@ -982,7 +1003,17 @@ export async function runChatFlow(input: ChatFlowInput) {
|
||||
});
|
||||
}
|
||||
await idle(minimumRunCount(execution.task));
|
||||
expect(matchesRunCount(execution.task, runs.filter((run) => !isResetRun(run)).length)).toBe(true);
|
||||
if (execution.suite.id === "hiring-templates") {
|
||||
// Refresh the whole consistent evidence generation for the final guard.
|
||||
hiringEvidence = await refreshHiringEvidence!();
|
||||
runs = hiringEvidence.runs;
|
||||
}
|
||||
assertChatFlowRunCount({ suiteId: execution.suite.id, task: execution.task,
|
||||
// Hiring must account for every actual company run, including unexpected resets.
|
||||
runs: execution.suite.id === "hiring-templates" ? runs : runs.filter((run) => !isResetRun(run)),
|
||||
hiringEvidence, hiringApiState: execution.suite.id === "hiring-templates"
|
||||
? hiringEvidence!.turnApiState : undefined,
|
||||
});
|
||||
for (const run of runs.filter((run) => !isResetRun(run))) {
|
||||
expect(run.runtimeMode).toBe(execution.profile.expectedRuntimeMode);
|
||||
expect(run.status).toBe(
|
||||
|
||||
@@ -2,7 +2,7 @@ import { createHash } from "node:crypto";
|
||||
import { readFileSync } from "node:fs";
|
||||
import type { RunnerProfileFixture, RunnerTaskFixture } from "./types.js";
|
||||
|
||||
export const HIRING_TEMPLATE_GRADER_VERSION = "paperclip.hiring-templates.v1";
|
||||
export const HIRING_TEMPLATE_GRADER_VERSION = "paperclip.hiring-templates.v3";
|
||||
export const HIRING_TEMPLATE_SKILL_KEY = "paperclipai/paperclip/paperclip-create-agent";
|
||||
export const HIRING_TEMPLATE_READ_FILES = [
|
||||
"SKILL.md",
|
||||
@@ -17,7 +17,8 @@ export const HIRING_TEMPLATE_SOURCE_FILES = [
|
||||
"skills/paperclip-create-agent/references/baseline-role-guide.md",
|
||||
] as const;
|
||||
export const hiringTemplateDefinitionDigest = createHash("sha256").update(
|
||||
["cases", "scoring", "flow"].map(part => readFileSync(new URL(`./hiring-template-${part}.ts`, import.meta.url), "utf8")).join("\n"),
|
||||
["hiring-template-cases.ts", "hiring-template-scoring.ts", "hiring-template-flow.ts", "hiring-template-turn-accounting.ts", "chat-flow.ts"]
|
||||
.map(file => readFileSync(new URL(`./${file}`, import.meta.url), "utf8")).join("\n"),
|
||||
).digest("hex");
|
||||
|
||||
export function hiringTemplateProfile(profile: RunnerProfileFixture): RunnerProfileFixture {
|
||||
@@ -49,7 +50,7 @@ export function hiringTemplateScenario(nonce: string, projectName: string) {
|
||||
|
||||
export const hiringTemplateTasks: readonly RunnerTaskFixture[] = [{
|
||||
id: "hire-coder-template-reuse", label: "Hire with production templates and reuse the coder",
|
||||
groups: ["chat"], flow: "agent_chat", workMode: "standard", expectedRunCount: 5,
|
||||
groups: ["chat"], flow: "agent_chat", workMode: "standard", expectedRunCount: 7, minimumExpectedRunCount: 5,
|
||||
attemptTimeoutMs: { local: 15 * 60_000, daytona: 15 * 60_000 },
|
||||
expectedTerminalState: { issue: "in_review", run: "succeeded" },
|
||||
buildTitle: nonce => `Production hiring templates ${nonce}`,
|
||||
|
||||
@@ -6,7 +6,7 @@ import { HIRING_TEMPLATE_GRADER_VERSION, HIRING_TEMPLATE_READ_FILES, HIRING_TEMP
|
||||
HIRING_TEMPLATE_SOURCE_FILES, hiringTemplateDefinitionDigest, hiringTemplateScenario } from "./hiring-template-cases.js";
|
||||
import { gradeHiringTemplate, hiringTemplateHash, hiringTemplateSize, type HiringAgent, type HiringDocument,
|
||||
type HiringInstructionSnapshot, type HiringTemplateEvidence } from "./hiring-template-scoring.js";
|
||||
import type { RunnerApi } from "./api.js";
|
||||
import { pollUntil, type RunnerApi } from "./api.js";
|
||||
|
||||
export async function readHiringInstructions(api: Pick<RunnerApi, "get">, agentId: string): Promise<HiringInstructionSnapshot> {
|
||||
const bundle = await api.get<{ mode: string | null; entryFile: string; files: Array<{ path: string; binary?: boolean }> }>(`/api/agents/${agentId}/instructions-bundle`);
|
||||
@@ -62,6 +62,61 @@ export function renderHiringCoderExample(reference: string, agentName: string, c
|
||||
return example.replaceAll("{{agentName}}", agentName).replaceAll("{{companyName}}", companyName).replaceAll("{{managerTitle}}", managerTitle).replaceAll("{{issuePrefix}}", issuePrefix);
|
||||
}
|
||||
|
||||
/** Reads only existing public surfaces; this does not fabricate lifecycle evidence. */
|
||||
export async function readHiringTurnApiState(api: Pick<RunnerApi, "get">, chatIssueId: string, allRuns: () => Promise<ChatRun[]>) {
|
||||
const [issue, comments, runs, wakes] = await Promise.all([
|
||||
api.get<ChatIssue>(`/api/issues/${chatIssueId}`),
|
||||
api.get<unknown[]>(`/api/issues/${chatIssueId}/comments`),
|
||||
allRuns(),
|
||||
api.get<{ events: Array<{ kind: string; status?: string; finishedAt?: string | null; runId?: string | null }>; truncated: boolean }>(`/api/issues/${chatIssueId}/diagnostics/wakes`),
|
||||
]);
|
||||
return { issue, comments, runs, wakes };
|
||||
}
|
||||
|
||||
type HiringTurnApiState = Awaited<ReturnType<typeof readHiringTurnApiState>>;
|
||||
interface HiringObservation {
|
||||
agents: HiringAgent[]; tasks: ChatIssue[]; runs: ChatRun[]; readRuns: HiringTemplateEvidence["readRuns"];
|
||||
apiState?: HiringTurnApiState; finalApiState?: HiringTurnApiState;
|
||||
}
|
||||
/** Retry whole observations, never splice different ledger generations together. */
|
||||
export async function waitForSettledHiringObservation(
|
||||
load: () => Promise<HiringObservation>, options: { deadlineAt?: number; intervalMs?: number } = {},
|
||||
) {
|
||||
let previous: string | undefined;
|
||||
return pollUntil({ label: "hiring work and completion deliveries settle consistently",
|
||||
deadlineAt: options.deadlineAt ?? Date.now() + 90_000, intervalMs: options.intervalMs ?? 1_000,
|
||||
load, accept: observation => {
|
||||
const before = observation.apiState, after = observation.finalApiState;
|
||||
const stable = before && after && JSON.stringify(before) === JSON.stringify(after)
|
||||
&& JSON.stringify(observation.runs) === JSON.stringify(after.runs);
|
||||
const completionRuns = observation.runs.filter(run => run.contextSnapshot?.wakeReason === "chat_task_completed");
|
||||
const completionsObserved = observation.tasks.filter(task => task.status === "done").every(task => {
|
||||
const callbacks = completionRuns.filter(run => Array.isArray(run.contextSnapshot?.chatCompletionUpdates)
|
||||
&& run.contextSnapshot.chatCompletionUpdates.filter((update: { id?: string }) => update.id === task.id).length === 1);
|
||||
return callbacks.length === 1 && after?.comments.some(value => {
|
||||
const comment = value as Record<string, unknown>;
|
||||
return comment.createdByRunId === callbacks[0]!.id && comment.issueId === after.issue.id
|
||||
&& comment.authorAgentId === callbacks[0]!.agentId && !comment.authorUserId
|
||||
&& typeof comment.body === "string" && comment.body.trim().length > 0;
|
||||
});
|
||||
});
|
||||
const pending = !after || after.wakes.truncated || after.wakes.events.some(wake => {
|
||||
if (wake.kind !== "wake_request") return false;
|
||||
if (wake.status === "coalesced") return !completionsObserved || !observation.runs.some(run => run.id === wake.runId
|
||||
&& ["succeeded", "failed", "cancelled"].includes(run.status));
|
||||
return !["completed", "skipped", "failed", "cancelled"].includes(wake.status ?? "");
|
||||
}) || observation.runs.some(run => ["queued", "running", "scheduled_retry"].includes(run.status));
|
||||
// Await both known callbacks, even before the pending outbox enqueues a
|
||||
// wake. Reply attribution proves the server acknowledged each delivery.
|
||||
const signature = stable && !pending && completionsObserved && after.issue.conversationState === "waiting"
|
||||
? JSON.stringify(observation) : undefined;
|
||||
const settled = Boolean(signature && signature === previous);
|
||||
previous = signature;
|
||||
return settled;
|
||||
},
|
||||
});
|
||||
}
|
||||
|
||||
export async function runHiringTemplateFlow(context: {
|
||||
input: ChatFlowInput; issue(): ChatIssue; turn(message: string, count: number): Promise<void>;
|
||||
tasks(): Promise<ChatIssue[]>; allRuns(): Promise<ChatRun[]>;
|
||||
@@ -80,16 +135,27 @@ export async function runHiringTemplateFlow(context: {
|
||||
};
|
||||
let source: Awaited<ReturnType<typeof readHiringTemplateSources>> | undefined;
|
||||
let scenario: ReturnType<typeof hiringTemplateScenario> | undefined;
|
||||
async function refresh() {
|
||||
async function loadObservation() {
|
||||
const observed = await Promise.all([api.get<HiringAgent[]>(`${company}/agents`), tasks(), allRuns()]);
|
||||
[evidence.agents, evidence.tasks, evidence.runs] = observed;
|
||||
evidence.readRuns = await Promise.all(evidence.runs.map(async run => {
|
||||
const [agents, taskRows, runs] = observed;
|
||||
const apiState = evidence.chatIssueId ? await readHiringTurnApiState(api, evidence.chatIssueId, allRuns) : undefined;
|
||||
const readRuns = await Promise.all(runs.map(async run => {
|
||||
const events = await collectRunEvents<{ seq?: number; eventType?: string; payload?: unknown; createdAt?: string }>(
|
||||
(afterSeq, limit) => api.get(`/api/heartbeat-runs/${run.id}/events?afterSeq=${afterSeq}&limit=${limit}`),
|
||||
);
|
||||
return { runId: run.id, agentId: run.agentId, events };
|
||||
}));
|
||||
// Bracket event collection with a fresh ledger/wake/comment observation.
|
||||
const finalApiState = evidence.chatIssueId ? await readHiringTurnApiState(api, evidence.chatIssueId, allRuns) : undefined;
|
||||
return { agents, tasks: taskRows, runs, readRuns, apiState, finalApiState };
|
||||
}
|
||||
async function refresh(settle = false) {
|
||||
const observed = settle && evidence.chatIssueId
|
||||
? await waitForSettledHiringObservation(loadObservation) : await loadObservation();
|
||||
Object.assign(evidence, { agents: observed.agents, tasks: observed.tasks, runs: observed.runs,
|
||||
readRuns: observed.readRuns, turnApiState: observed.finalApiState });
|
||||
}
|
||||
|
||||
try {
|
||||
await api.patch(`${company}/budgets`, { budgetMonthlyCents: 1_000 });
|
||||
await api.patch(`/api/agents/${f.agent.id}/budgets`, { budgetMonthlyCents: 1_000 });
|
||||
@@ -125,7 +191,7 @@ export async function runHiringTemplateFlow(context: {
|
||||
await turn(scenario.statusPrompt(initial[0]!.identifier ?? initial[0]!.id, second.identifier ?? second.id), 5);
|
||||
evidence.hiredInstructionsAfterReuse = await readHiringInstructions(api, hired.id);
|
||||
evidence.hiredSkillsAfterReuse = await readHiringSkillSelections(api, f.company.id, hired.id);
|
||||
await refresh();
|
||||
await refresh(true);
|
||||
const result = gradeHiringTemplate(evidence);
|
||||
for (const check of result.checks) input.check?.(`hiringTemplates.${check.dimension}.${check.id}`, check.passed, check.detail);
|
||||
await input.evidence("hiring-template.json", { schema: HIRING_TEMPLATE_GRADER_VERSION, definitionDigest: hiringTemplateDefinitionDigest,
|
||||
@@ -133,9 +199,10 @@ export async function runHiringTemplateFlow(context: {
|
||||
const failed = result.checks.filter(check => !check.passed);
|
||||
if (!result.outcomePassed) throw new Error(`Hiring-template workflow outcome failed: ${failed.filter(check => check.dimension === "outcome").map(check => check.id).join(", ")}`);
|
||||
if (result.comparisonStatus === "uncomparable") throw new Error(`Hiring-template source coverage is uncomparable: ${failed.filter(check => check.dimension === "coverage").map(check => check.id).join(", ")}`);
|
||||
return { evidence, refresh: async () => { await refresh(true); return evidence; } };
|
||||
} finally {
|
||||
let observationError: string | undefined;
|
||||
try { await refresh(); } catch (error) { observationError = String(error); }
|
||||
try { await refresh(Boolean(scenario)); } catch (error) { observationError = String(error); }
|
||||
await input.evidence("hiring-template.json", { schema: HIRING_TEMPLATE_GRADER_VERSION, definitionDigest: hiringTemplateDefinitionDigest,
|
||||
budgetGuard: { companyMonthlyCents: 1_000, leadMonthlyCents: 1_000 },
|
||||
scenario, source, evidence, result: gradeHiringTemplate(evidence), observationError });
|
||||
|
||||
@@ -1,6 +1,7 @@
|
||||
import { createHash } from "node:crypto";
|
||||
import { HIRING_TEMPLATE_READ_FILES, HIRING_TEMPLATE_SKILL_KEY } from "./hiring-template-cases.js";
|
||||
import type { ChatIssue, ChatRun } from "./chat-flow.js";
|
||||
import { gradeHiringTemplateTurns } from "./hiring-template-turn-accounting.js";
|
||||
|
||||
const record = (value: unknown): Record<string, any> =>
|
||||
value && typeof value === "object" && !Array.isArray(value) ? value as Record<string, any> : {};
|
||||
@@ -22,6 +23,8 @@ export interface HiringTemplateEvidence {
|
||||
connectionId: string; binding: unknown; tasks: ChatIssue[]; runs: ChatRun[];
|
||||
first?: HiringDocument; firstAfterReuse?: HiringDocument; second?: HiringDocument;
|
||||
readRuns: HiringReadRun[];
|
||||
/** Independent public chat/comment/run observations for strict lifecycle accounting. */
|
||||
turnApiState?: unknown;
|
||||
}
|
||||
|
||||
function expectedFixture(inputs: readonly string[], marker: string, separator: "-" | "_") {
|
||||
@@ -130,13 +133,9 @@ export function gradeHiringTemplate(e: HiringTemplateEvidence) {
|
||||
check("two-worker-tasks", "outcome", ids.every(Boolean) && new Set(ids).size === 2 && e.tasks.length === 2
|
||||
&& ids.every(id => e.tasks.some(t => t.id === id && t.status === "done" && t.assigneeAgentId === hired?.id
|
||||
&& t.projectId === e.projectId && !t.parentId)), "Both independent tasks are completed by the same coder in the chosen project.");
|
||||
check("five-successful-turns", "outcome", e.runs.length === 5 && e.runs.every(r => r.status === "succeeded" && r.runtimeMode === "native")
|
||||
&& e.runs.filter(r => r.agentId === e.leadId && r.contextSnapshot?.issueId === e.chatIssueId).length === 3
|
||||
&& ids.every(id => {
|
||||
const runs = e.runs.filter(r => r.contextSnapshot?.issueId === id);
|
||||
return runs.length === 1 && runs[0]?.agentId === hired?.id
|
||||
&& record(runs[0]?.contextSnapshot?.aiConnection).connectionId === e.connectionId;
|
||||
}), "Five turns include one actual coder execution for each task, attributed to the managed account.");
|
||||
const turnAccounting = gradeHiringTemplateTurns({ evidence: e, apiState: e.turnApiState });
|
||||
check("bounded-work-and-completion-turns", "outcome", turnAccounting.passed,
|
||||
"Exactly three requested lead turns and two coder executions, plus at most two strictly attributed completion notifications; every actual run remains counted.");
|
||||
check("initial-json-artifact", "outcome", documentMatches(e.first, expectedFixture(e.inputs, e.marker, "-"))
|
||||
&& e.first?.createdByAgentId === hired?.id, "Independent computation checks each original input/value and worker authorship.");
|
||||
check("reused-json-artifact", "outcome", documentMatches(e.second, expectedFixture(e.inputs, `REUSE${e.marker}`, "_"))
|
||||
@@ -161,8 +160,10 @@ export function gradeHiringTemplate(e: HiringTemplateEvidence) {
|
||||
&& e.hiredInstructions.entryFile === "AGENTS.md" && instruction.trim() === expectedCoder, "Saved hired instructions use the source revision's coder example with its company/name placeholders filled.");
|
||||
check("hired-instructions-durable", "coverage", Boolean(e.hiredInstructions) && sameJson(e.hiredInstructions, e.hiredInstructionsAfterReuse), "The same saved instruction bundle survives the reused worker execution.");
|
||||
check("hired-skills-durable", "coverage", Boolean(e.hiredSkills) && sameJson(e.hiredSkills, e.hiredSkillsAfterReuse), "The saved skill selections survive the reused worker execution.");
|
||||
check("completion-action-attribution", "coverage", turnAccounting.actionEvidence.status !== "uncomparable",
|
||||
"Notification actions require exact canonical/native identities; missing cross-namespace mapping is uncomparable, not proof of extra work.");
|
||||
return {
|
||||
checks, readReceipts: receipts,
|
||||
checks, readReceipts: receipts, turnAccounting,
|
||||
outcomePassed: checks.filter(c => c.dimension === "outcome").every(c => c.passed),
|
||||
comparisonStatus: checks.filter(c => c.dimension === "coverage").every(c => c.passed) ? "comparable" : "uncomparable",
|
||||
instructionSizes: { ceo: Object.fromEntries(Object.entries(e.leadInstructions?.files ?? {}).map(([file, content]) => [file, hiringTemplateSize(content)])), coder: hiringTemplateSize(instruction) },
|
||||
|
||||
@@ -0,0 +1,274 @@
|
||||
import assert from "node:assert/strict";
|
||||
import { it } from "vitest";
|
||||
import { gradeHiringTemplateTurns, HIRING_TEMPLATE_TURN_ACCOUNTING_VERSION,
|
||||
type HiringTemplateTurnPredicateId } from "./hiring-template-turn-accounting.js";
|
||||
import { createHiringTemplateTurnFixture as fixture, hiringTurnFixtureTimestamp as timestamp } from "./hiring-template-turn-fixture.js";
|
||||
|
||||
type Fixture = ReturnType<typeof fixture>;
|
||||
type FixtureContext = Fixture["evidence"]["runs"][number]["contextSnapshot"];
|
||||
const e = (f: Fixture) => f.evidence;
|
||||
const notification = (f: Fixture) => e(f).runs.find(r => r.id === "notify-1")!;
|
||||
const worker = (f: Fixture) => e(f).runs.find(r => r.id === "worker-first")!;
|
||||
const request = (f: Fixture) => e(f).runs.find(r => r.id === "lead-first")!;
|
||||
const sync = (f: Fixture) => { f.apiState.runs = structuredClone(e(f).runs); return f; };
|
||||
const fails = (f: { evidence: unknown; apiState: unknown }, predicate?: HiringTemplateTurnPredicateId) => {
|
||||
const result = gradeHiringTemplateTurns(f);
|
||||
assert.equal(result.passed, false);
|
||||
assert.equal(result.predicates.length, 12);
|
||||
if (predicate) assert.equal(result.predicates.find(p => p.id === predicate)?.passed, false, predicate);
|
||||
};
|
||||
|
||||
// Public positive and negative lifecycle cases ported from the published v1
|
||||
// retained-evidence sidecar. Its original-grade/provenance wrapper stays separate.
|
||||
for (const [count, label] of [[0, "five exact work turns"], [1, "six runs with one batched completion"], [2, "seven runs with distinct completions"]] as const) {
|
||||
it(`accepts ${label} while counting every run toward costs`, () => {
|
||||
const result = gradeHiringTemplateTurns(fixture(count));
|
||||
assert.equal(result.passed, true);
|
||||
assert.equal(result.version, HIRING_TEMPLATE_TURN_ACCOUNTING_VERSION);
|
||||
assert.equal(result.predicates.length, 12);
|
||||
assert.ok(result.predicates.every(predicate => predicate.passed));
|
||||
assert.deepEqual(result.counts, { requiredWorkTurns: 5, maximumCompletionTurns: 2, maximumTotalTurns: 7,
|
||||
requestedLeadTurns: 3, coderTurns: 2, completionTurns: count, unclassifiedTurns: 0,
|
||||
snapshotRunCount: 5 + count, actualRunCount: 5 + count, costAccountingRunCount: 5 + count });
|
||||
assert.deepEqual(Object.keys(result).sort(), ["actionEvidence", "counts", "passed", "predicates", "version"]);
|
||||
});
|
||||
}
|
||||
|
||||
const negatives: Array<[string, (f: Fixture) => void, HiringTemplateTurnPredicateId]> = [
|
||||
["an eighth run", f => { e(f).runs.push({ ...structuredClone(notification(f)), id: "extra" }); }, "exact-five-required-work-turns"],
|
||||
["a duplicate run ID", f => { e(f).runs.push(structuredClone(notification(f))); }, "complete-public-run-ledger"],
|
||||
["an arbitrary chat wake", f => { notification(f).contextSnapshot.wakeReason = "heartbeat_timer"; }, "exact-five-required-work-turns"],
|
||||
["a completion label on a coder run", f => { notification(f).agentId = "coder"; }, "exact-five-required-work-turns"],
|
||||
["an extra requested lead turn", f => { e(f).runs.push({ ...structuredClone(request(f)), id: "lead-extra" }); }, "exact-five-required-work-turns"],
|
||||
["an extra coder execution", f => { e(f).runs.push({ ...structuredClone(worker(f)), id: "worker-extra" }); }, "exact-five-required-work-turns"],
|
||||
["a missing requested lead turn", f => { e(f).runs = e(f).runs.filter(r => r.id !== "lead-status"); }, "exact-five-required-work-turns"],
|
||||
["a missing coder execution", f => { e(f).runs = e(f).runs.filter(r => r.id !== "worker-first"); }, "exact-five-required-work-turns"],
|
||||
["a missing production lead identity", f => { e(f).agents[0].id = "other"; }, "known-fixture-context"],
|
||||
["a failed notification", f => { notification(f).status = "failed"; }, "successful-native-without-retries"],
|
||||
["a legacy notification runtime", f => { notification(f).runtimeMode = "legacy"; }, "successful-native-without-retries"],
|
||||
["a retried notification", f => { notification(f).retryOfRunId = "prior"; }, "successful-native-without-retries"],
|
||||
["hidden process retries", f => { notification(f).processLossRetryCount = 1; }, "successful-native-without-retries"],
|
||||
["a scheduled retry", f => { notification(f).scheduledRetryAttempt = 1; }, "successful-native-without-retries"],
|
||||
["a continuation run", f => { notification(f).continuationAttempt = 1; }, "successful-native-without-retries"],
|
||||
["a run in another company", f => { notification(f).companyId = "other"; }, "resolved-company-account-and-identity"],
|
||||
["a wrong managed account", f => { notification(f).contextSnapshot.aiConnection.connectionId = "other"; }, "resolved-company-account-and-identity"],
|
||||
["a wrong provider", f => { worker(f).contextSnapshot.aiConnection.provider = "other"; }, "resolved-company-account-and-identity"],
|
||||
["a wrong auth method", f => { worker(f).contextSnapshot.aiConnection.method = "host"; }, "resolved-company-account-and-identity"],
|
||||
["a wrong responsible-user binding", f => { notification(f).contextSnapshot.aiConnection.responsibleUserId = "other"; }, "resolved-company-account-and-identity"],
|
||||
["a wrong run responsible user", f => { request(f).responsibleUserId = "other"; }, "resolved-company-account-and-identity"],
|
||||
["an accepted foreign identity", f => { notification(f).identityHistory[0].responsibleUserId = "other"; }, "resolved-company-account-and-identity"],
|
||||
["an identity receipt for another run", f => { notification(f).identityHistory[0].runId = "other"; }, "resolved-company-account-and-identity"],
|
||||
["a missing instruction receipt", f => { request(f).identityHistory.pop(); }, "three-distinct-requested-chat-turns"],
|
||||
["a rejected instruction receipt", f => { request(f).identityHistory[1].status = "rejected"; }, "three-distinct-requested-chat-turns"],
|
||||
["a wrong instruction comment", f => { request(f).identityHistory[1].messageId = "other"; }, "three-distinct-requested-chat-turns"],
|
||||
["a forged request author", f => { f.apiState.comments[0].authorUserId = "other"; }, "three-distinct-requested-chat-turns"],
|
||||
["an additional user request", f => { f.apiState.comments.push({ ...f.apiState.comments[0], id: "request-extra" }); }, "three-distinct-requested-chat-turns"],
|
||||
["two runs for one request comment", f => { e(f).runs.find(r => r.id === "lead-reuse")!.contextSnapshot.wakeCommentId = "request-1"; }, "three-distinct-requested-chat-turns"],
|
||||
["a different requested chat generation", f => { request(f).contextSnapshot.conversationSessionGeneration = 1; }, "three-distinct-requested-chat-turns"],
|
||||
["a worker assigned to another agent", f => { e(f).tasks[0].assigneeAgentId = "other"; }, "one-coder-execution-per-known-task"],
|
||||
["a task outside the selected project", f => { e(f).tasks[0].projectId = "other"; }, "one-coder-execution-per-known-task"],
|
||||
["a notification-created task", f => { e(f).tasks[0].originRunId = notification(f).id; }, "no-notification-created-extra-tasks"],
|
||||
["a task created by the status-only turn", f => { e(f).tasks[0].originRunId = "lead-status"; }, "task-origins-are-first-two-requested-turns"],
|
||||
["an extra task without an execution", f => { e(f).tasks.push({ ...e(f).tasks[0], id: "extra-task" }); }, "no-notification-created-extra-tasks"],
|
||||
["a wake without delivery IDs", f => { delete notification(f).contextSnapshot.chatCompletionDeliveryIds; }, "bounded-server-completion-receipts"],
|
||||
["a wake without task updates", f => { delete notification(f).contextSnapshot.chatCompletionUpdates; }, "bounded-server-completion-receipts"],
|
||||
["mismatched receipt cardinality", f => { notification(f).contextSnapshot.chatCompletionDeliveryIds!.push("extra-delivery"); }, "bounded-server-completion-receipts"],
|
||||
["a duplicate delivery across runs", f => { e(f).runs.find(r => r.id === "notify-2")!.contextSnapshot.chatCompletionDeliveryIds = notification(f).contextSnapshot.chatCompletionDeliveryIds; }, "bounded-server-completion-receipts"],
|
||||
["a duplicate task notification with a fresh delivery ID", f => { e(f).runs.find(r => r.id === "notify-2")!.contextSnapshot.chatCompletionUpdates = structuredClone(notification(f).contextSnapshot.chatCompletionUpdates); }, "bounded-server-completion-receipts"],
|
||||
["an unknown notified task", f => { notification(f).contextSnapshot.chatCompletionUpdates![0].id = "unknown"; }, "bounded-server-completion-receipts"],
|
||||
["a pre-completion wake", f => { notification(f).startedAt = timestamp(10); }, "bounded-server-completion-receipts"],
|
||||
["a false completedAt receipt", f => { notification(f).contextSnapshot.chatCompletionUpdates![0].completedAt = timestamp(10); }, "bounded-server-completion-receipts"],
|
||||
["a false task identifier", f => { notification(f).contextSnapshot.chatCompletionUpdates![0].identifier = "OTHER-1"; }, "bounded-server-completion-receipts"],
|
||||
["a wrong task link", f => { notification(f).contextSnapshot.chatCompletionUpdates![0].url = "/issues/OTHER-1"; }, "bounded-server-completion-receipts"],
|
||||
["a false task status", f => { notification(f).contextSnapshot.chatCompletionUpdates![0].status = "in_progress"; }, "bounded-server-completion-receipts"],
|
||||
["a task without a saved result", f => { notification(f).contextSnapshot.chatCompletionUpdates![0].hasSavedDocuments = false; }, "bounded-server-completion-receipts"],
|
||||
["a stale completion generation", f => { notification(f).contextSnapshot.conversationSessionGeneration = 1; }, "bounded-server-completion-receipts"],
|
||||
["a manually invoked completion label", f => { notification(f).invocationSource = "on_demand"; }, "bounded-server-completion-receipts"],
|
||||
["a mixed request/completion wake", f => { notification(f).contextSnapshot.wakeCommentId = "request-1"; }, "bounded-server-completion-receipts"],
|
||||
["a missing completion dispatch identity", f => { notification(f).identityHistory = []; }, "bounded-server-completion-receipts"],
|
||||
["a missing completion reply", f => { f.apiState.comments = f.apiState.comments.filter(c => c.id !== "reply-1"); }, "completion-runs-have-attributed-chat-replies"],
|
||||
["a reply attributed to a different run", f => { f.apiState.comments.find(c => c.id === "reply-1")!.createdByRunId = "lead-first"; }, "completion-runs-have-attributed-chat-replies"],
|
||||
["a reply by another agent", f => { f.apiState.comments.find(c => c.id === "reply-1")!.authorAgentId = "coder"; }, "completion-runs-have-attributed-chat-replies"],
|
||||
["a reply in another chat", f => { f.apiState.comments.find(c => c.id === "reply-1")!.issueId = "other"; }, "completion-runs-have-attributed-chat-replies"],
|
||||
["a reply before task completion", f => { f.apiState.comments.find(c => c.id === "reply-1")!.createdAt = timestamp(10); }, "completion-runs-have-attributed-chat-replies"],
|
||||
];
|
||||
for (const [label, mutate, predicate] of negatives) it(`rejects ${label}`, () => {
|
||||
const f = fixture(); mutate(f); fails(sync(f), predicate);
|
||||
});
|
||||
it("rejects extra runs present only in the final public ledger and counts their cost", () => {
|
||||
const f = fixture(); f.apiState.runs.push({ ...structuredClone(notification(f)), id: "extra" });
|
||||
fails(f, "complete-public-run-ledger");
|
||||
assert.equal(gradeHiringTemplateTurns(f).counts.costAccountingRunCount, 8);
|
||||
});
|
||||
it("rejects an account discrepancy between retained ledgers", () => {
|
||||
const f = fixture(); f.apiState.runs[0]!.contextSnapshot.aiConnection.method = "host";
|
||||
fails(f, "complete-public-run-ledger");
|
||||
});
|
||||
it("does not require a prior grade, artifact outcome, source coverage or report provenance", () => {
|
||||
const f = fixture();
|
||||
assert.equal(gradeHiringTemplateTurns({ ...f, result: { outcomePassed: false, comparisonStatus: "uncomparable" } } as typeof f).passed, true);
|
||||
assert.deepEqual(Object.keys(f).sort(), ["apiState", "evidence"]);
|
||||
});
|
||||
it("does not modify observations or expose opaque run content", () => {
|
||||
const f = fixture(); notification(f).nativeSessionId = "PRIVATE_SESSION";
|
||||
notification(f).hiddenReasoning = "PRIVATE_REASONING"; notification(f).credential = "PRIVATE_CREDENTIAL";
|
||||
const before = JSON.stringify(f), result = gradeHiringTemplateTurns(f);
|
||||
assert.equal(JSON.stringify(f), before);
|
||||
assert.equal(result.version, HIRING_TEMPLATE_TURN_ACCOUNTING_VERSION);
|
||||
assert.doesNotMatch(JSON.stringify(result), /PRIVATE_|"company"|"board"|"account"|"notify-1"|"lead-first"/);
|
||||
});
|
||||
|
||||
const missingObservations = [
|
||||
"evidence.agents", "evidence.tasks", "evidence.runs", "evidence.first", "evidence.second", "evidence.binding", "evidence.connectionId",
|
||||
"apiState.issue", "apiState.runs", "apiState.comments", "apiState.issue.companyId", "apiState.issue.conversationUserId",
|
||||
"apiState.issue.conversationAgentId", "apiState.issue.conversationSessionGeneration",
|
||||
"evidence.runs.0.identityHistory", "evidence.runs.0.contextSnapshot.aiConnection", "evidence.runs.0.startedAt", "evidence.runs.0.finishedAt",
|
||||
"evidence.tasks.0.completedAt", "evidence.tasks.0.originRunId", "apiState.comments.0.createdAt", "apiState.comments.3.createdByRunId",
|
||||
];
|
||||
for (const path of missingObservations) it(`fails closed without ${path}`, () => {
|
||||
const f = fixture(), keys = path.split(".");
|
||||
let target = f as unknown as Record<string, unknown>;
|
||||
for (const key of keys.slice(0, -1)) target = target[key] as Record<string, unknown>;
|
||||
delete target[keys.at(-1)!];
|
||||
fails(f);
|
||||
});
|
||||
for (const [label, value] of [["null", null], ["undefined", undefined], ["string", "bad"], ["number", 42], ["boolean", true], ["array", []]] as const) {
|
||||
it(`rejects ${label} evidence without throwing`, () => fails({ evidence: value, apiState: fixture().apiState }));
|
||||
it(`rejects ${label} API state without throwing`, () => fails({ evidence: fixture().evidence, apiState: value }));
|
||||
}
|
||||
it("fails closed on unstringifiable accounting fields", () => {
|
||||
for (const value of [1n, Object.assign({}, { nested: {} })]) {
|
||||
const f = fixture();
|
||||
if (typeof value === "object") value.nested = value;
|
||||
notification(f).contextSnapshot = { ...notification(f).contextSnapshot, chatCompletionUpdates: value } as unknown as FixtureContext;
|
||||
fails(sync(f), "complete-public-run-ledger");
|
||||
}
|
||||
});
|
||||
it("retains the observed public cost count when evidence getters throw", () => {
|
||||
const f = fixture();
|
||||
const evidence = new Proxy(f.evidence, { get() { throw new Error("Malformed observation"); } });
|
||||
const result = gradeHiringTemplateTurns({ evidence, apiState: f.apiState });
|
||||
assert.equal(result.passed, false);
|
||||
assert.equal(result.counts.costAccountingRunCount, 7);
|
||||
});
|
||||
|
||||
it("rejects a notification write_document despite valid tasks and replies", () => {
|
||||
const f = fixture();
|
||||
for (const event of f.evidence.readRuns[0]!.events) {
|
||||
const payload = event.payload.prpEvent.payload as Record<string, any>;
|
||||
if (payload.name) payload.name = "write_document";
|
||||
if (payload.item?.name) payload.item.name = "write_document";
|
||||
}
|
||||
fails(f, "completion-turns-only-report-actions");
|
||||
assert.equal(gradeHiringTemplateTurns(f).actionEvidence.status, "violated");
|
||||
});
|
||||
it("rejects a generic call_api document PUT even though it completes successfully", () => {
|
||||
const f = fixture();
|
||||
for (const event of f.evidence.readRuns[0]!.events) {
|
||||
const item = (event.payload.prpEvent.payload as Record<string, any>).item;
|
||||
if (item?.input) item.input.operationId = "PUT /api/issues/{id}/documents/{key}";
|
||||
if (item?.result) item.result.apiOperationId = "PUT /api/issues/{id}/documents/{key}";
|
||||
}
|
||||
fails(f, "completion-turns-only-report-actions");
|
||||
assert.equal(gradeHiringTemplateTurns(f).actionEvidence.status, "violated");
|
||||
});
|
||||
for (const defect of ["missing-stream", "missing-terminal", "missing-start", "unmatched-result", "separate-namespace",
|
||||
"api-error", "different-operation", "unknown-tool", "duplicate-result", "sequence-gap"] as const) {
|
||||
it(`does not certify notification action evidence with ${defect}`, () => {
|
||||
const f = fixture(), ledger = f.evidence.readRuns[0]!, events = ledger.events;
|
||||
const start = events[1]!.payload.prpEvent.payload as Record<string, any>;
|
||||
const result = events[2]!.payload.prpEvent.payload as Record<string, any>;
|
||||
if (defect === "missing-stream") f.evidence.readRuns = [];
|
||||
if (defect === "missing-terminal") ledger.events = events.slice(0, -1);
|
||||
if (defect === "missing-start") ledger.events = events.filter(e => e.eventType !== "item.started");
|
||||
if (defect === "unmatched-result") result.item.tool_use_id = "other";
|
||||
if (defect === "separate-namespace") { start.item.id = result.item.id = result.item.tool_use_id = "host-request"; }
|
||||
if (defect === "api-error") result.item.result.ok = false;
|
||||
if (defect === "different-operation") result.item.result.apiOperationId = "GET /api/issues/{id}";
|
||||
if (defect === "unknown-tool") (events[0]!.payload.prpEvent.payload as Record<string, any>).name = "unknown";
|
||||
if (defect === "duplicate-result") { ledger.events.splice(3, 0, structuredClone(events[2]!)); ledger.events.forEach((e, i) => e.seq = i + 1); }
|
||||
if (defect === "sequence-gap") events[2]!.seq++;
|
||||
fails(f, "completion-turns-only-report-actions");
|
||||
});
|
||||
}
|
||||
it("allows successful native readonly reads without claiming source identity", () => {
|
||||
const f = fixture();
|
||||
for (const ledger of f.evidence.readRuns) ledger.events = ledger.events.filter(event => !event.eventType.startsWith("item."))
|
||||
.map((event, i) => { const payload = event.payload.prpEvent.payload as Record<string, any>;
|
||||
if (payload.name) Object.assign(payload, { name: "Read File", operation: "read", readOnly: true, transport: "builtin" });
|
||||
return { ...event, seq: i + 1 }; });
|
||||
assert.equal(gradeHiringTemplateTurns(f).passed, true);
|
||||
});
|
||||
|
||||
for (const [label, value] of [["non-2xx API result", { status: 500 }], ["explicit API error", { error: "failed" }],
|
||||
["API is_error flag", { is_error: true }], ["API isError flag", { isError: true }]] as const) {
|
||||
it(`rejects ${label}`, () => {
|
||||
const f = fixture();
|
||||
Object.assign((f.evidence.readRuns[0]!.events[2]!.payload.prpEvent.payload as Record<string, any>).item.result, value);
|
||||
fails(f, "completion-turns-only-report-actions");
|
||||
});
|
||||
}
|
||||
it("does not trust a readonly hint on an unattributed call_api", () => {
|
||||
const f = fixture(), ledger = f.evidence.readRuns[0]!;
|
||||
ledger.events = ledger.events.filter(event => !event.eventType.startsWith("item."));
|
||||
ledger.events.forEach((event, i) => {
|
||||
event.seq = i + 1;
|
||||
const payload = event.payload.prpEvent.payload as Record<string, any>;
|
||||
if (payload.name) Object.assign(payload, { operation: "read", readOnly: true });
|
||||
});
|
||||
fails(f, "completion-turns-only-report-actions");
|
||||
});
|
||||
for (const [transport, operation, readOnly] of [["dynamic", "unknown", null], ["mcp", "execute", false]] as const) {
|
||||
it(`rejects ${transport} write_document with its actual production ${operation} shape`, () => {
|
||||
const f = fixture();
|
||||
for (const event of f.evidence.readRuns[0]!.events) {
|
||||
const payload = event.payload.prpEvent.payload as Record<string, any>;
|
||||
if (payload.name) Object.assign(payload, { name: "write_document", transport, operation, readOnly });
|
||||
if (payload.item?.name) payload.item.name = "write_document";
|
||||
}
|
||||
fails(f, "completion-turns-only-report-actions");
|
||||
assert.equal(gradeHiringTemplateTurns(f).actionEvidence.status, "violated");
|
||||
});
|
||||
}
|
||||
it("rejects an attempted mutation even if it fails and a benign read follows", () => {
|
||||
const f = fixture(), ledger = f.evidence.readRuns[0]!;
|
||||
const mutation = structuredClone(ledger.events.slice(0, 4));
|
||||
for (const event of mutation) {
|
||||
const payload = event.payload.prpEvent.payload as Record<string, any>;
|
||||
if (payload.executionId) payload.executionId = "mutation";
|
||||
if (payload.item) { payload.item.id = "mutation"; if (payload.item.tool_use_id) payload.item.tool_use_id = "mutation"; }
|
||||
if (payload.item?.input) payload.item.input.operationId = "POST /api/companies/{companyId}/issues";
|
||||
if (payload.item?.result) Object.assign(payload.item.result, { apiOperationId: "POST /api/companies/{companyId}/issues", ok: false, status: 403 });
|
||||
}
|
||||
ledger.events.unshift(...mutation);
|
||||
ledger.events.forEach((event, i) => event.seq = i + 1);
|
||||
fails(f, "completion-turns-only-report-actions");
|
||||
assert.equal(gradeHiringTemplateTurns(f).actionEvidence.status, "violated");
|
||||
});
|
||||
it("requires an accepted disposition and succeeded terminal in the action stream", () => {
|
||||
for (const defect of ["accepted", "terminal"] as const) {
|
||||
const f = fixture(), events = f.evidence.readRuns[0]!.events;
|
||||
const payload = events[defect === "accepted" ? 4 : 5]!.payload.prpEvent.payload as Record<string, any>;
|
||||
if (defect === "accepted") payload.result.schema = "wrong";
|
||||
else payload.runTerminalState = "failed";
|
||||
fails(f, "completion-turns-only-report-actions");
|
||||
}
|
||||
});
|
||||
|
||||
it("admits a correlated native finish without forcing Done versus yielded chat disposition", () => {
|
||||
for (const disposition of ["done", "yielded"]) {
|
||||
const f = fixture(), ledger = f.evidence.readRuns[0]!;
|
||||
const start = structuredClone(ledger.events[0]!), finish = structuredClone(ledger.events[3]!);
|
||||
for (const event of [start, finish]) Object.assign(event.payload.prpEvent.payload,
|
||||
{ executionId: "finish", name: "paperclip_finish", transport: "dynamic" });
|
||||
const proposal = { seq: 0, eventType: "run.result.proposed", payload: { prpEvent: {
|
||||
schema: "paperclip.prp.event.v1", schemaVersion: 1, sourceKind: "runner", itemId: "finish",
|
||||
payload: { schema: "paperclip.run_result.v1", reportedWorkDisposition: disposition },
|
||||
} } };
|
||||
ledger.events.splice(4, 0, start, proposal as unknown as typeof start, finish);
|
||||
ledger.events.forEach((event, i) => event.seq = i + 1);
|
||||
assert.equal(gradeHiringTemplateTurns(f).passed, true);
|
||||
(proposal.payload.prpEvent as Record<string, unknown>).itemId = "unrelated";
|
||||
fails(f, "completion-turns-only-report-actions");
|
||||
}
|
||||
});
|
||||
@@ -0,0 +1,317 @@
|
||||
/** Strict public lifecycle accounting for the bounded hiring-template journey.
|
||||
* Three requested CEO turns and two coder executions remain exact. Up to two
|
||||
* server completion turns are separately admitted and still count toward costs.
|
||||
* Outcome artifacts and source/read coverage are graded by their own checks.
|
||||
*/
|
||||
export const HIRING_TEMPLATE_TURN_ACCOUNTING_VERSION = "paperclip.hiring-template-turn-accounting.v2";
|
||||
const predicateIds = [
|
||||
"known-fixture-context",
|
||||
"complete-public-run-ledger",
|
||||
"exact-five-required-work-turns",
|
||||
"successful-native-without-retries",
|
||||
"resolved-company-account-and-identity",
|
||||
"three-distinct-requested-chat-turns",
|
||||
"one-coder-execution-per-known-task",
|
||||
"task-origins-are-first-two-requested-turns",
|
||||
"bounded-server-completion-receipts",
|
||||
"completion-runs-have-attributed-chat-replies",
|
||||
"no-notification-created-extra-tasks",
|
||||
"completion-turns-only-report-actions",
|
||||
] as const;
|
||||
export type HiringTemplateTurnPredicateId = typeof predicateIds[number];
|
||||
export interface HiringTemplateTurnPredicate { id: HiringTemplateTurnPredicateId; passed: boolean }
|
||||
export interface HiringTemplateTurnCounts {
|
||||
requiredWorkTurns: 5; maximumCompletionTurns: 2; maximumTotalTurns: 7;
|
||||
requestedLeadTurns: number; coderTurns: number; completionTurns: number; unclassifiedTurns: number;
|
||||
snapshotRunCount: number; actualRunCount: number; costAccountingRunCount: number;
|
||||
}
|
||||
export interface HiringTemplateTurnAccountingResult {
|
||||
version: typeof HIRING_TEMPLATE_TURN_ACCOUNTING_VERSION;
|
||||
counts: HiringTemplateTurnCounts;
|
||||
predicates: HiringTemplateTurnPredicate[];
|
||||
passed: boolean;
|
||||
actionEvidence: HiringCompletionActionEvidence;
|
||||
}
|
||||
export interface HiringCompletionActionEvidence {
|
||||
status: "verified" | "violated" | "uncomparable"; notificationRuns: number;
|
||||
canonicalExecutions: number; matchedNativeApiCalls: number; unknownExecutions: number;
|
||||
}
|
||||
type JsonRecord = Record<string, unknown>;
|
||||
const object = (value: unknown): JsonRecord =>
|
||||
value && typeof value === "object" && !Array.isArray(value) ? value as JsonRecord : {};
|
||||
const rows = (value: unknown): JsonRecord[] => Array.isArray(value) ? value.map(object) : [];
|
||||
const present = (value: unknown): value is string => typeof value === "string" && value.trim().length > 0;
|
||||
const unique = (values: readonly unknown[]) => values.every(present) && new Set(values).size === values.length;
|
||||
const sameSet = (a: readonly unknown[], b: readonly unknown[]) =>
|
||||
a.length === b.length && unique(a) && unique(b) && a.every(v => b.includes(v));
|
||||
const date = (value: unknown) => typeof value === "string" ? Date.parse(value) : NaN;
|
||||
const zero = (value: unknown) => value == null || value === 0;
|
||||
const noList = (value: unknown) => value == null || Array.isArray(value) && value.length === 0;
|
||||
const context = (run: JsonRecord) => object(run.contextSnapshot);
|
||||
const accepted = (run: JsonRecord) => rows(run.identityHistory).filter(i => i.status === "accepted");
|
||||
const dispatchIdentity = (run: JsonRecord, cause: string, user: unknown) => accepted(run).some(i => i.cause === cause
|
||||
&& i.responsibleUserId === user && !i.messageId && Number.isFinite(date(i.acceptedAt)));
|
||||
|
||||
// Compare only public accounting fields; malformed projections fail closed.
|
||||
function accountingProjection(run: JsonRecord): string | null {
|
||||
try {
|
||||
return JSON.stringify({ id: run.id, companyId: run.companyId,
|
||||
agentId: run.agentId, status: run.status, runtimeMode: run.runtimeMode,
|
||||
responsibleUserId: run.responsibleUserId, invocationSource: run.invocationSource,
|
||||
triggerDetail: run.triggerDetail, startedAt: run.startedAt, finishedAt: run.finishedAt,
|
||||
retryOfRunId: run.retryOfRunId, processLossRetryCount: run.processLossRetryCount,
|
||||
scheduledRetryAttempt: run.scheduledRetryAttempt, continuationAttempt: run.continuationAttempt,
|
||||
context: Object.fromEntries(["source", "issueId", "wakeReason", "wakeSource", "wakeTriggerDetail", "wakeCommentId",
|
||||
"wakeCommentIds", "conversationSessionGeneration", "chatCompletionDeliveryIds", "chatCompletionUpdates"]
|
||||
.map(key => [key, context(run)[key]])),
|
||||
account: Object.fromEntries(["connectionId", "responsibleUserId", "provider", "method", "mode"]
|
||||
.map(key => [key, object(context(run).aiConnection)[key]])),
|
||||
identities: rows(run.identityHistory).map(i => Object.fromEntries(["runId", "companyId", "responsibleUserId",
|
||||
"status", "cause", "messageId", "acceptedAt"].map(key => [key, i[key]]))) });
|
||||
} catch { return null; }
|
||||
}
|
||||
|
||||
const notificationReadApiOperations = new Set([
|
||||
"GET /api/issues/{id}", "GET /api/issues/{id}/comments", "GET /api/issues/{id}/documents", "GET /api/issues/{id}/documents/{key}",
|
||||
]);
|
||||
const discoveryNames = new Set(["search_api", "search_tasks"]);
|
||||
const forbiddenWorkTools = new Set(["write_document", "create_task", "hire_agent", "register_deliverable", "create_skill",
|
||||
"update_agent_instructions", "restore_agent_instructions", "create_project", "reassign_task", "set_dependencies", "set_task_title"]);
|
||||
const errored = (value: JsonRecord) => Boolean(value.error || value.is_error || value.isError);
|
||||
function eventEnvelope(event: JsonRecord) { return object(object(event.payload).prpEvent); }
|
||||
/** Only exact per-run IDs join provider executions to successful native actions.
|
||||
* ACPX host request IDs and provider stream IDs cannot be joined by order/name.
|
||||
*/
|
||||
export function gradeHiringCompletionActions(notifications: JsonRecord[], readRuns: unknown): HiringCompletionActionEvidence {
|
||||
const ledgers = rows(readRuns);
|
||||
const result: HiringCompletionActionEvidence = { status: "verified", notificationRuns: notifications.length,
|
||||
canonicalExecutions: 0, matchedNativeApiCalls: 0, unknownExecutions: 0 };
|
||||
const unresolved = () => { result.unknownExecutions++; if (result.status !== "violated") result.status = "uncomparable"; };
|
||||
const violation = () => { result.status = "violated"; };
|
||||
for (const run of notifications) {
|
||||
const matching = ledgers.filter(ledger => ledger.runId === run.id && ledger.agentId === run.agentId);
|
||||
if (matching.length !== 1) { unresolved(); continue; }
|
||||
const events = rows(matching[0]!.events), terminals = events.filter(event => event.eventType === "run.terminal");
|
||||
const acceptedResults = events.filter(event => event.eventType === "run.result.accepted");
|
||||
const terminalEnvelope = eventEnvelope(terminals[0] ?? {}), terminal = object(terminalEnvelope.payload);
|
||||
const acceptedEnvelope = eventEnvelope(acceptedResults[0] ?? {});
|
||||
if (!events.length || terminals.length !== 1 || acceptedResults.length !== 1
|
||||
|| terminalEnvelope.sourceKind !== "control_plane" || terminal.schema !== "paperclip.prp.terminal.v1"
|
||||
|| terminal.runTerminalState !== "succeeded" || terminal.turnTerminalState !== "completed"
|
||||
|| acceptedEnvelope.sourceKind !== "control_plane" || object(object(acceptedEnvelope.payload).result).schema !== "paperclip.run_result.v1"
|
||||
|| !events.every((event, i) => Number.isSafeInteger(event.seq) && event.seq === i + 1)) { unresolved(); continue; }
|
||||
const starts = new Map<unknown, JsonRecord>(), completed = new Map<unknown, JsonRecord>();
|
||||
const nativeStarts = new Map<unknown, JsonRecord>(), nativeResults = new Map<unknown, JsonRecord>();
|
||||
for (const event of events) {
|
||||
const envelope = eventEnvelope(event), payload = object(envelope.payload), item = object(payload.item);
|
||||
if (event.eventType?.toString().startsWith("tool.execution.")) {
|
||||
if (payload.schema !== "paperclip.tool.execution.v1" || !present(payload.executionId)) { unresolved(); continue; }
|
||||
if (event.eventType === "tool.execution.started") {
|
||||
if (starts.has(payload.executionId)) violation();
|
||||
starts.set(payload.executionId, payload);
|
||||
} else if (event.eventType === "tool.execution.completed") {
|
||||
if (completed.has(payload.executionId) || payload.status !== "completed" || errored(payload)
|
||||
|| payload.exitCode != null && payload.exitCode !== 0) violation();
|
||||
completed.set(payload.executionId, payload);
|
||||
} else if (["tool.execution.failed", "tool.execution.cancelled", "tool.execution.interrupted"].includes(String(event.eventType))) violation();
|
||||
}
|
||||
if (item.type !== "tool_use" && item.type !== "tool_result") continue;
|
||||
if (envelope.schema !== "paperclip.prp.event.v1" || envelope.schemaVersion !== 1 || envelope.sourceKind !== "runner"
|
||||
|| !present(item.id) || errored(item)) { unresolved(); continue; }
|
||||
if (event.eventType === "item.started" && item.type === "tool_use") {
|
||||
if (nativeStarts.has(item.id)) violation();
|
||||
nativeStarts.set(item.id, item);
|
||||
} else if (event.eventType === "item.completed" && item.type === "tool_result") {
|
||||
if (item.tool_use_id !== item.id || nativeResults.has(item.id) || errored(object(item.result))) violation();
|
||||
nativeResults.set(item.id, item);
|
||||
} else unresolved();
|
||||
}
|
||||
if (!sameSet([...starts.keys()], [...completed.keys()]) || !sameSet([...nativeStarts.keys()], [...nativeResults.keys()])) unresolved();
|
||||
for (const [id, start] of nativeStarts) {
|
||||
const receipt = object(nativeResults.get(id)?.result), input = object(start.input);
|
||||
if (start.name === "call_api") {
|
||||
if (present(input.operationId) && !notificationReadApiOperations.has(input.operationId)) violation();
|
||||
if (!present(input.operationId) || input.operationId !== receipt.apiOperationId || receipt.ok !== true
|
||||
|| !Number.isInteger(receipt.status) || Number(receipt.status) < 200 || Number(receipt.status) >= 300 || errored(receipt)) { unresolved(); continue; }
|
||||
if (!notificationReadApiOperations.has(input.operationId)) violation();
|
||||
} else if (start.name === "search_api") {
|
||||
if (!Array.isArray(receipt.results) || !Number.isInteger(receipt.total) || Number(receipt.total) < 0 || typeof receipt.guidance !== "string"
|
||||
|| receipt.nextCursor != null && typeof receipt.nextCursor !== "string") unresolved();
|
||||
} else if (start.name === "search_tasks") {
|
||||
if (!Array.isArray(receipt.tasks)) unresolved();
|
||||
} else if (forbiddenWorkTools.has(String(start.name))) violation();
|
||||
else unresolved();
|
||||
}
|
||||
for (const [id, start] of starts) {
|
||||
result.canonicalExecutions++;
|
||||
const finish = completed.get(id) ?? {}, native = nativeStarts.get(id);
|
||||
if (start.name !== finish.name) { unresolved(); continue; }
|
||||
if (forbiddenWorkTools.has(String(start.name))) { violation(); continue; }
|
||||
if (start.name === "paperclip_finish") {
|
||||
const proposals = events.filter(event => event.eventType === "run.result.proposed");
|
||||
const proposal = eventEnvelope(proposals[0] ?? {});
|
||||
if (start.transport !== "dynamic" || finish.transport !== "dynamic" || proposals.length !== 1
|
||||
|| proposal.sourceKind !== "runner" || proposal.itemId !== id
|
||||
|| object(proposal.payload).schema !== "paperclip.run_result.v1") unresolved();
|
||||
continue;
|
||||
}
|
||||
if (start.name === "ToolSearch" && start.transport === "builtin" && finish.transport === "builtin") continue;
|
||||
if (start.name === "call_api" || discoveryNames.has(String(start.name))) {
|
||||
if (!native || native.name !== start.name) { unresolved(); continue; }
|
||||
if (start.name === "call_api") result.matchedNativeApiCalls++;
|
||||
} else if (start.readOnly === true && finish.readOnly === true && ["read", "search", "list"].includes(String(start.operation))
|
||||
&& start.operation === finish.operation) continue;
|
||||
else unresolved();
|
||||
}
|
||||
// Every native action must also have an exact canonical provider identity.
|
||||
if ([...nativeStarts.keys()].some(id => !starts.has(id))) unresolved();
|
||||
}
|
||||
return result;
|
||||
}
|
||||
|
||||
function evaluate(evidence: unknown, apiState: unknown): HiringTemplateTurnAccountingResult {
|
||||
const e = object(evidence), api = object(apiState);
|
||||
const chat = object(api.issue), user = chat.conversationUserId, company = chat.companyId;
|
||||
const runs = rows(e.runs), publicRuns = rows(api.runs), comments = rows(api.comments);
|
||||
const agents = rows(e.agents), tasks = rows(e.tasks), lead = e.leadId;
|
||||
const coder = agents.find(a => a.id !== lead && a.name === e.hireName);
|
||||
const taskIds = [object(e.first).issueId, object(e.second).issueId];
|
||||
const taskById = new Map(tasks.map(t => [t.id, t]));
|
||||
const requests = comments.filter(c => c.authorUserId).sort((a, b) => date(a.createdAt) - date(b.createdAt));
|
||||
const requested: JsonRecord[] = [], workers: JsonRecord[] = [], notifications: JsonRecord[] = [], unknown: JsonRecord[] = [];
|
||||
for (const run of runs) {
|
||||
const c = context(run);
|
||||
if (run.agentId === lead && c.issueId === e.chatIssueId && c.wakeReason === "issue_commented") requested.push(run);
|
||||
else if (run.agentId === coder?.id && taskIds.includes(c.issueId) && c.wakeReason === "issue_assigned") workers.push(run);
|
||||
else if (run.agentId === lead && c.issueId === e.chatIssueId && c.wakeReason === "chat_task_completed") notifications.push(run);
|
||||
else unknown.push(run);
|
||||
}
|
||||
const predicates: HiringTemplateTurnPredicate[] = [];
|
||||
const check = (id: HiringTemplateTurnPredicateId, passed: unknown) => predicates.push({ id, passed: Boolean(passed) });
|
||||
const binding = object(e.binding);
|
||||
check("known-fixture-context", present(company) && present(user) && present(lead) && present(coder?.id)
|
||||
&& chat.id === e.chatIssueId && chat.conversationAgentId === lead
|
||||
&& typeof chat.conversationSessionGeneration === "number" && Number.isInteger(chat.conversationSessionGeneration) && chat.conversationSessionGeneration >= 0
|
||||
&& agents.length === 2 && agents.some(a => a.id === lead)
|
||||
&& unique(agents.map(a => a.id)) && agents.every(a => a.companyId === company)
|
||||
&& present(e.connectionId) && present(binding.provider) && present(binding.method) && binding.mode === "responsible_user"
|
||||
&& unique(taskIds) && unique(tasks.map(t => t.id)) && sameSet(taskIds, tasks.map(t => t.id)));
|
||||
check("complete-public-run-ledger", sameSet(runs.map(r => r.id), publicRuns.map(r => r.id))
|
||||
&& runs.every(r => {
|
||||
const projection = accountingProjection(r);
|
||||
return projection !== null && projection === accountingProjection(publicRuns.find(p => p.id === r.id) ?? {});
|
||||
}));
|
||||
check("exact-five-required-work-turns", requested.length === 3 && workers.length === 2 && unknown.length === 0
|
||||
&& notifications.length <= 2 && runs.length === 5 + notifications.length && runs.length <= 7);
|
||||
check("successful-native-without-retries", runs.length > 0 && runs.every(r => r.status === "succeeded"
|
||||
&& r.runtimeMode === "native" && Number.isFinite(date(r.startedAt)) && Number.isFinite(date(r.finishedAt))
|
||||
&& date(r.finishedAt) >= date(r.startedAt) && !r.retryOfRunId && zero(r.processLossRetryCount)
|
||||
&& zero(r.scheduledRetryAttempt) && zero(r.continuationAttempt)));
|
||||
check("resolved-company-account-and-identity", runs.length > 0 && runs.every(r => {
|
||||
const account = object(context(r).aiConnection), identities = accepted(r);
|
||||
return r.companyId === company && r.responsibleUserId === user && account.responsibleUserId === user
|
||||
&& account.connectionId === e.connectionId && ["provider", "method", "mode"].every(k => account[k] === binding[k])
|
||||
&& identities.length > 0 && identities.every(i => i.responsibleUserId === user
|
||||
&& i.runId === r.id && i.companyId === company);
|
||||
}));
|
||||
check("three-distinct-requested-chat-turns", requests.length === 3 && unique(requests.map(c => c.id))
|
||||
&& sameSet(requested.map(r => context(r).wakeCommentId), requests.map(c => c.id))
|
||||
&& requests.every(c => c.companyId === company && c.issueId === chat.id && c.authorUserId === user
|
||||
&& !c.authorAgentId && present(c.body) && Number.isFinite(date(c.createdAt)))
|
||||
&& requested.every(r => {
|
||||
const c = context(r), comment = requests.find(p => p.id === c.wakeCommentId);
|
||||
return r.invocationSource === "on_demand" && r.triggerDetail === "manual" && c.source === "issue.comment"
|
||||
&& c.wakeSource === "on_demand" && c.wakeTriggerDetail === "manual"
|
||||
&& c.conversationSessionGeneration === chat.conversationSessionGeneration
|
||||
&& noList(c.chatCompletionDeliveryIds) && noList(c.chatCompletionUpdates)
|
||||
&& sameSet(Array.isArray(c.wakeCommentIds) ? c.wakeCommentIds : [], [c.wakeCommentId])
|
||||
&& comment && date(r.startedAt) >= date(comment.createdAt)
|
||||
&& dispatchIdentity(r, "issue_commented", user)
|
||||
&& accepted(r).some(i => i.cause === "instruction" && i.messageId === c.wakeCommentId
|
||||
&& i.responsibleUserId === user && Number.isFinite(date(i.acceptedAt)));
|
||||
}));
|
||||
check("one-coder-execution-per-known-task", tasks.length === 2 && taskIds.every(id => {
|
||||
const task = taskById.get(id), matching = workers.filter(r => context(r).issueId === id), run = matching[0] ?? {};
|
||||
const c = context(run ?? {});
|
||||
return task && matching.length === 1 && task.companyId === company && task.assigneeAgentId === coder?.id
|
||||
&& task.projectId === e.projectId && !task.parentId && task.status === "done" && task.responsibleUserId === user
|
||||
&& Number.isFinite(date(task.completedAt)) && date(task.completedAt) >= date(run.startedAt)
|
||||
&& run.invocationSource === "assignment" && run.triggerDetail === "system"
|
||||
&& c.source === "paperclip_runner.create_task" && c.wakeSource === "assignment" && c.wakeTriggerDetail === "system"
|
||||
&& noList(c.chatCompletionDeliveryIds) && noList(c.chatCompletionUpdates)
|
||||
&& dispatchIdentity(run, "issue_assigned", user);
|
||||
}));
|
||||
check("task-origins-are-first-two-requested-turns", taskIds.every((id, index) => {
|
||||
const task = taskById.get(id), comment = requests[index];
|
||||
const run = requested.find(r => context(r).wakeCommentId === comment?.id);
|
||||
return task && run && task.createdByAgentId === lead && !task.createdByUserId && task.originRunId === run.id
|
||||
&& date(task.createdAt) >= date(run.startedAt);
|
||||
}));
|
||||
const deliveryIds = new Set<unknown>(), notifiedTasks = new Set<unknown>();
|
||||
check("bounded-server-completion-receipts", notifications.every(r => {
|
||||
const c = context(r), ids = Array.isArray(c.chatCompletionDeliveryIds) ? c.chatCompletionDeliveryIds : [];
|
||||
const updates = rows(c.chatCompletionUpdates);
|
||||
let valid = r.invocationSource === "automation" && r.triggerDetail === "system"
|
||||
&& c.wakeSource === "automation" && c.wakeTriggerDetail === "system" && !c.source
|
||||
&& c.conversationSessionGeneration === chat.conversationSessionGeneration
|
||||
&& !c.wakeCommentId && noList(c.wakeCommentIds)
|
||||
&& ids.length > 0 && ids.length <= 2 && ids.length === updates.length && unique(ids)
|
||||
&& ids.every(id => !deliveryIds.has(id)) && dispatchIdentity(r, "chat_task_completed", user);
|
||||
for (const update of updates) {
|
||||
const task = taskById.get(update.id);
|
||||
valid = Boolean(valid && task && !notifiedTasks.has(update.id) && update.status === "done"
|
||||
&& task.status === "done" && update.identifier === task.identifier && update.completedAt === task.completedAt
|
||||
&& present(task.identifier) && update.url === `/issues/${task.identifier}` && update.hasSavedDocuments === true
|
||||
&& date(r.startedAt) >= date(task.completedAt));
|
||||
notifiedTasks.add(update.id);
|
||||
}
|
||||
ids.forEach(id => deliveryIds.add(id));
|
||||
return valid;
|
||||
}));
|
||||
check("completion-runs-have-attributed-chat-replies", notifications.every(r => {
|
||||
const updates = rows(context(r).chatCompletionUpdates);
|
||||
const completedAt = Math.max(...updates.map(u => date(u.completedAt)));
|
||||
return updates.length > 0 && comments.some(c => c.companyId === company && c.issueId === chat.id
|
||||
&& c.authorAgentId === lead && !c.authorUserId && c.createdByRunId === r.id && present(c.body)
|
||||
&& (c.conversationSessionGeneration == null || c.conversationSessionGeneration === chat.conversationSessionGeneration)
|
||||
&& date(c.createdAt) >= completedAt && date(c.createdAt) >= date(r.startedAt));
|
||||
}));
|
||||
check("no-notification-created-extra-tasks", tasks.length === 2 && tasks.every(t =>
|
||||
!notifications.some(r => r.id === t.originRunId)));
|
||||
const actionEvidence = gradeHiringCompletionActions(notifications, e.readRuns);
|
||||
check("completion-turns-only-report-actions", actionEvidence.status === "verified");
|
||||
return {
|
||||
version: HIRING_TEMPLATE_TURN_ACCOUNTING_VERSION,
|
||||
counts: { requiredWorkTurns: 5, maximumCompletionTurns: 2, maximumTotalTurns: 7,
|
||||
requestedLeadTurns: requested.length, coderTurns: workers.length, completionTurns: notifications.length,
|
||||
unclassifiedTurns: unknown.length, snapshotRunCount: runs.length, actualRunCount: publicRuns.length,
|
||||
costAccountingRunCount: publicRuns.length },
|
||||
predicates, actionEvidence, passed: predicates.every(predicate => predicate.passed),
|
||||
};
|
||||
}
|
||||
|
||||
/** Accept untrusted evidence observations without needing an existing result. */
|
||||
export function gradeHiringTemplateTurns(input: { evidence: unknown; apiState: unknown }): HiringTemplateTurnAccountingResult {
|
||||
let snapshotRunCount = 0, actualRunCount = 0;
|
||||
try {
|
||||
const runs = object(object(input).apiState).runs;
|
||||
actualRunCount = Array.isArray(runs) ? runs.length : 0;
|
||||
} catch { /* An inaccessible public ledger is a failed observation. */ }
|
||||
try {
|
||||
const runs = object(object(input).evidence).runs;
|
||||
snapshotRunCount = Array.isArray(runs) ? runs.length : 0;
|
||||
} catch { /* Keep public cost accounting when the evidence is malformed. */ }
|
||||
try {
|
||||
const value = object(input);
|
||||
return evaluate(value.evidence, value.apiState);
|
||||
} catch {
|
||||
return {
|
||||
version: HIRING_TEMPLATE_TURN_ACCOUNTING_VERSION,
|
||||
counts: { requiredWorkTurns: 5, maximumCompletionTurns: 2, maximumTotalTurns: 7,
|
||||
requestedLeadTurns: 0, coderTurns: 0, completionTurns: 0, unclassifiedTurns: snapshotRunCount,
|
||||
snapshotRunCount, actualRunCount, costAccountingRunCount: actualRunCount },
|
||||
predicates: predicateIds.map(id => ({ id, passed: false })), passed: false,
|
||||
actionEvidence: { status: "uncomparable", notificationRuns: 0, canonicalExecutions: 0, matchedNativeApiCalls: 0, unknownExecutions: 0 },
|
||||
};
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,110 @@
|
||||
// Test-only public lifecycle facts shared by turn and outcome calibrations.
|
||||
interface FixtureIdentity {
|
||||
runId: string; companyId: string; responsibleUserId: string; status: string;
|
||||
cause: string; messageId: string | null; acceptedAt: string;
|
||||
}
|
||||
interface FixtureUpdate {
|
||||
id: string; identifier: string; status: string; completedAt: string;
|
||||
url: string; hasSavedDocuments: boolean;
|
||||
}
|
||||
interface FixtureRun {
|
||||
id: string; companyId: string; agentId: string; responsibleUserId: string;
|
||||
invocationSource: string; triggerDetail: string; status: string; runtimeMode: string;
|
||||
startedAt: string; finishedAt: string; retryOfRunId: string | null;
|
||||
processLossRetryCount: number; scheduledRetryAttempt: number; continuationAttempt: number;
|
||||
identityHistory: FixtureIdentity[];
|
||||
contextSnapshot: {
|
||||
source?: string; issueId: string; wakeReason: string; wakeSource: string;
|
||||
wakeTriggerDetail: string; conversationSessionGeneration: number;
|
||||
wakeCommentId?: string; wakeCommentIds: string[];
|
||||
aiConnection: { provider: string; method: string; mode: string; connectionId: string; responsibleUserId: string };
|
||||
chatCompletionDeliveryIds?: string[]; chatCompletionUpdates?: FixtureUpdate[];
|
||||
};
|
||||
[key: string]: unknown;
|
||||
}
|
||||
interface FixtureComment {
|
||||
id: string; companyId: string; issueId: string; authorUserId: string | null;
|
||||
authorAgentId: string | null; body: string; createdAt: string;
|
||||
createdByRunId?: string; conversationSessionGeneration?: number | null;
|
||||
}
|
||||
export const hiringTurnFixtureTimestamp = (seconds: number) => new Date(Date.UTC(2026, 9, 2, 0, 0, seconds)).toISOString();
|
||||
|
||||
/** Complete canonical/native read-action stream with one exact execution ID. */
|
||||
export function hiringNotificationActionEvents() {
|
||||
const id = "read-action";
|
||||
const payload = { schema: "paperclip.tool.execution.v1", executionId: id, transport: "dynamic",
|
||||
operation: "unknown", name: "call_api", readOnly: null, status: "running", exitCode: null };
|
||||
const items: Array<{ eventType: string; payload: Record<string, unknown> }> = [
|
||||
{ eventType: "tool.execution.started", payload },
|
||||
{ eventType: "item.started", payload: { kind: "tool", item: { id, type: "tool_use", name: "call_api",
|
||||
input: { operationId: "GET /api/issues/{id}/documents", pathParams: {} } } } },
|
||||
{ eventType: "item.completed", payload: { kind: "tool", item: { id, tool_use_id: id, type: "tool_result",
|
||||
result: { apiOperationId: "GET /api/issues/{id}/documents", ok: true, status: 200 } } } },
|
||||
{ eventType: "tool.execution.completed", payload: { ...payload, status: "completed" } },
|
||||
{ eventType: "run.result.accepted", payload: { result: { schema: "paperclip.run_result.v1" } } },
|
||||
{ eventType: "run.terminal", payload: { schema: "paperclip.prp.terminal.v1", runTerminalState: "succeeded", turnTerminalState: "completed" } },
|
||||
];
|
||||
return items.map((event, i) => ({ seq: i + 1, eventType: event.eventType, payload: { prpEvent: {
|
||||
schema: "paperclip.prp.event.v1", schemaVersion: 1, sourceKind: event.eventType.startsWith("run.") ? "control_plane" : "runner", payload: event.payload,
|
||||
} } }));
|
||||
}
|
||||
|
||||
export function createHiringTemplateTurnFixture(notificationCount = 2) {
|
||||
if (!Number.isInteger(notificationCount) || notificationCount < 0 || notificationCount > 2)
|
||||
throw new RangeError("The positive lifecycle fixture accepts zero, one or two notification turns.");
|
||||
const timestamp = hiringTurnFixtureTimestamp;
|
||||
const company = "company", user = "board", lead = "lead", coder = "coder", chat = "chat";
|
||||
const binding = { provider: "anthropic", method: "api_key", mode: "responsible_user" };
|
||||
function run(id: string, agentId: string, issueId: string, reason: string, start: number, finish: number, source?: string): FixtureRun {
|
||||
const invocationSource = reason === "issue_commented" ? "on_demand"
|
||||
: reason === "issue_assigned" ? "assignment" : "automation";
|
||||
const triggerDetail = reason === "issue_commented" ? "manual" : "system";
|
||||
return { id, companyId: company, agentId, responsibleUserId: user, invocationSource, triggerDetail,
|
||||
status: "succeeded", runtimeMode: "native", startedAt: timestamp(start), finishedAt: timestamp(finish),
|
||||
retryOfRunId: null, processLossRetryCount: 0, scheduledRetryAttempt: 0, continuationAttempt: 0,
|
||||
identityHistory: [{ runId: id, companyId: company, responsibleUserId: user, status: "accepted",
|
||||
cause: reason, messageId: null, acceptedAt: timestamp(start) }],
|
||||
contextSnapshot: { source, issueId, wakeReason: reason, wakeSource: invocationSource,
|
||||
wakeTriggerDetail: triggerDetail, conversationSessionGeneration: 0, wakeCommentIds: [],
|
||||
aiConnection: { ...binding, connectionId: "account", responsibleUserId: user } } };
|
||||
}
|
||||
const requested = [run("lead-first", lead, chat, "issue_commented", 1, 10, "issue.comment"),
|
||||
run("lead-reuse", lead, chat, "issue_commented", 21, 30, "issue.comment"),
|
||||
run("lead-status", lead, chat, "issue_commented", 51, 55, "issue.comment")];
|
||||
const comments: FixtureComment[] = requested.map((r, i) => {
|
||||
const id = `request-${i + 1}`;
|
||||
r.contextSnapshot.wakeCommentId = id;
|
||||
r.contextSnapshot.wakeCommentIds = [id];
|
||||
r.identityHistory.push({ runId: r.id, companyId: company, responsibleUserId: user, status: "accepted",
|
||||
cause: "instruction", messageId: id, acceptedAt: r.startedAt });
|
||||
return { id, companyId: company, issueId: chat, authorUserId: user, authorAgentId: null,
|
||||
body: `Fixture request ${i + 1}`, createdAt: timestamp([0, 20, 50][i]!) };
|
||||
});
|
||||
const workers = [run("worker-first", coder, "first-task", "issue_assigned", 4, 12, "paperclip_runner.create_task"),
|
||||
run("worker-reuse", coder, "second-task", "issue_assigned", 25, 40, "paperclip_runner.create_task")];
|
||||
const tasks = workers.map((r, i) => ({ id: r.contextSnapshot.issueId, companyId: company, projectId: "project",
|
||||
title: r.contextSnapshot.issueId, parentId: null as string | null, status: "done", identifier: `RUN-${i + 2}`, responsibleUserId: user,
|
||||
assigneeAgentId: coder, createdByAgentId: lead, createdByUserId: null as string | null, originRunId: requested[i]!.id,
|
||||
createdAt: timestamp([2, 22][i]!), completedAt: timestamp([11, 39][i]!) }));
|
||||
const notifications: FixtureRun[] = [];
|
||||
for (let i = 0; i < notificationCount; i++) {
|
||||
const r = run(`notify-${i + 1}`, lead, chat, "chat_task_completed", notificationCount === 1 ? 41 : [13, 41][i]!,
|
||||
notificationCount === 1 ? 45 : [18, 45][i]!);
|
||||
const covered = notificationCount === 1 ? tasks : [tasks[i]!];
|
||||
r.contextSnapshot.chatCompletionDeliveryIds = covered.map(t => `delivery-${t.id}`);
|
||||
r.contextSnapshot.chatCompletionUpdates = covered.map(t => ({ id: t.id, identifier: t.identifier,
|
||||
status: "done", completedAt: t.completedAt, url: `/issues/${t.identifier}`, hasSavedDocuments: true }));
|
||||
comments.push({ id: `reply-${i + 1}`, companyId: company, issueId: chat, authorAgentId: lead,
|
||||
authorUserId: null, createdByRunId: r.id, body: "The saved fixture is complete.",
|
||||
conversationSessionGeneration: null, createdAt: r.finishedAt });
|
||||
notifications.push(r);
|
||||
}
|
||||
const runs = [...requested, ...workers, ...notifications];
|
||||
const evidence = { leadId: lead, chatIssueId: chat, hireName: "Fixture Coder", projectId: "project", connectionId: "account", binding,
|
||||
agents: [{ id: lead, companyId: company, name: "Fixture CEO" }, { id: coder, companyId: company, name: "Fixture Coder" }],
|
||||
first: { issueId: tasks[0]!.id }, second: { issueId: tasks[1]!.id }, runs, tasks,
|
||||
readRuns: notifications.map(run => ({ runId: run.id, agentId: run.agentId, events: hiringNotificationActionEvents() })) };
|
||||
const apiState = { issue: { id: chat, companyId: company, conversationUserId: user,
|
||||
conversationAgentId: lead, conversationSessionGeneration: 0 }, comments, runs: structuredClone(runs) };
|
||||
return { evidence, apiState };
|
||||
}
|
||||
@@ -5,12 +5,14 @@ import { buildRunnerE2EProcessEnvironment } from "./harness-env.js";
|
||||
import { canonicalProviderEventsFromAcpxRuntimeEvent, canonicalProviderEventsFromCodex } from "../../packages/paperclip-runner/src/provider-events.js";
|
||||
import { loadDefaultAgentInstructionsBundle } from "../../server/src/services/default-agent-instructions.js";
|
||||
import { HIRING_TEMPLATE_READ_FILES, HIRING_TEMPLATE_SKILL_KEY, hiringTemplateInputs, hiringTemplateScenario } from "./hiring-template-cases.js";
|
||||
import { readHiringInstructions, readHiringTemplateSources, renderHiringCoderExample } from "./hiring-template-flow.js";
|
||||
import { readHiringInstructions, readHiringTemplateSources, renderHiringCoderExample, waitForSettledHiringObservation } from "./hiring-template-flow.js";
|
||||
import { gradeHiringTemplate, hiringTemplateHash, hiringTemplateReadReceipts, type HiringTemplateEvidence } from "./hiring-template-scoring.js";
|
||||
import type { RunnerApi } from "./api.js";
|
||||
import { createHiringTemplateTurnFixture } from "./hiring-template-turn-fixture.js";
|
||||
import { assertChatFlowRunCount } from "./chat-flow.js";
|
||||
|
||||
const beforeHire = "2026-10-01T12:00:00.000Z", hiredAt = "2026-10-01T12:01:00.000Z";
|
||||
const binding = { provider: "openai", method: "api_key", mode: "responsible_user" };
|
||||
const binding = { provider: "anthropic", method: "api_key", mode: "responsible_user" };
|
||||
const ceo = { "AGENTS.md": "You are the CEO. Lead the company." };
|
||||
const coder = "You are Casey, a software engineer at Fixture Company. Own software implementation and maintenance.";
|
||||
// Expected values are stated independently of the implementation under test.
|
||||
@@ -22,27 +24,28 @@ function commandEvents(file: string) {
|
||||
command: `cat /workspace/.agents/skills/paperclip-create-agent/${file}`, aggregatedOutput: "source bytes",
|
||||
} }).map(event => ({ eventType: event.eventType, createdAt: beforeHire, payload: { prpEvent: event } }));
|
||||
}
|
||||
function validEvidence(): HiringTemplateEvidence {
|
||||
function validEvidence(notificationCount = 0): HiringTemplateEvidence {
|
||||
const turnFixture = createHiringTemplateTurnFixture(notificationCount);
|
||||
const sourceFiles = [...HIRING_TEMPLATE_READ_FILES.map(file => `skills/paperclip-create-agent/${file}`),
|
||||
"skills/paperclip-create-agent/references/baseline-role-guide.md", "server/src/onboarding-assets/ceo/AGENTS.md"];
|
||||
const hashes = Object.fromEntries(sourceFiles.map(file => [file, hiringTemplateHash(file)]));
|
||||
const tasks = ["first", "second"].map(id => ({ id, companyId: "company", title: id, status: "done", assigneeAgentId: "coder", projectId: "project", parentId: null }));
|
||||
const runs = ["lead-1", "lead-2", "lead-3", "first", "second"].map(id => ({ id, companyId: "company", agentId: id.startsWith("lead") ? "lead" : "coder",
|
||||
status: "succeeded", runtimeMode: "native", contextSnapshot: { issueId: id.startsWith("lead") ? "chat" : id, aiConnection: { connectionId: "account" } } }));
|
||||
const tasks = turnFixture.evidence.tasks;
|
||||
const runs = turnFixture.evidence.runs;
|
||||
const document = (issueId: string, reference: string, outputs: string[]) => ({ issueId, key: "fixture", latestRevisionId: `${issueId}-revision`, createdByAgentId: "coder",
|
||||
body: JSON.stringify({ reference, entries: hiringTemplateInputs.map((input, index) => ({ input, value: outputs[index] })) }) });
|
||||
const first = document("first", "HIREfixture", values);
|
||||
const first = document("first-task", "HIREfixture", values);
|
||||
return { leadId: "lead", chatIssueId: "chat", hireName: "Casey", marker: "HIREfixture", projectId: "project", inputs: hiringTemplateInputs,
|
||||
expectedCeoFiles: ceo, leadInstructions: { mode: "managed", entryFile: "AGENTS.md", files: ceo },
|
||||
expectedSourceHashes: hashes, servedSourceHashes: { ...hashes }, assignedSkills: [HIRING_TEMPLATE_SKILL_KEY], coderExample: coder,
|
||||
hiredInstructions: { mode: "managed", entryFile: "AGENTS.md", files: { "AGENTS.md": coder } },
|
||||
hiredInstructionsAfterReuse: { mode: "managed", entryFile: "AGENTS.md", files: { "AGENTS.md": coder } },
|
||||
hiredSkills: [], hiredSkillsAfterReuse: [],
|
||||
agents: [{ id: "lead", name: "CEO", adapterConfig: { model: "model" } },
|
||||
{ id: "coder", name: "Casey", role: "engineer", reportsTo: "lead", adapterType: "paperclip_runner", createdAt: hiredAt,
|
||||
agents: [{ ...turnFixture.evidence.agents[0], id: "lead", name: "CEO", adapterConfig: { model: "model" } },
|
||||
{ ...turnFixture.evidence.agents[1], id: "coder", name: "Casey", role: "engineer", reportsTo: "lead", adapterType: "paperclip_runner", createdAt: hiredAt,
|
||||
adapterConfig: { model: "model" }, runtimeConfig: { aiConnection: binding } }],
|
||||
connectionId: "account", binding, tasks, runs, first, firstAfterReuse: { ...first }, second: document("second", "REUSEHIREfixture", reuseValues),
|
||||
readRuns: [{ runId: "lead-1", agentId: "lead", events: HIRING_TEMPLATE_READ_FILES.flatMap(commandEvents) }] };
|
||||
connectionId: "account", binding, tasks, runs, first, firstAfterReuse: { ...first }, second: document("second-task", "REUSEHIREfixture", reuseValues),
|
||||
turnApiState: turnFixture.apiState,
|
||||
readRuns: [...turnFixture.evidence.readRuns, { runId: "lead-1", agentId: "lead", events: HIRING_TEMPLATE_READ_FILES.flatMap(commandEvents) }] };
|
||||
}
|
||||
function fails(evidence: HiringTemplateEvidence, id: string) {
|
||||
expect(gradeHiringTemplate(evidence).checks.find(check => check.id === id)?.passed, id).toBe(false);
|
||||
@@ -63,7 +66,7 @@ describe("production hiring template oracle", () => {
|
||||
fails({ ...e, first: { ...e.first!, createdByAgentId: "lead" } }, "initial-json-artifact");
|
||||
});
|
||||
|
||||
it("requires the real hired identity, account, two distinct tasks and exactly five successful turns", () => {
|
||||
it("requires the real hired identity, account, two distinct tasks and exactly five required successful work turns", () => {
|
||||
const e = validEvidence();
|
||||
fails({ ...e, agents: [...e.agents, { ...e.agents[1]!, id: "replacement" }] }, "one-coder-hire");
|
||||
for (const wrong of [{ role: "qa" }, { reportsTo: "somebody" }, { adapterType: "codex_local" }, { name: "Another coder" }]) {
|
||||
@@ -74,10 +77,10 @@ describe("production hiring template oracle", () => {
|
||||
for (const patch of [{ assigneeAgentId: "lead" }, { parentId: "chat" }, { projectId: "other" }, { status: "backlog" }]) {
|
||||
fails({ ...e, tasks: [e.tasks[0]!, { ...e.tasks[1]!, ...patch }] }, "two-worker-tasks");
|
||||
}
|
||||
fails({ ...e, runs: e.runs.slice(1) }, "five-successful-turns");
|
||||
fails({ ...e, runs: [...e.runs, { ...e.runs[0]!, id: "extra" }] }, "five-successful-turns");
|
||||
fails({ ...e, runs: e.runs.map(r => r.id === "second" ? { ...r, contextSnapshot: { issueId: "second", aiConnection: { connectionId: "other" } } } : r) }, "five-successful-turns");
|
||||
fails({ ...e, runs: e.runs.map(r => r.id === "lead-3" ? { ...r, agentId: "coder" } : r) }, "five-successful-turns");
|
||||
fails({ ...e, runs: e.runs.slice(1) }, "bounded-work-and-completion-turns");
|
||||
fails({ ...e, runs: [...e.runs, { ...e.runs[0]!, id: "extra" }] }, "bounded-work-and-completion-turns");
|
||||
fails({ ...e, runs: e.runs.map(r => r.id === "worker-reuse" ? { ...r, contextSnapshot: { issueId: "second-task", aiConnection: { connectionId: "other" } } } : r) }, "bounded-work-and-completion-turns");
|
||||
fails({ ...e, runs: e.runs.map(r => r.id === "lead-status" ? { ...r, agentId: "coder" } : r) }, "bounded-work-and-completion-turns");
|
||||
fails({ ...e, firstAfterReuse: { ...e.first!, latestRevisionId: "modified" } }, "original-preserved");
|
||||
});
|
||||
|
||||
@@ -110,7 +113,7 @@ describe("production hiring template oracle", () => {
|
||||
|
||||
it("separates successful workflow outcomes from missing or mismatched source coverage", () => {
|
||||
const e = validEvidence();
|
||||
for (const patch of [{ readRuns: [] }, { assignedSkills: [] }, { expectedSourceHashes: {} }, { servedSourceHashes: {} },
|
||||
for (const patch of [{ readRuns: e.readRuns.filter(run => run.runId !== "lead-1") }, { assignedSkills: [] }, { expectedSourceHashes: {} }, { servedSourceHashes: {} },
|
||||
{ servedSourceHashes: { ...e.servedSourceHashes, "skills/paperclip-create-agent/SKILL.md": hiringTemplateHash("other checkout") } }]) {
|
||||
expect(gradeHiringTemplate({ ...e, ...patch })).toMatchObject({ outcomePassed: true, comparisonStatus: "uncomparable" });
|
||||
}
|
||||
@@ -120,7 +123,7 @@ describe("production hiring template oracle", () => {
|
||||
});
|
||||
|
||||
it("requires a completed pre-hire lead read rather than an echoed path, failed read or listing", () => {
|
||||
const e = validEvidence(), good = e.readRuns[0]!;
|
||||
const e = validEvidence(), good = e.readRuns.find(run => run.runId === "lead-1")!;
|
||||
for (const command of ["echo /workspace/.agents/skills/paperclip-create-agent/SKILL.md", "ls /workspace/.agents/skills/paperclip-create-agent/SKILL.md",
|
||||
"cat /workspace/.agents/skills/wrong-skill/SKILL.md", "cat $SKILL/SKILL.md", "cat /workspace/.agents/skills/paperclip-create-agent/SKILL.md > /dev/null",
|
||||
...["cat --help", "cat --version", "head --help", "tail --version", "head -n 0", "tail -c 0", "sed -n ''", "sed -n '1q'", "sed -n 'q'", "sed --help", "sed -n '1,200w /tmp/other'", "cat -n"].map(prefix => `${prefix} /workspace/.agents/skills/paperclip-create-agent/SKILL.md`),
|
||||
@@ -167,13 +170,71 @@ describe("production hiring template oracle", () => {
|
||||
});
|
||||
});
|
||||
|
||||
describe("executable hiring lifecycle count guards", () => {
|
||||
const task = { expectedRunCount: 7, minimumExpectedRunCount: 5 };
|
||||
it("grades and admits five required work turns with zero, batched or distinct completion turns", () => {
|
||||
for (const count of [0, 1, 2]) {
|
||||
const evidence = validEvidence(count);
|
||||
const result = gradeHiringTemplate(evidence);
|
||||
expect(result).toMatchObject({ outcomePassed: true, comparisonStatus: "comparable", turnAccounting: {
|
||||
passed: true, counts: { requestedLeadTurns: 3, coderTurns: 2, completionTurns: count,
|
||||
actualRunCount: 5 + count, costAccountingRunCount: 5 + count, unclassifiedTurns: 0 },
|
||||
} });
|
||||
expect(assertChatFlowRunCount({ suiteId: "hiring-templates", task, runs: evidence.runs,
|
||||
hiringEvidence: evidence, hiringApiState: evidence.turnApiState })?.passed).toBe(true);
|
||||
}
|
||||
});
|
||||
it("rejects an arbitrary same-count wake in both executable paths", () => {
|
||||
const evidence = validEvidence(2);
|
||||
const notification = evidence.runs.find(run => run.contextSnapshot?.wakeReason === "chat_task_completed")!;
|
||||
notification.contextSnapshot!.wakeReason = "heartbeat_timer";
|
||||
(evidence.turnApiState as { runs: unknown[] }).runs = structuredClone(evidence.runs);
|
||||
fails(evidence, "bounded-work-and-completion-turns");
|
||||
expect(() => assertChatFlowRunCount({ suiteId: "hiring-templates", task, runs: evidence.runs,
|
||||
hiringEvidence: evidence, hiringApiState: evidence.turnApiState })).toThrow();
|
||||
});
|
||||
it("rejects unrelated notification document writes in both executable paths", () => {
|
||||
const evidence = validEvidence(2);
|
||||
const ledger = evidence.readRuns.find(run => run.runId === "notify-1")!;
|
||||
for (const event of ledger.events) {
|
||||
const payload = (event.payload as Record<string, any>).prpEvent.payload;
|
||||
if (payload.name) payload.name = "write_document";
|
||||
if (payload.item?.name) payload.item.name = "write_document";
|
||||
}
|
||||
fails(evidence, "bounded-work-and-completion-turns");
|
||||
expect(() => assertChatFlowRunCount({ suiteId: "hiring-templates", task, runs: evidence.runs,
|
||||
hiringEvidence: evidence, hiringApiState: evidence.turnApiState })).toThrow();
|
||||
expect(gradeHiringTemplate(evidence).turnAccounting.actionEvidence.status).toBe("violated");
|
||||
});
|
||||
it("requires public lifecycle observations and preserves source/template coverage failures", () => {
|
||||
const evidence = validEvidence(2);
|
||||
fails({ ...evidence, turnApiState: undefined }, "bounded-work-and-completion-turns");
|
||||
expect(() => assertChatFlowRunCount({ suiteId: "hiring-templates", task, runs: evidence.runs,
|
||||
hiringEvidence: evidence })).toThrow();
|
||||
const changedInstructions = { ...evidence.hiredInstructions!, files: { "AGENTS.md": `${coder} Changed punctuation.` } };
|
||||
const result = gradeHiringTemplate({ ...evidence, readRuns: evidence.readRuns.filter(run => run.runId !== "lead-1"), hiredInstructions: changedInstructions,
|
||||
hiredInstructionsAfterReuse: structuredClone(changedInstructions) });
|
||||
expect(result).toMatchObject({ outcomePassed: true, comparisonStatus: "uncomparable", turnAccounting: { passed: true } });
|
||||
expect(result.checks.filter(check => check.dimension === "coverage" && !check.passed).map(check => check.id))
|
||||
.toEqual(["production-source-reads", "supplied-coder-instructions"]);
|
||||
});
|
||||
it("keeps every non-hiring chat count guard unchanged", () => {
|
||||
const runs = validEvidence(2).runs;
|
||||
for (const suiteId of ["agent-chat", "agent-chat-hardening", "agent-chat-qualification", "context-integrity"]) {
|
||||
expect(() => assertChatFlowRunCount({ suiteId, task: { expectedRunCount: 2 }, runs: runs.slice(0, 2) })).not.toThrow();
|
||||
expect(() => assertChatFlowRunCount({ suiteId, task: { expectedRunCount: 2 }, runs: runs.slice(0, 3) })).toThrow();
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe("production hiring fixture wiring and source observations", () => {
|
||||
it("keeps two explicit local native cells, production permissions and five-turn scope", () => {
|
||||
it("keeps two explicit local native cells, production permissions and bounded five-work-turn scope", () => {
|
||||
const cells = runnerMatrix.filter(cell => cell.suite.id === "hiring-templates");
|
||||
expect(cells.map(cell => cell.id)).toEqual(["hiring-templates.runner-codex.local.hire-coder-template-reuse", "hiring-templates.runner-acpx-claude.local.hire-coder-template-reuse"]);
|
||||
expect(runnerSuites.find(suite => suite.id === "hiring-templates")?.manualOnly).toBe(true);
|
||||
for (const cell of cells) {
|
||||
expect(cell.task.expectedRunCount).toBe(5);
|
||||
expect(cell.task.expectedRunCount).toBe(7);
|
||||
expect(cell.task.minimumExpectedRunCount).toBe(5);
|
||||
expect(cell.task.attemptTimeoutMs?.local).toBe(15 * 60_000);
|
||||
expect(buildRunnerE2EProcessEnvironment({}, [cell]).PAPERCLIP_RUNNER_API_TOOLS_ENABLED).toBe("true");
|
||||
const payload = cell.profile.buildAgent({ executionId: "fixture", workspacePath: "/workspace", environmentId: "env", environmentFixtureId: "local",
|
||||
@@ -222,3 +283,92 @@ describe("production hiring fixture wiring and source observations", () => {
|
||||
expect(() => renderHiringCoderExample("No example", "Casey", "Company", "CEO", "FIX")).toThrow();
|
||||
});
|
||||
});
|
||||
|
||||
describe("settled hiring observation", () => {
|
||||
function observation(count = 2) {
|
||||
const f = createHiringTemplateTurnFixture(count);
|
||||
const apiState = { ...f.apiState, issue: { ...f.apiState.issue, title: "Hiring fixture", status: "in_review", conversationState: "waiting" },
|
||||
wakes: { events: [], truncated: false } };
|
||||
return { agents: f.evidence.agents, tasks: f.evidence.tasks, runs: f.evidence.runs,
|
||||
readRuns: [], apiState, finalApiState: structuredClone(apiState) };
|
||||
}
|
||||
it("retries the whole snapshot when a completion appears between ledger reads", async () => {
|
||||
const racing = observation(), settled = observation();
|
||||
racing.runs = racing.runs.filter(run => run.id !== "notify-2");
|
||||
racing.apiState.runs = structuredClone(racing.runs);
|
||||
let reads = 0;
|
||||
const value = await waitForSettledHiringObservation(async () => ++reads === 1 ? racing : settled,
|
||||
{ deadlineAt: Date.now() + 1000, intervalMs: 0 });
|
||||
expect(reads).toBe(3);
|
||||
expect(value.runs).toHaveLength(7);
|
||||
});
|
||||
it("awaits pending completion wakes and then two stable observations", async () => {
|
||||
const pending = observation(), settled = observation();
|
||||
const wakes = { events: [{ kind: "wake_request", status: "queued", finishedAt: null }], truncated: false };
|
||||
pending.apiState.wakes = pending.finalApiState.wakes = wakes as typeof pending.apiState.wakes;
|
||||
let reads = 0;
|
||||
await waitForSettledHiringObservation(async () => ++reads === 1 ? pending : settled,
|
||||
{ deadlineAt: Date.now() + 1000, intervalMs: 0 });
|
||||
expect(reads).toBe(3);
|
||||
});
|
||||
it("awaits an outbox completion not yet represented by a wake or run", async () => {
|
||||
const awaitingCallback = observation(0), settled = observation(1);
|
||||
let reads = 0;
|
||||
const value = await waitForSettledHiringObservation(async () => ++reads === 1 ? awaitingCallback : settled,
|
||||
{ deadlineAt: Date.now() + 1000, intervalMs: 0 });
|
||||
expect(reads).toBe(3);
|
||||
expect(value.runs).toHaveLength(6);
|
||||
});
|
||||
it("never admits truncated diagnostics, active runs or inconsistent observations", async () => {
|
||||
for (const defect of ["truncated", "active", "inconsistent"] as const) {
|
||||
const bad = observation();
|
||||
if (defect === "truncated") bad.apiState.wakes.truncated = bad.finalApiState.wakes.truncated = true;
|
||||
if (defect === "active") bad.runs[0]!.status = "running";
|
||||
if (defect === "inconsistent") bad.finalApiState.comments.pop();
|
||||
await expect(waitForSettledHiringObservation(async () => bad,
|
||||
{ deadlineAt: Date.now() + 10, intervalMs: 0 })).rejects.toThrow(/Timed out/);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
it("requires attributable callback replies and does not ignore unresolved coalesced or unknown wakes", async () => {
|
||||
const f = createHiringTemplateTurnFixture(2);
|
||||
const state = { ...f.apiState, issue: { ...f.apiState.issue, title: "Fixture", status: "in_review", conversationState: "waiting" },
|
||||
wakes: { events: [] as Array<{ kind: string; status: string; runId?: string }>, truncated: false } };
|
||||
const valid = { agents: f.evidence.agents, tasks: f.evidence.tasks, runs: f.evidence.runs,
|
||||
readRuns: f.evidence.readRuns, apiState: state, finalApiState: structuredClone(state) };
|
||||
for (const defect of ["reply", "coalesced", "unknown", "scheduled-retry", "duplicate-callback"] as const) {
|
||||
const bad = structuredClone(valid);
|
||||
if (defect === "reply") bad.apiState.comments = bad.finalApiState.comments = [];
|
||||
if (defect === "coalesced") bad.apiState.wakes.events = bad.finalApiState.wakes.events = [{ kind: "wake_request", status: "coalesced", runId: "missing" }];
|
||||
if (defect === "unknown") bad.apiState.wakes.events = bad.finalApiState.wakes.events = [{ kind: "wake_request", status: "other" }];
|
||||
if (defect === "scheduled-retry") bad.runs[0]!.status = "scheduled_retry";
|
||||
if (defect === "duplicate-callback") bad.runs.push(structuredClone(bad.runs.find(run => run.id === "notify-1")!));
|
||||
await expect(waitForSettledHiringObservation(async () => bad,
|
||||
{ deadlineAt: Date.now() + 10, intervalMs: 0 })).rejects.toThrow(/Timed out/);
|
||||
}
|
||||
});
|
||||
|
||||
it("admits completed and accounted coalesced wake bookkeeping after callbacks", async () => {
|
||||
const f = createHiringTemplateTurnFixture(2);
|
||||
const state = { ...f.apiState, issue: { ...f.apiState.issue, title: "Fixture", status: "in_review", conversationState: "waiting" },
|
||||
wakes: { events: [{ kind: "wake_request", status: "completed", runId: "notify-1" },
|
||||
{ kind: "wake_request", status: "coalesced", runId: "notify-2" }], truncated: false } };
|
||||
const observation = { agents: f.evidence.agents, tasks: f.evidence.tasks, runs: f.evidence.runs,
|
||||
readRuns: f.evidence.readRuns, apiState: state, finalApiState: structuredClone(state) };
|
||||
let reads = 0;
|
||||
await waitForSettledHiringObservation(async () => { reads++; return observation; },
|
||||
{ deadlineAt: Date.now() + 1000, intervalMs: 0 });
|
||||
expect(reads).toBe(2);
|
||||
});
|
||||
|
||||
it("classifies absent notification action identity as uncomparable separately from unchanged source coverage", () => {
|
||||
const evidence = validEvidence(2);
|
||||
evidence.readRuns = evidence.readRuns.filter(run => run.runId === "lead-1");
|
||||
const result = gradeHiringTemplate(evidence);
|
||||
expect(result.comparisonStatus).toBe("uncomparable");
|
||||
expect(result.turnAccounting.actionEvidence.status).toBe("uncomparable");
|
||||
expect(result.checks.filter(check => check.dimension === "coverage" && !check.passed).map(check => check.id))
|
||||
.toEqual(["completion-action-attribution"]);
|
||||
expect(result.checks.find(check => check.id === "production-source-reads")?.passed).toBe(true);
|
||||
});
|
||||
Reference in new issue
Block a user