mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 11:13:44 +02:00
Retain matched OpenCode routing results and original handoff defects
Co-Authored-By: Paperclip <noreply@paperclip.ing>
This commit is contained in:
1 parent
5f868ab2f4
commit
c596787a98
3 files changed
+1467
-33
No files matched your search
File diff suppressed because it is too large.
Load diff
@@ -1,37 +1,47 @@
|
||||
# OpenCode skill routing and clickable delivery — 2026-10-02
|
||||
|
||||
**TL;DR: the two matched pairs are pending; the correction has no measured performance result yet.** Prior evidence shows original OpenCode work skipped the operational skill until status recovery, while explicit storage read the skill/reference and saved a document but gave a code-formatted path. The original failures and worse clickable handoff remain recorded in the [prior repair report](2026-10-02-legacy-document-skill-repair.md).
|
||||
**TL;DR: 0 newly failing cells, 0 newly passing cells, 2 unchanged passes, and 0 pending pairs in this matched trial. Both document cases pass in both variants, so this does not establish that the metadata/link change caused the earlier failures to recover. Handoff remains imperfect in the original case:** candidate uses a clickable link with the wrong `PAP` company prefix; baseline supplies only a bare prefix-less slug path. Neither defect is scored by the preserved original document-storage oracle. The explicit clickable-link cases deliver the correct `RUN` link in both variants.
|
||||
|
||||
This approved narrow correction expands the operational skill's stock discovery description to include Paperclip task/heartbeat work and document/file delivery. The existing early API recipe asks for a clickable Markdown link, and the reference provides an actual canonical UI link example. The full skill body remains on disk; no default-manual pointer, adapter policy override, full-body injection or native completion tool is added.
|
||||
The narrow correction expands the operational skill's stock discovery description to include Paperclip task/heartbeat work and document/file delivery. The early API recipe asks for a clickable Markdown link, and the reference shows a canonical UI-link example. The full skill stays on disk; the eight-word manual and shared prompts remain reduced. No full-body injection, adapter policy override or native tool guidance is added to legacy runs.
|
||||
|
||||
| Classic OpenCode case | Pre-correction skill | Corrected skill |
|
||||
| --- | --- | --- |
|
||||
| Original assigned skill, verbatim request | Pending | Pending |
|
||||
| Explicit Paperclip storage with clarified clickable UI link | Pending | Pending |
|
||||
| Classic OpenCode case | Pre-correction skill | Corrected skill | Independent evidence and limitation |
|
||||
| --- | --- | --- | --- |
|
||||
| Original assigned skill, verbatim request | Pass | Pass | Exactly one saved document with skill-only marker and persisted revision in each. Both handoff paths are deficient and outside this case's link-free oracle. |
|
||||
| Explicit Paperclip storage with clarified clickable UI link | Pass | Pass | Saved revision/content and exact clickable `/RUN/issues/RUN-1#document-context-integrity-output` in both. |
|
||||
|
||||
Both variants use the same original request and a future explicit request naming the clickable company-prefixed Paperclip UI link its oracle already checks. Bare paths and code-formatted paths fail calibration; saved content/revision and correct relative/same-app absolute links remain independent evidence. Historical API-link grades are not retroactively changed. Explicit results from this clarified fixture are not a direct repeat of the earlier usable-link wording.
|
||||
Original machine results, source/configuration proof, tool-read sequence, costs and evidence hashes are preserved in the [safe evidence projection](2026-10-02-opencode-skill-routing-link-qualification.json). All four cells clean up successfully; none uses automatic disposition recovery or an extra attempt. No provider was rerun to improve these results.
|
||||
|
||||
The tiny manual/shared prompts, native Codex fix #14920, models, effort, auth,
|
||||
permissions, tools, fixtures and behavioral graders are fixed in both variants.
|
||||
Only the two operational skill source files differ. First-assignment skill
|
||||
selection and reference reads will be reported separately from recovery; a model
|
||||
can still omit stock skill selection despite the improved metadata.
|
||||
## Sources, admission and inspectable reports
|
||||
|
||||
The bounded comparison selects only these two classic OpenCode cells per
|
||||
variant, four expected provider turns total, on separate frozen concurrency
|
||||
targets. Exact-source credential-free admission must pass before provider
|
||||
credentials. Existing hard stops and cleanup remain intact. No broad matrix or
|
||||
native provider rerun is part of this correction.
|
||||
- Candidate `fe9dc1e3c518825242ed889ab9c8352986f8c2ed`, frozen branch `codex/opencode-stock-routing-qualified`: [workflow](https://github.com/paperclipai/paperclip/actions/runs/37069547401), [published dashboard and screenshots](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37069547401-1/index.html).
|
||||
- Baseline `0d7ecfa96d72fba79b7f0a25052b42c0686c0488`, frozen branch `codex/opencode-pre-routing-skill`: [workflow](https://github.com/paperclipai/paperclip/actions/runs/37069552374), [published dashboard and screenshots](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37069552374-1/index.html).
|
||||
- Trusted default-branch workflows ran concurrently on separate targets. Each of four measured cells passed its own exact-source, credential-free 587-assertion prerequisite before provider credentials. Candidate source fingerprint is `aeb792549146cc29eda779e83ae22d336276e31162f738a1dde85a82e7833902`; baseline is `0315602ac6788cf0d65561a73cc85d790768ef90c6815257077a2381f961b96e`.
|
||||
- Git tree comparison verifies 8,242 identical tracked files. Only the two operational skill sources and an explicit baseline provenance receipt differ. Models, effort, tools, auth, permissions, budgets, manual/shared prompts and behavioral fixtures/graders are fixed; both fixture configuration hashes match. Suite/source digest differences correctly include the changed skill bytes.
|
||||
- Local source preparation passed 918 E2E support tests, eight canonical inventory calibrations, skill validation, E2E typecheck, two-cell discovery and both exact-source prerequisites. Both capability inventories use the same declared source list. Initial setup mistakes remain retained and did not reach providers.
|
||||
|
||||
Discovery delivery is verified for existing classic Codex/Claude/OpenCode and
|
||||
ACP Codex skill mounts. ACP Claude's names/root-only advertisement is a separate
|
||||
metadata follow-up. Hermes/Pi are not live-qualified here; custom ACP and OpenClaw
|
||||
have different delivery contracts. Hiring templates are owned by a separate
|
||||
implementation and are unchanged by this qualification.
|
||||
The explicit fixture now asks for a clickable company-prefixed Paperclip UI document link, matching its intended oracle. Bare paths and code-formatted paths fail calibration; correct relative/same-app absolute links pass. The original assigned-skill request and procedure remain verbatim. Earlier API-link grades are not retroactively changed, and these explicit results are not a direct repeat of the older “usable link” cohort.
|
||||
|
||||
Local preparation passes 918 E2E support tests and eight canonical inventory
|
||||
calibrations. Both capability inventories now use one declared skill-source
|
||||
list, including issue-documents.md; their existing heading parsing scopes are
|
||||
preserved. Skill validation passes. Final exact-source type/discovery/preflight,
|
||||
source freeze and protected runner dispatch remain pending. All prior failures,
|
||||
automatic recovery and costs remain retained; both PRs remain draft.
|
||||
## Which guidance was read before delivery
|
||||
|
||||
In the candidate original assignment, OpenCode loads the assigned Context integrity output skill, then operational Paperclip before any document write. The returned Paperclip skill is marked truncated, but the early saved-document recipe and reference pointer are visible. It reads `issue-documents.md`, then writes the public task document through the API. This differs from the previously failing candidate, which first loaded Paperclip only in a separate disposition-recovery run after its local-only output.
|
||||
|
||||
The current baseline original assignment first loads the assigned output skill and writes a workspace Markdown file. It later loads Paperclip and reads the issue-document reference, then saves a public document before completing the **same** assignment. Thus its successful delivery did not require improved metadata; stock selection timing varies between trials. Both explicit variants load Paperclip and read the reference before the document API call. Four rendered final-state screenshots were inspected, and public document revisions/content and completion comments agree with the retained results.
|
||||
|
||||
The candidate original comment has an actual Markdown href `/PAP/issues/RUN-1#document-context-integrity-output`, despite the measured company prefix being `RUN`. Baseline's original comment has a bare `/issues/runner-e2e-…#document-task-document` path without a Markdown anchor. The candidate's literal prefix matches the reference's example, making example copying a plausible cause, but this trial does not isolate that cause. Authenticated click navigation was not replayed after cleanup; neither original link is presented as a verified usable handoff. A minimal follow-up can show an issue-derived prefix instead of a literal company example. These original machine passes must not be presented as fully correct links.
|
||||
|
||||
## Timing, usage and cost
|
||||
|
||||
Four provider turns were expected and four assignment runs occurred. All four ledger receipts contain token usage and reported LLM cost. Local runtime remains unmetered; actual external billing is not established.
|
||||
|
||||
| Case | Baseline provider seconds | Candidate provider seconds | Baseline cell seconds | Candidate cell seconds | Baseline reported USD | Candidate reported USD |
|
||||
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
|
||||
| Original | 25.870 | 67.897 | 46.285 | 89.455 | 0.0058472545 | 0.0041383700 |
|
||||
| Explicit clickable storage | 73.984 | 82.751 | 94.804 | 102.901 | 0.0049436470 | 0.0066440790 |
|
||||
|
||||
Reported LLM totals are **$0.0107909015 baseline** and **$0.0107824490 candidate**. The candidate original case takes longer in this trial; these single observations, different cache/token receipts and unmetered runtime do not establish general speed, cost or quality equivalence.
|
||||
|
||||
## Preserved failures and scope
|
||||
|
||||
The [original combined prompt-removal report](2026-10-02-stock-harness-live-comparison.md) still records two new classic Claude/OpenCode document-delivery failures, two OpenCode ordered-case improvements, and separate unchanged credential/chat failures. The [first recipe repair report](2026-10-02-legacy-document-skill-repair.md) still records Claude's improvement, OpenCode's unresolved local-only output, and the candidate's worse non-clickable explicit handoff. New successes do not erase those earlier observations or establish robustness.
|
||||
|
||||
This comparison qualifies only classic OpenCode's two selected journeys. Existing classic Codex/Claude/OpenCode and ACP Codex skill mounts were audited; ACP Claude's names/root-only metadata and unsupported custom ACP delivery remain separate. Hermes/Pi/OpenClaw are not live-qualified here. Native completion descriptions and hiring templates have separate changes and measurements. Both PRs remain draft; no merge or further paid breadth is part of this report.
|
||||
@@ -4,10 +4,14 @@ Created: 2026-10-02. Status: item 1 merged for native Codex app-server in
|
||||
[PR #14920](https://github.com/paperclipai/paperclip/pull/14920). Item 2's default
|
||||
hire manual is reduced to identity only, and common legacy startup/resume
|
||||
instructions are in [PR #14948](https://github.com/paperclipai/paperclip/pull/14948).
|
||||
GitHub live qualification has measured document-delivery failures; PR #14948 remains
|
||||
draft. A matched prior-instruction comparison and its completed interrupted-cell
|
||||
recovery are recorded in the [live report](2026-10-02-stock-harness-live-comparison.md).
|
||||
Additional carriers and native instructions remain open.
|
||||
GitHub live qualification retains earlier document-delivery failures; the latest
|
||||
focused OpenCode comparison passes both cases in both variants but still exposes
|
||||
deficient original-case handoff links. PR #14948 remains draft. The original
|
||||
[live report](2026-10-02-stock-harness-live-comparison.md), completed
|
||||
[skill repair](2026-10-02-legacy-document-skill-repair.md) and latest
|
||||
[stock-selection/link comparison](2026-10-02-opencode-skill-routing-link-qualification.md)
|
||||
retain their separate sources and limitations. Additional carriers and fixed
|
||||
native instructions remain open; hiring templates are being reduced separately.
|
||||
|
||||
Goal: keep the agent's stock harness behavior and add only what it needs to work
|
||||
with Paperclip. Apply this across legacy adapters, the new Runner, and their
|
||||
@@ -438,11 +442,13 @@ will address them. Record intentional behavior explicitly rather than as a bug.
|
||||
| 2026-10-02 | Final bounded recovery completed; no further model reruns. | Both matched cohorts have all 24 results, 15 pass and nine fail. Two classic skill deliveries regress; two OpenCode ordered cases pass only with reduced instructions. Overall parity is not behavioral equivalence. All 48 retained result/receipt projections are hashed and sanitized; original interruptions and partial unknown spend remain recorded. |
|
||||
| 2026-10-02 | Dotta approved the measured legacy delivery repair and narrow follow-up qualification. | Early operational skill PUT/receipt/link guidance plus generic issue-document reference; eight-word manual retained. Focused original + explicit Paperclip-storage cases on classic Claude/OpenCode, with only the two skill sources varied. All four pairs completed: Claude original Fail → Pass, Claude explicit Pass → Pass, both OpenCode cases Fail → Fail. Explicit OpenCode handoff worsened beneath the unchanged machine grade (clickable API URL → code-formatted path). [Repair report](2026-10-02-legacy-document-skill-repair.md); no full matrix rerun. Claude chat-memory (F15) and ACP credential/receipt failures (F16) remain separate and unresolved. |
|
||||
| 2026-10-02 | Native tool-description comparison completed all six paired cases. | Zero newly failing cases, three unchanged completion passes, three unchanged blocker failures. Claude/OpenCode blocker API matchers pass in both variants; UI matcher wrongly demanded marker-only replies. Codex's exact-action punctuation failure is unchanged. Original failures retained; corrected 10-case browser calibration and separate retained-DOM replay pass all six visible replies. Exact Codex action failures remain. No native fixed-prompt removal measured. |
|
||||
| 2026-10-02 | Dotta approved minimal operational skill selection/link correction after retained OpenCode diagnosis. | Candidate `fe9dc1e3c` and baseline `0d7ecfa96d` have two matched Pass → Pass cases, zero new failures/passes and no pending pairs. All four exact-source gates pass 587 checks, all four provider runs and cleanup pass. Candidate original loads Paperclip/reference before delivery, but uses the wrong PAP prefix; baseline original saves publicly later within the same assignment and gives a bare path. Explicit clickable UI links are correct in both. [Complete report](2026-10-02-opencode-skill-routing-link-qualification.md); prior failures retained, no causal or broad quality claim. |
|
||||
|
||||
- F18: Native blocker browser assertion required a marker-only reply despite asking for owner/action/reason. Correct marker-plus-explanation checks symmetrically, calibrate contradictory and future-condition replies, retain original verdicts.
|
||||
- F19: Original prompt-removal cohorts loaded Paperclip; OpenCode skill truncation omitted the late API recipe. The early repair fixed Claude in one paired trial. In the repaired original OpenCode assignment, Paperclip was first loaded only during disposition recovery after a local-file write. Shared legacy operational-skill delivery/selection needs review before more recipe expansion.
|
||||
- F20: Explicit repaired OpenCode reads the early recipe/reference and saves a public document/revision, but hands off a code-formatted path rather than a clickable anchor. Baseline provides a clickable API URL, rejected by the UI-only oracle. Preserve both grades; clarify future clickable UI-link fixture and review the smallest link example separately.
|
||||
- F22: Legacy operational skill mounting is already mandatory, but its discovery description omitted ordinary task/heartbeat work and document delivery. The approved narrow fix expands stock metadata selection and adds a real Markdown link example in the existing reference. Original + clarified-explicit OpenCode paired qualification is pending; no full-body injection or native-tool leakage. ACP Claude names-only metadata, custom ACP unsupported delivery, OpenClaw wrappers and Pi HOME differences remain distinct follow-ups.
|
||||
- F22: Legacy operational skill mounting is already mandatory, but its discovery description omitted ordinary task/heartbeat work and document delivery. The approved narrow fix expands stock metadata selection and adds a real Markdown link example in the existing reference. Original + clarified-explicit OpenCode pairs both pass in both variants, so improvement causality is not established. Candidate loads the operational skill/reference before public delivery; baseline original writes locally first, then saves publicly in the same assignment. Original candidate copies the example PAP prefix into its link while baseline supplies a bare slug path; neither is scored by the original link-free oracle. No full-body injection or native-tool leakage. ACP Claude names-only metadata, custom ACP unsupported delivery, OpenClaw wrappers and Pi HOME differences remain distinct follow-ups.
|
||||
- F23: The new original OpenCode candidate copies `/PAP/issues/RUN-1#document-context-integrity-output` from the literal example despite actual prefix RUN. It passes durable-document grading but has an incorrect-prefix handoff. Baseline original also lacks a clickable canonical link. A reference-only issue-derived link construction is proposed; no further paid rerun.
|
||||
- F21: Skill heading insertion shifted both generated capability inventories. General CI caught stale metadata after the frozen repair campaign’s narrower build admitted providers. Regenerate canonical derived files and require stale-manifest/inventory checks before future provider admission.
|
||||
|
||||
For each completed item, add the chosen behavior, changed paths, verification
|
||||
|
||||
Reference in new issue
Block a user