66 KiB
Copilot 1.0.88 rich ACP capability audit
Current checkpoint (2026-09-30): Copilot profile v7 remains unqualified. Paid controller/Product harness source is 40064d28522de25fea85c1297f35b41bb8a8897a; runtime source is 5b8e4454ef0bf12d0bb068c2e41d8c9df9356a1c. All three target runtime builds are complete. Local attached-async attempt 40064-copilot-async-01 passed all eight matchers on gpt-5.6-luna, including exact one-command identity, trusted child exit and independent completion observations before terminal settlement, one run and one exact final marker. Done was visible, and full end integrity, owned process/semaphore retirement and temporary-root cleanup passed. The exact local native-permission-deny-write case 40064-copilot-deny-01 also passed all ten checks, with no target side effect through process retirement, correlated cancellation and successful cleanup/end integrity. These two local cases pass; the provider remains unqualified. Native per-run USD is unknown; included credits stayed 16→16 after async and reached 17 after denial, with possible dashboard lag, and extra usage remained disabled at $0. Result SHA-256: d1d074bc9df4644794dfa8b020fdeba8a9b0de4618e1f55d1845f7726a2ece33; reconciliation SHA-256: 8e9b9233ed987e9e522b17a271922fd409d311e04ee5c3422a4b5bd1fa2fce68. The prior exact head passed CI and Greptile for draft #14676. The earlier v6 denial pass and attached-command failures remain historical, including the doubled final marker and missing settlement observations; they do not qualify v7. See the comparative capability report for exact source identities and remaining local/Daytona, billing and cleanup gates.
Message identity implementation (2026-09-30, deterministic evidence and one paid async case; full qualification pending): the pinned 1.0.88
ACP mapper discards assistant.message_start and removes native messageId
from text deltas. The separate candidate keeps the original executable and
every embedded distribution asset, patches the hash-pinned JavaScript mapper,
and supplies its complete verified distribution through a per-spawn owned
entry shim. The shim installs the module guard before importing Copilot;
ambient distribution/version/module paths do not select executable content.
The original executable, inner archive, upstream app, patched app and complete
platform closure have independent hashes. Unknown upstream bytes are refused.
The mapping carries real message IDs on the ordered standard ACP stream, including empty starts; completed native messages do not echo their full text again. Distinct native messages remain distinct transcript items. Earlier messages close as commentary, and only the last message supplies final output. An empty final message clears preceding text. Reasoning and unidentified native session notices are not promoted into the final answer. This does not deduplicate equal text or infer a message boundary from a tool call.
A credential-free loopback probe through the actual pinned ARM64 process
observed three equal-text messages with three native IDs, three empty starts,
and attached-shell settlement. A separate actual ACPX transport fixture preserved
empty boundaries across three warm turns. All three platform distributions were
materialized; this is not native x64/Linux execution proof or paid qualification.
The complete candidate inventories
retain platform closure hashes. Private bounded probe receipts are under
runtime-6-6-9-eof-preparation/copilot-message-identity-candidate/ in the existing
qualification artifact collection. Copilot v7 binds this candidate to a new profile and warm identity; v6 paid
failures and missing settlement observations remain unchanged. The 1.0.89 comparison moves ACP transport into its native runtime and
does not establish that upgrading alone fixes message identity.
Initially audited 2026-09-28 against repository base c65fc9e3c81c41aafe421aa90a00514b84343285;
updated 2026-09-29. Status remains candidate, not qualified. On frozen source
bd4cc29c3017ed5e1484423e842fd48e3b2f49f3, profile v4 local hello, question/answer,
semantic plan, controller restart and file-edit cases passed. Daytona hello passed
one native gpt-5.6-luna run with six matchers on immutable image
sha256:bff4c3f291087a0eeae37e4c20dd51857b92833eaf73ba3aca4157f37de1e109.
Earlier question-marker, recovery, startup-timeout and PostgreSQL failures remain
retained; later success does not establish causes for unresolved earlier failures.
The earlier 2026-09-29 admission identity was profile v5, binding shared ACPX patch
79aad2d688b03362e8cfcbf7a08f78a8383869f6882a9c9ed66f1d18efb94f2b.
The native Copilot binary and permission/instruction policy are unchanged. The
new digest rejects old warm sessions; v4 paid receipts remain historical evidence,
not a claim of v5 qualification. The new native-protection Product cases below
have not run on the new source.
Daytona sandbox-specific analytics stabilized across three reads at a provisional $0.0039682144 for the closed attempt interval; invoice finality and native per-run Copilot USD remain unknown. The Product list-price estimate was $0.0040321867. One exact owned sandbox was absent after the full 360-second cleanup observation; 45 local PIDs retired and semaphore counts returned 255→266→255. The conditional 720-second infrastructure bound was $0.053496. These are distinct accounting claims, not a zero-cost assertion. The allocated provider budget remains $25 within the shared $100 ceiling; no new paid attempt follows automatically from registration. Mac x64, the remaining Daytona workflows, broader permissions/background variants and complete rich-field projection remain unqualified.
Connection policy audit, 2026-09-29
The pinned server accepts permission-changing session/set_config_option,
session/set_mode, and leading-slash CLI commands. Entering autopilot enables
allow-all; loading an autopilot conversation enables it again. A clean launch
configuration alone therefore cannot establish safe restored-session policy.
copilot-policy.ts and the optional ACPX protocolGuardFactory enforce policy
on each actual stdio connection, before bytes are delivered to the SDK or written
to the provider. Copilot admission requires the full new/load/resume result to
report the exact agent-mode URI and allow_all: off. Missing, duplicated,
unknown-authority, or custom-agent configuration fails closed. A connection may
select only its explicitly admitted model; a prompt requires proof of that
model. The guard rejects leading-slash commands, native mode/config controls,
and unsafe mode/config drift. Its authority is fresh for every connection and
remains invalid after a violation. It does not restore trust from cached ACPX
status. The stream errors, aborts outstanding response delivery, and terminates
the provider on violation; the existing runner owns bounded process cleanup.
Other providers do not install this guard.
The native ask_user and exit_plan_mode tools are absent from both agent and
plan mode in the pinned deterministic tool catalogs. Forced model tool calls
return explicit tool-unavailable results rather than producing a blocking
request. No --no-ask-user or tool-exclusion flag was used. Agent mode is the
only admitted production mode. If any native input-request event nevertheless
arrives, the guard terminates the connection; it never presents a form whose
answer cannot be delivered. This narrows the blocking-input question for the
admitted pinned mode without claiming an upstream question responder exists.
The offline conformance record
retains every attempt and its evidence digest, including the isolated comparison release. A protected file outside the
working directory was denied before its contents reached the fixture model.
A durable conversation was closed and loaded, then requested permission for a
shell write; reject_once prevented the file. The initial attempt to load an
empty conversation failed with resource-not-found and remains retained. The
second attempt seeds one text turn before close/load; it is not a retry of a
measured behavior failure. A separate
process-death attempt first performs one allowed seed append, then SIGKILLs and
replaces the exact provider using the same private session state. Load succeeds,
the seed append still occurs exactly once, and a fresh denied write remains
absent. Both native processes use permission request ID 0, proving why response
authority must belong to the current connection. These are local native
close/load and process-death probes, not final packaged Product qualification.
Production-pin probes use the verified 1.0.88 ARM64 binary; an explicitly selected
1.0.89 comparison has separately verified archive integrity. All use an explicitly constructed
credential-free environment, COPILOT_OFFLINE=true, and a synthetic OpenAI-style
model at an ephemeral loopback HTTP server. The synthetic gpt-4.1 ID does not
name an authenticated GitHub model selection. Counts include the setup turn.
There are zero paid provider calls and zero model spend. Raw evidence, including
the failed empty-session attempt, stays in the private
copilot-policy-20260929 artifact directory.
The prior Product question failure remains a model-behavior failure: the retained
snapshot contains the exact requested marker in the task instructions, the
answered Cobalt choice, and warm-session continuation. Copilot instead supplied
[terminal marker] to paperclip_finish. Neither the grader nor the terminal
marker requirement changed. A final-runtime rerun remains required.
Final v3 Product question attempt
The final v3 question receipt
retains the single reserved attempt at source 3d0c45920, with the exact frozen
controller distribution, native assets, pack, Node, daemon and sidecar hashes.
The first turn produced the structured question and the board answered Cobalt.
Continuation failed during native session.open recovery, before continuation
usage was reported, with an unclassified sidecar rejection. The 72.086-second
case failed; cleanup passed and no retry occurred. This is a recovery/admission
failure, distinct from the earlier literal-marker behavior failure. Neither is
removed from qualification evidence.
One run reports GitHub/unpriced usage: 28,545 input, 13,818 cached input and 525
output tokens. Provider USD and upstream model-request count remain unknown.
The refreshed account counter moved from 5 to 6 of 1,500 included credits;
additional usage remained disabled with a $0 budget and $0 account cash charges.
This account-level delta is not a per-turn USD allocation. The Product aggregate's
zero reported cost with unpriced/incomplete coverage is not an authoritative
zero-cost receipt. No additional live attempt is authorized by this result.
A credential-free native recovery reproduction
closes the provider, deletes the prior registered instruction copy, and reproduces
ENOENT at bindAcpxAgentFiles. Reopening with the current registered copy
succeeds with the same native session. A second explicit fixture turn proves
actual native session/load and reaches end_turn. The fixture uses exactly two
loopback responses, a synthetic credential and test-only model metadata;
it makes no paid calls and does not qualify authenticated model selection. This
supports the stale durable runtime-context diagnosis; the failed Product run did
not retain the underlying wire error. Cold rotated restoration now takes the
current authenticated runtime context. Live reconnect and pending warm-transition
receipts keep their existing context; replacing that context in an active provider
is a separate lifecycle operation.
Native instruction delivery
The initial candidate was profile v4. Its declaration binds native personal-file
instruction delivery, including replacement under the provider lifetime lease before native launch. The isolated
COPILOT_HOME/copilot-instructions.md is written atomically with mode 0600 under
the protected provider home, and empty instructions replace any stale content.
Concurrent contenders rejected by the lease and unauthenticated admissions do
not mutate the instruction file. The host awaits each write before releasing its
lease, including cancellation, so a late write cannot overwrite a successor.
Profile v3 identities are rejected before native launch; historical evidence
below remains labeled with its original profile and source.
The native model-request probe
retains a separate instruction-delivery failure: Copilot 1.0.88 ignores generic
ACP _meta.systemPrompt on new sessions, and ACPX does not include it on load.
Actual loopback model requests contained neither the Paperclip prompt nor the
registered directory guidance. Passing a refreshed descriptor alone is insufficient.
A credential-free follow-up wrote the composed text to the isolated provider
home's copilot-instructions.md, using Copilot's native personal-instruction
mechanism. The first request's system message contained the old directory and
original entry. After closing the provider, deleting the old directory and
reopening with fresh instructions, actual session/load and a second explicit
turn completed. The second model request contained the current directory and
updated entry, with neither old value in its system message. Both attempts used
two synthetic loopback responses and no paid calls. This probe is not a final
built-production or authenticated qualification result. A follow-up
source-level v4 probe
uses the actual production sandbox writer and confirms the same new/load model
request behavior. The post-lease source probe
reconfirms this after moving refresh under lifetime ownership, and retains an
earlier zero-call lease-admission failure. The final frozen pack still requires
independent verification.
Native detached work is explicitly unsupported
The broader 2026-09-29 deterministic probe reproduces early completion for
bash with mode: async and detach: true. After an allowed finite two-second
command, Copilot sends a completed tool update naming a detached shell ID, then
session.idle and end_turn before the marker file exists. A separate diagnostic
keeps observing for three seconds after terminal and proves the marker appears
late. Both attempts retain their failed settlement assertion.
The preceding standard tool-call notification includes rawInput.detach: true,
but the observed permission request carries only the command. The Copilot guard
rejects that update before the SDK can deliver any permission answer, terminates
the provider, and reports COPILOT_DETACHED_WORK_UNSUPPORTED with instructions to
run attached or use a managed runtime service. It also rejects detach carried
only in a permission request. Attached asynchronous work retains the normal
permission path. This is policy denial, not completed background settlement.
A native offline probe invokes the exact production TypeScript tool-update policy, kills the provider before permission dispatch, and observes no marker through the command's bounded delay. It makes one loopback fixture request and sends zero permission responses. Separate actual patched-ACPX stream tests prove the same rejection before SDK delivery. The native probe intentionally tests only tool admission: its synthetic model does not satisfy full production model admission. Neither layer's result is misrepresented as final-pack live evidence.
The bounded schema audit finds detach on bash, absent on read_bash,
stop_bash, and task; write_bash and PowerShell variants are not advertised
in this ARM64 catalog. Embedded JavaScript references write_bash but delegates
schemas to native Rust. The guard is independent of tool name, but this audit
cannot prove every platform/feature-flag variant or shell-created daemon
lifetime. Comprehensive background qualification remains blocked. Empty native
background-task notices and fixed delays cannot establish settlement.
An isolated exact-integrity 1.0.89 comparison passes the same attached async scenario and reproduces the detached late marker. Production remains pinned to 1.0.88. As checked on 2026-09-29, upstream #4743 is closed and explicitly concerns attached async commands, distinguishing deliberate detached services. These results do not establish that issue remains unfixed. #4537 remains open; the narrow denied-write/read/recovery probes did not reproduce its permission bypass. The original detached failures remain recorded as failures of the attempted runner settlement contract; they are not relabeled as successful settlement.
Evidence and scope
The audit inspected the exact @github/copilot@1.0.88 platform archives, the
native executable's embedded app.js, schemas/session-events.schema.json, CLI
help/config/environment output, and real ACP wire traffic. The embedded schema
SHA-256 is d8cb713c05d5278a68dde5c5d4482574f836e922ae13aa06f82474c209a7c6e9.
The complete event/field inventory
contains all 150 native event types and their data field names, including every
event we do not subscribe to. Fields in that file describe the native schema;
they are not a claim that every event was observed on the wire.
Primary external references: ACP server documentation,
CLI command reference,
denial bypass report #4537,
and background completion report #4743.
The exact pinned implementation takes precedence over moving documentation.
The harness priorities report
defines the requested product outcome. Existing Codex app-server contracts and
conformance tests are the comparison baseline, especially turn/steer,
turn/interrupt, thread/read, file changes, user input, and scoped approvals.
Distribution and authority
Launch the verified platform copilot executable directly with --acp --stdio.
Do not execute the mutable npm-loader.js, install hooks, or an ambient PATH
binary. materialize-copilot-binary.mjs validates exact package/version, executable
mode and bytes, then writes a native-closure manifest. Runtime descriptor leases
must independently verify the trusted closure digest before every launch.
| Platform | Executable SHA-256 | Bytes | Closure SHA-256 |
|---|---|---|---|
| macOS ARM64 | a9ff8babb10b7e443182ae96a8bc50a9c826ef1c773e1344c396eb5bf7f512c3 |
152595280 | fb3b367a45cd76122fe931521fa2a18adf234ba944fc302db9e10e005e57037e |
| macOS x64 | 85eb919f6b9b9dd833ce5e326cbf974b3ee2d4a9ac525c59d4ec9c9ec085715b |
165041200 | 05f3497b336b3efdec347beb2e3b80b02cfa95f811fafddc25d0b029ab95d711 |
| Linux x64 | 0059754cf78c3f3bf2c9d4564dfa7e9e25f3a3f8f411f2f0cdad9363f5662748 |
169544512 | 1a675c5b54ae4d94f08718a318451e0499708ded388b4cfd98acec6b4311ccbd |
All three archives were verified against the npm SHA-512 integrity value before
hashing the executable. Archive pins are retained in the materializer. The macOS
ARM64 executable has live evidence; Linux x64 now has the initialize-only image
proof described below. macOS x64 remains execution-unverified. The binary contains its JavaScript/native runtime and
extracts it into COPILOT_PKG_CACHE_HOME; this must be a fresh per-spawn private
lease directory, never a writable cache shared across executions. The native
distribution verifier supplied by the foundation owns that isolation boundary.
copilot-profile.ts supplies private HOME/XDG directories, COPILOT_HOME,
COPILOT_CACHE_HOME and an extraction-cache binding; update disabling; no built-in
MCP servers; no remote/remote-export or shell startup environment; and secret
environment stripping for child shells/MCP. Only explicitly bound
COPILOT_GITHUB_TOKEN may authenticate production use. The offline fixture uses
an intentionally separate, credential-free loopback provider, not this production
credential path. Private configuration disables hooks, memory and automatic IDE
attachment and has no trusted folders.
Never set COPILOT_ALLOW_ALL=true: this exact string also trusts workspace
hooks, plugins and MCP configuration. Paperclip's full-auto policy answers the
individual permission callback. It must not grant ambient configuration trust.
Mode/config changes, including autopilot and the provider's allow_all option,
must not bypass the admitted Paperclip policy. Slash-command discovery includes
commands capable of changing permissions, cwd, remote/export, MCP and schedules;
these are not authority to offer an unrestricted command UI. Adversarial
workspace/config qualification remains required before release.
The pinned ACP handler maps allow_always to native approve-for-session for
commands, writes, reads, MCP and other supported tools. Path approval is also
session-scoped; URL approval is session-scoped to an origin pattern. Factory
permissions omit that option and reject fabricated permanent approval. This
scope is confirmed in source, not inferred from the option label. Read/write
session grants are broader than one file and must be described accurately.
Negotiation and event contract
The observed initialize result advertises protocol 1; loadSession: true;
HTTP/SSE MCP; image and embedded-context input; no audio input; and session list
and close. It does not advertise steering, forking, goals, or a question/plan
extension responder. The source adapter supports model/reasoning/config changes,
but the offline session only returns mode and allow_all options. There is no
verified GitHub model ID from the historical offline probe. Require an explicitly selected model
and exact effective-model verification; never silently use the fixture's
gpt-4.1 or substitute another model. Authenticated discovery on 2026-09-28 subsequently advertised and accepted
gpt-5.6-luna; subsequent canonical and Product cases verified inference on that exact model.
The native event extension is real and is negotiated with:
{"clientCapabilities":{"_meta":{"github.com/copilot":{"events":["subagent.started","session.workspace_file_changed","assistant.usage"]}}}}
The notifications are:
{"method":"github.com/copilot/sessionEvent","params":{"sessionId":"...","type":"subagent.started","timestamp":"...","data":{},"agentId":"..."}}
The source caps subscription names at 128, payloads at 32 KiB, and pending
notification sends at 256. Oversized/unserializable payloads carry dataOmitted;
backpressure can drop events. There is no event ID or reliable replay contract.
skill.context_delivered and skill.context_delivered_ref are explicitly blocked
even when subscribed. These are provider restrictions, not missing runner parsing.
copilot-events.ts requests 22 event types and projects only bounded declared
fields after matching the admitted session and an active receiver. The pinned
notification has no originating turn ID; arrival during a later turn cannot
establish attribution. It records source method/type,
provider timestamp and subagent identity. Payloads cannot authorize filesystem
reads, workspace rebinding, permission changes, native-input replies or terminal
settlement. Inline binary assets are content-address verified; only metadata is
forwarded until a provider-session artifact resolver can upload them safely.
copilot-extension-adapter.ts retains every safe projected field in bounded,
session-scoped provider notices with method/event/session provenance and an
explicit unknown originating-turn attribution. Typed delegation, compaction and
artifact projections are withheld because they would falsely attribute delayed
session notifications to the receiving turn. Native correlated ACP turn events
remain separate. Follow-up: a versioned provider correlation contract is required
before promoting these notices into typed turn activity. Notices have readable
summaries; secret-shaped string values
are scrubbed without erasing numeric token counters. These display events never
create a usage charge, input-resolution acknowledgment, registered artifact, or
turn terminal event. The provider registry installs this factory and initialize
capability metadata in the provider branch.
Comparison against Codex app-server
“Source” means verified in the pinned implementation; “wire” means observed in the real ARM64 executable against the deterministic offline model. Product UI and Daytona claims require the separate live product qualification.
| Capability / Codex benchmark | Copilot native and ACP exposure | Runner and user-visible surface | Evidence / remaining gap |
|---|---|---|---|
| Authentication | Initialize advertises copilot-login, including terminal-auth command, args and label. |
Explicit company-bound COPILOT_GITHUB_TOKEN; missing binding fails before executable admission. Terminal login is not launched. |
Packaged initialize and host cleanup/retry test; terminal-auth remains intentionally unused because it would introduce ambient interactive identity. |
| Prompt attachments | Initialize advertises images and embedded context, and explicitly denies audio input. | Current runner prompt contract sends text. Image/context input blocks are not forwarded. | Observed initialize; P1 add validated attachment inputs. Audio is confirmed unsupported in this ACP advertisement. |
Active turn steering (turn/steer) |
Native SDK steering exists. ACP session/prompt unconditionally aborts the active session before sending a new prompt. |
Unsupported active steering; do not impersonate it with concurrent prompts. | Source; P1 add a versioned upstream ACP steering method. |
| Ordered follow-ups | Native pending-message controls and pending_messages.modified; notification has no queue body. |
Lifecycle activity only. Scheduler can start a subsequent completed-turn prompt, but that is not native queue delivery. | Source; P1 require queue acknowledgment and ordering contract. |
Interruption (turn/interrupt) |
Standard session/cancel, active prompt abort and process shutdown. |
Shared cancellation and bounded cleanup. | Source; authenticated cancellation/process-tree test pending. |
| Session recovery/history | session/load, list and close; native history and rewind richer. |
Exact identity/warm continuation through shared ACPX host; never replay approvals or mutations. | Live Product controller restart preserved the pending interaction and reused the same provider session. Provider-death restoration and history loading remain unqualified. |
| Session list / explicit close | Initialize advertises both methods. | Runner owns its selected-session registry and process cleanup; it does not call Copilot's list or explicit close methods. | Observed initialize; P2 company-scoped history/session management before consuming these interfaces. |
| Fork / history paging | Native CLI/SDK capabilities exist; no ACP fork advertised. | Unsupported. | Confirmed absent from initialize advertisement, not proof native harness lacks it; P2 upstream extension. |
| Tools / correlation | Standard tool_call/tool_call_update; parent identity in _meta["github.com/copilot"].agentId. HTTP/SSE MCP supported. |
Shared tool activity, authenticated runner-owned MCP bridge. | Real create/bash/read_bash traffic and canonical authenticated get-task-context pass; Product semantic question and plan calls observed. Full adversarial company-boundary coverage remains pending. |
| MCP transport selection | Both HTTP and SSE are advertised. | The assigned Paperclip gateway uses the controlled HTTP bridge. Arbitrary SSE endpoint configuration is not exposed. | Observed initialize; SSE remains unused, P2 only if a governed connection requires it. |
| Scoped approvals | session/request_permission, actual options allow_once/allow_always/reject_once. |
Shared durable permissions; only received decisions offered, policy enforced. | Real-service wire ID 0 denied before file creation, with no side effect through cleanup. Durable Product restrictive-mode recovery and wider tool denial remain unqualified. |
| Structured questions | Native ask_user callback and user_input.requested; current ACP adapter does not wire the responder. |
Emits capability-gap notice if native notification arrives; cannot claim answer delivery. | Actual agent-mode tool list omits ask_user without a suppression flag. Paperclip semantic questions are available: restart case passes, while the separate question case failed its exact marker. Pinned offline agent/plan catalogs now prove both tools unavailable; the production guard admits agent mode and terminates any unexpected native blocking request. |
| Plan approval | Native exit_plan_mode callback; notification contains plan content/actions but lacks qualified ACP responder. |
Capability-gap notice only; never synthesize plan acceptance. | Agent-mode tool list omits exit_plan_mode. Paperclip semantic plan/revision approval passes through the UI; Native plan-mode input is confirmed unavailable offline; production mode controls remain disabled. |
| Plan progress | Standard plan from todos SQL; native session.plan_changed has operation only. |
Existing ACP plan/activity; native operation preserved, planContentAvailable:false. |
Source; plan document reads require native interface. P1. |
| Models / reasoning / config | Source session/set_model, config options for model, reasoning, mode, custom agents, allow_all. |
Explicit model admission. Mode/governance changes must remain policy-gated. | Authenticated catalog, exact set_model/config echo and real inference verified for gpt-5.6-luna. Other models and config-mode changes remain unqualified. |
| Native notification attribution | Copilot passthrough identifies the session, optional subagent and timestamp, but no originating turn. | Bounded redacted session notices retain all admitted details; typed delegation/artifact/compaction projection is withheld. | Delayed same-session regression; P1 require a provider-origin turn correlation contract. Arrival time and optional payload IDs cannot establish it. |
| Usage | Standard prompt usage and context usage; native assistant usage, AI-unit checkpoint. | Token/counter metadata with source; multiplier and nano-AI-units distinct from USD. | Real token counters retain GitHub provenance; authoritative per-turn USD is unavailable. External included-credit snapshots are separate, with additional cash billing disabled. CLI requested-cost coverage fails closed when unknown. Never double-count passthrough. |
| Subagent activity | Native started/configured/completed/failed, model and tool IDs, token/call/duration stats. | Bounded structured activity retaining attribution and model-selection details. | Source and unit fixtures; live UI attribution pending. |
| Task file changes / diffs | Standard tools carry locations/diff content for create/edit/str_replace/apply_patch. | Shared ACP tool activity retains bounded rawOutput, inputUpdated and the first validated relative location. Structured tool content diffs/images, rawInput and secondary locations are dropped; no complete diff presentation is claimed. |
Real denied-create wire contains a diff, but wire presence is not runner/UI preservation. P1 add typed, bounded diff/image content and all validated locations with tool/session provenance; Product file-edit/validation and downloadable-file presentation pass; rich diff rendering remains unqualified. |
| Provider workspace files | session.workspace_file_changed.path is relative to provider session workspace files, not task cwd. |
Validated reference tagged provider_session_workspace, resolution required. |
Source + traversal tests; P1 safe file retrieval/upload. |
| Images / binary artifacts | Prompt image input; native content-addressed binary_asset base64. | Hash/length-validated metadata references; bytes not blindly read from disk or emitted in activity. | Source + unit digest tests; P1 durable artifact storage; >32 KiB payload provider omission remains. |
| Background settlement | Standard prompt waits for idle in tested attached async-shell case; lossy native idle/receipt also exist. | ACP terminal result remains authoritative; raw event cannot end turn. | Offline attached and real-service finite detached commands completed before end_turn; marker verified through cleanup. A deterministic detach:true case now reproduces end_turn before the finite command completes; see the 2026-09-29 blocker. Native detach is now rejected before permission delivery; comprehensive background qualification remains blocked pending unproven tool/platform variants. |
| Compaction/context | Native compaction lifecycle/token counts/context git metadata. | Safe bounded counters/status, immutable workspace binding. | Source + projection tests; raw summary/private custom instructions omitted. |
| Goals / remote/schedules | Native autopilot/objectives/remote/schedule facilities; no qualified ACP goal protocol. | Unsupported through this profile; remote disabled. | Source; P2 separate governance review before control exposure. |
Unused event and field accounting
The inventory is exhaustive for the pinned native event schema: 22 selectively
subscribed events, 15 standard-ACP projection events, 111 native passthrough events
not subscribed, and 2 provider-blocked events. Its fields array names every
native data field; fieldCoverage records each field's disposition and follow-up.
Preserving one standard ACP projection does not imply all native fields survive.
Reasons and priorities are explicit:
| Group | Reason and follow-up |
|---|---|
| Standard text/reasoning/tool/plan/config events | Standard ACP already transports the user-facing content. Additional native metadata is not assumed preserved. P2 compare native field inventory against normalization before adding fields; avoid duplicate output and raw reasoning. |
| Standard ACP parent tool content and locations | Confirmed shared normalization gap: structured content diff/image blocks, rawInput, and locations after the first safe relative path are not carried into canonical tool events. Only bounded rawOutput, inputUpdated, tool lifecycle/identity, and that first path survive. P1 introduce validated diff/image artifact schemas and preserve every safe location with attribution; retain secret redaction and workspace containment rather than forwarding raw provider objects. The retained offline diff proves harness exposure only. |
| Native user/plan/elicitation requests and completions | No qualified ACP responder/correlation acknowledgment. The 2026-09-29 agent/plan probes prove these two native tools unavailable on the pinned offline surface; the connection guard fails closed if any native request appears. Add a versioned upstream responder before displaying an answerable UI. |
| Native permission authorization internals, sandbox decisions, recovery | Standard permission callback is the policy decision boundary. P1 collect sanitized denial diagnostics; never let native carried-forward/assent events authorize actions. |
| Native usage diagnostics omitted from subscribed assistant.usage | Quota snapshots, reasoning summaries, fusion/RTE payloads and upstream service/cache diagnostics are not normalized. P2 type/redact useful performance details; quotas and model multipliers cannot substitute for verifiable dollar spend. |
| Native compaction summaries/checkpoint paths/custom instructions | Avoid copying private instruction/summary bodies or treating provider paths as task paths. P2 add explicit safe metadata schema/artifact retrieval where useful. |
| Native model-cache checkpoint data / premium request total | Complex cache state not normalized; premium request totals are not USD. P1 document billing provenance before integration. |
| Native artifact metadata/bytes | Base64 is verified then omitted; arbitrary nested metadata has no current UI contract. P1 safe artifact upload with content type/size/path policy, not a guessed cwd join. |
| Native hook/extension/skill events | Ambient hooks/extensions are disabled and runner-owned skills have their own attribution. P2 typed activity if a verified owned extension needs it. Two context-delivered types are blocked upstream. |
| Native MCP auth, headers, dynamic lists, reconnect and tool events | Only runner-owned bridge is admitted. P1 sanitized bridge availability diagnostics; do not surface OAuth/headers as unrestricted interactions. |
| Native remote/handoff, schedule, canvas, fusion/factory, memory/indexed-search, UI-ephemeral events | No corresponding admitted ACP control/resource or typed product surface. P2 investigate useful artifact/subagent projections separately; disabled remote authority stays disabled. |
| Native raw user/system messages, assistant lifecycle/retries, streaming internals, tool progress | Standard ACP provides primary conversation/tool lifecycle; duplicate/private bodies intentionally not subscribed. P1 audit lost meaningful progress/retry metadata with sanitized bounded samples. |
| Native external-tool/sampling/limits callbacks | No qualified ACP responder. P0 prove no unresolved request on admitted model/tools; otherwise keep release unqualified. |
| Native capability/model/session lifecycle/config notices | Initial handshake and normalized session/config are admission authority. P1 detect capability/model drift and fail closed rather than treat a notice as authorization. |
Subagent subscribed fields are preserved within the declared text bounds;
truncated display strings carry an explicit truncation marker. Empty native
pending_messages.modified and session.background_tasks_changed events have no
queue/task list to preserve; their projection explicitly says refresh unavailable.
Native context repository/git-root strings, completion receipt finalTool,
compaction summary internals, artifact metadata, nested cache state, and native
question/plan contents are individually marked in the inventory. No exposed
native field should be read as silently supported merely because its event name
appears in the subscription.
Deterministic verification and retained wire
The reusable scripts/probe-copilot-acp.py verifies the pinned executable before
running it with a fresh environment and a deterministic loopback OpenAI-compatible
fixture. COPILOT_OFFLINE=true; no credential is inherited. It bounds fake model
calls and process lifetime and removes its owned workspace. The fixture model ID
does not verify any GitHub model's availability.
python3 packages/paperclip-runner/scripts/probe-copilot-acp.py --package-root /path/to/copilot-darwin-arm64/package --scenario deny-write
python3 packages/paperclip-runner/scripts/probe-copilot-acp.py --package-root /path/to/copilot-darwin-arm64/package --scenario attached-shell
node --test packages/paperclip-runner/scripts/materialize-copilot-binary.test.mjs packages/paperclip-runner/scripts/build-copilot-distribution.test.mjs
pnpm --filter @paperclipai/paperclip-runner exec vitest run src/drivers/acpx/copilot-events.test.ts src/drivers/acpx/copilot-profile.test.ts src/drivers/acpx/copilot-evidence.test.ts
pnpm --filter @paperclipai/paperclip-runner exec vitest run src/drivers/acpx/copilot-extension-adapter.test.ts src/drivers/acpx/copilot-registry.test.ts
Retained real-binary evidence:
- Denial wire: request ID 0; reject_once; failed tool update; target file absent after end_turn. This narrow case did not reproduce #4537.
- Attached-shell settlement wire: two-second async attached command; marker written; output consumed through read_bash; idle followed by end_turn. This narrow case did not reproduce #4743.
- All three materialized platform executables match pinned digests. Only ARM64 wire behavior was exercised.
These are local offline conformance probes, not product E2E or live GitHub qualification. No screenshots were produced because no product UI was exercised. The later live sections record explicit authentication, exact-model inference, file validation, semantic tools, plan approval and controller-restart evidence. Remaining qualification blockers include the failed standalone question marker, every native blocking question/plan mode, restrictive permissions across native tools and configuration, provider-death recovery, active cancellation, multi-company isolation, rich artifact/diff UI, macOS x64 execution and Linux x64 Daytona E2E. Authoritative per-turn USD remains unavailable; external included-credit reconciliation is distinct from cost-limit coverage. Retry with another pinned release if a blocking interaction or settlement/denial case fails; do not suppress the interaction to obtain a pass.
Build-owned native distribution
buildPinnedCopilotDistribution({ outputRoot }) in
scripts/build-copilot-distribution.mjs downloads the exact platform npm archive
from registry.npmjs.org, verifies its pinned SHA-512 integrity, and admits only
the four expected regular tar members. Traversal, links, PAX overrides, duplicate
entries, bad checksums, hidden trailers, and oversized input fail admission. The
binary is independently checked against the pinned SHA-256 and exact size before
it is materialized; no install script, npm launcher, or downloaded executable runs
during this build. outputRoot is the selected pack's exact
provider-assets/copilot/<platform>-<arch> directory.
The factory resolves those assets from the runner's verified package authority, including the descriptor-loaded sidecar path, then the native verifier makes a private executable lease and fresh extraction cache. Callers cannot choose a runtime binary or distribution root.
On 2026-09-28 the strict archive reader verified the actual pinned archives for
all three platforms. A fresh macOS ARM64 registry download completed the full
builder, returning the profile digest above and closure
sha256:fb3b367a45cd76122fe931521fa2a18adf234ba944fc302db9e10e005e57037e.
The temporary output was removed after verification. This packaging proof used
no model credentials, executed no provider turn, and incurred $0 model spend;
it does not qualify either local product behavior or Daytona execution.
The Copilot branch connects all three closed registries: profile installation selects the pinned native verifier, profile extensions advertise only the 22 selected native event types and create the Copilot adapter, and candidate packs select the verified archive builder. Registry conformance checks the complete subagent field projection through the shared turn binder, attribution, canonical schema validation, meaningful display details, and stale/cross-session rejection. Admission error classification distinguishes missing authentication, account or organization denial, and unavailable explicit models using fixed safe messages; unrelated runner integrity errors keep their original classification.
Complete provider-pack proof
The retained packaged-launch evidence
records clean source revision 5272fc6398d42344d1888a3f97ca6909684eefbf,
provider-pack digest sha256:4ba17b2455b0ab92cfe6ee223f77708379fe219792ccb66dbdef70060e6e22e1,
the profile and native closure digests, and the exact protocol-1 initialize response.
The complete pack was built with standalone Node 24.19.0. Its packaged
verifyAcpxProfileInstallation registry acquired a private native command lease,
launched Copilot 1.0.88, preserved numeric request ID 0, and observed clean EOF
settlement. The native terminal-login command path is sanitized in the fixture.
The smoke uses COPILOT_OFFLINE=true and an explicitly configured loopback
metadata-only provider; that server received zero requests, and no prompt was
sent. This low-level packaging test deliberately does not claim production
authentication or model availability. The real host separately rejects absent,
blank, or NUL-containing explicit credentials before opening a command lease;
ambient GH_TOKEN and GITHUB_TOKEN cannot satisfy admission. A host regression
verifies that this failure releases ownership and permits a subsequent explicitly
bound retry without spawning a provider during the test. The controller now mints
a provider/session binding from explicit credential names; the sidecar rejects
unbound ambient credentials and removes the binding before native launch.
Caller-supplied binding markers cannot override the controller-generated value.
The final runner spawn allowlist now preserves the selected credential and
binding through local and remote launch specifications; regression tests cover
the complete controller-to-launcher-to-sidecar boundary. Pending direct product
backends reject before driver construction. Qualification remains available
through the existing host-controlled runnerd CLI.
Probe attempts are accounted for: an initial smoke client closed stdin before
initialize completed and was corrected; a bare unauthenticated initialize then
hit the 20-second deadline; a metadata-fixture initialize passed; the final pack
was rebuilt with the authentication preflight and passed again. After rebasing
onto foundation 5aeebb20c, the complete pack was rebuilt and the fifth initialize
probe passed, with numeric ID 0, clean EOF, and zero fixture HTTP requests. All
five attempts were local with no credentials or inference. After the final
foundation f80c312cd and Copilot review fixes, the complete pack was rebuilt
and a sixth initialize probe passed with the same results. After foundation
7721662f2 fixed the final spawn boundary, the pack was rebuilt from the source
above and a seventh initialize probe passed. All seven probes used $0 model and
infrastructure spend. There was no Daytona
deployment. The candidate remains unqualified.
Final focused checks at the source revision above passed: 200 Copilot/provider-host,
environment, backend-admission and durable-control-plane tests (including four
retained-evidence cases), 12 strict
builder/materializer, candidate-registry and probe-cleanup tests, all six Daytona
image-content tests, and the runner
TypeScript build including generated schema checks and verified sidecar bundles.
The evidence update itself passed the four evidence cases again. The complete
pack remains inspectable at /tmp/paperclip-copilot-launch-boundary-pack-20260928
on the build host; the sanitized tracked fixture provides the portable proof.
The exact Docker resolution command, seeded from the tracked lockfile, produced
650e23d20e967bcfbfced888e131199b9a06e66a1ba4f64cfb68383b59def4a8, matching
the reviewed image pin; the subsequent frozen runner install passed. Generated
lock changes remain uncommitted. These checks do not substitute for live GitHub
or product/Daytona qualification.
/path/to/standalone/node packages/paperclip-runner/scripts/build-provider-pack.mjs /absolute/provider-pack --candidate-providers=copilot
/absolute/provider-pack/node_modules/node/bin/node packages/paperclip-runner/scripts/copilot-provider-pack-smoke.mjs /absolute/provider-pack
Both Copilot builder scripts are explicit Daytona image hash inputs. The selected
copilot candidate also changes image identity, while manifest qualification stays
pending. Linux x64 binaries are pinned and buildable; live Linux/Daytona behavior
still requires the qualification cases above.
Authenticated qualification preparation (2026-09-28)
The sanitized authenticated discovery
uses the previously verified native pack and a dedicated, short-lived personal
Copilot Requests token, explicitly bound as COPILOT_GITHUB_TOKEN. The token has
no repository write authority. No credential, account identifier, provider
session ID, or private path is retained in that fixture.
Three metadata-only sessions were created and closed: catalog discovery,
session/set_model, then session/set_model plus session/set_config_option
with exact gpt-5.6-luna echo. The native catalog advertises 21 model IDs and
marks Luna enabled. No session/prompt request was sent. This verifies auth and
model selection only. The reproducible metadata probe is
scripts/discover-copilot-acp.mjs; it refuses any outbound method outside its
closed initialize/new/select/close list and denies unexpected inbound requests.
The initial paid qualification reservation is $2 within the provider's $25
allocation and shared $100 budget. The account dashboard baseline is 0 of 1,500
included AI credits; additional paid usage is disabled with a $0 cash budget.
Included credit consumption must still be reconciled and reported separately
from cash charges. The sanitized first live proof
records one get-task-context turn completed in 33.383 seconds: one authenticated
semantic call and successful result, completed terminal, complete transcript and
mock-only boundary all pass the unchanged canonical scorer. The first attempt
failed before prompting because a stale Rust binary omitted credential binding;
the fresh source-built release binary fixed admission. The live attempt's
post-turn package-provenance lookup failed, so the same retained artifact was
scored offline after regenerating the exact source tarball. Both original
failure records remain inspectable; recovery sent no additional model prompt.
The sibling launcher now validates that tarball before paid work.
The receipt reports 24,258 input tokens, 11,781 cached input tokens and 441 output
tokens. Its $0.00326022 catalog estimate is not a GitHub charge. The runner's
providerRequests: 1 is a terminal receipt count, not an observed count of
upstream HTTP requests. GitHub still displayed 0/1,500 credits after this small
turn; lag or rounding prevents exact credit attribution. Additional paid usage remains disabled with a $0 budget; that billing boundary
does not imply zero consumption of included credits.
The pinned CLI documents --max-ai-credits as a soft cap, minimum 30 credits.
Its ACP startup/session implementation does not propagate that option into
sessionLimits, unlike the native interactive/server paths; ACP exposes no
budget configuration option. Treat provider-enforced per-session credit limits
as unavailable through this pinned integration (P1 follow-up: upstream ACP limit
support). The 60-second deadline, $0.50 declared envelope and $2 reservation do
not become a per-response hard cap. All profiles remain pending until the full
local and Daytona qualification succeeds.
First local Product result and accounting defect (2026-09-28)
The retained local Product proof
records extended-harnesses.runner-acpx-copilot.local.hello-complete from committed
source bcc9c638a25b91b84065f12633f083bd4f7a689f. The first attempt passed all six
behavior matchers in 38.436 seconds, with one provider turn, no automatic retry,
and successful cleanup. Browser evidence shows the exact completion marker once,
the task marked Done, the Copilot agent, and expandable tool activity. The active
turn deadline was 120 seconds within the existing $2 reservation. This is one
basic Product case, not full local or Daytona qualification.
The original result is retained unchanged with digest
sha256:92c8aff3960e443be0c009c891d49d285476c3d6b67999fe97bb323076c4dc8b.
It exposed a shared accounting defect: model-family inference labeled the biller
openai, while missing provider cost became costUsd: 0, costStatus: reported
and billing.complete: true. Those fields are invalid accounting evidence. The
correct biller is GitHub, provider USD cost is unknown, and the retained behavioral
pass must not be interpreted as accounting qualification. The shared server fix
was incorporated before the following file and question cases. The original
result also lacks source SHA fields; the independent launcher manifest records
the exact committed source above.
After this case, GitHub displayed 1 of 1,500 included AI credits consumed across the successful protocol and Product turns together. Exact per-run credit use remains unknown. Additional usage is disabled, with $0 of the $0 cash budget spent; included-credit consumption is reported separately. The Product receipt contains 31,332 input, 15,451 cached input and 337 output tokens, without a verified USD receipt.
The shared eval CLI now preserves completed provider outcomes while returning a
separate nonzero accounting failure when cost coverage is unknown or its bound
is exceeded. The sibling scorer retains semantic assertions under
accounting_failure; roster and campaign orchestration stop subsequent queued
cases. External dashboard reconciliation does not override the CLI budget
result. Profiles remain pending until the outstanding qualification matrix passes.
Merged-source local Product coverage (2026-09-28)
The retained merged-source proof
records both successful behavior and failures without rewriting original results.
The file and question cases used source ee9536001fbe733b2386dd3379730a4e0be59488,
an immutable runnerd whose Rust tree matches that source, and verified native
assets. A separate complete compiled pack has digest
sha256:809ea6bef2fc1122ef214840d119f37854598007ff7449189027294b0802a681.
Targeted validation passed 67 provider/contract tests, 11 packaging/script tests,
93 Product harness tests, and the TypeScript/verified-sidecar build. The packaged
offline registry launch also passed without a prompt or fixture HTTP request.
| Local case | Retained result | Accounting and remaining limits |
|---|---|---|
| File edit and validation | 7/7 matchers passed in 59.696s. Native shell output proves the file was edited and compared successfully; the downloadable file, exact marker and Done status were visually checked. | Correct GitHub biller, unpriced, absent USD field, incomplete billing. Native raw output contains exit code 0, but normalized typed exitCode is absent; follow-up P2 is preserving this structured result. |
| Question and continuation | 4/6 matchers passed in 102.939s. Cobalt was selected through the real question card and continuation reused the same provider session. The final output was literally [terminal marker], so the exact-marker checks failed. |
Original run-log events and the browser screenshot confirm provider behavior, not public redaction. No retry or grader relaxation. Continuation is GitHub/unpriced. The paused first run incorrectly said reported with no cost and zero counters; separate fix c276496e4 keeps absent cost unpriced and passed 13 accounting tests. |
| Plan approval | Failed after 9.531s during embedded PostgreSQL bootstrap, before any provider call. | Host semaphore exhaustion was independently confirmed. The failed attempt is retained and does not qualify plan behavior. |
| Controller restart | Not executed. | Held before launch because the same host resource exhaustion affected other Product and DB checks. |
The two inference cases each had a 120-second active deadline, a 300-second outer deadline, a $2 reservation and zero automatic retries. GitHub's displayed included credits moved from 1 to 2 after the file case and from 2 to 3 after the question case. Additional paid usage stayed disabled with a $0 budget and $0 cash charge. Display deltas are not exact per-request credit receipts; provider USD remains unknown. At that historical checkpoint, five provider turns were observed; the underlying HTTP model-request count is unavailable. No additional paid attempt is running. Later plan, restart and native risk-probe results follow below. This checkpoint remains unchanged as historical evidence; its failed question result is not erased.
Final shared-source packaging checkpoint
The final pack proof
uses committed source 8aa867b64d5fc2fd62cff110bd000addf5dc54de on foundation
f063fbf2b, including the paused-cost and bounded native-copy fixes. Pack digest
is sha256:627992ea80e8be7154d07cbc9925b781519199e25690b02d4a044f44342e6bd8.
Two clean resolutions from the tracked manifest graph and lock produced identical
dependency bytes, SHA256 aa97f89ba8a7c63573114dda54895523d316df67956c65eeccbc89d3166bb1b4;
the frozen filtered install passed and no generated lock is committed. The
packaged registry probe again initialized numeric request 0 and closed cleanly,
with no credential, prompt or fixture HTTP request. TypeScript and sidecar build,
68 provider/contract tests, and 11 builder/script tests passed at this checkpoint.
The refreshed immutable runnerd has the matching Rust tree and digest
sha256:e6a9fb5170b76a49b8411834b3706e8edf8f1a1ae85ad13b55368158aa7f67a0.
This packaging proof does not rerun or supersede the live results above. Additional
Product starts are held while host PostgreSQL semaphore capacity is restored.
Real-service permission and detached-command probes
Two subsequent single-turn probes used exact gpt-5.6-luna through the final
pack's verified native command lease, with isolated configuration, no ambient
credentials or MCP servers, and a dedicated explicitly bound token. Probe source
47eed960b380e8c8054eb19985aaefeabc6336c3 is retained in
scripts/qualify-copilot-acp.mjs. Five credential-free probe tests include actual
JSON-RPC framing, numeric request ID 0, native option identity and process reaping.
Each live probe had a $2 reservation, 120-second prompt and 180-second outer
deadline, an awake supervisor, and zero retries.
- Denied write: Copilot requested a native file edit at 7.177s. The client returned its exact
reject_onceoption for request ID 0. The turn ended at 7.187s and all 79 filesystem observations through 12.837s remained absent. Native exit and owned process-group cleanup passed. - Detached command: the native tool call explicitly requested
mode: async,detach: true. Only the exact finite three-second marker command receivedallow_once. Native output confirmed the detached shell exited 0 at 10.083s; the marker was present before the 10.675s terminal response and remained correct through cleanup.
These are actual GitHub-service results, distinct from the earlier loopback fixtures. They qualify these two narrow file/command oracles only. They do not prove all tools, an arbitrarily long detached process, governed Product approval surfaces or Daytona execution. Authoritative USD remains absent. GitHub's display was still 3/1,500 included credits after denial; display granularity or delay prevents a zero-use claim. The post-command dashboard also remained at 3/1,500 included credits, with additional usage disabled and $0 cash charges. Exact per-probe credit use remains unknown. The overall candidate remains pending.
Semantic plan after host capacity recovered
A separately authorized Product plan attempt
passed all six matchers in 45.625 seconds at source
c06fc5fccc88f5816450434493451b9d2d339125, using the final provider pack and updated
immutable daemon. The original PostgreSQL startup failure is retained separately;
this was one explicit new attempt with no automatic retry. Browser review shows
the complete Plan revision 1, a confirmation targeting that exact revision,
the approval message and the exact terminal marker once.
This is Paperclip semantic planning. It does not enable Copilot's unwired native
exit_plan_mode responder. The approved continuation deliberately opened a fresh
session after the adapter configuration changed and forceFreshSession was
requested, so this case does not prove warm-session reuse. Both paused and
completed runs now correctly report GitHub and unpriced with no USD field,
including the paused run's zero normalized token counters. Cleanup passed and
the retained fixture process audit found no remaining owned processes.
GitHub subsequently displayed 5/1,500 included credits, an aggregate increase of 2 from the snapshot before the denied-write, detached-command and two-turn plan batch. Display delay and granularity prevent allocation among those calls. Additional usage stayed disabled with a $0 budget and $0 cash charge. Provider USD remains unknown; this external reconciliation is not an authoritative per-turn cost receipt.
Controller restart and pending question recovery
The retained restart proof
passed all six matchers in 50.875 seconds at source
19ca0f558d3aebddedc6ff14836ff3e71498e68e, using the same final pack and immutable
daemon. The exact pending Cobalt/Amber interaction remained visible after server
restart. The board selected Cobalt; the continuation reused the same persisted
provider session and emitted the exact terminal marker once. Screenshots show
the recovered question, selected answer and Done task. This proves controller
reconnect with a preserved session, not reconstruction after provider death.
Both paused and completed receipts are GitHub/unpriced with no USD field. The
continuation reports 49,436 input, 44,075 cached input and 418 output tokens.
Cleanup passed, the bounded supervisor exited successfully and an exact owned-root
process audit found no retained processes. This was one explicit attempt with
two provider turns, no automatic retry, a 120-second active deadline, a 300-second
outer deadline and a $2 reservation. The earlier standalone question's literal
[terminal marker] failure remains unchanged and continues to block its cell.
Eleven provider turns have now been observed across the canonical, Product and native risk probes; the number of upstream HTTP model requests is unavailable. GitHub displayed 5/1,500 included credits both before and after restart, with additional usage disabled and $0 cash charge. Display delay and precision mean that the unchanged counter cannot prove zero included-credit consumption. This external reconciliation remains separate from unknown provider USD. No further paid prompt is authorized or running from this branch. The profile remains pending.
Final provider checks after these evidence updates pass 74 focused TypeScript
tests and 24 packaging/discovery/risk-probe tests. The discovery regressions cover
rejecting the mutable auto model selector, releasing leases on early setup
failures, and rejecting pending RPC calls immediately after native exit or
malformed output. These probe-script changes do not alter the final runtime pack.
Review tightened the denial probe to require the native rawInput.fileName to
resolve to the exact marker target. An unrelated edit or a command merely
mentioning the marker cannot satisfy the oracle. The original paid denial names
that exact target; offline replay of its retained wire and marker observations
passes the corrected oracle. The fixture records the original evidence and oracle
script digests. No additional provider prompt was sent.
Linux image initialize-only proof
The Linux x64 pack proof
uses image ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5457769683fd310223d3b0d4f1ed9a6cf341bdb16514746b3aaeabca2e888fee
from the same source 8aa867b64d5fc2fd62cff110bd000addf5dc54de. Its platform-specific
pack digest is sha256:b11984fe02bd5d5a36b01ed559d2f4c765663df09a46117d3092b1183be08ba4.
The verified executable matches the Linux pin, initializes Copilot 1.0.88/ACP 1
with request ID 0, and exits 0 after stdin EOF. The maintained smoke script ran
under network-none, a read-only root and private tmpfs, with no provider
credentials and zero fixture requests or inference. Missing-token production
preflight still rejects. Two earlier operator invocations used the wrong image
repository or Node path and failed before provider launch; both are retained.
This is Linux packaging evidence, not an authenticated Daytona Product run or
model qualification. The profile remains pending.
Bounded native tool observations and protection evaluations
The sidecar and direct TypeScript driver share one projector under drivers/acpx,
which projects an allowlist of active-turn ACP tool/permission fields into
existing provider.notice.recorded details and provenance: validated relative target,
request/tool identities, command SHA-256, explicit mode/detach, and linked shell
start/completion with provider-reported exit code. Session passthrough notifications
remain uncorrelated and cannot supply this evidence. Native ACP exposes an execution
kind, not a trustworthy bash name in the title. Strings use semantic redaction;
ambiguous identities, exhausted bounds or failed projection make evidence incomplete.
Projection cannot change permission delivery or terminal authority. This is an additive
observation change: Copilot profile v4 is unchanged, while source/pack provenance
changes. Older sessions lacking these notices cannot pass the new evidence oracles.
The manual local copilot-protection Product suite registers an exact browser-denial
case with explicit cancellation and a finite attached-command settlement case. Both
remain unqualified until separately built and run. The denial case expects an unfinished
task/cancelled run; full restrictive workflow completion is still a gap because later
MCP permissions must not be auto-approved from provider-controlled titles.
A separate credential-free real-binary fixture on 2026-09-29 exercised Copilot1.0.88
(executable SHA-256 a9ff8babb10b7e443182ae96a8bc50a9c826ef1c773e1344c396eb5bf7f512c3)
against a loopback synthetic model. It started an eight-second attached async command;
the model immediately returned terminal text while the process was still live. The
binary itself issued a read-shell update, and process absence plus the marker were
observed before the ACP prompt result. The model never called read_bash. The private
copilot-v4-protection-preparation/adversarial-attached-v2 receipt retains monotonic
ordering and process identity. This verifies one pinned finite attached scenario and
is distinct from detached-policy rejection or comprehensive background qualification.