fix(runner): restore task runtime parity (#12685)

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task view shows a running agent and lets an operator guide that
agent.
> - The merged runner stack lost parts of the accepted task experience.
> - Native event errors could hide current reasoning from the operator.
> - Queued message steering had no server route on `master`.
> - This pull request restores the task-runtime behavior and keeps the
runner experimental gate.
> - The benefit is a visible and steerable native run with durable
fallback behavior.

## Linked Issues or Issue Description

**What happened?**

The task view could stop showing current runner reasoning. The steering
action also failed because the server route was absent. Runner
instruction files were not declared as supported.

**Expected behavior**

The task view must show current provider activity. It must use the live
log when durable native events are empty or unavailable. The operator
must be able to steer a queued message into the active native turn.

**Steps to reproduce**

1. Enable the Paperclip Runner experimental setting.
2. Start a native runner task.
3. Open the task view while the run emits reasoning.
4. Queue a message and select the steering action.

**Paperclip version or commit**

The regression reproduces on `24a674f8858060e77ea1beb50689d26473e91431`.

**Additional context**

Related closed work: Refs #12592.

## What Changed

- Restored the queued-comment steering route for active native sessions.
- Added durable and queue-bound steering acknowledgements for safe
retries.
- Restored runner instruction bundle support.
- Added live-log fallback when native events are empty or unavailable.
- Restored the compact live reasoning ticker in the task view.
- Added a visible temporary-unavailable state when both activity sources
fail.
- Kept the unified Paperclip Runner experimental gate unchanged.

## Verification

- `pnpm --filter @paperclipai/ui typecheck`
- `pnpm --filter @paperclipai/server typecheck`
- `pnpm -r typecheck`
- `pnpm check:token-gates`
- `pnpm build`
- Seven focused test files passed with 140 tests.
- The final steering regression file passed with 12 tests.
- The broad local test run reached unrelated workspace, port, and shared
database failures. The changed-area tests remained green.

## Risks

- The steering route changes queue and run records in one transaction.
Tests cover stale targets, unavailable sessions, lost responses, and
wrong-queue acknowledgements.
- Native events remain the primary transcript source. The live log is
used only when event data is absent or its poll fails.
- The experimental gate still hides and rejects the runner when the
setting is off.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex, GPT-5, with tool use, code execution, and subagent review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes — no
documentation change is required for this regression repair
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
This commit is contained in:
Dotta authored and GitHub committed 2026-09-01 21:17:54 -05:00
1 parent b4f302d040
commit 72b9f92d76
14 files changed
+884 -34

No files matched your search

@@ -13,6 +13,11 @@ export interface NativeRunTranscriptSource {
runtimeMode?: "legacy" | "native";
}
export interface NativeRunTranscriptError {
message: string;
failedAt: string;
}
function isLive(status: string): boolean {
return status === "queued" || status === "running";
}
@@ -30,15 +35,17 @@ export function useNativeRunTranscripts(runs: readonly NativeRunTranscriptSource
[nativeRunsKey],
);
const [eventsByRun, setEventsByRun] = useState<Map<string, HeartbeatRunEvent[]>>(new Map());
const [errorsByRun, setErrorsByRun] = useState<Map<string, NativeRunTranscriptError>>(new Map());
const cursorByRunRef = useRef(new Map<string, number>());
useEffect(() => {
let cancelled = false;
let timer: number | null = null;
const refresh = async () => {
const refresh = async (runsToRefresh: readonly NativeRunTranscriptSource[]) => {
const updates = new Map<string, HeartbeatRunEvent[]>();
await Promise.all(nativeRuns.map(async (run) => {
const errors = new Map<string, NativeRunTranscriptError>();
await Promise.all(runsToRefresh.map(async (run) => {
try {
let cursor = cursorByRunRef.current.get(run.id) ?? 0;
const incoming: HeartbeatRunEvent[] = [];
@@ -56,8 +63,12 @@ export function useNativeRunTranscripts(runs: readonly NativeRunTranscriptSource
}
if (incoming.length > 0) updates.set(run.id, incoming);
cursorByRunRef.current.set(run.id, cursor);
} catch {
} catch (error) {
// Keep the last durable cursor; the next poll retries this run only.
errors.set(run.id, {
message: error instanceof Error ? error.message : "Native run activity could not be loaded",
failedAt: new Date().toISOString(),
});
}
}));
@@ -75,13 +86,27 @@ export function useNativeRunTranscripts(runs: readonly NativeRunTranscriptSource
}
return next;
});
setErrorsByRun((previous) => {
const next = new Map<string, NativeRunTranscriptError>();
for (const runId of retainedIds) {
const error = errors.get(runId);
if (error) next.set(runId, previous.get(runId) ?? error);
}
return next;
});
if (nativeRuns.some((run) => isLive(run.status))) {
timer = window.setTimeout(refresh, EVENT_POLL_INTERVAL_MS);
if (nativeRuns.some((run) => isLive(run.status)) || errors.size > 0) {
const retryRuns = nativeRuns.filter(
(run) => isLive(run.status) || errors.has(run.id),
);
timer = window.setTimeout(
() => void refresh(retryRuns),
EVENT_POLL_INTERVAL_MS,
);
}
};
void refresh();
void refresh(nativeRuns);
return () => {
cancelled = true;
if (timer !== null) window.clearTimeout(timer);
@@ -96,5 +121,5 @@ export function useNativeRunTranscripts(runs: readonly NativeRunTranscriptSource
return transcripts;
}, [eventsByRun, nativeRuns]);
return { transcriptByRun };
return { transcriptByRun, errorsByRun };
}