mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 00:54:38 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task view shows a running agent and lets an operator guide that agent. > - The merged runner stack lost parts of the accepted task experience. > - Native event errors could hide current reasoning from the operator. > - Queued message steering had no server route on `master`. > - This pull request restores the task-runtime behavior and keeps the runner experimental gate. > - The benefit is a visible and steerable native run with durable fallback behavior. ## Linked Issues or Issue Description **What happened?** The task view could stop showing current runner reasoning. The steering action also failed because the server route was absent. Runner instruction files were not declared as supported. **Expected behavior** The task view must show current provider activity. It must use the live log when durable native events are empty or unavailable. The operator must be able to steer a queued message into the active native turn. **Steps to reproduce** 1. Enable the Paperclip Runner experimental setting. 2. Start a native runner task. 3. Open the task view while the run emits reasoning. 4. Queue a message and select the steering action. **Paperclip version or commit** The regression reproduces on `24a674f8858060e77ea1beb50689d26473e91431`. **Additional context** Related closed work: Refs #12592. ## What Changed - Restored the queued-comment steering route for active native sessions. - Added durable and queue-bound steering acknowledgements for safe retries. - Restored runner instruction bundle support. - Added live-log fallback when native events are empty or unavailable. - Restored the compact live reasoning ticker in the task view. - Added a visible temporary-unavailable state when both activity sources fail. - Kept the unified Paperclip Runner experimental gate unchanged. ## Verification - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm -r typecheck` - `pnpm check:token-gates` - `pnpm build` - Seven focused test files passed with 140 tests. - The final steering regression file passed with 12 tests. - The broad local test run reached unrelated workspace, port, and shared database failures. The changed-area tests remained green. ## Risks - The steering route changes queue and run records in one transaction. Tests cover stale targets, unavailable sessions, lost responses, and wrong-queue acknowledgements. - Native events remain the primary transcript source. The live log is used only when event data is absent or its poll fails. - The experimental gate still hides and rejects the runner when the setting is off. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, with tool use, code execution, and subagent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes — no documentation change is required for this regression repair - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge
89 lines
2.4 KiB
TypeScript
89 lines
2.4 KiB
TypeScript
// @vitest-environment jsdom
|
|
|
|
import { act } from "react";
|
|
import { createRoot, type Root } from "react-dom/client";
|
|
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
|
|
import { useNativeRunTranscripts } from "./useNativeRunTranscripts";
|
|
|
|
const eventsMock = vi.hoisted(() => vi.fn());
|
|
|
|
vi.mock("@/api/heartbeats", () => ({
|
|
heartbeatsApi: { events: eventsMock },
|
|
}));
|
|
|
|
function Probe() {
|
|
const { errorsByRun } = useNativeRunTranscripts([
|
|
{ id: "native-run", status: "succeeded", runtimeMode: "native" },
|
|
]);
|
|
return (
|
|
<div data-testid="errors">
|
|
{[...errorsByRun.keys()].join(",")}
|
|
</div>
|
|
);
|
|
}
|
|
|
|
function MultiRunProbe() {
|
|
useNativeRunTranscripts([
|
|
{ id: "failed-run", status: "succeeded", runtimeMode: "native" },
|
|
{ id: "healthy-run", status: "succeeded", runtimeMode: "native" },
|
|
]);
|
|
return null;
|
|
}
|
|
|
|
describe("useNativeRunTranscripts", () => {
|
|
let container: HTMLDivElement;
|
|
let root: Root;
|
|
|
|
beforeEach(() => {
|
|
vi.useFakeTimers();
|
|
eventsMock.mockReset();
|
|
container = document.createElement("div");
|
|
document.body.appendChild(container);
|
|
root = createRoot(container);
|
|
});
|
|
|
|
afterEach(() => {
|
|
act(() => root.unmount());
|
|
container.remove();
|
|
vi.useRealTimers();
|
|
});
|
|
|
|
it("exposes event transport failures and retries terminal runs until recovery", async () => {
|
|
eventsMock
|
|
.mockRejectedValueOnce(new Error("event endpoint unavailable"))
|
|
.mockResolvedValue([]);
|
|
|
|
await act(async () => {
|
|
root.render(<Probe />);
|
|
await Promise.resolve();
|
|
});
|
|
expect(container.textContent).toBe("native-run");
|
|
|
|
await act(async () => {
|
|
await vi.advanceTimersByTimeAsync(2_000);
|
|
});
|
|
expect(eventsMock).toHaveBeenCalledTimes(2);
|
|
expect(container.textContent).toBe("");
|
|
});
|
|
|
|
it("retries only terminal runs whose event request failed", async () => {
|
|
eventsMock.mockImplementation((runId: string) => (
|
|
runId === "failed-run"
|
|
? Promise.reject(new Error("event endpoint unavailable"))
|
|
: Promise.resolve([])
|
|
));
|
|
|
|
await act(async () => {
|
|
root.render(<MultiRunProbe />);
|
|
await Promise.resolve();
|
|
});
|
|
expect(eventsMock).toHaveBeenCalledTimes(2);
|
|
|
|
await act(async () => {
|
|
await vi.advanceTimersByTimeAsync(2_000);
|
|
});
|
|
expect(eventsMock).toHaveBeenCalledTimes(3);
|
|
expect(eventsMock.mock.calls.at(-1)?.[0]).toBe("failed-run");
|
|
});
|
|
});
|