Files
PaperClipAI/tests/runner-e2e/everyday-cases.ts
T
DottaandPaperclip aa8fc86331 feat(connections): prefer native apps and ask users to choose external providers (#13941)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents governed access to external services.
> - Native connections should remain the first choice for a supported
app.
> - Other apps may be available through an external MCP provider.
> - The user must know which external provider handles the connection
and choose it before setup.
> - This pull request adds ranked alternatives and server-authored
instructions to connection search.
> - Agents can follow the returned instructions while Paperclip
validates saved choices and access.

## Linked Issues or Issue Description

Related: #13879, which fixed inline MCP provider setup. This PR adds
discovery and provider selection on top of that work.

**Subsystem affected**

Cross-cutting: shared connection contracts, server search and intent
services, native runtime, CLI, inline setup UI, and evals.

**Problem or motivation**

An agent cannot offer a clear external-provider choice when Paperclip
has no native connection for an app. Adding provider-specific branches
to the core prompt would make those instructions harder to maintain.

**Proposed solution**

Prefer a native connection. Otherwise return verified alternatives in
Composio, Arcade, Executor, Zapier order. Include an external-service
disclosure, a question with None, and the next instruction in the search
result. Validate the saved human choice before creating a selected
fallback setup card. Reuse existing provider accounts and verify
underlying app access separately.

**Alternatives considered**

Do not silently choose a provider. Do not claim that broad execution
tools prove support for every app. Reuse existing questions and
connection intents rather than add another connection model.

**Roadmap alignment**

Extends the existing MCP Tool Gateway & Apps and Agent evals & feedback
capabilities. The MCP aggregators experiment remains the gate. No
duplicate provider-routing PR was found in the public search.

## What Changed

- Add a dated support index and authorized cached-tool evidence for
external routes.
- Return provider questions and next-step instructions from
`connections_search`.
- Preserve pending choices and declines across continuation. Validate
company, task, agent, human, app, and current route eligibility.
- Carry the selected app into new setup and account reuse, validate
explicit provider requests against persisted human messages, and
distinguish provider readiness from app authorization.
- Sync native, MCP, REST, and CLI contracts. Keep core agent
instructions provider-neutral.
- Add production-component Storybooks, focused database tests, and three
real-agent browser eval cases.
- Record the plan, observed failures, fixes, passing evidence, and
acceptance limits.

## Verification

- Latest head `586f0e6cd`: 54 checks passed, 2 skipped; Greptile 5/5 and
all review threads resolved.

- After rebasing on master `18dac1e1e`: 64 focused shared, validator,
route-contract, and database tests passed; server typecheck passed.
- Embedded-browser test drive on the rebased head: native Jira card,
HubSpot external-provider question, Arcade account reuse, one actual MCP
read against a local synthetic fixture, reload persistence, and None
preventing further calls. A real OpenAI-backed agent performed discovery
and continuation.
- UX observation: the agent initially combined mutually exclusive
request fields; the server rejected it and the agent recovered without
changing access. This extra retry remains visible in the transcript.
- After rebase: 23 focused eval grader/catalog tests, affected
TypeScript checks, token gates, production UI build, and Storybook build
passed. Full local tests are intentionally excluded at the maintainer's
request.
- Before rebase: four browser/real-agent attempts passed: native Jira,
None, and reuse of the second provider on two Codex profiles.
- Browser evals used an isolated deterministic MCP fixture through the
real Paperclip gateway. They do not prove production compatibility with
all four providers.
- Review `Apps / Connections / Provider choice` in Storybook. Choose
Arcade, continue through Access, and verify the app name,
external-service disclosure, and URL configuration.
- The detailed verification report is
`doc/connections/2026-09-23-aggregator-routing-verification.md`.

## Risks

- The public support index is finite and can age. Account capability and
app authorization still require verification after selection.
- Existing installed-tool permissions remain in effect. Provider choice
is not a new execution permission boundary. An early Mini attempt
skipped search; clearer provider-neutral instructions made the targeted
rerun pass. This is not a measured reliability rate.
- Explicit requests skip provider confirmation only when a clear
persisted human message or saved provider choice supports them. Other
phrasing falls back to confirmation; the agent query alone is not
consent. These routes do not add tool permissions.
- No database migration or legacy Composio broker is added.
Real-provider acceptance remains separate from fixture proof.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser tools. The exact deployment variant and
context-window size are not exposed in this session. Product evals
separately used the repository's primary Codex and Codex Mini profiles;
those agents supplied test behavior, not independent provider
compatibility proof.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 16:36:49 -05:00

121 lines
6.8 KiB
TypeScript

import type { RunnerProfileFixture, RunnerTaskFixture } from "./types.js";
/** No fixture completion/API directions: the normal execution prompt owns that contract. */
export function productionStoryProfile(
profile: RunnerProfileFixture,
): RunnerProfileFixture {
return {
...profile,
buildAgent(input) {
const payload = profile.buildAgent(input);
return {
...payload,
name: `Studio Lead ${input.executionId}`,
role: "ceo",
title: "Studio Lead",
capabilities:
"Builds small projects, delegates implementation, reviews delivered work, and hires teammates when requested.",
instructionsBundle: {
entryFile: "AGENTS.md",
files: {
"AGENTS.md":
"You lead a small software studio. Help the user build useful small projects. Respect their requirements and verify the delivered work.",
},
},
};
},
};
}
export const SLUGIFY_REQUIREMENTS = `Build a small dependency-free Python command-line tool in slugify.py. It accepts one positional string and prints its slug: trim whitespace, lowercase, replace each sequence outside ASCII a-z and 0-9 with one hyphen, then trim hyphens. Expose slugify(text) for reuse. Include a README and automated tests. Deliver the source and tests as a downloadable ZIP. Work in the project workspace.`;
export const SLUGIFY_REVISION = `Add an optional --separator argument that accepts either - or _. Keep - as the default and preserve the earlier behavior. Deliver an updated ZIP, keeping the previous download available.`;
export const LATE_REQUIREMENT = `Also support --max-length as a positive integer. Truncate the final slug to that length, then remove any trailing separator. In particular, input " Queue--Ready!! " with --max-length 7 must print "queue-r". Preserve the default behavior.`;
const definitions = [
["provider-decline", "Decline external providers", "Connect HubSpot so you can read my recent contacts. If I choose None for now, explain that you could not retrieve the contacts and finish without connecting.", 2],
["provider-second", "Choose and reuse the second external provider", "Connect HubSpot and read my recent contacts. Report the contact names and verification code returned by the service.", 2],
["provider-native", "Prefer a built-in connection over external providers", "Connect Jira so you can read my recent issues. If I choose Not now, explain that you could not retrieve the issues and finish without connecting.", 2],
[
"build-revise",
"Build, download, and revise a project",
SLUGIFY_REQUIREMENTS,
2,
],
[
"delegate-feedback",
"Delegate implementation and preserve late feedback",
`Have Riley Builder implement the following as one child task. You review the downloaded result when it is ready. Keep implementation with Riley and post a progress note linking the child while work is underway. ${SLUGIFY_REQUIREMENTS}`,
4,
],
[
"agent-review-handoff",
"Delegate work through an agent review handoff",
`Have Riley Builder implement the following as one child task. The child must keep its original Riley assignee throughout. Before Riley finishes, require a native needs_review report with exactly one attention request: kind review, ownerClass agent, targetAgentId set to your exact lead agent id, and a summary naming you as the reviewer. The child must remain in_review while waiting. When the durable review wake arrives, inspect the child task context, approve the review through the native resolve_review tool with decision accept, then finish the parent task. Do not patch the child status, reassign the child, self-approve the child from the parent run, or bypass the review interaction. ${SLUGIFY_REQUIREMENTS}`,
3,
],
[
"hire-reuse",
"Hire one teammate, then reuse that agent",
`Hire exactly one agent named Morgan QA, reporting to you, using the same available AI connection and native runner configuration as you. Have Morgan implement the following in one child task, then review the result. ${SLUGIFY_REQUIREMENTS}`,
6,
],
[
"service-approve",
"Use a connection after approval",
`Use the connected page service to find recent pages and create a short Markdown briefing document on this task. Include the titles and verification code returned by the service.`,
2,
],
[
"service-decline",
"Respect a declined tool action",
`Use the connected page service to find recent pages and create a short Markdown briefing document on this task. Include the titles and verification code returned by the service. If I decline the action, do not try again or request another connection. Instead, give me a brief explanation that you could not retrieve the data. That explanation is the complete allowed fallback; no briefing is required after a decline.`,
2,
],
[
"connection-decline",
"Respect Not now on a new connection",
`Please connect Notion so you can read my recent pages and write a short briefing. If I choose Not now, do not try again or use another service. Instead, give me a brief explanation that you could not retrieve the pages. That explanation is the complete allowed fallback; no briefing is required after a decline.`,
2,
],
[
"recover-controller",
"Recover work after the server restarts",
SLUGIFY_REQUIREMENTS,
3,
],
[
"stop-redirect",
"Stop a task and send a new direction once",
SLUGIFY_REQUIREMENTS,
2,
],
[
"create-skill-studio",
"Create and edit a company skill",
"Create one company skill using this complete SKILL.md content:\n---\nname: release-readiness-checklist\ndescription: A bounded checklist for validating a release before handoff.\n---\n\n# Release Readiness Checklist\n\n1. Verify checks.\n2. Review evidence.\n3. Record the handoff.\n\nUse request key create-skill-e2e-001. After creating it, report the created skill and finish the task.",
1,
],
] as const;
export const everydayTasks: readonly RunnerTaskFixture[] = definitions.map(
([id, label, prompt, expectedRunCount]) => ({
id,
label,
groups: [],
workMode: "standard",
flow: "everyday_workflow",
expectedRunCount,
attemptTimeoutMs: { local: 12 * 60_000, daytona: 30 * 60_000 },
expectedTerminalState: { issue: "done", run: "succeeded" },
buildTitle: (nonce) => `${label} ${nonce}`,
buildPrompt: () => prompt,
buildVisibleMarker: (nonce) => `STUDIO_${nonce}`,
buildMatchers: () => [], // The workflow records independent artifact and lifecycle checks.
}),
);
/** Only these stories execute downloaded Python ZIPs in the pinned oracle. */
export function requiresEverydayArtifactOracle(caseId: string): boolean {
return ["build-revise", "delegate-feedback", "agent-review-handoff", "hire-reuse", "recover-controller", "stop-redirect"].includes(caseId);
}