Files
PaperClipAI/tests/runner-e2e/EVERYDAY-WORKFLOWS.md
T
DottaandPaperclip 6c1a75da49 feat(connections): make AgentMail a default connection with inline setup (#14772)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents access to external services.
> - AgentMail needs both a saved key and an inbox assigned to the agent.
> - Chat requests offered a setup link instead of an inline card and
could treat a saved key as complete.
> - Inbox setup also hid address conflicts behind a generic server error
and a separate review step.
> - This pull request makes AgentMail a default connection, adds the
inline card, reduces setup to two steps, and shows conflicts beside the
address.
> - Shared native dropdown styles also give every caret a consistent
inset.

## Linked Issues or Issue Description

**What happened?**

AgentMail requests in chat did not show a usable inline connection card.
Manual setup required extra screens, ignored saved account keys, and
could trap new-address setup in a locked inbox dropdown. Agent selectors
omitted the avatar from the selected value. A taken address could
produce an HTTP 403 from AgentMail and appear as an internal server
error. Native dropdown arrows also touched the right edge of their
fields.

**Expected behavior**

Make AgentMail available as a default connection. Ask for the API key
inline, with a direct link to its provider page. Default human access to
the company and agent access to the requesting agent. Resume the agent
only after an assigned inbox is active. Manual setup should ask for an
agent and email address, then finish. Address checks should run as the
user types. Taken addresses should show clickable alternatives. A domain
dropdown beside the name should prefer a verified custom domain. Setup
should suggest authorized saved AgentMail keys and show agent avatars in
the picker and selected value.

**Steps to reproduce**

1. Ask an agent to connect AgentMail when it has no assigned inbox.
2. Check that an inline API-key card appears and links to the provider's
API-key page.
3. Open AgentMail setup, choose an agent, and request an address that is
already taken.
4. Correct the inline error, refresh, and finish setup with the same
request ID.
5. Inspect native dropdown carets in light, dark, disabled, and
right-to-left states.

Uses the bounded provider-error parser merged in #14768. Related work:
#13256 introduced AgentMail; #14725 expanded connection search.

## What Changed

- Stop recurring email queries for tasks that have no email thread.
Share the query between the thread provider and activity view. Keep
email-task updates and invalidation-based discovery.
- Make AgentMail available without the experimental chat setting. Keep
the catalog, setup and management routes, agent Channels tab, task email
feed, receiving worker, and agent tools available by default. Other
experimental chat providers stay gated.

- Make the email address and copy icon a single clickable action with
the shared Copied! confirmation. Add View inbox linking directly to the
matching AgentMail console inbox, with the address encoded as one URL
path segment.

- Reorganize inbox Settings around the copyable email address, usage
instructions, and receiving status. Move reconnect credentials into a
disclosure and separate the Disconnect action. Add production Settings
stories for active, paused, unassigned-address, revoked, webhook,
long-address, mobile, and reconnect states. Show repair controls when
the inbox has an error. Keep usage instructions tied to an active inbox
with an address.

- Add AgentMail channel intents and an inline key field with the direct
API-key URL.
- Keep setup and retry state tied to the interaction. Require an active
inbox for completion. Preserve company and agent access checks.
- Reduce manual setup to agent selection and email selection. Put the
domain dropdown beside the address and default to a verified custom
domain. Preserve explicit choices across reloads. Keep receiving
settings under Advanced options.
- Check the initial address and edits after a 350 ms pause. Abort
superseded requests and ignore stale responses. Show clickable
suggestions and retain known creation conflicts across reloads.
- Add a company-scoped, manager-only address check using the saved
credential. Search the visible inbox list instead of fetching an
uncreated inbox: live AgentMail retains negative lookups that can break
subsequent access-key creation. Unlisted addresses remain unknown;
creation is authoritative.
- Suggest labeled saved AgentMail keys in both manual setup and the
inline card. Filter by company, provider, active credential, and
current-user grants on the server. Prefer an account key and preserve
the selected key or an explicit new-key choice across refresh. Use
verified scope metadata and bounded concurrent checks for legacy keys.
Never return secret values.
- Catch an inbox-only key before the email step. Allow its existing
inbox only after an explicit choice. Recover old locked drafts at the
key picker. Save the replacement key before retiring an empty draft,
then use a new setup URL so refresh preserves the switched account; stop
if cleanup fails. Preserve already allocated addresses and their
original accounts.
- Use the shared AgentSelect in email setup. Show the canonical agent
avatar in each option and the selected value, including other consumers
of the shared component. Add regression coverage for legacy and current
Lucide agent-mention icon formats.
- Start each catalog Add connection with a fresh setup identity. Honor
Finish setup's exact draft/account/address instead of resuming an
unrelated browser draft. Return Cancel and Done to Connectors and Email
settings to the inbox. Group the task/thread explanation in a How it
Works card.
- Route AgentMail catalog removal through the email inbox control API,
including unfinished drafts. Refresh both the catalog and inbox views.
- Render each inbox management tab separately. Access uses the saved
account grants and agent controls; Conversations and Activity use the
shared persisted email feed. Activity lifecycle actions use the email
API. Reconnect returns to inbox Settings. Conversation failures show a
retry instead of a false empty state. Email delivery recovery stays in
the task.
- Map documented provider address conflicts to a field error. Preserve
actionable messages for other failures.
- Preserve non-secret draft fields across refresh, scoped to the
requested agent. Never save API keys in browser storage. Resume partial
inbox creation with the original agent, address, and request ID.
- Show an already-created address with explicit retry and new-address
recovery instead of locked inputs. Preserve the original inbox and
resumable draft when choosing another address. Distinguish runtime-key
404 errors and log safe provider status/operation/code.
- Apply final agent access once within email setup authorization for a
new account whose original installs are unchanged. Preserve later
permission edits and reused account installs. Support in-place retry of
progress loading.
- Let a failed inline setup change keys after retiring an empty draft.
Persist its replacement setup identity without storing secrets. Recover
a server-saved account when refresh interrupts the save response, while
preserving intentional account changes.
- Render the production setup in Storybook and add error, recovery, and
mobile states.
- Inset native select carets in shared CSS. Preserve custom icons,
listboxes, keyboard behavior, and forced-color controls.
- Add browser regression coverage and an AgentMail Product E2E case with
persisted-state and rendered-card evidence.

## Verification

- Full `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and
`git diff --check` passed after the default-availability change.
- All 485 focused tests passed. These cover setup, management, catalog
and route gates, connection intents, email authorization, Cursor
execution, and the OpenAPI contract. All 39 email integration tests run
with the experimental chat setting off.
- The shared polling change passed four behavioral tests, UI typecheck
and build, and token gates.
- `tests/e2e/agentmail.spec.ts` passed with the actual server setting
off. This full-stack browser test uses simulated provider responses. It
covers catalog entry, saved keys, editable address and domain controls,
creation, conflicts, retry, all management tabs, clipboard feedback, the
provider link, and task email rendering.
- In the live local browser, Add connection reached the editable email
step with the saved account key. The verified custom domain was selected
by default. Both domain choices worked. The existing inbox Settings page
remained available. Both active inboxes completed new mail checks with
the setting off. No new provider inbox or email message was created for
this pass.
- Earlier live provider acceptance covered creation on a verified custom
domain, Finish connecting on the reported draft, successful mail checks
after refresh, and catalog removal of disposable draft and active
connections. Clicking the email address copied the exact address and
showed Copied!. View inbox opened the same inbox in AgentMail’s console.
No email messages were sent.
- Production setup and Settings Storybook builds and interactions
passed. Settings states include active, paused, unassigned, revoked,
webhook, long-address, mobile, and reconnect. Receiving and
revoked-access stories had zero accessibility violations.
- Full local `pnpm test:run` on an earlier revision completed with
14,709 passing, 87 skipped, and four transient failures. All four failed
cases passed in focused reruns without product changes. That serial full
local command was not repeated after each follow-up. The latest-head
full CI suite is the final test gate.
- CI found an obsolete browser assertion that hid every channel when the
flag was off. Updated it to keep AgentMail and the Channels surface
visible while preserving the GitHub chat route gates. All 11 provider
browser tests passed locally after scoping the Channels selector to the
agent sidebar. Two initial local attempts stopped at temporary Postgres
initialization. The passing run used a separate disposable database on
the existing local Postgres server; it was removed after the test.
- Updated the remaining sidebar and aggregator discovery assertions for
default AgentMail availability. Ordinary task fixtures now return no
email thread. All 128 sidebar/task-page tests and all 42 aggregator
tests passed locally.
- Latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`: full CI
passed, with 54 successful checks including Snyk and two intentional
Storybook skips. The CI run is
https://github.com/paperclipai/paperclip/actions/runs/37020833647. A
fresh Greptile review scored 5/5 with no unresolved threads. Live model
evaluations and inbound/outbound email delivery were not run.

## Risks

- AgentMail no longer needs experimental opt-in. Setup still requires a
human to connect an account and assign an inbox. Inline setup creates an
inbox after a human submits a new or saved key. Company access, agent
access, inbox assignment, and completion checks remain enforced.
- AgentMail read APIs cannot prove global address availability. The
visible-list check is bounded to 100 entries and cannot see inboxes
outside the key’s scope. The UI reports this limitation, suggests
alternatives without claiming they are free, and keeps final creation
conflicts inline. Lookup outages show an error without preventing the
authoritative creation attempt.
- Native select CSS affects the whole app. Custom-icon selects and
multi-row lists are excluded. Forced-color mode keeps the browser caret.
- Saved-key discovery uses stored verified scope metadata and checks
authorized legacy credentials concurrently within a shared three-second
deadline. Provider outages mark legacy choices unavailable; users can
still enter another key. Final use rechecks authorization and provider
access.
- No database migration or transport default change. Live connection
remains the default.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, and code execution. The
exact served model ID and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; full-suite
limitation documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (latest head
`b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 10:01:15 -05:00

15 KiB

Everyday Paperclip workflow evals

This manual suite tests useful work through the production browser, public API, native runner, and normal agent instructions. It complements the tightly scripted runner contract fixtures. It does not add a scheduled or default paid run: --all and generic profile selectors exclude it. Select the suite or an exact execution ID explicitly.

Stories and assertions

Story Cases Required evidence
Build a small project and revise it build-revise Download both ZIPs through the UI; independently execute the delivered CLI and import its function; test the revision; retrieve the original bytes again.
Delegate and incorporate late feedback delegate-feedback One child assigned to Riley; send feedback while the child runs; find it in the child history and independently test --max-length in the delivered ZIP. The worker must not execute on the parent.
Hire a teammate and use them again hire-reuse One Morgan QA reporting to the lead, native runner and the same encrypted connection bindings, real child execution, then a second usable delivery from that same agent.
Decide on an installed service action service-approve, service-decline Assign an authenticated local MCP fixture with Ask first; match its tool action and connection ID; no provider call before approval; exactly one after approval and a verified document; none after decline.
Decline a new connection connection-decline Start without service connections; match a Notion connection intent; click Not now; verify the saved rejection, no new connection or repeated request, and an explanation followed by Done.
Request an email address agentmail-setup Enable Chat connectors, ask for an email address, and require a durable AgentMail card for the requesting agent and user. Reload and verify one password field, the direct API-key link, and no access selectors or modal. Decline and verify no connection or repeated request.
Continue work after a controller restart recover-controller Observe saved source, persist a user message, restart the isolated controller, and independently test the delivered result.
Stop work and change direction stop-redirect Click Stop, send one new request, reload, observe exactly one stored user message and the new answer, and reach Done.
Create and edit a company skill create-skill-studio Create one skill through the runner, verify its persisted library entry and activity-feed card, open Skill Studio, save an edit, and verify the edit after returning. Local Codex, local ACPX Claude, and warm Daytona cells are explicit.

For normal completion, all story tasks must reach Done, with no active run, pending completion confirmation, or scheduled recovery. Runs must prove native identity and native terminal contracts. A workspace-contention cancellation is not provider execution only when the persisted pre-dispatch record explicitly says providerWorkStarted: false and no process/session/runner identity exists. Other unexplained cancellations remain failures. Twelve total run records bound each story, including contention and recovery.

Arbitrary runner-process termination is not part of the model scorecard. The historical recover-runner, recover-runner-safe, and recover-runner-uncertain attempts remain available as diagnostics, with their original grades and costs. The first two did not establish a safe restart boundary, and the uncertainty case measures a deterministic safety rule. None supports ranking models. See controlled recovery tests.

Matrix and running

The local matrix has fourteen cases on native Codex gpt-5.6-sol, native ACPX Claude claude-sonnet-5, and native Codex gpt-5.4-mini: 42 cells. The two core profiles also declare build/revise, delegation, controller-restart, and skill-creation cases on Daytona: eight cells. Remote runner-process killing is not supported. For remote controller restart, a verified first download supplies the persistence checkpoint; the controller is interrupted during a subsequent revision with another queued requirement.

pnpm test:e2e:runner -- --list --suite everyday-workflows
pnpm test:e2e:runner -- --suite everyday-workflows --environment local --max-parallel 2
pnpm test:e2e:runner -- --id everyday-workflows.runner-codex-mini.local.build-revise
pnpm test:e2e:runner -- --suite everyday-workflows --environment daytona --max-parallel 2

Before project stories or the Python calibration tests, start Docker on the harness host and fetch the pinned oracle image. CI prepares and verifies this same pinned image before the paid project-story cells; artifact checks run on the harness host. The workflow verifies the exact repository digest after the pull. This is required for local and Daytona stories.

docker pull python@sha256:9d2e5553305c7c7b0097999bb17187c69b921ccd6bc9d40e4bb5ebe652c00285
python3 tests/runner-e2e/everyday-artifact.py --preflight

The harness checks this prerequisite before it creates the task. It does not pull an image during a model attempt or fall back to host execution.

Use the credential and immutable Daytona image setup in README.md. Provider calls cost money. Each cell owns an isolated instance and project. There are no real third-party mutations in the service fixture; it exercises production connection, transport, tool approval, and document delivery paths.

Deterministic checks and calibration

pnpm test:e2e:runner:typecheck
pnpm test:e2e:runner:unit
python3 -m unittest discover -s tests/runner-e2e -p test_everyday_artifact.py

The independent oracle rejects wrong output, ignored late feedback, trailing separator bugs, invalid argument acceptance, duplicate source modules, archive path traversal, and symlinks. Passing agent-authored tests cannot override it. Lifecycle calibration rejects legacy execution, missing runner identity, unexpected crashes, workers on the parent, and answers left in review.

ZIP evaluation runs delivered Python in a Docker container with a read-only project mount and root filesystem, no network, a non-root user, no Linux capabilities, and bounded CPU, memory, process count, output, and duration. Only the extracted delivery enters the container. The container is removed after grading. Calibration includes attempts to read a host file and reach a host loopback service.

Evalbook evidence and qualification

Each packaged attempt retains snapshots/everyday-workflow.json, downloaded ZIPs, assertions, actual task comments and run records, timing, accounting, source provenance, and screenshots. The story records a digest of its harness sources. Infrastructure failures and failed attempts must remain inspectable.

Import packaged results with paperclip-evals/evals/everyday-workflows/import_results.py. It uses the canonical Runner Evalbook generator and the built Runner Lab viewer. It does not invent provider transcripts, tool counts, model observations, or cost estimates. The selected model is checked against persisted native execution inputs; that is distinct from provider-side model identity verification.

Initial live results are diagnostic. They are not a reliability estimate or a model ranking. Before promotion, freeze both source revisions and harness digest, run at least three independent local repetitions, qualify the eight remote cells against a verified image, and review every failure. Keep model quality, lifecycle correctness, infrastructure availability, and latency separate.

Revised evaluation contract (14 September, second campaign)

That campaign used 32 cells: the original local stories plus two local Codex text-only safe-replacement probes, and the unchanged six remote cells. The old recover-runner results remain historical; recover-runner-uncertain is a new case that expects a visible Blocked safety stop, preserved source and queued input, and no unverified provider replay. Its Retry control is inspected, not claimed to restore work. Successful manual recovery remains unqualified.

recover-runner-safe interrupts a text-only Codex turn and queues new direction. A pass requires the server's durable verified_safe_replacement evidence and the new answer. No safety proof is injected or fabricated. If that premise cannot be verified in a live probe, report it as an unqualified recovery boundary, not an established product defect. Claude has no catalog cell for this Codex-specific replacement proof. Deterministic native-safe-replacement tests cover its proof and admission gates independently of model behavior.

Delegation now submits feedback through the existing child task composer and records the delivered comment ID. The child must consume the message and deliver the revised program. This does not require a lead to relay a parent comment. The separate issue-update-comment-wakeup route tests exercise exact supported mention routing, including access, dependency, identity, and duplicate-wake gates.

The approval case provisions an authenticated local service through the public API, with a random server-held credential that never enters the agent environment or browser trace. Approval/decline interactions still use the browser. Provider captures distinguish rejected unauthenticated requests from accepted calls. The old public-endpoint attempts remain boundary evidence, not an isolation promise.

Stop now waits for the owned runner to exit, records project file hashes, and checks them again after the new response. This proves stability over that interval, not indefinite monitoring. Hiring and declined-access policy changes are deferred by user decision; their old results must not be presented as new campaign runs.

Decline correction (14 September, third campaign)

That campaign used 35 cells (29 local, six remote). service-decline tests rejection of a protected action on an already installed service; its former "connection request" title was misleading. connection-decline separately tests Not now on new Notion setup. Both permit a brief explanation as the complete fallback, so Done is expected after that explanation. Neither test requires completion after refusing work that is still required.

The installed-service decline fixture now uses the same server-held credential as approval. The harness requires one pending interaction, validates its kind and connection/provider identity before clicking, and waits for the exact interaction's saved decision. Wrong interactions fail decision-request-matches-story with a screenshot; they are not evidence of an ignored decline. Both decline stories check a new explanation after the decision and reject repeated requests.

Historical attempts remain unchanged. This campaign resumes the previously deferred decline cases; hiring remains deferred. Notion setup is declined in the UI, so this test neither authenticates to nor reads real Notion data.

Decision screenshots are included in the evidence package. Before capturing the final screen, the harness waits for the thread and latest persisted agent comment to render, then scrolls that comment into view. A Done header alone is not proof that the final response was visible.

Recovery scope correction (14 September)

This correction reduced the catalog to 30 cells: 24 local and six remote. Forced runner crash probes are retired from paid selection. Their original attempt IDs remain in Evalbook's Diagnostics history and Latest pages; they are excluded from the main matrix without changing grades or deleting evidence. Reported spend still includes all attempts.

recover-controller and stop-redirect retain concrete supported journeys: restart the controller while preserving the runner, or use Stop and submit a new direction. Their assertions verify pending input, saved work, and the next usable result. Neither claims recovery from an arbitrary provider-process crash.

A future user-facing crash-recovery case needs a reproducible recoverable fault, an identified supported recovery action, and evidence through the final usable result. A missing test premise must be reported as unexercised, not a model failure. Do not introduce a new paid case just to replace a retired row.

Skill creation (16 September)

create-skill-studio adds five cells: three local profiles and the two core profiles on Daytona. The current catalog has 35 cells: 27 local and eight remote. The test opens the created skill from its task-feed card, checks the canonical skill identity in Studio, saves an edit, and returns to the same skill in the task sidebar. A model's authored document heading is not used as the identity check.

External-provider fallback

Three explicit local cases cover aggregator routing with the normal production agent guidance. They add nine local cells. The AgentMail setup case adds three local cells; the suite now has 50 cells total.

The AgentMail case uses real model discovery and production interaction/UI paths, but declines before sending credentials to AgentMail. Its independent grader is calibrated against missing, duplicate, misaddressed, hidden, and malformed cards. The email integration suite separately proves credential/inbox creation, access defaults, assignment checks, and completion using a fixture provider. Neither test qualifies live AgentMail delivery. Run a bounded local model cell with:

pnpm test:e2e:runner -- --id everyday-workflows.runner-codex-mini.local.agentmail-setup --max-automatic-retries 0
  • provider-native: Jira is supported natively and by aggregators. Require the Jira connection card directly, decline it in the browser, and verify no provider question, connection creation, or repeated request.
  • provider-decline: HubSpot has no built-in connector in this fixture. Require Composio, Arcade, Zapier, and None in that order with external-service disclosure. Restart the controller, reload the question, choose None through the UI, and verify one saved answer, no connection changes, and no fabricated result.
  • provider-second: Install a deterministic Arcade gateway with a read-only HubSpot action through public APIs. Choose Arcade in the browser after restart. Require no calls before selection, exactly one call afterwards, the independently generated contact marker in the agent response, and no duplicate connection.

These use the existing native profiles and 12-minute local attempt deadline; expected provider runs are two per case. There are no real third-party mutations. The fixture server is closed and the harness cleans its disposable instance. Provider-choice screenshots, persisted interactions, gateway call counts, source revision, harness digest, and existing usage/cost evidence accompany each attempt. connection-routing-evidence.test.ts calibrates the grader against undisclosed routing, incorrect ordering, early calls, duplicate questions, and fabricated reads. Run a single everyday-workflows.runner-codex-mini.local.provider-decline cell first; do not treat these fixtures as live provider compatibility tests.