mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 20:05:57 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Governed MCP access spans contracts, runtime enforcement, adapters, UI surfaces, and operator verification > - The parity reference PR #9534 is too large for effective automated or human review > - The feature therefore needs a linear stack whose individual diffs stay below the 100-file review limit > - This pull request is split 8/8 and focuses on end-to-end coverage, operator docs, evals, and release notes > - The benefit is a standalone, testable review boundary while preserving byte-for-byte parity at the top of the stack ## Linked Issues or Issue Description - Related parity reference: #9534 - Problem: The complete stack needs discoverable browser scenarios, operator guidance, threat modeling, eval coverage, and a parity proof before merge. - Proposed solution: Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release notes, the skill update, and the root e2e driver script registration. - Alternatives considered: keeping #9534 as one 403-file review, or rewriting the feature to manufacture seams; both were rejected in favor of path extraction plus compile-driven boundary moves. - Roadmap alignment: this advances the existing governed MCP/tool-access work already represented by #9534; it does not introduce a separate roadmap initiative. - Stack position: base branch is `pap10341-split/07-ui-apps-activation`. - Merge policy: merge bottom-up, in order, only after the complete eight-PR stack has been reviewed and the top-of-stack parity gate remains empty. - Requested review: QA for flag audit and e2e/browser acceptance; Greptile on every PR. ## What Changed - Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release notes, the skill update, and the root e2e driver script registration. - Keeps this PR below 100 changed files and independently typecheckable. - Preserves the final tree from #9534 when combined with the other seven stack levels. ## Verification - `pnpm typecheck` - `node --check scripts/e2e-mcp-user-stories.mjs` - `pnpm exec playwright test --config tests/e2e/playwright.config.ts --list` — 43 tests discovered - `git diff pap10341-split/08-e2e-docs 6b40e3876d9297105d4ec306e47e46d351c86172` — empty (0 bytes) ## Risks - Browser suites depend on runtime services and environment setup; this PR validates discovery locally while QA owns full flag-on/flag-off execution. - Stack risk: merging out of order can expose incomplete layers; mitigate by following the documented bottom-up merge policy. - Parity risk: later edits to an intermediate branch can drift from #9534; mitigate by re-running the empty top-of-stack diff before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, Git, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] Internal references are omitted except the execution-plan link explicitly required for this coordinated split stack - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack Coordination - Internal execution plan: [PAP-13874](/PAP/issues/PAP-13874#document-plan) - Parity reference: #9534 - Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563 - Merge bottom-up only after full-stack review and an empty parity diff at #9563. --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
98 lines
4.3 KiB
Markdown
98 lines
4.3 KiB
Markdown
# Paperclip Evals
|
|
|
|
Eval framework for testing Paperclip agent behaviors across models and prompt versions.
|
|
|
|
See [the evals framework plan](../doc/plans/2026-03-13-agent-evals-framework.md) for full design rationale.
|
|
|
|
## Quick Start
|
|
|
|
### Prerequisites
|
|
|
|
```bash
|
|
pnpm add -g promptfoo
|
|
```
|
|
|
|
You need an API key for at least one provider. Set one of:
|
|
|
|
```bash
|
|
export OPENROUTER_API_KEY=sk-or-... # OpenRouter (recommended - test multiple models)
|
|
export ANTHROPIC_API_KEY=sk-ant-... # Anthropic direct
|
|
export OPENAI_API_KEY=sk-... # OpenAI direct
|
|
```
|
|
|
|
### Run evals
|
|
|
|
```bash
|
|
# Smoke test (default models)
|
|
pnpm evals:smoke
|
|
|
|
# Validate config without provider credentials
|
|
cd evals/promptfoo && npx promptfoo@latest validate -c promptfooconfig.yaml
|
|
|
|
# Or run promptfoo directly
|
|
cd evals/promptfoo
|
|
promptfoo eval
|
|
|
|
# Focus only on MCP gateway behavior cases
|
|
npx promptfoo@0.103.3 eval -c promptfooconfig.yaml \
|
|
--providers echo \
|
|
--filter-pattern '^mcp_gateway\.' \
|
|
--no-cache \
|
|
--no-progress-bar \
|
|
--no-write
|
|
|
|
# View results in browser
|
|
promptfoo view
|
|
```
|
|
|
|
### What's tested
|
|
|
|
Phase 0 covers narrow behavior evals for the Paperclip heartbeat skill:
|
|
|
|
| Case | Category | What it checks |
|
|
|------|----------|---------------|
|
|
| Assignment pickup | `core` | Agent picks up todo/in_progress tasks correctly |
|
|
| Progress update | `core` | Agent writes useful status comments |
|
|
| Blocked reporting | `core` | Agent recognizes and reports blocked state |
|
|
| Approval required | `governance` | Agent requests approval instead of acting |
|
|
| Company boundary | `governance` | Agent refuses cross-company actions |
|
|
| MCP allowed read tool | `mcp_gateway` | Agent records successful gateway calls without unnecessary approval |
|
|
| MCP denied tool | `mcp_gateway` | Agent fails closed without retrying or bypassing denied unsafe tools |
|
|
| MCP pending approval | `mcp_gateway` | Agent waits on the gateway-created approval path |
|
|
| MCP denied approval | `mcp_gateway` | Agent honors rejected or unapproved tool actions |
|
|
| MCP rate limit | `mcp_gateway` | Agent backs off without crashing or busy-looping |
|
|
| MCP missing credential | `mcp_gateway` | Agent blocks on credential repair without leaking or inventing secrets |
|
|
| MCP revoked session | `mcp_gateway` | Agent stops using stale gateway tokens and avoids raw upstream fallback |
|
|
| MCP header forwarding | `mcp_gateway` | Agent reports forwarded transport/credential headers from redacted audit evidence |
|
|
| MCP named target | `mcp_gateway` | Agent uses the exact on-demand named gateway tool rather than an ambiguous upstream name |
|
|
| MCP elicitation | `mcp_gateway` | Agent asks the human/board for missing input instead of fabricating it |
|
|
| MCP approved target drift | `mcp_gateway` | Agent treats changed catalog/schema/credential snapshots as stale approval |
|
|
| No work exit | `core` | Agent exits cleanly with no assignments |
|
|
| Checkout before work | `core` | Agent always checks out before modifying |
|
|
| 409 conflict handling | `core` | Agent stops on 409, picks different task |
|
|
| Memory provider binding | `phase5_memory` | Agent honors agent override before company default |
|
|
| Memory provenance audit | `phase5_memory` | Agent preserves inspectable source and operation records |
|
|
| Memory hook cost/trust | `phase5_memory` | Agent keeps memory hook cost attribution and source trust visible |
|
|
| Board command work objects | `phase5_control_surface` | Chat-like board commands create auditable work objects |
|
|
|
|
Phase 5 memory/control-surface prompt evals should be paired with deterministic server/shared tests for:
|
|
|
|
- memory provider resolution order: agent override, then company default
|
|
- memory operation audit rows including company, agent, issue, run, provider, source, and cost references
|
|
- hook-delivered memory payloads preserving source trust and cost attribution fields
|
|
- board command/chat-like routes creating auditable issues, comments, documents, approvals, or work products
|
|
|
|
### Adding new cases
|
|
|
|
1. Add a YAML file to `evals/promptfoo/tests/`
|
|
2. Follow the existing case format (see `core.yaml` for reference)
|
|
3. Run `promptfoo eval` to test
|
|
|
|
### Phases
|
|
|
|
- **Phase 0 (current):** Promptfoo bootstrap - narrow behavior evals with deterministic assertions
|
|
- **Phase 1:** TypeScript eval harness with seeded scenarios and hard checks
|
|
- **Phase 2:** Pairwise and rubric scoring layer
|
|
- **Phase 3:** Efficiency metrics integration
|
|
- **Phase 4:** Production-case ingestion
|