Published package evidence
A report you can interrogate, not just a score you have to trust.
This evidence came from a clean, isolated run of the live npm 0.4.7 package against RelayForge, a synthetic incident-routing MCP fixture.
Readiness
84/100
Grade B
10 tools · 1,113 estimated definition tokens
Input contracts
Object schemas use valid required semantics and documented properties.
Sensitive inputs
No credential-like values are requested from the model.
Destructive clarity
The destructive operation includes an explicit warning.
Output contracts
All ten tools still lack optional output schemas, so the result stays below A.
Execution matrix
One task. Three model families. Inspectable evidence.
Each selected model called the named read-only tool once and received one non-error result. Displayed call counts match the replay.
| Family | Model | Evidence | Calls / results | Tokens |
|---|---|---|---|---|
| OpenAI | openrouter/openai/gpt-5.6-sol | proven | 1 / 1 | 1,622 |
| Anthropic | openrouter/anthropic/claude-sonnet-5 | proven | 1 / 1 | 4,772 |
| openrouter/google/gemini-3.7-flash | proven | 1 / 1 | 1,717 |
Replay excerpt
The report keeps the call and result in order so the operator can judge the run.
- →call get_incident {"incident_id":"inc_demo_001"}
- ←result: Checkout latency above SLO · checkout-api · high · open
What this proves
A real, non-error MCP result occurred for every selected model.
What this does not prove
Universal task correctness, security, load readiness, or compatibility with every MCP client.