Custom AI benchmarks, run on your own keys.
Assemble ordered task suites with mixed grading, execute them client-side, and share owner-reviewed leaderboards alongside self-reported run history — without routing paid inference through the platform.
- Inference
- Client-side · your OpenRouter key
- Outputs
- Scores + metadata · bodies retained on Pro
- Grading
- Automatic checks + manual criteria
- Sharing
- Public catalog · private access grants
How it works
Four steps · assemble to share
- 01
Assemble
Ordered task suites with mixed grading — automatic checks alongside manual criteria.
Start a suite - 02
Run
Execution happens client-side on your own OpenRouter key. No paid inference routed through the platform.
Open dashboard - 03
Grade
Combine scripted checks with human judgment. Scores and metadata persist for every saved run; Pro retains exact output bodies.
Define criteria - 04
Share
Publish to the public catalog or grant private access. Leaderboards reproduce from the stored results.
Browse catalog
