Skip to main content
Benchmark marketplace

Custom AI benchmarks, run on your own keys.

Assemble ordered task suites with mixed grading, execute them client-side, and share owner-reviewed leaderboards alongside self-reported run history — without routing paid inference through the platform.

Inference
Client-side · your OpenRouter key
Outputs
Scores + metadata · bodies retained on Pro
Grading
Automatic checks + manual criteria
Sharing
Public catalog · private access grants

How it works

Four steps · assemble to share

See examples
  1. 01

    Assemble

    Ordered task suites with mixed grading — automatic checks alongside manual criteria.

    Start a suite
  2. 02

    Run

    Execution happens client-side on your own OpenRouter key. No paid inference routed through the platform.

    Open dashboard
  3. 03

    Grade

    Combine scripted checks with human judgment. Scores and metadata persist for every saved run; Pro retains exact output bodies.

    Define criteria
  4. 04

    Share

    Publish to the public catalog or grant private access. Leaderboards reproduce from the stored results.

    Browse catalog