Plans · Individual
Pricing
Free authors run ad-hoc models with provider defaults. Pro locks model and generation configuration, keeps full experiment history, adds parameterized dataset sweeps, and reopens the exact output bodies behind every run.
Bring your own key
Free
$0 · foreverAd-hoc models with provider defaults. One saved run per benchmark; output bodies stay session-only.
Pro
$8/mo · $80/yrControlled model targets, dataset sweeps, full run history, and reopenable output bodies.
Secure checkout and subscription management are provided by Stripe Managed Payments and Link.
What you get in each plan
Free → Pro
CapabilityFree planPro plan
Dataset sweeps
- Import bounded CSV or JSONL prompt datasetsFree planPro plan
- Slice score breakdowns across expanded runsFree planRun sharedPro plan
Core
- Create and share benchmark suitesFree planPro plan
- Run models with your own API keysFree planPro plan
- Save scores and grading criteriaFree planPro plan
LLM Judge Labs
- Author rubric judges and multi-model consensusFree planPro plan
- Run shared judge benchmarks with your own OpenRouter keyFree planPro plan
- Calibrate judges against reviewed saved examplesFree planPro plan
Model targeting
- Ad-hoc model with OpenRouter provider defaultsFree planPro plan
- Define ordered, controlled OpenRouter model targetsFree planPro plan
- Configure typed generation settings and provider routingFree planPro plan
History & outputs
- Keep one saved run per benchmark and ad-hoc modelFree planPro planFull history
- Reopen exact output bodies (text, JSON, Markdown, code, HTML, SVG)Free planSession onlyPro plan
Sharing
- Duplicate shared benchmarksFree planPro plan
