Step 1 — Ask
What do you want to test?
Type a category in plain language. BankBench checks what you already have, tells you what's missing, and drafts new tests.
Step 2 — Drafts
Review the drafts
Each draft is shaped like an Inspect test: what the test is, how it runs, and how we grade it. Uncheck any you don't want.
Step 3 — Quality
How much can you trust these tests?
Each test is scored on the five dimensions from the AI Evaluation Quality Scorecard (San Joaquin et al., Feb 2026). The grade rolls up across the whole suite.
Step 4 — Run
Pick the models, run the tests
Choose 3–5 models to compare. BankBench calls each one against every selected test.
Step 5 — Review
Results
Every model, every test, in one place. Full answers are hidden by default.
Library
Everything you've saved
Drafts you've approved, plus scenarios from past sessions. This is what the Scout checks against when you ask something new.
Results
Model comparison
Aggregated across every scored run, grouped by model.
Settings
API keys & judge model
Keys are saved only on this computer. They're sent directly to the provider you choose — nowhere else.