› › › ›

Step 1 — Ask

What do you want to test?

Type a category in plain language. BankBench checks what you already have, tells you what's missing, and drafts new tests.

Step 2 — Drafts

Review the drafts

Each draft is shaped like an Inspect test: what the test is, how it runs, and how we grade it. Uncheck any you don't want.

0 drafts selected

Step 3 — Quality

How much can you trust these tests?

Each test is scored on the five dimensions from the AI Evaluation Quality Scorecard (San Joaquin et al., Feb 2026). The grade rolls up across the whole suite.

Step 4 — Run

Pick the models, run the tests

Choose 3–5 models to compare. BankBench calls each one against every selected test.

Models
0 selected
$/Mtok

Step 5 — Review

Results

Every model, every test, in one place. Full answers are hidden by default.

Library

Everything you've saved

Drafts you've approved, plus scenarios from past sessions. This is what the Scout checks against when you ask something new.

Results

Model comparison

Aggregated across every scored run, grouped by model.

Settings

API keys & judge model

Keys are saved only on this computer. They're sent directly to the provider you choose — nowhere else.

One key, many models. Get one at openrouter.ai/keys.
NVIDIA NIM endpoints. Get one at build.nvidia.com.
A small, fast model is fine — it just needs to follow instructions.