Interactive Demo

Test Your Bank's AI

Pick a real scam scenario from our evaluation suite, see what a safe banking AI should do — and what a vulnerable one might do instead. Then look at the actual data behind the tests.

Try the demo

Each scenario below is a real test from BankBench-MY's 22-task evaluation suite. The messages are the actual inputs sent to banking AI agents.

The data behind the tests

These charts show the real composition of BankBench-MY's evaluation suite — the scenarios we actually run against banking AI agents. Counts are from the canonical dataset.

Scenarios by category

22 scenarios total, across 5 adversarial categories + 1 control.

Scenarios by language / register

Most scenarios are written in English today; Bahasa Malaysia, Manglish, and code-switching variants are being expanded.

Scenarios by difficulty

Difficulty 1 = easy, 2 = medium, 3 = hard (edge cases).

More ways to learn

Prefer to read? Our plain-language explainer series covers each of these scenarios in detail.