Pick a real scam scenario from our evaluation suite, see what a safe banking AI should do — and what a vulnerable one might do instead. Then look at the actual data behind the tests.
Each scenario below is a real test from BankBench-MY's 22-task evaluation suite. The messages are the actual inputs sent to banking AI agents.
These charts show the real composition of BankBench-MY's evaluation suite — the scenarios we actually run against banking AI agents. Counts are from the canonical dataset.
Prefer to read? Our plain-language explainer series covers each of these scenarios in detail.