A multilingual (EN / Bahasa Malaysia / Manglish) safety evaluation for banking-agent LLMs — does a banking chatbot leak OTPs and account data, process unauthorised transfers, or follow phishing links when conversation register shifts mid-conversation? Built on Inspect AI, developed under the Sinar fellowship. Status: ongoing work — some surfaces below are live, some are still in progress.
BankBench-MY's scenarios and scorer are the shared core. Everything else is a different way of putting that core to work — some finished, some still in progress.
The core: adversarial banking-agent scenarios in EN/BM/Manglish, an Inspect AI task + scorer, and a live agentic sandbox you can run them against.
Try it live →A/B/C/D graded, versioned cross-model comparisons — benchmark-lifecycle framing (drift, deprecation, saturation) instead of a one-off leaderboard number.
Not yet migratedThe eval set becomes training data — fine-tune toward the behavior BankBench-MY measures, then re-measure it. Track live status on the training-loop dashboard.
View progress →Plain-language explainers for consumers and a regulator-facing gap brief, built from the same findings above.
Not yet writtenFields, not project names — each maps to ongoing work adjacent to BankBench-MY, most of it not yet public.