BankBench-MY

A multilingual (EN / Bahasa Malaysia / Manglish) safety evaluation for banking-agent LLMs — does a banking chatbot leak OTPs and account data, process unauthorised transfers, or follow phishing links when conversation register shifts mid-conversation? Built on Inspect AI, developed under the Sinar fellowship. Status: ongoing work — some surfaces below are live, some are still in progress.

One eval core, four applied surfaces

BankBench-MY's scenarios and scorer are the shared core. Everything else is a different way of putting that core to work — some finished, some still in progress.

BankBench-MY meta-overview diagram — one shared eval core with four applied surfaces (BankBench itself, plus Scorecard, plus Model, plus Public Education) and a row of related domain fields below
live — sandbox

BankBench itself

The core: adversarial banking-agent scenarios in EN/BM/Manglish, an Inspect AI task + scorer, and a live agentic sandbox you can run them against.

Try it live →
coming soon

+ Scorecard

A/B/C/D graded, versioned cross-model comparisons — benchmark-lifecycle framing (drift, deprecation, saturation) instead of a one-off leaderboard number.

Not yet migrated
in progress

+ Model

The eval set becomes training data — fine-tune toward the behavior BankBench-MY measures, then re-measure it. Track live status on the training-loop dashboard.

View progress →
coming soon

+ Public Education

Plain-language explainers for consumers and a regulator-facing gap brief, built from the same findings above.

Not yet written

Related domains this work also draws on

Fields, not project names — each maps to ongoing work adjacent to BankBench-MY, most of it not yet public.

Mechanistic Interpretability Adversarial Sandbox & Honeypot Red-teaming Multi-Agent Collusion & Cross-Language Pressure Testing Meta-Evaluation & Benchmark Tooling AI Governance & Standards Mapping AI Safety Engineering Curriculum & Education