Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Bengali, spoken by roughly a quarter of a billion people, has recently gained a growing collection of large language model (LLM) evaluation benchmarks, most of which are direct translations of established English test sets. At the same time, benchmark data contamination (BDC) - the leakage of evaluation items into model training corpora - is now a well-documented threat to the validity of LLM evaluation, and recent evidence shows that contamination can cross language barriers: a model exposed to an English test set can obtain inflated scores on its translated counterpart. Whether Bengali benchmarks suffer from such contamination, and how stable their results are under semantically equivalent perturbations, has never been examined. We present the first systematic audit framework for the reliability of Bengali LLM benchmarks, together with a completed pilot validating the full pipeline on free-tier compute. The framework combines (i) likelihood-based membership signals (Min-K% Prob), (ii) generation-based tests (TS-Guessing-style option completion), (iii) ground-truth n-gram overlap against fully public pre-training corpora, and (iv) a perturbation suite measuring the stability of benchmark rankings. In the pilot (Qwen2.5-1.5B-Instruct on 500 BoolQ-bn items), accuracy is 0.568 (95% CI [0.524, 0.610]) - barely above the 0.50 chance level of the binary task - and a distractor perturbation shifts accuracy by -1.0 pp (McNemar p = 0.302). The complete model x benchmark audit grid is being executed with the released protocol; this manuscript reports the framework, benchmark inventory, protocol, and pilot. All code, data manifests, and per-item audit outputs are publicly released.</p>

Show More

Keywords

model benchmark bengali evaluation benchmarks

Related Articles

PORE

About

Connect