BanglaMATH: A Bangla benchmark dataset for testing LLM mathematical reasoning at grades 6, 7, and 8

More recent
The Three Books of Science

More or equal citational love
Continuum rich-get-richer processes: Mean field analysis with an application to firm size

More recent
Detecting sub-populations in online health communities: A mixed-methods exploration of breastfeeding messages in BabyCenter Birth Clubs

Less or equal citational love
A blind spot for large language models: Supradiegetic linguistic information

BanglaMATH: A Bangla benchmark dataset for testing LLM mathematical reasoning at grades 6, 7, and 8

Tabia Tanzin Prama, Christopher M. Danforth, and Peter Sheridan Dodds

Proceedings of The 3rd Workshop on Mathematical Natural Language Processing (MathNLP 2025), 2025

journal version | journal page | Google Scholar

Times cited: 5

Abstract:

Large Language Models (LLMs) have tremendous potential to play a key role in supporting mathematical reasoning, with growing use in education and AI research. However, most existing benchmarks are limited to English, creating a significant gap for low-resource languages. For example, Bangla is spoken by nearly 250 million people who would collectively benefit from LLMs capable of native fluency. To address this, we present BanglaMATH, a dataset of 1.7 k Bangla math word problems across topics such as Arithmetic, Algebra, Geometry, and Logical Reasoning, sourced from Bangla elementary school workbooks and annotated with details like grade level and number of reasoning steps. We have designed BanglaMATH to evaluate the mathematical capabilities of both commercial and open-source LLMs in Bangla, and we find that Gemini 2.5 Flash and DeepSeek V3 are the only models to achieve strong performance, with≥ 80% accuracy across three elementary school grades. Furthermore, we assess the robustness and language bias of these top-performing LLMs by augmenting the original problems with distracting information, and translating the problems into English. We show that both LLMs fail to maintain robustness and exhibit significant performance bias in Bangla. Our study underlines current limitations of LLMs in handling arithmetic and mathematical reasoning in low-resource languages, and highlights the need for further research on multilingual and equitable mathematical understanding.

This is the default HTML.
You can replace it with your own.
Include your own code without the HTML, Head, or Body tags.

BibTeX:

@inproceedings{prama2025a,
  author =	 {Prama, Tabia Tanzin and Danforth, Christopher M. and
                  Dodds, P.},
  title =	 {Bangla{MATH}: {A} {B}angla benchmark dataset for testing
                  {LLM} mathematical reasoning at grades 6, 7, and 8},
  booktitle =	 {Proceedings of The 3rd Workshop on Mathematical
                  Natural Language Processing (MathNLP 2025)},
  year =	 {2025},
  key =		 {},
  pages =	 {134–149},
  url =		 {https://aclanthology.org/2025.mathnlp-main.10/},
}

More recent
Detecting sub-populations in online health communities: A mixed-methods exploration of breastfeeding messages in BabyCenter Birth Clubs

BanglaMATH: A Bangla benchmark dataset for testing LLM mathematical reasoning at grades 6, 7, and 8

Tabia Tanzin Prama, Christopher M. Danforth, and Peter Sheridan Dodds

Proceedings of The 3rd Workshop on Mathematical Natural Language Processing (MathNLP 2025), 2025

journal version | journal page | Google Scholar

Times cited: 5

Abstract:

BibTeX:

Share this page:

Some of our Panometer’s online instruments:

Storywrangler: Track and compare Twitter n-grams from 2008 on in 100+ languages.

The Lexicocalorimeter: Measuring calories in and calories out with tweets.

The POTUSometer: Computational history, narrative control, ratios, and chronopathy—measuring how time flies and crawls.

Explore the Teletherm: the on-average coldest and warmest days of the year.

The Hedonometer: Measuring the happiness (and sadness) of all kinds of texts.

© Peter Sheridan Dodds, 7+13+5, 1995–