2 Sources
[1]
'Are you joking, mate?' AI doesn't get sarcasm in non-American varieties of English
UNSW Sydney provides funding as a member of The Conversation AU. In 2018, my Australian co-worker asked me, "Hey, how are you going?". My response - "I am taking a bus" - was met with a smirk. I had recently moved to Australia. Despite studying English for more than 20 years, it took me a while to
[2]
'Are you joking, mate?' AI doesn't get sarcasm in non-American varieties of English
In 2018, my Australian co-worker asked me, "Hey, how are you going?" My response -- "I am taking a bus" -- was met with a smirk. I had recently moved to Australia. Despite studying English for more than 20 years, it took me a while to familiarize myself with the Australian variety of the
Share
Copy Link
New research reveals that large language models have difficulty detecting sarcasm and sentiment in Australian, Indian, and British English, highlighting the need for more diverse language training in AI.
Researchers have developed a new tool called BESSTIE (Benchmark for Sentiment and Sarcasm in Three International English varieties) to evaluate the performance of large language models (LLMs) in detecting sentiment and sarcasm across different English varieties. The study, published in the Findings of the Association for Computational Linguistics 2025, highlights significant challenges faced by AI in understanding non-American English
1
2
.Dr. Siddharth Srivastava, the lead researcher, shares a personal anecdote that illustrates the complexity of language varieties. Despite studying English for over two decades, he found himself confused by Australian English upon moving to Australia. This experience mirrors the challenges faced by AI models, which are predominantly trained and tested on Standard American English
1
.
Source: The Conversation
BESSTIE is the first benchmark of its kind, focusing on three English varieties: Australian, Indian, and British. The researchers collected data from Google Maps reviews and Reddit posts, using language variety predictors to ensure a high probability of specific language varieties. The benchmark evaluates nine powerful, freely usable large language models, including RoBERTa, mBERT, Mistral, Gemma, and Qwen
1
2
.The study revealed several important insights:
Performance disparity: LLMs performed better on Australian and British English (native varieties) compared to Indian English (non-native variety)
1
2
.Sentiment vs. Sarcasm: AI models were more adept at detecting sentiment than sarcasm across all varieties
1
2
.Sarcasm detection challenges: The models struggled significantly with sarcasm, achieving only 62% accuracy for Australian English and about 57% for Indian and British English
1
2
.
Source: Tech Xplore
1
2
.Related Stories
The research underscores the importance of evaluating AI models in specific national contexts. As LLMs become increasingly prevalent worldwide, there's a growing recognition of the need to adapt these tools for diverse language varieties
1
2
.Dr. Srivastava and his team are currently working on a project to implement LLMs in hospital emergency departments to assist patients with varying English proficiencies. Additionally, initiatives like the University of Western Australia and Google's project to improve LLM efficacy for Aboriginal English demonstrate the increasing focus on language diversity in AI development
1
2
.The BESSTIE benchmark represents a significant step towards more inclusive and accurate AI language models. By highlighting the current limitations in processing non-American English varieties, this research paves the way for future improvements in AI's ability to understand and interpret diverse language patterns, ultimately leading to more effective and equitable AI applications across different cultures and regions.
Summarized by
Navi
[1]
1
Technology

2
Technology

3
Policy and Regulation
