Cultural Evaluation as the Next Phase of AI Benchmarking

 

“When a measure becomes a target, it ceases to be a good measure.”

- Goodhart’s Law

 

 

AI evaluation tends to treat high scores as definitive proof of broad competence. Today, AI systems are setting records – acing professional exams, outscoring standardized tests, and ranking at the top of global performance charts. But in early-grade classrooms across Sub-Saharan Africa, these numbers often mean very little. A system that performs well on abstract benchmarks can still falter in a multilingual, noisy classroom where children switch between languages mid-sentence and answer questions through a cultural frame the model has never encountered. 

Cultural evaluation is not a minor tweak to AI benchmarking; it is a necessary course correction. Without it, we risk confusing high test scores for real understanding. 

 

The gap between benchmarks and classrooms

Current testing methods often rely on standardized AI benchmarks, such as multiple-choice knowledge tests, math word-problem sets, coding tasks, and general reasoning exams. These tests are useful for comparing models under controlled conditions, but they can also encourage a false sense of confidence. However, competence is never culturally neutral. In culturally rooted language tasks, models that excel on general benchmarks often stumble when meaning depends on local context, shared experience, idiom, humor, history, or worldview. They may translate words accurately across many languages, yet miss the cultural logic woven into them. The fundamental limitation is not grammar; it is context.

Many tech systems claim to support dozens of African languages, but covering a language is not the same as understanding the culture that shapes it. Because current evaluation practices focus heavily on general reasoning in abstract scenarios, they overlook the kind of culturally grounded understanding that matters in real classrooms, thereby providing inaccurate guidance on a tool’s real-world validity. 

 

What a proverb reveals about AI understanding

A concrete example illustrates the problem. Consider the Ethiopian proverb, “ቆዳ ሲወደድ፥ በሴህን እረድ.” This conveys the idea that when leather becomes valuable, that is the right time to slaughter your ox. It is a proverb about timing, about recognizing when conditions are favorable and acting accordingly, similar to the English expression, “strike while the iron is hot.” It reflects a practical understanding of opportunity and resources deeply embedded in Ethiopian society.

When we tested DeepSeek R1 on this proverb, the model literally translated the words and concluded that it means, “meaningful gains require giving up something.” That interpretation completely misses the cultural core: the proverb is not about abstract sacrifice. It is about reading the moment. A student raised in the culture would grasp this distinction immediately; the model cannot.

This kind of failure motivated our development of ProverbEval (Azime et al., NAACL 2025), a benchmark built around proverb understanding in historically under-resourced languages. The premise is simple: if an AI system truly understands a language, it should be able to reason through the culturally grounded expressions that real speakers use. Proverbs serve as an excellent stress-test because their true meaning depends on shared cultural knowledge that cannot be deciphered through literal translation alone. Our finding confirmed a systemic pattern: models that score well on standard benchmarks often fail when meaning is rooted in lived human experience rather than formal machine logic.

 

When multiple choice hides the problem

Another critical pattern is the gap between recognizing the right answer and generating one independently. AI systems often perform well on multiple-choice questions because the correct answer is already visible among the options. In that setting, the model only has to select the most likely answer, and its performance can be shaped by clues in the wording, the order of the options, or even a tendency to favor particular answer labels. Zheng et al. (2023) show that this makes multiple-choice performance less reliable as evidence of real understanding. Our own findings in ProverbEval confirmed this: models that successfully picked the correct interpretation from a list were often unable to explain the proverb’s meaning when asked to respond freely. 

In primary education, this mismatch matters directly. True learning means narrating, explaining, and reasoning, not just guessing between A, B, C, or D. If evaluation frameworks rely primarily on multiple-choice formats, they will continue to produce reassuringly high scores that mask the kinds of failures that show up in real classrooms. A high average accuracy score can hide who is being left behind.

 

Beyond text: Why multimodal evaluation matters

The mismatch deepens when we move beyond text. In real classrooms, communication rarely happens through written words alone. Students interact through speech, images, gestures, and shared cultural references. Yet, most evaluation frameworks still test AI systems primarily through text inputs, overlooking the multimodal reality of how children actually learn.

This is the exact challenge we set out to address with Afri-MCQA (Tonja et al., 2026), a multilingual cultural question-answering benchmark spanning 15 African languages across 12 countries. The benchmark evaluates models across text, speech, and visual inputs grounded in local contexts. For example, a question might show an image of a tandar ƙasa (a traditional Hausa clay pan) and ask the model to identify it in both English and Hausa, and via both written text and spoken audio.

What we found was striking: models that performed well on text-based tasks often struggled when the exact same questions were posed through spoken local languages, or when the reasoning required culturally grounded visual understanding. Simply put, fluency in multilingual text does not guarantee competence in multimodal cultural reasoning. A system might correctly answer a written question about Gonder’s Fasil Ghebbi castle in English yet fail when asked the same question in spoken Amharic. This issue is not that the language is unknown to the model; it is that the model had never encountered that specific cultural and geographic knowledge in an audio format.

These findings point in a clear direction: the next generation of AI evaluation needs to be multimodal and speech-aware, reflecting how language is used in everyday settings. Without this shift, AI systems will continue to appear capable in controlled evaluations while failing in the actual environments where they are deployed. 

 

What cultural evaluation should look like

If we are serious about deploying AI systems that work in African classrooms, evaluation must mirror everyday cultural experiences. Benchmarks should reflect the kinds of knowledge people actually use --  local practices, community activities, culturally unique ways of expressing meaning. This means mathematical problems must be set in the context of real-world financial activities and social practices, as we explored in our work on socio-cultural localization of math word problems (Azime et al., 2025). It means translation systems must be evaluated not just for linguistic accuracy, but for cultural alignment and contextual appropriateness, as the CAMMT benchmark (Villa-Cueva et al., 2025) has begun to do. It also means evaluation frameworks must test how systems handle real spoken dialects, drawing on large-scale multilingual African speech resources like the WAXAL corpus (Diack et al., 2026). 

The common thread is clear: evaluation must move beyond controlled, text-only settings and into the settings these systems will be used in classrooms, homes, and communities where language, culture, and context intersect. While no single benchmark can fully capture this reality, culturally grounded evaluation is where we must begin.

 

The question that matters

We now have the tools to build this kind of evaluation. ProverbEval tests whether models understand culturally embedded meaning. Afri-MCQA tests whether that understanding holds across modalities and languages. Our math localization framework tests if educational content can be adapted to reflect the lived realities of students. None of these benchmarks existed just two years ago. Together, they represent the beginning of an evaluation infrastructure built by African researchers, grounded in African contexts, and designed to ask the vital questions that standard benchmarks ignore.

The core question is no longer whether an AI model can process multimodal inputs. It is whether we are measuring what actually matters, and whether the people building these systems are truly listening to the communities where these tools are deployed. We have started asking the right questions. It is time for the rest of the field to catch up.

 

The opinions expressed are those of the authors alone.


Israel Azime

My research focuses on effective domain adaptation and evaluation methods for large language models (LLMs) in low-resource human and data languages. The overarching goal of my work is to advance the development of AI systems that are accurate, interpretable, robust, and culturally grounded, while remaining human-centered and contextually aware. I investigate how LLMs can better understand, reason, and communicate across linguistic, cultural, and structural boundaries, bridging the gap between human meaning and machine representation. Ultimately, my work aims to create socially aware and context-sensitive AI systems that perform reliably across diverse languages and data modalities.

Next
Next

Whose Voice Is the Model Listening For?