Why the data behind voice technology matters
A few weeks ago, ICTWorks raised a valid alarm: voice-only interfaces are often failing small businesses in low- and middle-income countries (LMICs). At CLEAR Global, we’ve spent years documenting that the problem isn’t the interface, but the foundational data. Voice technology is currently built on a digital divide that favors a handful of dominant languages and excludes the linguistic realities of the four billion people who speak the rest. The starting point for this is the enormous difference between the availability of language data for dominant versus low-resource languages.
For voice to fulfill its promise, we must start investing in gathering quality data that reflects the real world.
Without quality language data, the technology will never catch up
Automatic speech recognition, text-to-speech systems, and voice assistants all depend on vast amounts of recorded speech to learn the sounds, rhythms, and patterns of a language. When that data is scarce, unrepresentative, or collected under controlled conditions that bear no resemblance to how people actually speak, the resulting tools perform poorly in the real world.
The reasons why entire societies have been historically overlooked are the same ones that now keep them away from technology. Overlooking the need for quality language data will only widen the digital divide and, as a consequence, people’s access to opportunities.
The oral reality: why voice is still the way forward
This is a question of basic access and right to information. As many as 7,000 languages are spoken today, two-thirds of which do not have a written form. In many low- and middle-income countries, English, French or another colonial language is the teaching medium. But that doesn’t mean that the majority or even most people are comfortable reading and writing them. For example, in parts of Mozambique, according to the Ministry of State Administration, 81% of women are not literate, yet only 20% understand Portuguese.
For these communities, text-based interfaces, which are often built in colonial languages, are a non-starter. Multilingual voice technology has the power to overcome language and literacy barriers, allowing users to navigate digital spaces and access vital information in the audio form that is easiest for them.
Why thinking “technology will get us there” is a mirage
Organizations often believe the claim that Big Tech “supports” hundreds of languages. In reality, high-quality voice technology is only available for a handful of languages. Voice tools and models are typically directionally interesting projects, but they deteriorate rapidly beyond the languages that tech companies consider to be commercially viable.
Even when the “right” language is supported, the tech often fails because:
- It ignores variants: models for Portuguese are usually geared toward European or Brazilian variants and fail users in Angola or Mozambique.
- It lacks diversity: existing voice datasets are often narrow, dominated by young, male voices and urban accents. This excludes rural women who may use different terms or accents.
- It can’t code-switch: in the real world, speakers frequently mix local languages with a lingua franca. Most standard models cannot recognize or reproduce this.
Building data for the real world
If we want voice technology to work for a female farmer in Niger or a small business owner in Bangladesh, the data must reflect their lives. This requires a shift from off-the-shelf solutions to fit-for-purpose technology.
CLEAR Global and others focus on building foundational data; a repository of spoken and written datasets for underserved languages. To be effective, this data must:
- Include natural speech from community members of different ages and genders to minimize bias.
- Vary audio quality and include background noise to reflect real-life conditions rather than laboratory settings.
- Ensure informed consent, particularly in conflict zones where voice data could identify an individual speaker.
Perhaps most importantly, the people that speak these languages should be involved in all stages of the development of the data and technology. This means inclusion efforts that go far beyond the validation of decisions that have already been taken. Involving communities at every stage, from defining what data is collected to deciding how models are trained and deployed, ensures that the resulting tools are culturally appropriate, practically useful, and trusted by the people who will use them. Without this foundation, even well-intentioned projects risk repeating historical patterns in which outside institutions have defined, categorized, and controlled the languages and identities of marginalized communities.
The investment gap
Building quality voice functionalities is not an optional “add-on” that can be rushed in the later stages of a tech project. It requires a collective approach where organizations, tech developers, and communities work together to make cohesive and culturally fit products.
We cannot wait for the perfect tool to arrive from Silicon Valley. Impact is ultimately determined by usage, and access alone is not sufficient. If we want people to thrive, we must stop building on flawed foundations and start investing in language technology that makes it possible to hear them.


Aimee Ansari
Chief Executive, CLEAR Global