Playbook for voice data collection for low-resource languages

The TWB Voice Playbook is a practical guide to planning and managing voice data collection projects for low-resource languages. It is aimed at both new and experienced teams and covers the full process, from setting up the project to publishing your dataset.  The playbook draws on CLEAR Global’s experience developing and running TWB Voice, a […]

Chichewa synthetic voice dataset, TTS models, ASR models

Below is a curated collection of open resources for text-to-speech (TTS), automatic speech recognition (ASR), and synthetic voice datasets in the Chichewa language.  Text-to-Speech (TTS) Models Explore our collection of Chichewa TTS models, including XTTS and other multilingual models fine-tuned for natural-sounding Chichewa speech. View Chichewa TTS Models on Hugging Face Automatic Speech Recognition (ASR) Models […]

Hausa synthetic voice dataset, TTS models, ASR models

Below is a curated collection of open resources for text-to-speech (TTS), automatic speech recognition (ASR), and synthetic voice datasets in the Hausa language.  Text-to-Speech (TTS) Models Explore our collection of Hausa TTS models, including XTTS and other multilingual models fine-tuned for natural-sounding Hausa speech. View Hausa TTS Models on Hugging Face Automatic Speech Recognition (ASR) […]

Dholuo synthetic voice dataset, TTS models, ASR models

Below is a curated collection of open resources for text-to-speech (TTS), automatic speech recognition (ASR), and synthetic voice datasets in the Dholuo language. Text-to-Speech (TTS) Models Explore our collection of Dholuo TTS models, including XTTS and other multilingual models fine-tuned for natural-sounding Dholuo speech. View Dholuo TTS Models on Hugging Face Automatic Speech Recognition (ASR) […]

Marma TTS and text data resources

Text data This dataset contains sentences in the Marma language (ISO code: rmz), with both original and normalized forms. The dataset is designed to support language technology development for the Marma language, a Tibeto-Burman language spoken primarily by the Marma people in Bangladesh and Myanmar. This dataset was created as part of a project funded […]

Language technology for humanitarian action

Language technology for humanitarian action How collaborating on language technology development could reduce the digital language gap   Up to four billion people globally face digital or language exclusion. Developing language technology for humanitarian and development use represents a significant opportunity to reduce exclusion across the aid sector and beyond. This brief proposes a collective […]