Shuwa Arabic voice dataset

Voice datasets are structured collections of audio recordings paired with corresponding text transcriptions, metadata, and annotations. These datasets serve as the foundation for training and evaluating speech recognition systems, text-to-speech engines, and other voice-enabled applications. High-quality voice datasets are essential for developing accurate and robust speech technologies.  Voice models are machine learning systems trained on […]

Kanuri TTS and ASR models

Voice datasets are structured collections of audio recordings paired with corresponding text transcriptions, metadata, and annotations. These datasets serve as the foundation for training and evaluating speech recognition systems, text-to-speech engines, and other voice-enabled applications. High-quality voice datasets are essential for developing accurate and robust speech technologies. Voice models are machine learning systems trained on […]

Playbook for voice data collection for low-resource languages

The TWB Voice Playbook is a practical guide to planning and managing voice data collection projects for low-resource languages. It is aimed at both new and experienced teams and covers the full process, from setting up the project to publishing your dataset.  The playbook draws on CLEAR Global’s experience developing and running TWB Voice, a […]

Chichewa synthetic voice dataset, TTS models, ASR models

Below is a curated collection of open resources for text-to-speech (TTS), automatic speech recognition (ASR), and synthetic voice datasets in the Chichewa language.  Text-to-Speech (TTS) Models Explore our collection of Chichewa TTS models, including XTTS and other multilingual models fine-tuned for natural-sounding Chichewa speech. View Chichewa TTS Models on Hugging Face Automatic Speech Recognition (ASR) Models […]

Hausa synthetic voice dataset, TTS models, ASR models

Below is a curated collection of open resources for text-to-speech (TTS), automatic speech recognition (ASR), and synthetic voice datasets in the Hausa language.  Text-to-Speech (TTS) Models Explore our collection of Hausa TTS models, including XTTS and other multilingual models fine-tuned for natural-sounding Hausa speech. View Hausa TTS Models on Hugging Face Automatic Speech Recognition (ASR) […]

Dholuo synthetic voice dataset, TTS models, ASR models

Below is a curated collection of open resources for text-to-speech (TTS), automatic speech recognition (ASR), and synthetic voice datasets in the Dholuo language. Text-to-Speech (TTS) Models Explore our collection of Dholuo TTS models, including XTTS and other multilingual models fine-tuned for natural-sounding Dholuo speech. View Dholuo TTS Models on Hugging Face Automatic Speech Recognition (ASR) […]

Marma TTS and text data resources

Text data This dataset contains sentences in the Marma language (ISO code: rmz), with both original and normalized forms. The dataset is designed to support language technology development for the Marma language, a Tibeto-Burman language spoken primarily by the Marma people in Bangladesh and Myanmar. This dataset was created as part of a project funded […]

Human-Centred Technology Design in Humanitarian Action

Guides on co-creating digital tools with crisis-affected people On this page, you will find two essential guides on Human-Centred Technology Design in Humanitarian Action. These resources, developed by CLEAR Global, are part of an initiative funded by the UK Humanitarian Innovation Hub (UKHIH) aimed at establishing participatory and community driven approaches for developing digital technology […]