graviti
产品公开数据集应用市场解决方案知识库关于我们
633
1
2
MAGICDATAMandarinChineseReadSpeechCorpus_1
来自MagicHub
概要
讨论
代码
活动
MAGICDATA Mandarin Chinese Read Speech Corpus
aedf50b·
Aug 19, 2021 3:05 AM
·1Commits

Overview

MAGICDATA Mandarin Chinese Read Speech Corpus was developed by MAGIC DATA Technology Co., Ltd. and freely published for non-commercial use. The contents and the corresponding descriptions of the corpus include:

The corpus contains 755 hours of speech data, which is mostly mobile recorded data. 1080 speakers from different accent areas in China are invited to participate in the recording. The sentence transcription accuracy is higher than 98%. Recordings are conducted in a quiet indoor environment. The database is divided into training set, validation set, and testing set in a ratio of 51: 1: 2. Detail information such as speech data coding and speaker information is preserved in the metadata file. The domain of recording texts is diversified, including interactive Q&A, music search, SNS messages, home command and control, etc. Segmented transcripts are also provided. The corpus aims to support researchers in speech recognition, machine translation, speaker recognition, and other speech-related fields. Therefore, the corpus is totally free for academic use. The corpus is a subset of a much bigger data ( 10566.9 hours Chinese Mandarin Speech Corpus ) set which was recorded in the same environment. Please feel free to contact us via business@magicdatatech.com for more details.

语料库包含755小时的语音数据,主要是移动终端的录音数据,邀请来自中国不同口音区域的1080名演讲者参与录制,句子转录准确率高达98%以上。 录音在安静的室内环境中进行。数据库分为训练集,验证集和测试集,比例为51:1:2。诸如语音数据编码和说话者信息的细节信息被保存在元数据文件中。 录音文本十分多样化,包括互动问答,音乐搜索,SNS信息,家庭指挥和控制等,还提供了分段的转录文本。

欢迎business@magicdatatech.com

数据预览
查看数据
🎉感谢MagicHub的贡献
数据集信息
应用场景NLP
标注类型暂无
任务类型ASR
LicenseCustom
更新时间2021-08-19 03:05:49
数据概要
数据格式Audio
数据数量609.55K
已标注数量611550
文件大小81GB
版权归属方
MagicHub
标注方
未知
了解更多和支持