LibriSpeech-Gloss: A Large-Scale Sign Language Gloss Dataset using TxtSLProcess Framework
Main Article Content
Abstract
Sign language datasets often have limited vocabulary and are usually word-based, which restricts their use in real-world applications. To address this, this work introduces LibriSpeech-Gloss, a large-scale sign language gloss dataset generated from the LibriSpeech-PC corpus using an extended ver
sion of our previously published TxtSLProcess text-to-gloss toolkit. In our prior work, TxtSLProcess was introduced as part of the EqualWe framework for real-time Speech-to-Sign Language conversion. In this work, we extend TxtSLProcess with an expanded negation map, comprehensive verb normalization with learned suffix rules and enhanced contraction handling to support large-scale offline dataset generation. The LibriSpeech-Gloss dataset is generated for all LibriSpeech-PC partitions including train-clean-100, train-clean-360, train-other-500, dev-clean, dev-other, test-clean and test-other covering approximately 981 hours of aligned English-to-gloss sentence pairs totaling 266,655 utterances across seven evaluation partitions. The dataset provides general-domain coverage with diverse sentence structures and significantly exceeds existing sign language gloss datasets in scale and vocabulary diversity, making it avaluable resource for advancing sign language translation research.
Article Details
Upon receipt of accepted manuscripts, authors will be invited to complete a copyright license to publish the paper. At least the corresponding author must send the copyright form signed for publication. It is a condition of publication that authors grant an exclusive licence to the the INFOCOMP Journal of Computer Science. This ensures that requests from third parties to reproduce articles are handled efficiently and consistently and will also allow the article to be as widely disseminated as possible. In assigning the copyright license, authors may use their own material in other publications and ensure that the INFOCOMP Journal of Computer Science is acknowledged as the original publication place.
References
[1] Shree, B., & Chawla, S. (2026). A framework for speech-to-sign language conversion for hearing. Proceedings of Sixth Doctoral Symposium on Computational Intelligence: DoSCI 2025, Vol. 3, 395.
[2] Loper, E., & Bird, S. (2002). NLTK: The natural language toolkit. Proceedings of the ACL-02 Workshop on Effective Tools and Methodologies for Teaching Natural Language Processing and Computational Linguistics, 63–70.
[3] El-Gayyar, M. M., Ibrahim, A. S., & Wahed, M. E. (2016). Translation from Arabic speech to Arabic sign language based on cloud computing. Egyptian Informatics Journal, 17(3), 295–303.
[4] Shahriar, R., Zaman, A. G. M., Ahmed, T., Khan, S. M., & Maruf, H. M. (2017). A communication platform between Bangla and sign language. IEEE R10 Humanitarian Technology Conference, 1–4.
[5] Panayotov, V., Chen, G., Povey, D., & Khudanpur, S. (2015). LibriSpeech: An ASR corpus based on public domain audio books. ICASSP, 5206–5210.
[6] Meister, A., Novikov, M., Karpov, N., Bakhturina, E., Lavrukhin, V., & Ginsburg, B. (2023). LibriSpeech-PC: Benchmark for punctuation and capitalization in ASR models. IEEE ASRU Workshop, 1–7.
[7] Camgoz, N. C., Hadfield, S., Koller, O., Ney, H., &Bowden, R.(2018). Neural sign language translation. CVPR, 7784–7793.
[8] Koller, O. (2020). Quantitative survey of sign language recognition. arXiv preprint arXiv:2008.09918.
[9] Duarte, A., Palaskar, S., Ventura, L., Ghadiyaram, D., DeHaan, K., Metze, F., Torres, J., & Giro-i Nieto, X. (2021). How2Sign: A large-scale multi modal dataset. CVPR, 2735–2744.
[10] Stoll, S., Camgoz, N. C., Hadfield, S., & Bowden, R. (2020). Text2Sign: Neural machine translation and GANs. International Journal of Computer Vision, 128(4), 891–908.
[11] Sayed, S. A., Seoud, R. A. A. A., & Naby, H. Y. A. (2024). CNNs for Arabic speech recognition. Journal of Electrical and Computer Engineering, 2024.
[12] Othman, A., & Jemni, M. (2012). ASLG-PC12 corpus. LREC Workshop, 151–154.
[13] Elnashar, A., Hamdan, K., Al Seiari, S., Shanableh, Y., & Barlas, G. (2024). Bi-directional translation methods. Procedia Computer Science, 239, 1879–1886.
[14] Naik, S., Golivadekar, M., Sreemathy, R., & Turuk, M. P. (2025). Text-to-gloss translation model. Springer ICIVC, 343–357.
[15] Amiruzzaman, S., Amiruzzaman, M., Batchu, R. M., Dracup, J., Pham, A., Crocker, B., Ngo, L., & Dewan, M. A. A. (2026). Bidirectional ASL
translation. Computers, 15(1), 20.
[16] Nguyen-Xuan, S., Le, G. D., & Nguyen, H. (2025). Gloss annotation scheme. Springer, 30–45.