NEXT WORD PREDICTION IN YORÙBÁ TEXT: A LONG SHORT TERM MEMORY APPROACH
Main Article Content
Abstract
Yorùbá is a language spoken by over 50 million people and being a tonal and diacritic language poses significant challenge to user on digital space , the lack of tools that support seamless data entry in the language increase the hurdles of communicating with the language on digital space , these has forced younger generations to transition from their mother tongue to a high resource language that has massive support thereby leading to wearing off of the Yorùbá cultural identity in digital space gradually.
This study aim is to develop three LSTM next word based text prediction for Yorùbá language to enhance digital typing and promote linguistic accuracy on social media networks.
The study uses the OpenSLR 86 corpus for model development. This study implements and compares three deep learning models namely the Vanilla LSTM, the Stacked LSTM, and the attention LSTM model.
The findings show that the vanilla LSTM outperforms all other complex models, achieving the highest top k accuracy and lowest perplexity score of 251 as against other models .
These model developed in this study will not only reduces the cognitive load of Yorùbá speakers but will preserve Yorùbá cultural identity across digital space.
Article Details
Upon receipt of accepted manuscripts, authors will be invited to complete a copyright license to publish the paper. At least the corresponding author must send the copyright form signed for publication. It is a condition of publication that authors grant an exclusive licence to the the INFOCOMP Journal of Computer Science. This ensures that requests from third parties to reproduce articles are handled efficiently and consistently and will also allow the article to be as widely disseminated as possible. In assigning the copyright license, authors may use their own material in other publications and ensure that the INFOCOMP Journal of Computer Science is acknowledged as the original publication place.
References
[1] Ajao, J. F., Babatunde, R. S., Asapetu, S. O., and Yusuff, S. R. Implementation of Yorùbá unicode generation for an indigenous keyboard. Techno-science Journal for Community Development in Africa, 1(1):71–80, 2020.
[2] Aliyu, E. O. A deep learning approach for Yoruba language next word generation. In 2024 IEEE 5th International Conference on Electro-Computing Technologies for Humanity (NIGERCON), pages 1–4. IEEE, 2024.
[3] Brigato, L. and Iocchi, L. A close look at deep learning with small data. In 2020 25th Interna-tional Conference on Pattern Recognition (ICPR), pages 2490–2497. IEEE, 2021.
[4] Das, M. and Alphonse, P. J. A. A comparative study on tf-idf feature weighting method and its analysis using unstructured dataset. arXiv preprint arXiv:2308.04037, 2023.
[5] Gutkin, A., Demirs¸ahin, I., Kjartansson, O., Rivera, C., and Túbo. sún, K. Developing an Open-Source Corpus of Yoruba Speech. In Proceed-ings of Interspeech 2020, pages 404–408, Shang-hai, China, October 2020. International Speech and Communication Association (ISCA).
[6] Heo, D., Rim, D. N., and Choi, H. N-gram prediction and word difference represen-tations for language modeling. arXiv preprint arXiv:2409.03295, 2024.
[7] Jimoh, T. A., De Wille, T., and Nikolov, N. S. Bridging gaps in natural language processing for Yorubá: A systematic review of a decade of progress and prospects. Natural Language Pro-cessing Journal, page 100194, 2025.
[8] Kik, A., Adamec, M., Aikhenvald, A. Y., Ba-jzekova, J., Baro, N., Bowern, C., et al. Lan-guage and ethnobiological skills decline pre-cipitously in papua new guinea, the world’s most linguistically diverse nation. Proceed-ings of the National Academy of Sciences, 118(22):e2100096118, 2021.
[9] Mienye, I. D., Swart, T. G., and Obaido, G. Re-current neural networks: A comprehensive review of architectures, variants, and applications. Infor-mation, 15(9):517, 2024.
[10] Minaee, S., Mikolov, T., Nikzad, N., Chenaghlu, M., Socher, R., Amatriain, X., and Gao, J. Large language models: A survey. arXiv preprint arXiv:2402.06196, 2024.
[11] Noh, S. H. Analysis of gradient vanishing of rnns and performance comparison. Information, 12(11):442, 2021.
[12] Olorunfemi, T. O., Azubuike, O. L., and Ojo, O. E. On-screen keyboard with dictionary mapping for Yorùbá language. The Journal of Computer Sci-ence and Its Applications, 27(1):104–115, 2020.
[13] Oluokun, S. O., Ayobami, A. J., and Adefunso,
A. Enhancing Yoruba text autocompletion with an attention-augmented recurrent neural network. International Journal of Research and Innovation in Applied Science, 10(10):1734–1752, 2025.
[14] Oyekanmi, E. O., Al-Turjman, F., and Fadare, O. Toward the realization of a high performing con-tinuous yorùbá speech to text translation using a mobile phone. In Artificial Intelligence Learn-ing Facilitators, pages 153–191. Auerbach Pub-lications, 2025.
[15] Raeini, M. G. The evolution of language models: From n-grams to llms, and beyond. Natural Lan-guage Processing Journal, 12:100168, 2025.