NEXT WORD PREDICTION IN YORÙBÁ TEXT: A LONG SHORT TERM MEMORY APPROACH

Main Article Content

Ayodeji Obaude
Ezekiel Oyekanmi
Joshua Tom

Abstract


Yorùbá is a language spoken by over 50 million people and being a tonal and diacritic language poses significant challenge to user on digital space , the lack of  tools that support seamless data entry in the language  increase the hurdles of communicating with the language on digital space , these has forced younger generations to transition from their mother tongue to a high resource language that has massive support  thereby leading to wearing off of the Yorùbá cultural identity in digital space gradually.
This study aim is to develop three LSTM  next word based text prediction for Yorùbá language to enhance digital typing and promote linguistic accuracy on social media networks.
The study uses the OpenSLR 86 corpus for model development. This study implements and compares three deep learning models namely the Vanilla LSTM, the Stacked LSTM, and the attention LSTM model.
The findings show that the vanilla LSTM outperforms all other complex models, achieving the highest top k accuracy and lowest perplexity score  of 251 as against other models .
These model  developed in this study will not only  reduces the cognitive load of Yorùbá speakers but will preserve Yorùbá  cultural identity across digital space.

Article Details

How to Cite
Obaude, A., Oyekanmi, E., & Tom, J. (2026). NEXT WORD PREDICTION IN YORÙBÁ TEXT: A LONG SHORT TERM MEMORY APPROACH. INFOCOMP Journal of Computer Science, 25(1), e5467. https://doi.org/10.18760/v25.5467
Section
Machine Learning and Computational Intelligence

References

[1] Ajao, J. F., Babatunde, R. S., Asapetu, S. O., and Yusuff, S. R. Implementation of Yorùbá unicode generation for an indigenous keyboard. Techno-science Journal for Community Development in Africa, 1(1):71–80, 2020.

[2] Aliyu, E. O. A deep learning approach for Yoruba language next word generation. In 2024 IEEE 5th International Conference on Electro-Computing Technologies for Humanity (NIGERCON), pages 1–4. IEEE, 2024.

[3] Brigato, L. and Iocchi, L. A close look at deep learning with small data. In 2020 25th Interna-tional Conference on Pattern Recognition (ICPR), pages 2490–2497. IEEE, 2021.

[4] Das, M. and Alphonse, P. J. A. A comparative study on tf-idf feature weighting method and its analysis using unstructured dataset. arXiv preprint arXiv:2308.04037, 2023.

[5] Gutkin, A., Demirs¸ahin, I., Kjartansson, O., Rivera, C., and Túbo. sún, K. Developing an Open-Source Corpus of Yoruba Speech. In Proceed-ings of Interspeech 2020, pages 404–408, Shang-hai, China, October 2020. International Speech and Communication Association (ISCA).

[6] Heo, D., Rim, D. N., and Choi, H. N-gram prediction and word difference represen-tations for language modeling. arXiv preprint arXiv:2409.03295, 2024.

[7] Jimoh, T. A., De Wille, T., and Nikolov, N. S. Bridging gaps in natural language processing for Yorubá: A systematic review of a decade of progress and prospects. Natural Language Pro-cessing Journal, page 100194, 2025.

[8] Kik, A., Adamec, M., Aikhenvald, A. Y., Ba-jzekova, J., Baro, N., Bowern, C., et al. Lan-guage and ethnobiological skills decline pre-cipitously in papua new guinea, the world’s most linguistically diverse nation. Proceed-ings of the National Academy of Sciences, 118(22):e2100096118, 2021.

[9] Mienye, I. D., Swart, T. G., and Obaido, G. Re-current neural networks: A comprehensive review of architectures, variants, and applications. Infor-mation, 15(9):517, 2024.

[10] Minaee, S., Mikolov, T., Nikzad, N., Chenaghlu, M., Socher, R., Amatriain, X., and Gao, J. Large language models: A survey. arXiv preprint arXiv:2402.06196, 2024.

[11] Noh, S. H. Analysis of gradient vanishing of rnns and performance comparison. Information, 12(11):442, 2021.

[12] Olorunfemi, T. O., Azubuike, O. L., and Ojo, O. E. On-screen keyboard with dictionary mapping for Yorùbá language. The Journal of Computer Sci-ence and Its Applications, 27(1):104–115, 2020.

[13] Oluokun, S. O., Ayobami, A. J., and Adefunso,

A. Enhancing Yoruba text autocompletion with an attention-augmented recurrent neural network. International Journal of Research and Innovation in Applied Science, 10(10):1734–1752, 2025.

[14] Oyekanmi, E. O., Al-Turjman, F., and Fadare, O. Toward the realization of a high performing con-tinuous yorùbá speech to text translation using a mobile phone. In Artificial Intelligence Learn-ing Facilitators, pages 153–191. Auerbach Pub-lications, 2025.

[15] Raeini, M. G. The evolution of language models: From n-grams to llms, and beyond. Natural Lan-guage Processing Journal, 12:100168, 2025.