AUT Journal of Modeling and Simulation

AUT Journal of Modeling and Simulation

Persian Typographical Error Type Detection using Many-to-Many Deep Neural Networks on Algorithmically-Generated Misspellings

Document Type : Research Article

Authors
School of Electrical and Computer Engineering, College of Engineering, University of Tehran, Tehran, Iran
10.22060/miscj.2026.25563.5482
Abstract
Spelling correction is a challenge in the field of natural language processing. The objective of spelling correction tasks is to recognize and rectify spelling errors automatically. The development of applications that can diagnose and correct Persian spelling errors has become more important in order to improve the quality of text. Typographical Error Type Detection in Persian is a relatively understudied area. Therefore, this paper presents an approach for detecting typographical errors in Persian texts. Our work includes the presentation of a publicly available dataset called FarsTypo, which comprises 3.4 million words arranged in chronological order and tagged with their corresponding part-of-speech. These words cover a wide range of topics and linguistic styles. We develop an algorithm designed to apply Persian-specific errors to a subset of these words. By leveraging FarsTypo, we establish a baseline and conduct a comparison of various methodologies employing different architectures. We propose a Deep Sequential Neural Network for token-level error type classification. The model incorporates three key components: joint word-level and character-level embeddings to handle both in-vocabulary and out-of-vocabulary words, bidirectional LSTM layers to capture contextual dependencies in text, and a TimeDistributed classifier for predicting 51 error categories. Experimental results demonstrate that the proposed model achieves 97.62% accuracy, 98.83% precision, and 98.61% recall on the constructed dataset. In addition, the model shows higher inference efficiency compared to industrial spell-checking systems. The results indicate that separating error detection from correction and leveraging combined word–character sequential modeling improves typographical error type detection.
Keywords
Subjects


Articles in Press, Accepted Manuscript
Available Online from 11 August 2026