DOI Number : 10.5614/itbj.ict.res.appl.2013.7.2.5
A Proposed Arabic Handwritten Text Normalization Method

Tarik Abu-Ain1, Siti Norul Huda Sheikh Abdullah1, Khairuddin Omar1, Ashraf Abu-Ein2, Bilal Bataineh1 & Waleed Abu-Ain1

1Pattern Recognition Research Group, Center for Artificial Intelligence Technology, Faculty of Information Science and Technology, Universiti Kebangsaan Malaysia, 43600, Bangi, Selangor, Malaysia.
2Computer Engineering Department, Al-Balqa' Applied University, Faculty of Engineering Technology, 15008 Amman, 11134 Jordan.


Text normalization is an important technique in document image analysis and recognition. It consists of many preprocessing stages, which include slope correction, text padding, skew correction, and straight the writing line. In this side, text normalization has an important role in many procedures such as text segmentation, feature extraction and characters recognition. In the present article, a new method for text baseline detection, straightening, and slant correction for Arabic handwritten texts is proposed. The method comprises a set of sequential steps: first components segmentation is done followed by components text thinning; then, the direction features of the skeletons are extracted, and the candidate baseline regions are determined. After that, selection of the correct baseline region is done, and finally, the baselines of all components are aligned with the writing line.  The experiments are conducted on IFN/ENIT benchmark Arabic dataset. The results show that the proposed method has a promising and encouraging performance.

Keywords: Arabic handwriting, baseline detection, preprocessing, slant correction, sub-word extraction, handwritten text normalization

