Toward enhanced Arabic speech recognition using part of speech tagging

AbuZeina, Dia; Al-Khatib, Wasfi; Elshafei, Moustafa; Al-Muhtaseb, Husni

doi:10.1007/s10772-011-9121-5

Toward enhanced Arabic speech recognition using part of speech tagging

Published: 12 October 2011

Volume 14, pages 419–426, (2011)
Cite this article

International Journal of Speech Technology Aims and scope Submit manuscript

Dia AbuZeina¹,
Wasfi Al-Khatib¹,
Moustafa Elshafei¹ &
…
Husni Al-Muhtaseb¹

234 Accesses
6 Citations
Explore all metrics

Abstract

One major source of suboptimal performance in automatic continuous speech recognition systems is misrecognition of small words. In general, errors resulting from small words are much more than errors resulting from long words. Therefore, compounding some words (small or long) to produce longer words is welcome by speech recognition decoders. In this paper, we present a novel approach to artificially generate compound words using part of speech tagging. For this purpose, we consider two cases in Arabic speech where two words are pronounced without a silence period in between: a noun followed by an adjective, and a preposition followed by any word. To collect the candidate compound words, we use Stanford Arabic tagger to tag all words in our baseline transcription corpus. Then, compound words are generated whenever any of the two cases occur in a sequence of two words. The unique compound words are then added to the expanded pronunciation dictionary, as well as to the language model. Using Sphinx 3, we test the proposed method for a 5.4 hours speech corpus of modern standard Arabic. The results show a significant improvement, as the word error rate is reduced by 2.39%.

This is a preview of subscription content, log in via an institution to check access.

Access this article

Log in via an institution

Price excludes VAT (USA)
Tax calculation will be finalised during checkout.

Instant access to the full article PDF.

Institutional subscriptions

Diacritics Effect on Arabic Speech Recognition

Article 10 July 2019

The impact of phonological rules on Arabic speech recognition

Article 24 July 2017

Arabic Speech Processing: State of the Art and Future Outlook

References

AbuZeina, D., Al-Khatib, W., Elshafei, M., & Al-Muhtaseb, H. (2011). Cross-word Arabic pronunciation variation modeling for speech recognition. International Journal of Speech Technology.
Albared, M., Omar, N., Ab Aziz, M., Zakree, M., & Nazri, A. (2010). Automatic part of speech tagging for Arabic: an experiment using Bigram hidden Markov model. In RSKT’10 proceedings of the 5th international conference on rough set and knowledge technology.
Google Scholar
Alghamdi, M., Almuhtasib, H., & Elshafei, M. (2004). Arabic phonological rules. Journal of King Saud University: Computer and Information Sciences, 16, 1–25.
Google Scholar
Ali, M., Elshafei, M., Alghamdi, M., Almuhtaseb, H., & Alnajjar, A. (2009). Arabic phonetic dictionaries for speech recognition. Journal of Information Technology Research, 2(4), 67–80.
Article Google Scholar
Al Shamsi, F., & Guessoum, A. (2006). A hidden Markov model-based PS tagger for Arabic.
Google Scholar
Berton, A., Fetter, P., & Regel-Brietzmann, P. (1996). Compound words in large-vocabulary German speech recognition systems. In Proceedings of the international conference of speech and language processing (ICSLP).
Google Scholar
Diab, M., Hacioglu, K., & Jurafsky, D. (2004). Automatic tagging of Arabic text: from raw text to base phrase chunks. In 5th meeting of the North American chapter of the Association for Computational Linguistics/Human Language Technologies conference.
Google Scholar
Elshafei, A. M. (1991). Toward an Arabic text-to-speech system. Arabian Journal for Science and Engineering, 16(4B), 565–583.
MathSciNet Google Scholar
Elshafei, M., Almuhtasib, H., & Alghamdi, M. (2002). Techniques for high quality text-to-speech. Information Sciences, 140(3–4), 255–267.
Article MATH Google Scholar
El Hadj, Y., Al-Sughayeir, I., & Al-Ansari, A. (2009). Arabic part-of-speech tagging using the sentence structure. In Proceedings of the second international conference on Arabic language resources and tools. The MEDAR Consortium.
Finke, M., & Waibel, A. (1997). Speakingmode dependent pronunciation modeling in large vocabulary conversational speech recognition. In EUROSPEECH-1997 (pp. 2379–2382).
Google Scholar
IPA for Arabic (2011). http://en.wikipedia.org/wiki/Wikipedia:IPA_for_Arabic.
Jelinek, F. (1999). Statistical methods for speech recognition. In Language, speech and communication series. Cambridge: MIT Press.
Google Scholar
Lei, X., Wang, W., & Stolcke, A. (2009). Data-driven lexicon expansion for Mandarin broadcast news and conversation speech recognition. In IEEE international conference on acoustics, speech and signal processing. ICASSP 2009 (pp. 4329–4332). 19–24 April 2009. doi:10.1109/ICASSP.2009.4960587.
Google Scholar
MITCogNet (2010). http://mitpdev.mit.edu/library/erefs/arbib/images/figures/A248_fig001.gif.
Plötz, T. Advanced stochastic protein sequence analysis. Ph.D. Thesis, Bielefeld University, June 2005.
Stanford Log-linear Part-of-Speech Tagger (2011). http://nlp.stanford.edu/software/tagger.shtml.
Siegler, M. A., & Stern, R. M. (1995). On the effects of speech rate in large vocabulary speech recognition systems. In Proc. IEEE int. conf. acoust. speech signal process (Vol. 1, pp. 612–615). Detroit.
Google Scholar
Saon, G., & Padmanabhan, M. (2001). Data-driven approach to designing compound words for continuous speech recognition. IEEE Transactions on Speech and Audio Processing, 9(4):327–332
Article Google Scholar
Sloboda, T., & Waibel, A. (1996). Dictionary learning for spontaneous speech recognition. In Proc. ICSLP-96 (pp. 2328–2331). Philadelphia, PA, USA.
Google Scholar
Zhang, J., Gao, J., & Zhou, M. (2000). Extraction of Chinese compound words: an experimental study on a very large corpus. In Proceedings of the second workshop on Chinese language processing: held in conjunction with the 38th annual meeting of the Association for Computational Linguistics, 8 October, Hong Kong.
Google Scholar

Download references

Author information

Authors and Affiliations

King Fahd University of Petroleum and Minerals, Dhahran, Saudi Arabia
Dia AbuZeina, Wasfi Al-Khatib, Moustafa Elshafei & Husni Al-Muhtaseb

Authors

Dia AbuZeina
View author publications
You can also search for this author in PubMed Google Scholar
Wasfi Al-Khatib
View author publications
You can also search for this author in PubMed Google Scholar
Moustafa Elshafei
View author publications
You can also search for this author in PubMed Google Scholar
Husni Al-Muhtaseb
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Dia AbuZeina.

Rights and permissions

Reprints and permissions

About this article

Cite this article

AbuZeina, D., Al-Khatib, W., Elshafei, M. et al. Toward enhanced Arabic speech recognition using part of speech tagging. Int J Speech Technol 14, 419–426 (2011). https://doi.org/10.1007/s10772-011-9121-5

Download citation

Received: 31 July 2011
Accepted: 28 September 2011
Published: 12 October 2011
Issue Date: December 2011
DOI: https://doi.org/10.1007/s10772-011-9121-5

Keywords

Access this article

Log in via an institution

Price excludes VAT (USA)
Tax calculation will be finalised during checkout.

Instant access to the full article PDF.

Institutional subscriptions

Toward enhanced Arabic speech recognition using part of speech tagging

Abstract

Access this article

Similar content being viewed by others

Diacritics Effect on Arabic Speech Recognition

The impact of phonological rules on Arabic speech recognition

Arabic Speech Processing: State of the Art and Future Outlook

References

Author information

Authors and Affiliations

Corresponding author

Rights and permissions

About this article

Cite this article

Keywords

Navigation

Toward enhanced Arabic speech recognition using part of speech tagging

Abstract

Access this article

Similar content being viewed by others

Diacritics Effect on Arabic Speech Recognition

The impact of phonological rules on Arabic speech recognition

Arabic Speech Processing: State of the Art and Future Outlook

References

Author information

Authors and Affiliations

Corresponding author

Rights and permissions

About this article

Cite this article

Share this article

Keywords

Search

Navigation