Semantic-Aware Deep Neural Attention Network for Machine Translation Detection

Shi, Yangbin; Lu, Jun; Gu, Shuqin; Wang, Qiang; Zheng, Xiaolin

doi:10.1007/978-981-16-7512-6_6

Yangbin Shi^7,8,
Jun Lu⁸,
Shuqin Gu⁸,
Qiang Wang⁸ &
…
Xiaolin Zheng⁷

Part of the book series: Communications in Computer and Information Science ((CCIS,volume 1464))

Included in the following conference series:

China Conference on Machine Translation

282 Accesses

Abstract

Web crawling is an important way to collect a massive training corpus for building a high-quality machine translation system. However, a large amount of data collected comes from machine-translated texts rather than native speakers or professional translators, severely reducing the benefit of data scale. Traditional machine translation detection methods generally require human-crafted feature engineering and are difficult to distinguish the fine-grained semantic difference between real text and pseudo text from a modern neural machine translation system. To address this problem, we propose two semantic-aware models based on the deep neural network to automatically learn semantic features of text for monolingual scenarios and bilingual scenarios, respectively. Specifically, our models incorporate the global semantic from BERT and the local semantic from convolutional neural network together for monolingual detection and further explores the semantic consistency relationship for bilingual detection. The experimental results on the Chinese-English machine translation detection task show that our models achieve 83.12% \(F_{1}\) in the monolingual detection and 85.53% \(F_{1}\) in the bilingual detection respectively, which is better than the strong BERT baselines by 2.2–3.2%.

Supported in part by the National Key R&D Program of China (No. 2018YFB1403001).

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 54.99; Price excludes VAT (USA)

Softcover Book: USD 69.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Notes

1.
http://www.statmt.org/wmt18/translation-task.html.
2.
To our knowledge, all the four machine translators are NMT systems.
3.
https://www.nltk.org/.
4.
https://pytorch.org/.
5.
http://www.statmt.org/wmt17/translation-task.html.

References

Aharoni, R., Koppel, M., Goldberg, Y.: Automatic detection of machine translated text and translation quality estimation. In: Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), vol. 2, pp. 289–295 (2014)
Google Scholar
Antonova, A., Misyurev, A.: Building a web-based parallel corpus and filtering out machine-translated text. In: Proceedings of the 4th Workshop on Building and Using Comparable Corpora: Comparable Corpora and the Web, pp. 136–144. Association for Computational Linguistics (2011)
Google Scholar
Arase, Y., Zhou, M.: Machine translation detection from monolingual web-text. In: Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), vol. 1, pp. 1597–1607 (2013)
Google Scholar
Bahdanau, D., Cho, K., Bengio, Y.: Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473 (2014)
Biçici, E., Yuret, D.: Instance selection for machine translation using feature decay algorithms. In: Proceedings of the Sixth Workshop on Statistical Machine Translation, pp. 272–283. Association for Computational Linguistics (2011)
Google Scholar
Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., Specia, L.: SemEval-2017 task 1: semantic textual similarity-multilingual and cross-lingual focused evaluation. arXiv preprint arXiv:1708.00055 (2017)
Chen, B., Kuhn, R., Foster, G., Cherry, C., Huang, F.: Bilingual methods for adaptive training data selection for machine translation. In: Proceedings of AMTA, pp. 93–103 (2016)
Google Scholar
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: BERT: pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)
Eetemadi, S., Lewis, W., Toutanova, K., Radha, H.: Survey of data-selection methods in statistical machine translation. Mach. Transl. 29, 189–223 (2015). https://doi.org/10.1007/s10590-015-9176-1
Article Google Scholar
Junczys-Dowmunt, M., et al.: Marian: fast neural machine translation in C++. In: Proceedings of ACL 2018, System Demonstrations, pp. 116–121. Association for Computational Linguistics, Melbourne, July 2018. http://www.aclweb.org/anthology/P18-4020
Kim, Y.: Convolutional neural networks for sentence classification. arXiv preprint arXiv:1408.5882 (2014)
Lison, P., Tiedemann, J.: OpenSubtitles 2016: extracting large parallel corpora from movie and tv subtitles (2016)
Google Scholar
Ma, M., Nirschl, M., Biadsy, F., Kumar, S.: Approaches for neural-network language model adaptation. In: Proceedings of Interspeech, Stockholm, Sweden, pp. 259–263 (2017)
Google Scholar
Moore, R.C., Lewis, W.: Intelligent selection of language model training data (2010)
Google Scholar
Papineni, K., Roukos, S., Ward, T., Zhu, W.J.: BLEU: a method for automatic evaluation of machine translation. In: Proceedings of the ACL, pp. 311–318 (2002). https://doi.org/10.3115/1073083.1073135. http://dx.doi.org/10.3115/1073083.1073135
Peris, Á., Chinea-Ríos, M., Casacuberta, F.: Neural networks classifier for data selection in statistical machine translation. Prague Bull. Math. Linguist. 108(1), 283–294 (2017)
Article Google Scholar
Poncelas, A., Shterionov, D., Way, A., Wenniger, G.M.D.B., Passban, P.: Investigating backtranslation in neural machine translation. arXiv preprint arXiv:1804.06189 (2018)
Rarrick, S., Quirk, C., Lewis, W.: MT detection in web-scraped parallel corpora. In: Proceedings of the Machine Translation Summit (MT Summit XIII) (2011)
Google Scholar
Resnik, P., Smith, N.A.: The web as a parallel corpus. Comput. Linguist. 29(3), 349–380 (2003)
Article Google Scholar
Sennrich, R., Haddow, B., Birch, A.: Improving neural machine translation models with monolingual data. arXiv preprint arXiv:1511.06709 (2015)
Shao, Y.: HCTI at SemEval-2017 task 1: use convolutional neural network to evaluate semantic textual similarity. In: Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pp. 130–133 (2017)
Google Scholar
Sutskever, I., Vinyals, O., Le, Q.V.: Sequence to sequence learning with neural networks. In: Advances in Neural Information Processing Systems, pp. 3104–3112 (2014)
Google Scholar
Vaswani, A., et al.: Attention is all you need. In: Advances in Neural Information Processing Systems, pp. 5998–6008 (2017)
Google Scholar
Zeiler, M.D.: ADADELTA: an adaptive learning rate method. arXiv preprint arXiv:1212.5701 (2012)
Zens, R., Och, F.J., Ney, H.: Phrase-based statistical machine translation. In: Jarke, M., Lakemeyer, G., Koehler, J. (eds.) KI 2002. LNCS (LNAI), vol. 2479, pp. 18–32. Springer, Heidelberg (2002). https://doi.org/10.1007/3-540-45751-8_2
Chapter Google Scholar

Download references

Author information

Authors and Affiliations

College of Computer Science and Technology, Zhejiang University, Hangzhou, China
Yangbin Shi & Xiaolin Zheng
Machine Intelligence Technology Lab, Alibaba Group, Hangzhou, China
Yangbin Shi, Jun Lu, Shuqin Gu & Qiang Wang

Authors

Yangbin Shi
View author publications
You can also search for this author in PubMed Google Scholar
Jun Lu
View author publications
You can also search for this author in PubMed Google Scholar
Shuqin Gu
View author publications
You can also search for this author in PubMed Google Scholar
Qiang Wang
View author publications
You can also search for this author in PubMed Google Scholar
Xiaolin Zheng
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Xiaolin Zheng .

Editor information

Editors and Affiliations

Xiamen University, Xiamen, China
Jinsong Su
The University of Edinburgh, Edinburgh, UK
Rico Sennrich

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Shi, Y., Lu, J., Gu, S., Wang, Q., Zheng, X. (2021). Semantic-Aware Deep Neural Attention Network for Machine Translation Detection. In: Su, J., Sennrich, R. (eds) Machine Translation. CCMT 2021. Communications in Computer and Information Science, vol 1464. Springer, Singapore. https://doi.org/10.1007/978-981-16-7512-6_6

Download citation

DOI: https://doi.org/10.1007/978-981-16-7512-6_6
Published: 30 October 2021
Publisher Name: Springer, Singapore
Print ISBN: 978-981-16-7511-9
Online ISBN: 978-981-16-7512-6
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics