The present research analyses the performance of two free open-source neural machine translation (NMT) systems —Google Translate and DeepL— in the (ES>EN) translation of somatisms such as tomar el pelo and meter la pata, their nominal variants (tomadura/tomada de pelo and metedura/metida de pata), and other lower-frequency variants such as meter la pata hasta el corvejón, meter la gamba and metedura/metida de gamba. The machine translation outcomes will be contrasted and classified depending on whether these idioms are presented in their continuous or discontinuous form (Anastasiou 2010), i.e., whether different n-grams split the idiomatic sequence (or not), which may pose some difficulties for their automatic detection and translation. Overall, the insights gained from this study will prove useful in determining for which of the different scenarios either Google Translate or DeepL delivers a better performance under the challenge of phraseological variation and discontinuity.
The present research introduces the tool gApp, a Python-based text preprocessing system for the automatic identification and conversion of discontinuous multiword expressions (MWEs) into their continuous form in order to enhance neural machine translation (NMT). To this end, an experiment with semi-fixed verb–noun idiomatic combinations (VNICs) will be carried out in order to evaluate to what extent gApp can optimise the performance of the two main free open-source NMT systems —Google Translate and DeepL— under the challenge of MWE discontinuity in the Spanish into English directionality. In the light of our promising results, the study concludes with suggestions on how to further optimise MWE-aware NMT systems.
Reseña de la obra: CORPAS PASTOR, G. y SEGHIRI DOMÍNGUEZ, M. (eds.) (2016). Corpus-based Approaches to Translation and Interpreting: From Theory to Applications. (Studien zur romanischen Sprachwissenschaft und interkulturellen Kommunikation, 106). Frankfurt: Peter Lang. ISBN 9783631609569 / E-ISBN 9783653060553. DOI 10.3726/9783653060553.
scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.