European Union parallel multilingual texts are ideal for training machine translation
the verdict
SUPPORTED
the evidence backs this
refutedsupported
the weight of evidence
2 sources for · 0 against
Official documentation and research note that parallel texts and translation memories from European Union institutions serve as important linguistic resources for training automated machine translation systems and conducting cross-language research.
ns and verifications Data policy Transparency See all Europa Science Experience Scientific excellence for EU policy More about us Meet the passionate scientists of the Joint Research Centre
EAC-Translation Memory
Introduction
Languages / File format
Text types / Domain
Statistics on the corpus
Conditions for Use
Further Translation Memories available here
Download the EAC Translation Memory
Referring to this resource
Acknowledgements and Contact
Introduction
In October 2012, the European Union's (EU) Directorate General for Education and Culture (
DG EAC ) released a translation memory (TM), i.e. a collection of sentences and their professionally produced translations, in twenty-six languages. The data gets distributed via the web pages of the EC's Joint Research Centre (JRC). Here we describe
this resource, which bears the name EAC Translation Memory, short EAC-TM.
view details
Translation Memories are
parallel texts , i.e. texts and their manually produced translations. They are also referred to as bi-texts. A
translation memory is a collection of small text segments and their translations (referred to as translation units, TU). These TUs can be sentences or parts of sentences. Translation memories are used to support translators by ensuring that pieces
of text that have already been translated do not need to be translated again.
Both translation memories and parallel texts are important linguistic resources that
can be used for a variety of purposes , including:
training automatic systems for statistical machine translation (SMT);
producing monolingual or multilingual lexical and semantic resources such as dictionaries and ontologies;
training and testing multilingual information extraction software;
checking translation consistency automatically;
testing and benchmarking alignment software (for sentences, words, etc.).
The value of a parallel corpus grows with its size and with the number of languages for which translations exist. While parallel corpora for some
We present a new, unique and freely available parallel corpus containing European Union (EU) documents of mostly legal nature. It is available in all 20 official EUanguages, with additional documents being available in the languages of the EU candidate countries. The corpus consists of almost 8,000 documents per language, with an average size of nearly 9 million words per language. Pair-wise paragraph alignment information produced by two different aligners (Vanilla and HunAlign) is available for all 190+ language pair combinations. Most texts have been manually classified according to the EUROVOC subject domains so that the collection can also be used to train and test multi-label classification algorithms and keyword-assignment software. The corpus is encoded in XML, according to the Text Encoding Initiative Guidelines. Due to the large number of parallel texts in many languages, the JRC-Acquis is particularly suitable to carry out all types of cross-language research, as well as to test and benchmark text analysis software across different languages (for instance for alignment, sentence splitting and term extraction).
Everything we examined (2)
This check searched the claim as stated. It did not run a separate search for evidence against it.