How comparable are parallel corpora? Measuring the distribution of general vocabulary and connectives
Abstract
In this paper, we question the homogeneity of a large parallel corpus by measuring the similarity between various sub-parts. We compare results obtained using a general measure of lexical similarity based on χ2 and by counting the number of discourse connectives. We argue that discourse connectives provide a more sensitive measure, revealing differences that are not visible with the general measure. We also provide evidence for the existence of specific characteristics defining translated texts as opposed to non-translated ones, due to a universal tendency for explicitation.
Date Issued
2011
Publication Type
Conference Item
Language(s)
en
Author(s)
Cartoni, Bruno | |
Popescu-Belis, Andrei | |
Meyer, Thomas |
Additional Credits
ISBN
978-1-937284-015
Access(Rights)
restricted