Identifying Multi-Word Expressions from Parallel Corpora with Kernel Methods and Crowdsourcing
Conference Poster
Publication Date:
2014
abstract:
We propose a new methodology for the identification of MWEs from parallel multilingual corpora. Our approach is inspired by one of the most significant properties that characterize the majority of MWEs, which goes under the name of non-translatability: an MWE cannot be translated
from one language to another on a word by word basis (Sag et al., 2002; Monti, 2012). The methodology envisions a three-stage process. The first phase makes use of automatic kernel methods for the identification of possible candidate pairs of expressions which has high recall
and low precision. In the second phase, a "word by word" automatic translation system will filter out those candidate pairs which are literal translations (and therefore not MWEs). In the third phase, a crowdsourcing system is used to further validate the list of final candidates
from one language to another on a word by word basis (Sag et al., 2002; Monti, 2012). The methodology envisions a three-stage process. The first phase makes use of automatic kernel methods for the identification of possible candidate pairs of expressions which has high recall
and low precision. In the second phase, a "word by word" automatic translation system will filter out those candidate pairs which are literal translations (and therefore not MWEs). In the third phase, a crowdsourcing system is used to further validate the list of final candidates
Iris type:
4.3 Poster
Keywords:
Multi-word expression; kernel methods; crowdsourcing
List of contributors:
Monti, Johanna; Sangati, Federico; van Cranenburgh Andreas,
Full Text:
Book title:
Parseme Frankfurt, 8-9 September 2014