Preordering using a target-language parser via cross-language syntactic projection for statistical machine translation

I Goto, M Utiyama, E Sumita, S Kurohashi - ACM Transactions on Asian …, 2015 - dl.acm.org
ACM Transactions on Asian and Low-Resource Language Information Processing …, 2015dl.acm.org
When translating between languages with widely different word orders, word reordering can
present a major challenge. Although some word reordering methods do not employ source-
language syntactic structures, such structures are inherently useful for word reordering.
However, high-quality syntactic parsers are not available for many languages. We propose a
preordering method using a target-language syntactic parser to process source-language
syntactic structures without a source-language syntactic parser. To train our preordering …
When translating between languages with widely different word orders, word reordering can present a major challenge. Although some word reordering methods do not employ source-language syntactic structures, such structures are inherently useful for word reordering. However, high-quality syntactic parsers are not available for many languages. We propose a preordering method using a target-language syntactic parser to process source-language syntactic structures without a source-language syntactic parser. To train our preordering model based on ITG, we produced syntactic constituent structures for source-language training sentences by (1) parsing target-language training sentences, (2) projecting constituent structures of the target-language sentences to the corresponding source-language sentences, (3) selecting parallel sentences with highly synchronized parallel structures, (4) producing probabilistic models for parsing using the projected partial structures and the Pitman-Yor process, and (5) parsing to produce full binary syntactic structures maximally synchronized with the corresponding target-language syntactic structures, using the constraints of the projected partial structures and the probabilistic models. Our ITG-based preordering model is trained using the produced binary syntactic structures and word alignments. The proposed method facilitates the learning of ITG by producing highly synchronized parallel syntactic structures based on cross-language syntactic projection and sentence selection. The preordering model jointly parses input sentences and identifies their reordered structures. Experiments with Japanese--English and Chinese--English patent translation indicate that our method outperforms existing methods, including string-to-tree syntax-based SMT, a preordering method that does not require a parser, and a preordering method that uses a source-language dependency parser.
ACM Digital Library
Showing the best result for this search. See all results