Towards Automatically Aligning German Compounds with English Word Groups in an Example-Based Translation System.

Saved in:
Bibliographic Details
Title: Towards Automatically Aligning German Compounds with English Word Groups in an Example-Based Translation System.
Language: English
Authors: Jones, Daniel, Alexa, Melina
Peer Reviewed: N
Page Count: 6
Publication Date: 1994
Document Type: Reports - Research
Descriptors: English, Foreign Countries, German, Language Patterns, Language Processing, Lexicology, Machine Translation, Vocabulary
Abstract: As part of the development of a completely sub-symbolic machine translation system, a method for automatically identifying German compounds was developed. Given a parallel bilingual corpus, German compounds are identified along with their English word groupings by statistical processing alone. The underlying principles and the design process are described here. Design began with small-scale word-alignment experiments, using 2,543 English words and 1,898 German words that yielded unique lexical items in each language. A technique for decreasing reliance on one-to-one word correspondences was then applied, resulting in a distinct ability to capture relationships between compounds and non-compounded expressions. Statistical analysis of these relationships provides data on which to base machine translation operations. It is concluded that the method used is effective on identifying cross-language lexical fertility to establish translation units for re-combination within the example-based translation process. (MSE)
Entry Date: 1995
Accession Number: ED377711
Database: ERIC
Description
Abstract:As part of the development of a completely sub-symbolic machine translation system, a method for automatically identifying German compounds was developed. Given a parallel bilingual corpus, German compounds are identified along with their English word groupings by statistical processing alone. The underlying principles and the design process are described here. Design began with small-scale word-alignment experiments, using 2,543 English words and 1,898 German words that yielded unique lexical items in each language. A technique for decreasing reliance on one-to-one word correspondences was then applied, resulting in a distinct ability to capture relationships between compounds and non-compounded expressions. Statistical analysis of these relationships provides data on which to base machine translation operations. It is concluded that the method used is effective on identifying cross-language lexical fertility to establish translation units for re-combination within the example-based translation process. (MSE)