Bilingual corresponding corpus: R&D and application

Author: Wang Kefei
Publisher:
Publish Date: 2005-12-01
Features: The characteristics of our bilingual corpus are as follows:
1) In terms of design, the entire corpus is divided into several sub-corpora, such as the encyclopedia corpus, specialized corpus, translation text corpus, and corresponding sentence corpus.
2) The corpus is both divisible and combinable: it can be used for individual research, such as the encyclopedia corpus and corresponding sentence corpus being beneficial for bilingual dictionary compilation, the specialized corpus for automatic translation research, and the translation text corpus for translation style and translation teaching research; when combined, it can be used for research on word frequency, collocation, corresponding words, sentence patterns, style, and genre based on large-scale corpus data.
3) This corpus distinguishes four types of data: foreign language (English/Japanese) original texts, Chinese original texts, foreign language (English/Japanese) translated texts, and Chinese translated texts, allowing for monolingual research, bilingual comparative research, and comparative research between original and translated texts.
4) Using this bilingual corpus, a large number of corresponding words and phrases can be searched, enriching the compilation of bilingual dictionaries (English/Japanese-Chinese and Chinese-English/Japanese).

📌 Related Posts