TAILIEUCHUNG - Báo cáo khoa học: "A Geometric View on Bilingual Lexicon Extraction from Comparable Corpora"

We present a geometric view on bilingual lexicon extraction from comparable corpora, which allows to re-interpret the methods proposed so far and identify unresolved problems. This motivates three new methods that aim at solving these problems. Empirical evaluation shows the strengths and weaknesses of these methods, as well as a signiﬁcant gain in the accuracy of extracted lexicons. and polysemy problems. | A Geometric View on Bilingual Lexicon Extraction from Comparable Corpora E. Gaussiery . Renders I. Matveeva C. Gouttey H. Dejeany Xerox Research Centre Europe 6 Chemin de Maupertuis 38320 Meylan France Dept of Computer Science University of Chicago 1100 E. 58th St. Chicago IL 60637 USA matveeva@ Abstract We present a geometric view on bilingual lexicon extraction from comparable corpora which allows to re-interpret the methods proposed so far and identify unresolved problems. This motivates three new methods that aim at solving these problems. Empirical evaluation shows the strengths and weaknesses of these methods as well as a significant gain in the accuracy of extracted lexicons. 1 Introduction Comparable corpora contain texts written in different languages that roughly speaking talk about the same thing . In comparison to parallel corpora ie corpora which are mutual translations comparable corpora have not received much attention from the research community and very few methods have been proposed to extract bilingual lexicons from such corpora. However except for those found in translation services or in a few international organisations which by essence produce parallel documentations most existing multilingual corpora are not parallel but comparable. This concern is reflected in major evaluation conferences on crosslanguage information retrieval CLIR . CLEF1 which only use comparable corpora for their multilingual tracks. We adopt here a geometric view on bilingual lexicon extraction from comparable corpora which allows one to re-interpret the methods proposed thus far and formulate new ones inspired by latent semantic analysis LSA which was developed within the information retrieval IR community to treat synonymous and polysemous terms Deerwester et al. 1990 . We will explain in this paper the motivations behind the use of such methods for bilingual lexicon extraction from comparable corpora and show how to

Minh Nguyệt 79 8 pdf

Upload

Bấm vào đây để xem trước nội dung

Tải xuống

TÀI LIỆU LIÊN QUAN

Guide to Geometric Algebra in Practice

458 73 0

Geometric Algebra and its Application to Mathematical Physics

1 49 0

Developing geometric thinking for children in early childhood education through geometric activities

6 65 0

Ebook Geometric modeling (2nd edition): Part 2

243 82 1

Lecture Engineering drawing and design - Lecture 6: Representation of features geometric tolerances

13 77 0

Multi-item fuzzy inventory problem with space constraint via geometric programming method

12 73 0

The necessary and sufficient conditions for a probability distribution belongs to the domain of geometric attraction of standard Laplace distribution

4 84 0

On chi-square type distributions with geometric degrees of freedom in relation to geometric sums

5 55 0

The use of geometric dynamic softwares by teachers in teaching mathematics at secondary school in Vietnam

6 34 1

Curriculum Geometric geodesy: Part I

189 12 1

TÀI LIỆU XEM NHIỀU

Một Case Về Hematology (1)

8 461864 55

Giới thiệu :Lập trình mã nguồn mở

14 22634 59

Tiểu luận: Tư tưởng Hồ Chí Minh về xây dựng nhà nước trong sạch vững mạnh

13 10884 529

Câu hỏi và đáp án bài tập tình huống Quản trị học

14 10064 446

Phân tích và làm rõ ý kiến sau: “Bài thơ Tự tình II vừa nói lên bi kịch duyên phận vừa cho thấy khát vọng sống, khát vọng hạnh phúc của Hồ Xuân Hương”

3 9518 104

Ebook Facts and Figures – Basic reading practice: Phần 1 – Đặng Tuấn Anh (Dịch)

249 8279 1125

Tiểu luận: Nội dung tư tưởng Hồ Chí Minh về đạo đức

16 8230 423

Mẫu đơn thông tin ứng viên ngân hàng VIB

8 7864 2220

Đề tài: Dự án kinh doanh thời trang quần áo nữ

17 6683 253

Vật lý hạt cơ bản (1)

29 5769 85

TỪ KHÓA LIÊN QUAN

TÀI LIỆU MỚI ĐĂNG

Giáo án mầm non chương trình đổi mới: Gia đình vui nhộn

4 312 1 26-04-2024

Động cơ đốt trong và máy kéo công nghiêp tập 1 part 7

23 258 0 26-04-2024

extremetech Hacking Firefox phần 7

46 187 0 26-04-2024

TƯƠNG QUAN GIỮA MÔ HỌC, GIẢI PHẪU VÀ HÌNH ẢNH CỦA CÁC KHỐI U PHẦN PHỤ

3 167 0 26-04-2024

MySQL Basics for Visual Learners PHẦN 9

15 183 0 26-04-2024

Posted prices versus bargaining in markets_7

23 155 0 26-04-2024

Đề tài: Tìm hiểu một số yêu cầu đặt ra với một phòng thu âm, để đảm bảo chất lượng âm thanh trong sản phẩm đa phương tiện

8 159 1 26-04-2024

HƯỚNG DẪN SỬ DỤNG PHẦN MỀM CAITA part 9

18 128 0 26-04-2024

Kỹ thuật nuôi cá rồng part 5

7 127 0 26-04-2024

Báo cáo nghiên cứu nông nghiệp " Field control of pest fruit flies in Vietnam "

14 116 0 26-04-2024

TÀI LIỆU HOT

Mẫu đơn thông tin ứng viên ngân hàng VIB

8 7864 2220

Giáo trình Tư tưởng Hồ Chí Minh - Mạch Quang Thắng (Dành cho bậc ĐH - Không chuyên ngành Lý luận chính trị)

152 5722 1368

Ebook Chào con ba mẹ đã sẵn sàng

112 3767 1231

Ebook Tuyển tập đề bài và bài văn nghị luận xã hội: Phần 1

62 5318 1136

Ebook Facts and Figures – Basic reading practice: Phần 1 – Đặng Tuấn Anh (Dịch)

249 8279 1125

Giáo trình Văn hóa kinh doanh - PGS.TS. Dương Thị Liễu

561 3498 643

Tiểu luận: Tư tưởng Hồ Chí Minh về xây dựng nhà nước trong sạch vững mạnh

13 10884 529

Giáo trình Sinh lí học trẻ em: Phần 1 - TS Lê Thanh Vân

122 3683 525

Giáo trình Pháp luật đại cương: Phần 1 - NXB ĐH Sư Phạm

274 4045 514

Bài tập nhóm quản lý dự án: Dự án xây dựng quán cafe

35 4127 480