TAILIEUCHUNG - Báo cáo khoa học: "On-line Language Model Biasing for Statistical Machine Translation"

The language model (LM) is a critical component in most statistical machine translation (SMT) systems, serving to establish a probability distribution over the hypothesis space. Most SMT systems use a static LM, independent of the source language input. While previous work has shown that adapting LMs based on the input improves SMT performance, none of the techniques has thus far been shown to be feasible for on-line systems. | On-line Language Model Biasing for Statistical Machine Translation Sankaranarayanan Ananthakrishnan Rohit Prasad and Prem Natarajan Raytheon BBN Technologies Cambridge MA 02138 . sanantha rprasad pnataraj @ Abstract The language model LM is a critical component in most statistical machine translation SMT systems serving to establish a probability distribution over the hypothesis space. Most SMT systems use a static LM independent of the source language input. While previous work has shown that adapting LMs based on the input improves SMT performance none of the techniques has thus far been shown to be feasible for on-line systems. In this paper we develop a novel measure of cross-lingual similarity for biasing the LM based on the test input. We also illustrate an efficient on-line implementation that supports integration with on-line SMT systems by transferring much of the computational load off-line. Our approach yields significant reductions in target perplexity compared to the static LM as well as consistent improvements in SMT performance across language pairs English-Dari and English-Pashto . 1 Introduction While much of the focus in developing a statistical machine translation SMT system revolves around the translation model TM most systems do not emphasize the role of the language model LM . The latter generally follows a n-gram structure and is estimated from a large monolingual corpus of target sentences. In most systems the LM is independent of the test input . fixed n-gram probabilities determine the likelihood of all translation hypotheses regardless of the source input. The views expressed are those of the author and do not reflect the official policy or position of s f i 5 Some previous work exists in LM adaptation for SMT. Snover et al. 2008 used a cross-lingual information retrieval CLIR system to select a subset of target documents comparable to the source document bias LMs estimated from these subsets were interpolated with a static

Kiều Trang 74 5 pdf

Upload

Bấm vào đây để xem trước nội dung

Tải xuống

TÀI LIỆU LIÊN QUAN

Online language teaching: The pedagogical challenges

20 79 0

English-English language graduation thesis: A study on studying Speaking skill online for the 2nd year English major students at Hai Phong Management and Technology University

73 50 1

Báo cáo khoa học: "Fast Online Lexicon Learning for Grounded Language Acquisition"

10 70 0

Báo cáo khoa học: "Dialect Classification for online podcasts fusing Acoustic and Language based Structural and Semantic Information"

4 50 0

Báo cáo khoa học: "Advanced Online Learning for Natural Language Processing"

1 118 0

Quantitative analysis of the effect of synchronous online discussions on oral and written language development for EFL university students in Vietnam

11 78 2

Advantages and disadvantages of online homework software: The case of “life” in Vietnam

7 39 1

Báo cáo khoa học: "User Participation Prediction in Online Forums"

10 54 0

Báo cáo khoa học: "Learning from evolving data streams: online triage of bug reports"

10 73 0

Báo cáo khoa học: "Automatically Generated Customizable Online Dictionaries"

7 42 0

TÀI LIỆU XEM NHIỀU

Một Case Về Hematology (1)

8 462337 61

Giới thiệu :Lập trình mã nguồn mở

14 25975 79

Tiểu luận: Tư tưởng Hồ Chí Minh về xây dựng nhà nước trong sạch vững mạnh

13 11341 542

Câu hỏi và đáp án bài tập tình huống Quản trị học

14 10546 466

Phân tích và làm rõ ý kiến sau: “Bài thơ Tự tình II vừa nói lên bi kịch duyên phận vừa cho thấy khát vọng sống, khát vọng hạnh phúc của Hồ Xuân Hương”

3 9838 108

Ebook Facts and Figures – Basic reading practice: Phần 1 – Đặng Tuấn Anh (Dịch)

249 8889 1161

Tiểu luận: Nội dung tư tưởng Hồ Chí Minh về đạo đức

16 8502 426

Mẫu đơn thông tin ứng viên ngân hàng VIB

8 8100 2279

Giáo trình Tư tưởng Hồ Chí Minh - Mạch Quang Thắng (Dành cho bậc ĐH - Không chuyên ngành Lý luận chính trị)

152 7727 1790

Đề tài: Dự án kinh doanh thời trang quần áo nữ

17 7245 268

TỪ KHÓA LIÊN QUAN

TÀI LIỆU MỚI ĐĂNG

Đóng mới oto 8 chỗ ngồi part 9

10 179 3 25-12-2024

Báo cáo nghiên cứu nông nghiệp " Field control of pest fruit flies in Vietnam "

14 189 4 25-12-2024

Báo cáo nghiên cứu khoa học " HÃY LÀM CHO HUẾ XANH HƠN VÀ ĐẸP HƠN "

6 180 3 25-12-2024

báo cáo hóa học:" Quality of data collection in a large HIV observational clinic database in sub-Saharan Africa: implications for clinical research and audit of care"

7 154 4 25-12-2024

Đề tài " Dự báo về tác động của Tổ chức Thương mại Thế giới WTO đối với các doanh nghiệp xuất khẩu vừa và nhỏ Việt Nam – Những giải pháp đề xuất "

72 184 2 25-12-2024

Báo cáo y học: "The Factors Influencing Depression Endpoints Research (FINDER) study: final results of Italian patients with depressio"

9 148 1 25-12-2024

Báo cáo " Bàn về hành vi pháp luật và hành vi đạo đức "

11 177 2 25-12-2024

ĐỀ TÀI " ĐÁNH GIÁ HIỆU QUẢ HOẠT ĐỘNG KINH DOANH NGOẠI HỐI CỦA NGÂN HÀNG THƯƠNG MẠI CỔ PHẦN XUẤT NHẬP KHẨU VIỆT NAM "

51 149 3 25-12-2024

Báo cáo nghiên cứu khoa học " Vai trò chính quyền địa phương trong phát triển kinh tế : khu chuyên doanh gốm sứ ( Trung Quốc ) và Bát Tràng ( Việt Nam )("

11 213 1 25-12-2024

báo cáo khoa học: "Malignant peripheral nerve sheath tumor arising from the greater omentum: Case report"

4 140 1 25-12-2024

TÀI LIỆU HOT

Mẫu đơn thông tin ứng viên ngân hàng VIB

8 8100 2279

Giáo trình Tư tưởng Hồ Chí Minh - Mạch Quang Thắng (Dành cho bậc ĐH - Không chuyên ngành Lý luận chính trị)

152 7727 1790

Ebook Chào con ba mẹ đã sẵn sàng

112 4406 1371

Ebook Tuyển tập đề bài và bài văn nghị luận xã hội: Phần 1

62 6281 1266

Ebook Facts and Figures – Basic reading practice: Phần 1 – Đặng Tuấn Anh (Dịch)

249 8889 1161

Giáo trình Văn hóa kinh doanh - PGS.TS. Dương Thị Liễu

561 3837 680

Giáo trình Sinh lí học trẻ em: Phần 1 - TS Lê Thanh Vân

122 3919 609

Giáo trình Pháp luật đại cương: Phần 1 - NXB ĐH Sư Phạm

274 4705 565

Tiểu luận: Tư tưởng Hồ Chí Minh về xây dựng nhà nước trong sạch vững mạnh

13 11341 542

Bài tập nhóm quản lý dự án: Dự án xây dựng quán cafe

35 4504 490