TAILIEUCHUNG - Báo cáo khoa học: "Improved Smoothing for N-gram Language Models Based on Ordinary Counts"

Kneser-Ney (1995) smoothing and its variants are generally recognized as having the best perplexity of any known method for estimating N-gram language models. Kneser-Ney smoothing, however, requires nonstandard N-gram counts for the lowerorder models used to smooth the highestorder model. For some applications, this makes Kneser-Ney smoothing inappropriate or inconvenient. In this paper, we introduce a new smoothing method based on ordinary counts that outperforms all of the previous ordinary-count methods we have tested, with the new method eliminating most of the gap between Kneser-Ney and those methods. . | Improved Smoothing for N-gram Language Models Based on Ordinary Counts Robert C. Moore Chris Quirk Microsoft Research Redmond WA 98052 USA bobmoore chrisq @ Abstract Kneser-Ney 1995 smoothing and its variants are generally recognized as having the best perplexity of any known method for estimating N-gram language models. Kneser-Ney smoothing however requires nonstandard N-gram counts for the lower-order models used to smooth the highest-order model. For some applications this makes Kneser-Ney smoothing inappropriate or inconvenient. In this paper we introduce a new smoothing method based on ordinary counts that outperforms all of the previous ordinary-count methods we have tested with the new method eliminating most of the gap between Kneser-Ney and those methods. 1 Introduction Statistical language models are potentially useful for any language technology task that produces natural-language text as a final or intermediate output. In particular they are extensively used in speech recognition and machine translation. Despite the criticism that they ignore the structure of natural language simple N-gram models which estimate the probability of each word in a text string based on the N 1 preceding words remain the most widely used type of model. The simplest possible N-gram model is the maximum likelihood estimate MLE which takes the probability of a word wn given the preceding context W1. wn-1 to be the ratio of the number of occurrences in a training corpus of the Ngram W1 . .wn to the total number of occurrences of any word in the same context C w1 . . . wn p w w1---w -1 . C wi w -iw One obvious problem with this method is that it assigns a probability of zero to any N-gram that is not observed in the training corpus hence numerous smoothing methods have been invented that reduce the probabilities assigned to some or all observed N-grams to provide a non-zero probability for N-grams not observed in the training corpus. The best methods for smoothing .

Lệ Khanh 66 4 pdf

Upload

Bấm vào đây để xem trước nội dung

Tải xuống

TÀI LIỆU LIÊN QUAN

Báo cáo khoa học: "Improved Smoothing for N-gram Language Models Based on Ordinary Counts"

4 57 0

TÀI LIỆU XEM NHIỀU

Một Case Về Hematology (1)

8 462292 61

Giới thiệu :Lập trình mã nguồn mở

14 24934 79

Tiểu luận: Tư tưởng Hồ Chí Minh về xây dựng nhà nước trong sạch vững mạnh

13 11287 542

Câu hỏi và đáp án bài tập tình huống Quản trị học

14 10511 466

Phân tích và làm rõ ý kiến sau: “Bài thơ Tự tình II vừa nói lên bi kịch duyên phận vừa cho thấy khát vọng sống, khát vọng hạnh phúc của Hồ Xuân Hương”

3 9791 108

Ebook Facts and Figures – Basic reading practice: Phần 1 – Đặng Tuấn Anh (Dịch)

249 8876 1160

Tiểu luận: Nội dung tư tưởng Hồ Chí Minh về đạo đức

16 8467 426

Mẫu đơn thông tin ứng viên ngân hàng VIB

8 8090 2279

Giáo trình Tư tưởng Hồ Chí Minh - Mạch Quang Thắng (Dành cho bậc ĐH - Không chuyên ngành Lý luận chính trị)

152 7473 1763

Đề tài: Dự án kinh doanh thời trang quần áo nữ

17 7189 268

TỪ KHÓA LIÊN QUAN

TÀI LIỆU MỚI ĐĂNG

THE ANTHROPOLOGY OF ONLINE COMMUNITIES BY Samuel M.Wilson and Leighton C. Peterson

19 211 4 27-11-2024

Báo cáo nghiên cứu nông nghiệp " Biofertiliser inoculant technology for the growth of rice in Vietnam: Developing technical infrastructure for quality assurance and village production for farmers "

12 132 2 27-11-2024

Bảng màu theo chữ cái – V

11 153 2 27-11-2024

Chương 10: Các phương pháp tính quá trình quá độ trong mạch điện tuyến tính

57 226 7 27-11-2024

Color Atlas of Ophthamology

165 132 2 27-11-2024

báo cáo hóa học:" Quality of data collection in a large HIV observational clinic database in sub-Saharan Africa: implications for clinical research and audit of care"

7 146 4 27-11-2024

Báo cáo nghiên cứu khoa học " NÂNG QUAN HỆ KINH TẾ THƯƠNG MẠI VIỆT NAM - TRUNG QUỐC LÊN TẦM CAO THỜI ĐẠI "

8 159 1 27-11-2024

Chủ đề 3 : SỰ CÂN BẰNG CỦA VẬT RẮN (4 tiết)

9 199 1 27-11-2024

CUỘC KHÁNG CHIẾN CHỐNG THỰC DÂN PHÁP KẾT THÚC (1953 - 1954)_5

11 133 1 27-11-2024

Lập trình Java cơ bản : Luồng và xử lý file part 8

5 133 1 27-11-2024

TÀI LIỆU HOT

Mẫu đơn thông tin ứng viên ngân hàng VIB

8 8090 2279

Giáo trình Tư tưởng Hồ Chí Minh - Mạch Quang Thắng (Dành cho bậc ĐH - Không chuyên ngành Lý luận chính trị)

152 7473 1763

Ebook Chào con ba mẹ đã sẵn sàng

112 4364 1369

Ebook Tuyển tập đề bài và bài văn nghị luận xã hội: Phần 1

62 6156 1259

Ebook Facts and Figures – Basic reading practice: Phần 1 – Đặng Tuấn Anh (Dịch)

249 8876 1160

Giáo trình Văn hóa kinh doanh - PGS.TS. Dương Thị Liễu

561 3790 680

Giáo trình Sinh lí học trẻ em: Phần 1 - TS Lê Thanh Vân

122 3909 609

Giáo trình Pháp luật đại cương: Phần 1 - NXB ĐH Sư Phạm

274 4618 562

Tiểu luận: Tư tưởng Hồ Chí Minh về xây dựng nhà nước trong sạch vững mạnh

13 11287 542

Bài tập nhóm quản lý dự án: Dự án xây dựng quán cafe

35 4454 490