TAILIEUCHUNG - Báo cáo khoa học: "Learning Sub-Word Units for Open Vocabulary Speech Recognition"

Large vocabulary speech recognition systems fail to recognize words beyond their vocabulary, many of which are information rich terms, like named entities or foreign words. Hybrid word/sub-word systems solve this problem by adding sub-word units to large vocabulary word based systems; new words can then be represented by combinations of subword units. Previous work heuristically created the sub-word lexicon from phonetic representations of text using simple statistics to select common phone sequences. . | Learning Sub-Word Units for Open Vocabulary Speech Recognition Carolina Parada1 Mark Dredze1 Abhinav Sethy2 and Ariya Rastrow1 1 Human Language Technology Center of Excellence Johns Hopkins University 3400 N Charles Street Baltimore MD USA carolinap@ mdredze@ ariya@ 2IBM . Watson Research Center Yorktown Heights NY USA asethy@ Abstract Large vocabulary speech recognition systems fail to recognize words beyond their vocabulary many of which are information rich terms like named entities or foreign words. Hybrid word sub-word systems solve this problem by adding sub-word units to large vocabulary word based systems new words can then be represented by combinations of subword units. Previous work heuristically created the sub-word lexicon from phonetic representations of text using simple statistics to select common phone sequences. We propose a probabilistic model to learn the subword lexicon optimized for a given task. We consider the task of out of vocabulary OOV word detection which relies on output from a hybrid model. A hybrid model with our learned sub-word lexicon reduces error by and absolute at a 5 false alarm rate on an English Broadcast News and MIT Lectures task respectively. 1 Introduction Most automatic speech recognition systems operate with a large but limited vocabulary finding the most likely words in the vocabulary for the given acoustic signal. While large vocabulary continuous speech recognition LVCSR systems produce high quality transcripts they fail to recognize out of vocabulary OOV words. Unfortunately OOVs are often information rich nouns such as named entities and foreign words and mis-recognizing them can have a disproportionate impact on transcript coherence. 712 Hybrid word sub-word recognizers can produce a sequence of sub-word units in place of OOV words. Ideally the recognizer outputs a complete word for in-vocabulary IV utterances and sub-word units for OOVs. Consider the word Slobodan the

Ngọc Tuấn 82 10 pdf

Upload

Bấm vào đây để xem trước nội dung

Tải xuống

TÀI LIỆU LIÊN QUAN

TÀI LIỆU XEM NHIỀU

Một Case Về Hematology (1)

8 462342 61

Giới thiệu :Lập trình mã nguồn mở

14 26076 79

Tiểu luận: Tư tưởng Hồ Chí Minh về xây dựng nhà nước trong sạch vững mạnh

13 11348 542

Câu hỏi và đáp án bài tập tình huống Quản trị học

14 10552 466

Phân tích và làm rõ ý kiến sau: “Bài thơ Tự tình II vừa nói lên bi kịch duyên phận vừa cho thấy khát vọng sống, khát vọng hạnh phúc của Hồ Xuân Hương”

3 9843 108

Ebook Facts and Figures – Basic reading practice: Phần 1 – Đặng Tuấn Anh (Dịch)

249 8891 1161

Tiểu luận: Nội dung tư tưởng Hồ Chí Minh về đạo đức

16 8506 426

Mẫu đơn thông tin ứng viên ngân hàng VIB

8 8101 2279

Giáo trình Tư tưởng Hồ Chí Minh - Mạch Quang Thắng (Dành cho bậc ĐH - Không chuyên ngành Lý luận chính trị)

152 7756 1792

Đề tài: Dự án kinh doanh thời trang quần áo nữ

17 7271 268

TỪ KHÓA LIÊN QUAN

TÀI LIỆU MỚI ĐĂNG

Đóng mới oto 8 chỗ ngồi part 9

10 179 3 28-12-2024

Data Structures and Algorithms - Chapter 8: Heaps

41 188 5 28-12-2024

Hướng dẫn chế độ dinh dưỡng cho người bệnh viêm khớp

5 168 2 28-12-2024

báo cáo hóa học:" Quality of data collection in a large HIV observational clinic database in sub-Saharan Africa: implications for clinical research and audit of care"

7 154 4 28-12-2024

CHƯƠNG 2: RỦI RO THÂM HỤT TÀI KHÓA

28 160 1 28-12-2024

Đề tài " Dự báo về tác động của Tổ chức Thương mại Thế giới WTO đối với các doanh nghiệp xuất khẩu vừa và nhỏ Việt Nam – Những giải pháp đề xuất "

72 187 2 28-12-2024

Báo cáo " Thẩm quyền quản lí nhà nước đối với hoạt động quảng cáo thực trạng và hướng hoàn thiện "

7 206 7 28-12-2024

Báo cáo nghiên cứu khoa học " NÂNG QUAN HỆ KINH TẾ THƯƠNG MẠI VIỆT NAM - TRUNG QUỐC LÊN TẦM CAO THỜI ĐẠI "

8 174 1 28-12-2024

Sáng kiến kinh nghiệm môn mỹ thuật

5 175 1 28-12-2024

CÂU HỎI TRẮC NGHIỆM HSLS NƯỚC TIỂU

9 177 0 28-12-2024

TÀI LIỆU HOT

Mẫu đơn thông tin ứng viên ngân hàng VIB

8 8101 2279

Giáo trình Tư tưởng Hồ Chí Minh - Mạch Quang Thắng (Dành cho bậc ĐH - Không chuyên ngành Lý luận chính trị)

152 7756 1792

Ebook Chào con ba mẹ đã sẵn sàng

112 4409 1371

Ebook Tuyển tập đề bài và bài văn nghị luận xã hội: Phần 1

62 6290 1266

Ebook Facts and Figures – Basic reading practice: Phần 1 – Đặng Tuấn Anh (Dịch)

249 8891 1161

Giáo trình Văn hóa kinh doanh - PGS.TS. Dương Thị Liễu

561 3841 680

Giáo trình Sinh lí học trẻ em: Phần 1 - TS Lê Thanh Vân

122 3920 609

Giáo trình Pháp luật đại cương: Phần 1 - NXB ĐH Sư Phạm

274 4712 565

Tiểu luận: Tư tưởng Hồ Chí Minh về xây dựng nhà nước trong sạch vững mạnh

13 11348 542

Bài tập nhóm quản lý dự án: Dự án xây dựng quán cafe

35 4510 490