TAILIEUCHUNG - Báo cáo khoa học: "Updating a Name Tagger Using Contemporary Unlabeled Data"

For many NLP tasks, including named entity tagging, semi-supervised learning has been proposed as a reasonable alternative to methods that require annotating large amounts of training data. In this paper, we address the problem of analyzing new data given a semi-supervised NE tagger trained on data from an earlier time period. We will show that updating the unlabeled data is sufﬁcient to maintain quality over time, and outperforms updating the labeled data. | Updating a Name Tagger Using Contemporary Unlabeled Data Cristina Mota L2F INESC-ID 1ST NYU Rua Alves Redol 9 1000-029 Lisboa Portugal cmota@ Ralph Grishman New York University Computer Science Department NeW York NY 10003 USA grishman@ Abstract For many NLP tasks including named entity tagging semi-supervised learning has been proposed as a reasonable alternative to methods that require annotating large amounts of training data. In this paper we address the problem of analyzing new data given a semi-supervised NE tagger trained on data from an earlier time period. We will show that updating the unlabeled data is sufficient to maintain quality over time and outperforms updating the labeled data. Furthermore we will also show that augmenting the unlabeled data with older data in most cases does not result in better performance than simply using a smaller amount of current unlabeled data. 1 Introduction Brill 2003 observed large gains in performance for different NLP tasks solely by increasing the size of unlabeled data but stressed that for other NLP tasks such as named entity recognition NER we still need to focus on developing tools that help to increase the size of annotated data. This problem is particularly crucial when processing languages such as Portuguese for which the labeled data is scarce. For instance in the first NER evaluation for Portuguese HAREM Santos and Cardoso 2007 only two out of the nine participants presented systems based on machine learning and they both argued they could have achieved significantly better results if they had larger training sets. Semi-supervised methods are commonly chosen as an alternative to overcome the lack of annotated resources because they present a good trade-off between amount of labeled data needed and performance achieved. Co-training is one of those methods and has been extensively studied in NLP Nigam and Ghani 2000 Pierce and Cardie 2001 Ng and Cardie 2003 Mota and Grishman 2008 . In .

Minh Dân 54 4 pdf

Upload

Bấm vào đây để xem trước nội dung

Tải xuống

TÀI LIỆU LIÊN QUAN

View Updating and Relational Theory

262 39 1

Editing and Updating Data in a Web Forms DataGrid

10 63 0

Ableton Live 6 PowerThe Comprehensive Guide

449 35 0

Báo cáo khoa học: "Updating a Name Tagger Using Contemporary Unlabeled Data"

4 45 0

Lecture Introduction to web engineering - Lec 31: Deleting and updating records in MySQL using PHP

32 97 0

Analysis of PSO technique to tune updating factors of PD based fuzzy logic controllers

5 73 0

Lecture Introduction to web engineering - Lec 31: Deleting and updating records in MySQL using PHP

32 67 3

Model updating of a scaled piping system and vibration attenuation via locally resonant bandgap formation

16 36 3

Updating Server Data Using a Web Service

6 59 0

Updating a DataSet with a Many-to-Many Relationship

19 58 0

TÀI LIỆU XEM NHIỀU

Một Case Về Hematology (1)

8 461867 55

Giới thiệu :Lập trình mã nguồn mở

14 22643 59

Tiểu luận: Tư tưởng Hồ Chí Minh về xây dựng nhà nước trong sạch vững mạnh

13 10892 529

Câu hỏi và đáp án bài tập tình huống Quản trị học

14 10066 446

Phân tích và làm rõ ý kiến sau: “Bài thơ Tự tình II vừa nói lên bi kịch duyên phận vừa cho thấy khát vọng sống, khát vọng hạnh phúc của Hồ Xuân Hương”

3 9519 104

Ebook Facts and Figures – Basic reading practice: Phần 1 – Đặng Tuấn Anh (Dịch)

249 8281 1125

Tiểu luận: Nội dung tư tưởng Hồ Chí Minh về đạo đức

16 8238 423

Mẫu đơn thông tin ứng viên ngân hàng VIB

8 7864 2220

Đề tài: Dự án kinh doanh thời trang quần áo nữ

17 6687 253

Vật lý hạt cơ bản (1)

29 5770 85

TỪ KHÓA LIÊN QUAN

TÀI LIỆU MỚI ĐĂNG

Động cơ đốt trong và máy kéo công nghiêp tập 1 part 7

23 258 0 27-04-2024

Sáng tạo trong thuật toán và lập trình với ngôn ngữ Pascal và C# Tập 2 - Chương 4

47 246 1 27-04-2024

Công nghiệp gang thép Việt Nam : Một giai đoạn phát triển và chuyển đổi chính sách mới part 5

6 194 0 27-04-2024

THE ANTHROPOLOGY OF ONLINE COMMUNITIES BY Samuel M.Wilson and Leighton C. Peterson

19 145 0 27-04-2024

Lịch sử Đội TNTP Hồ Chí Minh - CHƯƠNG III VÂNG LỜI BÁC DẠY, LÀM NGHÌN VIỆC TỐT, CHỐNG MỸ, CỨU NƯỚC, THIẾU NIÊN SĂN SÀNG

45 137 0 27-04-2024

Khurana et al. Journal of Orthopaedic Surgery and Research 2010, 5:23

7 133 0 27-04-2024

báo cáo hóa học:" Endoscopic decompression for intraforaminal and extraforaminal nerve root compression"

7 107 0 27-04-2024

HƯỚNG DẪN SỬ DỤNG PHẦN MỀM CAITA part 9

18 130 0 27-04-2024

Kỹ thuật nuôi cá rồng part 5

7 127 0 27-04-2024

Gastroenterology an illustrated colour text - part 10

10 89 0 27-04-2024

TÀI LIỆU HOT

Mẫu đơn thông tin ứng viên ngân hàng VIB

8 7864 2220

Giáo trình Tư tưởng Hồ Chí Minh - Mạch Quang Thắng (Dành cho bậc ĐH - Không chuyên ngành Lý luận chính trị)

152 5737 1368

Ebook Chào con ba mẹ đã sẵn sàng

112 3767 1231

Ebook Tuyển tập đề bài và bài văn nghị luận xã hội: Phần 1

62 5319 1136

Ebook Facts and Figures – Basic reading practice: Phần 1 – Đặng Tuấn Anh (Dịch)

249 8281 1125

Giáo trình Văn hóa kinh doanh - PGS.TS. Dương Thị Liễu

561 3499 643

Tiểu luận: Tư tưởng Hồ Chí Minh về xây dựng nhà nước trong sạch vững mạnh

13 10892 529

Giáo trình Sinh lí học trẻ em: Phần 1 - TS Lê Thanh Vân

122 3684 525

Giáo trình Pháp luật đại cương: Phần 1 - NXB ĐH Sư Phạm

274 4046 515

Bài tập nhóm quản lý dự án: Dự án xây dựng quán cafe

35 4128 480