TAILIEUCHUNG - Báo cáo khoa học: "Evaluating language understanding accuracy with respect to objective outcomes in a dialogue system"

It is not always clear how the differences in intrinsic evaluation metrics for a parser or classiﬁer will affect the performance of the system that uses it. We investigate the relationship between the intrinsic evaluation scores of an interpretation component in a tutorial dialogue system and the learning outcomes in an experiment with human users. Following the PARADISE methodology, we use multiple linear regression to build predictive models of learning gain, an important objective outcome metric in tutorial dialogue. We show that standard intrinsic metrics such as F-score alone do not predict the outcomes well. . | Evaluating language understanding accuracy with respect to objective outcomes in a dialogue system Myroslava O. Dzikovska and Peter Bell and Amy Isard and Johanna D. Moore Institute for Language Cognition and Computation School of Informatics University of Edinburgh United Kingdom @ Abstract It is not always clear how the differences in intrinsic evaluation metrics for a parser or classifier will affect the performance of the system that uses it. We investigate the relationship between the intrinsic evaluation scores of an interpretation component in a tutorial dialogue system and the learning outcomes in an experiment with human users. Following the PARADISE methodology we use multiple linear regression to build predictive models of learning gain an important objective outcome metric in tutorial dialogue. We show that standard intrinsic metrics such as F-score alone do not predict the outcomes well. However we can build predictive performance functions that account for up to 50 of the variance in learning gain by combining features based on standard evaluation scores and on the confusion matrix entries. We argue that building such predictive models can help us better evaluate performance of NLP components that cannot be distinguished based on F-score alone and illustrate our approach by comparing the current interpretation component in the system to a new classifier trained on the evaluation data. 1 Introduction Much of the work in natural language processing relies on intrinsic evaluation computing standard evaluation metrics such as precision recall and F-score on the same data set to compare the performance of different approaches to the same NLP problem. However once a component such as a parser is included in a larger system it is not always clear that improvements in intrinsic evaluation scores will translate into improved overall system performance. Therefore extrinsic or task-based evaluation can be used to

Tùng Linh 59 11 pdf

Upload

Bấm vào đây để xem trước nội dung

Tải xuống

TÀI LIỆU LIÊN QUAN

Báo cáo khoa học: "Evaluating language understanding accuracy with respect to objective outcomes in a dialogue system"

11 50 0

Báo cáo khoa học: "EVALUATING DISCOURSE PROCESSING ALGORITHMS"

11 44 0

Báo cáo khoa học: "Evaluating and Combining Approaches to Selectional Preference Acquisition"

8 67 0

Báo cáo khoa học: "Sentiment Summarization: Evaluating and Learning User Preferences"

9 72 0

Báo cáo khoa học: "Evaluating the Inferential Utility of Lexical-Semantic Resources"

9 41 0

Báo cáo khoa học: "PEAS, the first instantiation of a comparative framework for evaluating parsers of French"

4 44 0

Báo cáo khoa học: "Probing the lexicon in evaluating commercial MT systems Martin"

8 43 0

Báo cáo khoa học: "Evaluating Distributional Models of Semantics for Syntactically Invariant Inference"

11 38 0

Báo cáo khoa học: " A Framework for Evaluating Spoken Dialogue Agents"

10 37 0

Báo cáo khoa học: "Re-evaluating the Role of B LEU in Machine Translation Research"

8 65 0

TÀI LIỆU XEM NHIỀU

Một Case Về Hematology (1)

8 461887 55

Giới thiệu :Lập trình mã nguồn mở

14 22723 61

Tiểu luận: Tư tưởng Hồ Chí Minh về xây dựng nhà nước trong sạch vững mạnh

13 10906 530

Câu hỏi và đáp án bài tập tình huống Quản trị học

14 10083 447

Phân tích và làm rõ ý kiến sau: “Bài thơ Tự tình II vừa nói lên bi kịch duyên phận vừa cho thấy khát vọng sống, khát vọng hạnh phúc của Hồ Xuân Hương”

3 9540 104

Ebook Facts and Figures – Basic reading practice: Phần 1 – Đặng Tuấn Anh (Dịch)

249 8302 1127

Tiểu luận: Nội dung tư tưởng Hồ Chí Minh về đạo đức

16 8248 423

Mẫu đơn thông tin ứng viên ngân hàng VIB

8 7867 2220

Đề tài: Dự án kinh doanh thời trang quần áo nữ

17 6713 253

Giáo trình Tư tưởng Hồ Chí Minh - Mạch Quang Thắng (Dành cho bậc ĐH - Không chuyên ngành Lý luận chính trị)

152 5795 1391

TỪ KHÓA LIÊN QUAN

TÀI LIỆU MỚI ĐĂNG

Giáo án mầm non chương trình đổi mới: Gia đình vui nhộn

4 313 1 02-05-2024

CẤU TẠO HẠT NHÂN NGUYÊN TỬ-ĐỘ HỤT KHỐI-NĂNG LƯỢNG LIÊN KẾT-LK RIÊNG

12 270 0 02-05-2024

Bibliography on Medieval Women, Gender, and Medicine 1980-2009

82 211 0 02-05-2024

Trading Strategies Profit Making Techniques For Stock_8

23 176 1 02-05-2024

TƯƠNG QUAN GIỮA MÔ HỌC, GIẢI PHẪU VÀ HÌNH ẢNH CỦA CÁC KHỐI U PHẦN PHỤ

3 169 0 02-05-2024

Bơm máy nén quạt trong công nghiệp part 8

20 199 2 02-05-2024

MySQL Basics for Visual Learners PHẦN 9

15 186 0 02-05-2024

MySQL Database Usage & Administration PHẦN 9

37 143 0 02-05-2024

BÀI GIẢNG VỀ - MẠCH ĐIỆN II - Chương I: Phân tích mạch trong miền thời gian

38 143 0 02-05-2024

Hệ thống làm lạnh và điều hòa không khí

21 127 0 02-05-2024

TÀI LIỆU HOT

Mẫu đơn thông tin ứng viên ngân hàng VIB

8 7867 2220

Giáo trình Tư tưởng Hồ Chí Minh - Mạch Quang Thắng (Dành cho bậc ĐH - Không chuyên ngành Lý luận chính trị)

152 5795 1391

Ebook Chào con ba mẹ đã sẵn sàng

112 3772 1233

Ebook Tuyển tập đề bài và bài văn nghị luận xã hội: Phần 1

62 5334 1136

Ebook Facts and Figures – Basic reading practice: Phần 1 – Đặng Tuấn Anh (Dịch)

249 8302 1127

Giáo trình Văn hóa kinh doanh - PGS.TS. Dương Thị Liễu

561 3518 644

Tiểu luận: Tư tưởng Hồ Chí Minh về xây dựng nhà nước trong sạch vững mạnh

13 10906 530

Giáo trình Sinh lí học trẻ em: Phần 1 - TS Lê Thanh Vân

122 3695 525

Giáo trình Pháp luật đại cương: Phần 1 - NXB ĐH Sư Phạm

274 4071 516

Bài tập nhóm quản lý dự án: Dự án xây dựng quán cafe

35 4136 480