TAILIEUCHUNG - Báo cáo khoa học: "Learning with Unlabeled Data for Text Categorization Using Bootstrapping and Feature Projection Techniques"

A wide range of supervised learning algorithms has been applied to Text Categorization. However, the supervised learning approaches have some problems. One of them is that they require a large, often prohibitive, number of labeled training documents for accurate learning. Generally, acquiring class labels for training data is costly, while gathering a large quantity of unlabeled data is cheap. We here propose a new automatic text categorization method for learning from only unlabeled data using a bootstrapping framework and a feature projection technique. From results of our experiments, our method showed reasonably comparable performance compared with a supervised. | Learning with Unlabeled Data for Text Categorization Using Bootstrapping and Feature Projection Techniques Youngjoong Ko Dept. of Computer Science Sogang Univ. Sinsu-dong 1 Mapo-gu Seoul 121-742 Korea kyj@ Abstract A wide range of supervised learning algorithms has been applied to Text Categorization. However the supervised learning approaches have some problems. One of them is that they require a large often prohibitive number of labeled training documents for accurate learning. Generally acquiring class labels for training data is costly while gathering a large quantity of unlabeled data is cheap. We here propose a new automatic text categorization method for learning from only unlabeled data using a bootstrapping framework and a feature projection technique. From results of our experiments our method showed reasonably comparable performance compared with a supervised method. If our method is used in a text categorization task building text categorization systems will become significantly faster and less expensive. 1 Introduction Text categorization is the task of classifying documents into a certain number of pre-defined categories. Many supervised learning algorithms have been applied to this area. These algorithms today are reasonably successful when provided with enough labeled or annotated training examples. For example there are Naive Bayes McCallum and Nigam 1998 Rocchio Lewis et al. 1996 Nearest Neighbor CNN Yang et al. 2002 TCFP Ko and Seo 2002 and Support Vector Machine SVM Joachims 1998 . However the supervised learning approach has some difficulties. One key difficulty is that it requires a large often prohibitive number of labeled training data for accurate learning. Since a labeling task must be done manually it is a painfully time-consuming process. Furthermore since the application area of text categorization has diversified from newswire articles and web pages to E-mails and newsgroup postings it is also a difficult task to

TAILIEUCHUNG - Chia sẻ tài liệu không giới hạn
Địa chỉ : 444 Hoang Hoa Tham, Hanoi, Viet Nam
Website : tailieuchung.com
Email : tailieuchung20@gmail.com
Tailieuchung.com là thư viện tài liệu trực tuyến, nơi chia sẽ trao đổi hàng triệu tài liệu như luận văn đồ án, sách, giáo trình, đề thi.
Chúng tôi không chịu trách nhiệm liên quan đến các vấn đề bản quyền nội dung tài liệu được thành viên tự nguyện đăng tải lên, nếu phát hiện thấy tài liệu xấu hoặc tài liệu có bản quyền xin hãy email cho chúng tôi.
Đã phát hiện trình chặn quảng cáo AdBlock
Trang web này phụ thuộc vào doanh thu từ số lần hiển thị quảng cáo để tồn tại. Vui lòng tắt trình chặn quảng cáo của bạn hoặc tạm dừng tính năng chặn quảng cáo cho trang web này.