International Journal of Engineering، جلد ۳۴، شماره ۱۲، صفحات ۲۶۴۸-۲۶۵۷

عنوان فارسی
چکیده فارسی مقاله
کلیدواژه‌های فارسی مقاله

عنوان انگلیسی A Two-Level Semi-supervised Clustering Technique for News Articles
چکیده انگلیسی مقاله The web and social media are overcrowded with news pieces in terms of amount and diversity. Document clustering is a useful technique that is widely used in organizing and managing data into smaller groups. One of the factors influencing the quality of clustering is the way documents are represented. Some traditional methods of document representation depend on word frequencies and create sparse and large-sized document vectors. These methods cannot preserve proximity information between documents. In addition, neural network-based methods that preserve proximity information suffer from poor interpretability. Conceptual text representation methods have overcome the shortcomings of previous methods, but semi-supervised text clustering does not currently use concept-based document representation. This paper presents a two-level semi-supervised text clustering method that uses labeled and unlabeled data simultaneously to achieve higher clustering quality. In the first level, documents are represented based on the concepts extracted from the raw corpus. Second, the semi-supervised clustering process applies unlabeled data to capture the overall structure of the clusters and a small amount of labeled data to adjust the center of the clusters. Experiments on the Reuters-21578 data collection show that the proposed model is superior to other semi-supervised approaches in both text classification and text clustering.
کلیدواژه‌های انگلیسی مقاله News Clustering,Two-level clustering,Semi-supervised,word embedding,Document clustering

نویسندگان مقاله S. M. Sadjadi |
Faculty of Computer Engineering, Shahrood University of Technology, Shahrood, Iran

H. Mashayekhi |
Faculty of Computer Engineering, Shahrood University of Technology, Shahrood, Iran

H. Hassanpour |
Faculty of Computer Engineering, Shahrood University of Technology, Shahrood, Iran


نشانی اینترنتی https://www.ije.ir/article_137466_5d6eec012a034833f12edb2a3da1cce9.pdf
فایل مقاله فایلی برای مقاله ذخیره نشده است
کد مقاله (doi)
زبان مقاله منتشر شده en
موضوعات مقاله منتشر شده
نوع مقاله منتشر شده
برگشت به: صفحه اول پایگاه   |   نسخه مرتبط   |   نشریه مرتبط   |   فهرست نشریات