دورية أكاديمية

Using Fuzzy Logic to Leverage HTML Markup for Web Page Representation.

التفاصيل البيبلوغرافية
العنوان: Using Fuzzy Logic to Leverage HTML Markup for Web Page Representation.
المؤلفون: Garcia-Plaza, Alberto P., Fresno, Victor, Unanue, Raquel Martinez, Zubiaga, Arkaitz
المصدر: IEEE Transactions on Fuzzy Systems; Aug2017, Vol. 25 Issue 4, p919-933, 15p
مصطلحات موضوعية: FUZZY logic, HTML (Document markup language), WEBSITES, FUZZY systems, DOCUMENT clustering, INFORMATION retrieval, META tags (HTML)
مستخلص: The selection of a suitable document representation approach plays a crucial role in the performance of a document clustering task. Being able to pick out representative words within a document can lead to substantial improvements in document clustering. In the case of web documents, the HTML markup that defines the layout of the content provides additional structural information that can be further exploited to identify representative words. In this paper, we introduce a fuzzy term weighing approach that makes the most of the HTML structure for document clustering. We set forth and build on the hypothesis that a good representation can take advantage of how humans skim through documents to extract the most representative words. The authors of web pages make use of HTML tags to convey the most important message of a web page through page elements that attract the readers’ attention, such as page titles or emphasized elements. We define a set of criteria to exploit the information provided by these page elements, and introduce a fuzzy combination of these criteria that we evaluate within the context of a web page clustering task. Our proposed approach, called abstract fuzzy combination of criteria (AFCC), can adapt to datasets whose features are distributed differently, achieving good results compared with other similar fuzzy logic based approaches and TF-IDF across different datasets. [ABSTRACT FROM PUBLISHER]
Copyright of IEEE Transactions on Fuzzy Systems is the property of IEEE and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
قاعدة البيانات: Complementary Index
الوصف
تدمد:10636706
DOI:10.1109/TFUZZ.2016.2586971