Skip to main navigation Skip to search Skip to main content

A Clustering of an Important and Improved Text Documents Using Datamining Techniques

  • J. Ranjith
  • , S. Raghavendra*
  • , M. Sharada
  • , Sheo Kumar
  • , B. Sindhusaranya
  • *Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

One of the many things that need to be done because the amount of text data available on such devices is also escalating rapidly, and we need efficient ways to sort and analyze large textual datasets. One of the popular ways data mining techniques can be conducted on given data is by clustering, an effective method for grouping similar text documents so they can be processed for relevance, summarization and knowledge acquisition. This study aims to explore the use of sophisticated data mining approaches to improve clustering of textual documents, examining techniques that maximize both performance and accuracy in high-dimensional spaces of text categories. In this study we compare various clustering algorithms including but not limited to K Means, hierarchical clustering and density based algorithm with preprocessing techniques such as tokenization, stop-word removal and term frequency-inverse document frequency (TF-IDF) weighting. We also examine dimensionality reduction methods (Latent Semantic Analysis (LSA) and Principal Component Analysis (PCA)) to deal with the sparsely and complexity of textual data. The performance of the implication patterns in the clustering task on benchmark datasets is evaluated using quality metrics like silhouette score, entropy, and purity through empirical experiments. Superior clustering results show that preprocessing pipelines tailored for text, and hybrid approaches combining machine learning algorithms with text-specific enhancements, are particularly advantageous.

Original languageEnglish
Title of host publicationIntelligent Data Engineering and Analytics - Proceedings of the 13th International Conference on Frontiers in Intelligent Computing
Subtitle of host publicationTheory and Applications FICTA 2025
EditorsVikrant Bhateja, Preeti Patel, Jinshan Tang
PublisherSpringer Science and Business Media Deutschland GmbH
Pages519-528
Number of pages10
ISBN (Print)9783032189738
DOIs
Publication statusPublished - 2026
Externally publishedYes
Event13th International Conference on Frontiers in Intelligent Computing: Theory and Applications, FICTA 2025 - London, United Kingdom
Duration: 06-06-202507-06-2025

Publication series

NameSmart Innovation, Systems and Technologies
Volume479 SIST
ISSN (Print)2190-3018
ISSN (Electronic)2190-3026

Conference

Conference13th International Conference on Frontiers in Intelligent Computing: Theory and Applications, FICTA 2025
Country/TerritoryUnited Kingdom
CityLondon
Period06-06-2507-06-25

All Science Journal Classification (ASJC) codes

  • General Decision Sciences
  • General Computer Science

Fingerprint

Dive into the research topics of 'A Clustering of an Important and Improved Text Documents Using Datamining Techniques'. Together they form a unique fingerprint.

Cite this