TY - GEN
T1 - A Clustering of an Important and Improved Text Documents Using Datamining Techniques
AU - Ranjith, J.
AU - Raghavendra, S.
AU - Sharada, M.
AU - Kumar, Sheo
AU - Sindhusaranya, B.
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.
PY - 2026
Y1 - 2026
N2 - One of the many things that need to be done because the amount of text data available on such devices is also escalating rapidly, and we need efficient ways to sort and analyze large textual datasets. One of the popular ways data mining techniques can be conducted on given data is by clustering, an effective method for grouping similar text documents so they can be processed for relevance, summarization and knowledge acquisition. This study aims to explore the use of sophisticated data mining approaches to improve clustering of textual documents, examining techniques that maximize both performance and accuracy in high-dimensional spaces of text categories. In this study we compare various clustering algorithms including but not limited to K Means, hierarchical clustering and density based algorithm with preprocessing techniques such as tokenization, stop-word removal and term frequency-inverse document frequency (TF-IDF) weighting. We also examine dimensionality reduction methods (Latent Semantic Analysis (LSA) and Principal Component Analysis (PCA)) to deal with the sparsely and complexity of textual data. The performance of the implication patterns in the clustering task on benchmark datasets is evaluated using quality metrics like silhouette score, entropy, and purity through empirical experiments. Superior clustering results show that preprocessing pipelines tailored for text, and hybrid approaches combining machine learning algorithms with text-specific enhancements, are particularly advantageous.
AB - One of the many things that need to be done because the amount of text data available on such devices is also escalating rapidly, and we need efficient ways to sort and analyze large textual datasets. One of the popular ways data mining techniques can be conducted on given data is by clustering, an effective method for grouping similar text documents so they can be processed for relevance, summarization and knowledge acquisition. This study aims to explore the use of sophisticated data mining approaches to improve clustering of textual documents, examining techniques that maximize both performance and accuracy in high-dimensional spaces of text categories. In this study we compare various clustering algorithms including but not limited to K Means, hierarchical clustering and density based algorithm with preprocessing techniques such as tokenization, stop-word removal and term frequency-inverse document frequency (TF-IDF) weighting. We also examine dimensionality reduction methods (Latent Semantic Analysis (LSA) and Principal Component Analysis (PCA)) to deal with the sparsely and complexity of textual data. The performance of the implication patterns in the clustering task on benchmark datasets is evaluated using quality metrics like silhouette score, entropy, and purity through empirical experiments. Superior clustering results show that preprocessing pipelines tailored for text, and hybrid approaches combining machine learning algorithms with text-specific enhancements, are particularly advantageous.
UR - https://www.scopus.com/pages/publications/105039010547
UR - https://www.scopus.com/pages/publications/105039010547#tab=citedBy
U2 - 10.1007/978-3-032-18974-5_46
DO - 10.1007/978-3-032-18974-5_46
M3 - Conference contribution
AN - SCOPUS:105039010547
SN - 9783032189738
T3 - Smart Innovation, Systems and Technologies
SP - 519
EP - 528
BT - Intelligent Data Engineering and Analytics - Proceedings of the 13th International Conference on Frontiers in Intelligent Computing
A2 - Bhateja, Vikrant
A2 - Patel, Preeti
A2 - Tang, Jinshan
PB - Springer Science and Business Media Deutschland GmbH
T2 - 13th International Conference on Frontiers in Intelligent Computing: Theory and Applications, FICTA 2025
Y2 - 6 June 2025 through 7 June 2025
ER -