TY - GEN
T1 - An expemmental study of the effect of frequency of co-occurrence of features in clustering
AU - Pai, Radhika M.
AU - Ananthanarayana, V. S.
PY - 2007
Y1 - 2007
N2 - In this paper, an attempt has been made to explore the effect of frequency of co-occurrence of features on the accuracy of the clustering results. This has been achieved by incorporating the frequency component in the clustering algorithm. The frequency, we mean here is the number of times the sequence of features appear in the data set. We try to utilize this component in the algorithm and study its effect on the resultant accuracy. The algorithm we have used is the PC(pattern count)-tree based clustering algorithm. The PC-tree is a compact and complete representation of the data set. It is data order independent and incremental. It can be applied to changing data and changing knowledge. i.e. dynamic databases. This algorithm is based on a compact data structure called PC-tree. The node of the PC-tree has, in addition to other fields a count field, which keeps track of the count of the number of features shared by the pattern. In the literature, the PC-tree was used for clustering and the count field was used only to retrieve back the transactions. In this paper, we try to make use of this field in clustering. We have also used the partitioned PC-tree based algorithm and studied the effect of frequency on the accuracy. We have conducted extensive experiments with the OCR handwritten digit dataset, a real dataset and observed the effect of frequency on the clustering results. The results of all our experiments are tabulated.
AB - In this paper, an attempt has been made to explore the effect of frequency of co-occurrence of features on the accuracy of the clustering results. This has been achieved by incorporating the frequency component in the clustering algorithm. The frequency, we mean here is the number of times the sequence of features appear in the data set. We try to utilize this component in the algorithm and study its effect on the resultant accuracy. The algorithm we have used is the PC(pattern count)-tree based clustering algorithm. The PC-tree is a compact and complete representation of the data set. It is data order independent and incremental. It can be applied to changing data and changing knowledge. i.e. dynamic databases. This algorithm is based on a compact data structure called PC-tree. The node of the PC-tree has, in addition to other fields a count field, which keeps track of the count of the number of features shared by the pattern. In the literature, the PC-tree was used for clustering and the count field was used only to retrieve back the transactions. In this paper, we try to make use of this field in clustering. We have also used the partitioned PC-tree based algorithm and studied the effect of frequency on the accuracy. We have conducted extensive experiments with the OCR handwritten digit dataset, a real dataset and observed the effect of frequency on the clustering results. The results of all our experiments are tabulated.
UR - https://www.scopus.com/pages/publications/51549104346
UR - https://www.scopus.com/pages/publications/51549104346#tab=citedBy
U2 - 10.1109/ISSPA.2007.4555535
DO - 10.1109/ISSPA.2007.4555535
M3 - Conference contribution
AN - SCOPUS:51549104346
SN - 1424407796
SN - 9781424407798
T3 - 2007 9th International Symposium on Signal Processing and its Applications, ISSPA 2007, Proceedings
BT - 2007 9th International Symposium on Signal Processing and its Applications, ISSPA 2007, Proceedings
T2 - 2007 9th International Symposium on Signal Processing and its Applications, ISSPA 2007
Y2 - 12 February 2007 through 15 February 2007
ER -