Applied K means clustering to find best number of clusters with below listed text vectorization techniques
- Bag of Words
- Term Frequency — Inverse Document Frequency
- Average word to vector
- TF IDF weighted word to vector
Solution
https://www.kaggle.com/pradipdharam/10-k-means-clustering-on-amazonfinefoodreviews
Observations
‘Average word to vector’ and ‘tf-idf weighted word to vector’ performs better in order to find optimal number of clusters and volume.
Bag of Words:
· Optimal number of clusters found is 16
· Clusters are not getting formed properly.
· Data volume in each cluster is not consistent.
Cluster wise data volume:
Cluster wise data volume:
0 9977
11 5
15 3
4 3
7 1
14 1
6 1
13 1
5 1
12 1
3 1
10 1
2 1
9 1
1 1
8 1Name: cluster, dtype: int64
Term Frequency — Inverse Document Frequency (TF-IDF):
· Optimal number of clusters found is 18
· Clusters are not getting formed properly.
· Data volume in each cluster is not consistent.
Cluster wise data volume: 8 9967 6 11 3 5 1 2 0 2 5 1 12 1 4 1 11 1 7 1 10 1 2 1 17 1 9 1 13 1 16 1 14 1 15 1 Name: cluster, dtype: int64
Average Word to Vector:
· Optimal number of clusters found is 18
· Clusters are getting formed properly.
· Data volume in each cluster is consistent.
Cluster wise data volume: 10 1093 1 910 15 861 5 796 17 744 16 688 3 626 4 622 12 610 13 591 11 526 7 430 9 415 6 322 14 295 2 265 8 202 0 4 Name: cluster, dtype: int64
TF-IDF weighted Word to Vector:
· Optimal number of clusters found is 18
· Clusters are getting formed properly.
· Data volume in each cluster is consistent.
Cluster wise data volume: 7 1240 12 1125 14 1025 11 1007 3 796 8 678 13 467 10 439 15 410 1 408 17 397 2 389 6 336 9 283 4 279 16 268 0 238 5 215 Name: cluster, dtype: int64