K Means ++ Clustering on Amazon Fine Food Reviews

 


Applied K means clustering to find best number of clusters with below listed text vectorization techniques

  • Bag of Words
  • Term Frequency — Inverse Document Frequency
  • Average word to vector
  • TF IDF weighted word to vector

Solution

https://www.kaggle.com/pradipdharam/10-k-means-clustering-on-amazonfinefoodreviews

Observations

‘Average word to vector’ and ‘tf-idf weighted word to vector’ performs better in order to find optimal number of clusters and volume.

Bag of Words:

· Optimal number of clusters found is 16

· Clusters are not getting formed properly.

· Data volume in each cluster is not consistent.

Cluster wise data volume:

Cluster wise data volume: 
0 9977
11 5
15 3
4 3
7 1
14 1
6 1
13 1
5 1
12 1
3 1
10 1
2 1
9 1
1 1
8 1

Name: cluster, dtype: int64

 

Term Frequency — Inverse Document Frequency (TF-IDF):

· Optimal number of clusters found is 18

· Clusters are not getting formed properly.

· Data volume in each cluster is not consistent.

Cluster wise data volume: 
8     9967
6       11
3        5
1        2
0        2
5        1
12       1
4        1
11       1
7        1
10       1
2        1
17       1
9        1
13       1
16       1
14       1
15       1
Name: cluster, dtype: int64

 

Average Word to Vector:

· Optimal number of clusters found is 18

· Clusters are getting formed properly.

· Data volume in each cluster is consistent.

Cluster wise data volume: 
10    1093
1      910
15     861
5      796
17     744
16     688
3      626
4      622
12     610
13     591
11     526
7      430
9      415
6      322
14     295
2      265
8      202
0        4
Name: cluster, dtype: int64

 

TF-IDF weighted Word to Vector:

· Optimal number of clusters found is 18

· Clusters are getting formed properly.

· Data volume in each cluster is consistent.

Cluster wise data volume: 
7     1240
12    1125
14    1025
11    1007
3      796
8      678
13     467
10     439
15     410
1      408
17     397
2      389
6      336
9      283
4      279
16     268
0      238
5      215
Name: cluster, dtype: int64

 

Previous Post Next Post