{"id":9545,"date":"2026-07-27T12:30:28","date_gmt":"2026-07-27T10:30:28","guid":{"rendered":"https:\/\/www.kubicek.ai\/?post_type=lexicon&#038;p=9545"},"modified":"2026-07-27T13:23:35","modified_gmt":"2026-07-27T11:23:35","slug":"k-means-clustering","status":"publish","type":"lexicon","link":"https:\/\/www.kubicek.ai\/en\/lexicon\/k-means-clustering\/","title":{"rendered":"k-means clustering"},"content":{"rendered":"<p class=\"wp-block-paragraph\"><strong>k-means clustering<\/strong> is a fundamental unsupervised learning algorithm that divides data into a predetermined number k of groups so as to minimise the sum of squared distances of points from the centres of their groups. It proceeds iteratively in two alternating steps: first it assigns each point to the nearest centre, then it recomputes the centres as the mean of the assigned points; this repeats while assignments keep changing. The algorithm always converges, but only to a local optimum, so the result depends on the initial choice of centres \u2013 which is why the careful k-means++ initialisation and several independent runs are used. The choice of k is an external decision, estimated by the elbow method or the silhouette score. Crucial are the assumptions the method silently makes: clusters roughly spherical, of similar size and comparable density. It fails on elongated, nested or markedly unequal groups, and it is also sensitive to outliers and to feature scale, which must be normalised beforehand. In high dimensions Euclidean distance loses its discriminating power, so it is often combined with dimensionality reduction. Typical uses are customer segmentation or thematic grouping of embeddings.<\/p>\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n<p class=\"wp-block-paragraph\">Picture a school gym with three hundred people standing in it, and you want four groups. You place four leaders in the middle at random and say: &#8220;Everyone go to the nearest leader.&#8221; Four huddles form. Then you tell the leaders: &#8220;Move to the middle of your huddle.&#8221; Some people are now closer to a different leader, so they switch over. You repeat this and before long nobody moves any more \u2013 you have your groups. Two catches: you have to know in advance that you want four, not five. And if the people are standing in a long line along the wall, splitting them by distance from a centre gives a nonsensical result.<\/p>\n","protected":false},"featured_media":0,"template":"","class_list":["post-9545","lexicon","type-lexicon","status-publish","hentry"],"acf":[],"_links":{"self":[{"href":"https:\/\/www.kubicek.ai\/en\/wp-json\/wp\/v2\/lexicon\/9545","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.kubicek.ai\/en\/wp-json\/wp\/v2\/lexicon"}],"about":[{"href":"https:\/\/www.kubicek.ai\/en\/wp-json\/wp\/v2\/types\/lexicon"}],"wp:attachment":[{"href":"https:\/\/www.kubicek.ai\/en\/wp-json\/wp\/v2\/media?parent=9545"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}