{"id":9577,"date":"2026-07-27T12:30:27","date_gmt":"2026-07-27T10:30:27","guid":{"rendered":"https:\/\/www.kubicek.ai\/?post_type=lexicon&#038;p=9577"},"modified":"2026-07-27T13:34:49","modified_gmt":"2026-07-27T11:34:49","slug":"k-nearest-neighbors-k-nn","status":"publish","type":"lexicon","link":"https:\/\/www.kubicek.ai\/en\/lexicon\/k-nearest-neighbors-k-nn\/","title":{"rendered":"k-Nearest Neighbors (k-NN)"},"content":{"rendered":"<p class=\"wp-block-paragraph\"><strong>k-Nearest Neighbors (k-NN)<\/strong> is the simplest supervised learning algorithm and an example of so-called lazy learning: it essentially has no training phase, it merely stores the entire training set. A prediction for a new sample is derived only at query time \u2013 the k nearest training points are found according to a chosen metric, typically Euclidean or cosine, and the result is determined by their vote in classification or their average in regression, possibly weighted by distance. The choice of k is the key hyperparameter: a value of one gives a model extremely sensitive to noise and outliers, while too large a k erases local structure and leads to underfitting. Feature normalisation is a necessary condition, because otherwise the scale of one variable drowns out all the others. The main weaknesses are the cost of inference, which grows with dataset size, and the curse of dimensionality \u2013 in high dimensions distance stops being informative. Precisely for that reason, modern deployments combine it with embeddings and approximate nearest-neighbour search in a vector database.<\/p>\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n<p class=\"wp-block-paragraph\">It is the method you use when moving to a new neighbourhood and wanting to estimate the price of a house. You do not study economic models \u2013 you look at what the five nearest houses of similar size sold for and take the average. You need no preparation for that, just a list of sales. Two things make it treacherous. If you look at only the single nearest house, you may hit the one sold within a family for a token sum. And if you let the comparison include distance from the brook in centimetres and age in seconds, the large numbers will drown out everything else.<\/p>\n","protected":false},"featured_media":0,"template":"","class_list":["post-9577","lexicon","type-lexicon","status-publish","hentry"],"acf":[],"_links":{"self":[{"href":"https:\/\/www.kubicek.ai\/en\/wp-json\/wp\/v2\/lexicon\/9577","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.kubicek.ai\/en\/wp-json\/wp\/v2\/lexicon"}],"about":[{"href":"https:\/\/www.kubicek.ai\/en\/wp-json\/wp\/v2\/types\/lexicon"}],"wp:attachment":[{"href":"https:\/\/www.kubicek.ai\/en\/wp-json\/wp\/v2\/media?parent=9577"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}