{"id":9528,"date":"2026-07-27T12:30:28","date_gmt":"2026-07-27T10:30:28","guid":{"rendered":"https:\/\/www.kubicek.ai\/?post_type=lexicon&#038;p=9528"},"modified":"2026-07-27T13:10:14","modified_gmt":"2026-07-27T11:10:14","slug":"batch","status":"publish","type":"lexicon","link":"https:\/\/www.kubicek.ai\/en\/lexicon\/batch\/","title":{"rendered":"Batch"},"content":{"rendered":"<p class=\"wp-block-paragraph\">A <strong>batch<\/strong> is a group of training samples pushed through the model together, from which a single update of the parameters is computed. The batch size is one of the key hyperparameters and represents a trade-off between the quality of the gradient estimate and computational efficiency. A small batch gives a noisy but often beneficial gradient \u2013 the random scatter helps escape poor minima and acts as a regularizer \u2013 while making worse use of the accelerator&#8217;s parallel capacity. A large batch gives a smoother, more faithful estimate, permits a higher learning rate and keeps the hardware far busier, but enlarging it excessively harms generalization and beyond a certain point stops delivering any speed-up at all. When the desired batch will not fit in memory, gradient accumulation is used: the gradient is computed from several smaller portions in turn and the parameters are updated only once they have been summed. In the training of large language models the batch is measured not in samples but in tokens, and routinely runs into the millions, because that is the only way to keep thousands of accelerators working in parallel sensibly occupied.<\/p>\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n<p class=\"wp-block-paragraph\">Imagine you want to find out how the residents of a city vote. Asking each person individually and revising your estimate after every answer is exhausting, and the estimate lurches wildly depending on who you happen to run into. Counting all five hundred thousand people first and only then adjusting your estimate is absurdly slow \u2013 you would get an answer once a year. A sensible pollster asks a thousand people, forms a picture from them, adjusts the estimate and moves on to the next thousand. That thousand is the batch: big enough not to be a fluke, small enough to handle right away.<\/p>\n","protected":false},"featured_media":0,"template":"","class_list":["post-9528","lexicon","type-lexicon","status-publish","hentry"],"acf":[],"_links":{"self":[{"href":"https:\/\/www.kubicek.ai\/en\/wp-json\/wp\/v2\/lexicon\/9528","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.kubicek.ai\/en\/wp-json\/wp\/v2\/lexicon"}],"about":[{"href":"https:\/\/www.kubicek.ai\/en\/wp-json\/wp\/v2\/types\/lexicon"}],"wp:attachment":[{"href":"https:\/\/www.kubicek.ai\/en\/wp-json\/wp\/v2\/media?parent=9528"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}