What artificial intelligence and large language models are, in plain words

Basics6 min read

Behind most of today's coding assistants stands one kind of technology: the large language model, or LLM for short. It is worth understanding what it is, because that explains both its strength and its weaknesses.

Three concepts that are easy to confuse

How a model learns

During training, the model is given countless fragments of text and learns to predict what should come next. Repeated billions of times, this exercise has a surprising effect: the model absorbs grammar, style, many facts and also typical constructs in programming languages. After this stage, models are usually fine-tuned further so that they follow instructions more readily and respond in a useful and safe way.

Why it can write code

Code is text too, with its own grammar and recurring patterns. The model has seen a great deal of it, so it can match a typical solution to a description of a task. It does not "understand" a program the way a human does: it generates what statistically fits the instruction and the context of the conversation.

What this means in practice

A model is not a search engine

A search engine finds existing pages. A model generates a new answer based on what it has learned, which is why the answer can be both credible and false. It is good to keep this in mind with every answer, especially when it comes to facts, numbers and library names.

Once you know what a model is, move on to practice: prompt engineering for everyone.