What artificial intelligence and large language models are, in plain words

Behind most of today's coding assistants stands one kind of technology: the large language model, or LLM for short. It is worth understanding what it is, because that explains both its strength and its weaknesses.
Three concepts that are easy to confuse
- Artificial intelligence (AI) is a broad umbrella term: anything that imitates tasks requiring human intelligence, from a spam filter to image recognition.
- Machine learning is a way of building AI: instead of writing rules by hand, we show a program a great many examples and it draws out the patterns itself.
- A large language model (LLM) is a machine learning model trained on huge collections of text, and often code as well. It can produce the next pieces of text that fit an instruction.
How a model learns
During training, the model is given countless fragments of text and learns to predict what should come next. Repeated billions of times, this exercise has a surprising effect: the model absorbs grammar, style, many facts and also typical constructs in programming languages. After this stage, models are usually fine-tuned further so that they follow instructions more readily and respond in a useful and safe way.
Why it can write code
Code is text too, with its own grammar and recurring patterns. The model has seen a great deal of it, so it can match a typical solution to a description of a task. It does not "understand" a program the way a human does: it generates what statistically fits the instruction and the context of the conversation.
What this means in practice
- The model is excellent where there are patterns: typical functions, forms, simple apps, translation between programming languages.
- It can be unsure where patterns are scarce: niche libraries, very new tool versions, unusual business logic.
- It sounds confident even when it is wrong. More on this in the text on tokens, context and hallucinations.
- It does not know your project until you show it. It knows only what it was given in the conversation, plus what it learned earlier.
A model is not a search engine
A search engine finds existing pages. A model generates a new answer based on what it has learned, which is why the answer can be both credible and false. It is good to keep this in mind with every answer, especially when it comes to facts, numbers and library names.
Once you know what a model is, move on to practice: prompt engineering for everyone.


