What isan LLM?
The large language model behind ChatGPT, Copilot and the like, without the jargon.
Get in touch
Let's talk about your project.
First we check whether the project fits your business model. Then you get a proposal with phases and effort.
Talk about AI in your business or call: +49 151 1576 5566An LLM, short for large language model, is an AI model that understands text and writes text of its own. It sits inside chatbots and AI assistants that you talk to in plain language.
- An LLM only knows what was in its training data.
- When something is missing, it sometimes pieces together an answer that sounds plausible and is still wrong.
- It gets your company knowledge through RAG or through fine-tuning.
- Usage is billed in tokens, chunks of text of roughly three quarters of a word.
What does an LLM know, and where does it hit its limits in your business?
A large language model stores a surprising amount of factual knowledge, yet it struggles to retrieve it in a targeted way, can't show where a statement comes from and is hard to bring up to date. It has never seen your parts list from last quarter. Ask about it and the model either has to pass or makes something up (that's what people call a hallucination).
In day to day work that means an LLM shines at wording, summarizing and sorting. For facts from your own business it needs documents to rely on, plus a person who checks and approves before anything goes out.
How do you get your own knowledge into an LLM, and what does running it cost?
There are two routes. With RAG, a separate program looks up matching passages in your documents before every answer and the model writes from them without changing itself. With fine-tuning, you train the model further on your own data. In a study presented at EMNLP 2024, RAG came out ahead on factual knowledge throughout, while fine-tuning suits style, format and technical language better.
Billing works in tokens, counted separately for your input and the answer, and the answer costs considerably more. So the figure that holds up is the price per completed task, not the price per million tokens. If you'd rather not send data to the cloud, you can also run an open model on your own hardware.
Frequently asked questions
What is the difference between an LLM and RAG?
The LLM is the language model that writes. RAG is the machinery around it that hands it matching passages from a search before each answer, so it doesn't have to rely on memory alone.
What is a token?
A chunk of text that language models calculate with, roughly three quarters of a word. How many tokens a text produces depends on the model.
Can an LLM run on your own server?
Yes, with so-called open models whose weights have been published. How much hardware you need depends on the size of the model.
Terms you should know in the same context
RAG (retrieval augmented generation) · AI hallucination · Fine-tuning · Local AI · AI assistant · Back to the AI glossary A to Z
A language model writes well, but what it should know about your business is something you have to give it.

