Large Language Models
Everything we've answered about how large language models work: tokens, context windows, model size, and open vs. closed models.
12 questions in this cluster
“Large language model” is a term everyone uses and few people can define precisely, so this cluster starts there — what an LLM actually is in plain terms, and why larger models generally perform better even though parameter count matters less on its own than it used to. From that foundation, the mechanical vocabulary gets unpacked one piece at a time: what a token is and why it drives cost, what a context window controls, what temperature does to an output, and what an embedding is doing when it lets a model recognize semantic similarity between two pieces of text.
The more technical questions get equal space rather than being waved off as too advanced: what retrieval-augmented generation does to reduce hallucination, what test-time compute is and how it improves reasoning, how speculative decoding speeds up responses, and why some models need far more memory to run than others of a similar size. Open-source versus closed models closes out the cluster, since it’s the distinction that determines who can actually inspect, modify, or self-host a given model.
AI Agents Explained: A Complete Guide to Autonomous and Agentic AI
Read the full guide →What is an ai models parameter count and why does it matter less than it used to?
An AI model's parameter count is the number of adjustable internal values that determine its learned behavior, and while a larger count historically correlated with greater capability, this relationship has weakened as more efficient training techniques and better data let smaller models match or exceed older, larger ones.
What is speculative decoding and how does it speed up ai responses?
Speculative decoding speeds up AI responses by using a smaller, faster draft model to quickly guess several upcoming words at once, which the larger, more capable model then verifies in a single pass rather than generating each word one at a time, producing the same quality output considerably faster when the draft guesses turn out to be correct.
What is test time compute and how does it improve ai reasoning?
Test-time compute refers to the additional computational effort an AI model spends actually working through a problem at the moment it's asked, rather than during its original training, and increasing this effort — letting a model reason through more intermediate steps before answering — has proven to genuinely improve accuracy on complex reasoning tasks beyond what training alone achieves.
What is an embedding and how does it let ai models understand semantic similarity?
An embedding is a numerical representation of text, an image, or other data as a list of numbers positioned within a mathematical space, where items with similar meaning end up positioned closer together, letting AI models mathematically compare and measure semantic similarity between different pieces of content in a way that goes well beyond simple keyword matching.
What is retrieval augmented generation and why does it reduce hallucination?
Retrieval-augmented generation, commonly called RAG, is a technique where an AI model first retrieves relevant information from a specific external knowledge source before generating its response, reducing hallucination by grounding the model's answer in retrieved, verifiable source material rather than relying purely on potentially imprecise information learned during training.
What is temperature in ai model settings and how does it affect the output?
Temperature is a setting that controls how random or predictable an AI model's output is, with a lower temperature producing more consistent, conservative responses and a higher temperature producing more varied, creative, but potentially less reliable output, making it a genuinely useful parameter to adjust depending on whether a task calls for precision or creative variation.
Why do some ai models require significantly more memory to run than others of similar size?
AI models with a similar total parameter count can still require significantly different amounts of memory to actually run, since factors like numerical precision used for the model's weights, the specific architecture design, and whether techniques like quantization have been applied all meaningfully affect actual memory requirements beyond parameter count alone.
What Does 'Context Window' Mean for an AI Model?
A context window is the maximum amount of text — measured in tokens — that an AI model can consider at once, including the prompt, any attached documents, and its own prior conversation history.
What Is a Large Language Model, in Plain Terms?
A large language model (LLM) is a computer program trained on huge amounts of text so it can predict and generate human-like language, letting it answer questions, write text, summarize, and hold conversations.
What Is a 'Token' in AI and Why Does It Matter for Cost?
A token is a small chunk of text — often a word, part of a word, or punctuation mark — that AI models use as their basic unit of processing, and most AI providers charge based on the number of tokens a request and its response use.
What Is the Difference Between Open-Source and Closed AI Models?
Open-source AI models release their weights (and sometimes training details) for anyone to download, run, and modify, while closed models are only accessible through a provider's API or product, with the underlying model kept private.
Why Do Larger AI Models Generally Perform Better?
Larger AI models generally perform better because more parameters, more training data, and more compute together let a model capture more nuanced patterns in language, a relationship researchers describe with 'scaling laws' — though bigger is not unconditionally better.
Other topics in AI Models & Technology
AI Agents
Everything we've answered about AI agents: how they differ from chatbots, taking real-world actions, and the risks of account access.
AI Hallucination & Accuracy
Everything we've answered about why AI models get things wrong: hallucinated facts, fake citations, and how to fact-check AI output.
AI Training & Fine-Tuning
Everything we've answered about how AI models are trained: pretraining, fine-tuning, RLHF, system prompts, and knowledge cutoffs.
Related categories
AI Models & Companies
Sourced answers about specific AI products and the companies behind them — Gemini, Llama, Perplexity, Copilot, and how to choose between providers.
Prompting & Everyday AI Use
Sourced, practical answers about getting better results from AI tools — prompt engineering, AI-assisted writing, productivity workflows, and getting started.
AI Policy, Law & Safety
Sourced answers about AI regulation, copyright and intellectual property, AI safety and alignment, and data privacy.
AI Tools & Assistants
Direct, sourced answers about the AI assistants and generative tools people actually use day to day — ChatGPT, Claude, AI coding assistants, and AI image generators.