Learn / Category
Learn
AI basics
30 topics in this category.
Guides
1 guide
- Abliterated vs uncensored models
Uncensored names the result, a model that does not refuse; abliterated names one cheap way to get there, by editing out the single direction in the model's activations that carries refusal.
Definitions, A to Z
29 definitions
- AI hallucination
Confident output from a generative AI model that is false or invented, which NIST calls confabulation and OWASP counts as a cause of misinformation.
- AI safety
Keeping AI systems from causing harm, and where that ends and AI security begins.
- Chat template
The format, usually a Jinja template shipped with a model, that turns a list of system, user, assistant and tool messages into the single token sequence a chat model reads.
- Context window
The limit on how much text, counted in tokens, a language model can take into account at once, including the response it is writing.
- Embedding
A list of numbers an embedding model produces from a piece of text so that texts with similar meaning get nearby vectors, used to search by meaning and able to leak the text it came from.
- Federated learning
Training one shared machine learning model across many devices or organisations, each of which keeps its data and sends only model updates to a central server.
- Fine-tuning
Further training of an already pre-trained model on a smaller, targeted dataset, which is how models get customised and one way their safety training gets undone.
- GGUF
The single-file model format that llama.cpp and other GGML-based runners load, holding a model's weights together with metadata such as its tokenizer and chat template.
- In-context learning
A language model's ability to pick up a task from instructions or examples placed in its prompt, with no retraining.
- Instruction hierarchy
The ranking that tells a language model whose instructions win in a conflict, and why that ranking is a trained tendency rather than a security boundary.
- Instruction tuning
The second, much smaller training stage that teaches a pre-trained language model to follow instructions, and a stage researchers have poisoned with a hundred examples.
- Language model
A probability distribution over sequences of tokens, learned from text, that lets software predict and generate natural language.
- llama.cpp
The open-source C/C++ engine, built on the GGML library, that runs language models from GGUF files on local hardware and can serve them over an HTTP API.
- Logit bias
An API parameter that adds a fixed amount to a language model's score for chosen tokens before the next token is picked, and a known route for reading out a model's hidden scores.
- Mechanistic interpretability
Reverse engineering what a trained model's internal parts compute, which NIST records among the proposed ways to find backdoor features in a model.
- Model checkpoint
A saved snapshot of a model's parameters and optionally its optimiser state at a specific point in training, serialised to a file format such as safetensors or .pt.
- Model distillation
Training a smaller student model to reproduce a larger teacher model's behaviour by learning from the teacher's outputs.
- Model drift
The change in a deployed machine learning model's performance over time, usually a fall as the data it meets moves away from the data it learned from.
- Open-source AI
An AI system anyone may use, study, modify and share, released with its weights, its full training and running code, and enough detail about its training data to build a substantially equivalent system.
- Open-weight model
An AI model whose trained weights are published for anyone to download and run, which is less than open source and more than API access, and which puts the model's safeguards in the hands of whoever holds the file.
- Prompt
The complete input a caller sends to a language model in a single turn.
- Quantization
Storing a model's weights in fewer bits than it was trained with, so it needs less memory to run, at the cost of rounding every weight.
- Retrieval-augmented generation
A way of building language-model applications in which the system first searches a collection of documents, then hands the best matches to the model alongside the question.
- Safetensors
Hugging Face's file format for model weights, which stores only tensors and a JSON header so that loading a downloaded file cannot run code the way a pickle file can.
- System One model
A class of AI model that returns typed, probabilistic decisions instead of generated text, so software can act on the answer directly. The term was coined by TypeSafe AI in September 2026.
- System prompt
The instructions an application gives a language model before the conversation starts.
- Token
The unit of text a language model reads and writes, which may be a word, part of a word, a character or a byte, and the unit context windows are sized in.
- Vector database
A database that stores embeddings and finds the ones nearest a query, where retrieval-augmented generation keeps its documents and where one user's data can reach another's answers.
- Vibe coding
Building software by telling an AI what you want and accepting the code it writes without reading it.