Guides and definitions
Learn
Each topic says what a term means, in under a minute, then goes deeper. Topics are sorted into 4 categories.
AI agents
2 guides · 54 definitions
- AI agent attacks
Every named attack on an AI agent, grouped by the step of the agent's loop it targets.
- How to scope write permissions for LLM agents
An LLM agent that changes staging or production config should propose changes rather than apply them, hold only short-lived task-scoped credentials, and be checked by a policy the model cannot change.
- Agent Control Standard
An OWASP open specification for runtime governance of AI agents, where a separate Guardian Agent allows, denies, or modifies an agent's actions through middleware hooks.
- Agent data injection
An indirect prompt injection that corrupts the trusted metadata an AI agent relies on, so the agent acts on attacker-controlled data while still doing the user's task.
- Agent harness
The software layer that wraps a language model, manages its context and tool use, and runs the loop that turns a model into an agent. NIST calls it scaffolding software; Anthropic places it as the third layer in the four-layer framework.
AI basics
1 guide · 29 definitions
- Abliterated vs uncensored models
Uncensored names the result, a model that does not refuse; abliterated names one cheap way to get there, by editing out the single direction in the model's activations that carries refusal.
- AI hallucination
Confident output from a generative AI model that is false or invented, which NIST calls confabulation and OWASP counts as a cause of misinformation.
- AI safety
Keeping AI systems from causing harm, and where that ends and AI security begins.
- Chat template
The format, usually a Jinja template shipped with a model, that turns a list of system, user, assistant and tool messages into the single token sequence a chat model reads.
- Context window
The limit on how much text, counted in tokens, a language model can take into account at once, including the response it is writing.
AI governance
1 guide · 19 definitions
- How to use LLMs securely in regulated industries
To use an LLM securely in a regulated industry, treat the provider like any outside service that handles regulated data. Four questions about each use, not the brand of model, decide the controls.
- AI Office
The European Commission function that helps implement and supervise the EU AI Act, enforcing the rules for general-purpose AI models and, since the Digital Omnibus on AI, supervising some AI systems' providers directly.
- AI risk management
The NIST framework for identifying, assessing and treating the risks of AI systems, organised around four functions and seven trustworthiness characteristics.
- AI system impact assessment
A documented study of how an AI system and its foreseeable uses could affect individuals and societies, the subject of ISO/IEC 42005:2025 and required in its own form for Canadian federal automated systems that make or support administrative decisions about clients.
- Automation bias
The tendency to accept an automated system's output instead of checking for yourself, first studied in aviation and now named in the EU AI Act's human oversight rules.
AI security
9 guides · 78 definitions
- Guardrail bypass techniques
The named ways attackers get a language model past its safety controls, grouped by the control each one mainly gets around, plus the reconnaissance that comes first.
- How to detect shadow AI
Shadow AI can be found in proxy and firewall logs, identity-provider app consents, code repositories, cloud audit logs and scans for self-hosted models. Each source misses something, so check several and give staff a safe way to report what they use.
- How to red team an LLM application
Red team an LLM application by starting from the harm that would matter, attacking the whole deployed application where untrusted text gets in, probing by hand before automating, and repeating every attack.
- Input filtering vs output validation
Input filtering screens what goes into an LLM; output validation checks what comes out before a user or another system acts on it. Each misses what the other catches, and OWASP and MITRE ATLAS list both.
- Security risks of running LLMs locally
Running a language model locally keeps prompts away from a provider, but it puts a downloaded model file, an inference server and the model's behaviour on your own machine, and it leaves prompt injection and hallucinated output in place.