What matters in AI.

Subscribe

Learn / AI basics

Definition · AI basics

Mechanistic interpretability

Mechanistic interpretability is the study of a trained neural network's internal parts, its features, neurons and circuits, to work out the computations that turn inputs into outputs. In AI security, NIST records it among the current proposed approaches to finding backdoor features in a model, the hidden rules a poisoning attack can leave behind.

Last reviewed

Key points

  • Mechanistic interpretability reads a model from the inside, working out what its internal parts compute instead of testing what it outputs.
  • NIST records it among current proposed approaches to finding backdoor features in poisoned models. Elsewhere, NIST warns that weights may not be auditable the way open-source code is.
  • It does not work reliably yet. In a 2023 test on image classifiers with 12 implanted backdoors, even the best of nine tools helped people pick the trigger out of eight options less than half the time.
  • The clearest success leaned on the data. In a 2025 auditing game, all three teams given training data found a planted hidden objective, two using interpretability tools; the team with neither weights nor data did not.

Testing a model shows what it does on the inputs you try. A backdoor is built to pass that test: it behaves normally until an input carries the attacker’s trigger. Mechanistic interpretability tries to read the rule from the inside instead.

What it looks at

Bereska and Gavves describe the field as “reverse engineering the computational mechanisms and representations learned by neural networks into human-understandable algorithms and concepts”. It works bottom up, through “features, neurons, layers, and connections”, and asks which parts cause which output.

For backdoors, the appeal is finding the trigger without owning a sample of it. Tools that only explain inputs you already have “cannot help to identify failures that are not present in some readily available dataset”, Casper and colleagues note. Feature synthesis tools instead generate the input that most strongly activates a part of the model, which can surface a trigger nobody has seen.

Why it matters

NIST warns that organizations and researchers “may not be able to audit and identify vulnerabilities encoded into a model’s weights in the same way it is often possible to audit open-source software”. Controls such as dataset provenance and an AI bill of materials ask a different question: where a model and its data came from.

In its mitigations for model poisoning, NIST says new approaches are required and lists “methods from the field of mechanistic interpretability to identify backdoor features” among those currently proposed. Mechanistic interpretability is the research that would let the weights answer the question themselves. It does not do that reliably yet.

What it can find today

Less than the name suggests. Casper and colleagues implanted 12 trojans, their word for backdoors, in image classifiers in 2023. Most of 16 attribution tools, which highlight the pixels behind an answer, “struggle to beat a trivial edge-detector baseline” at locating a trigger in an image that contained it. Visualising inner neurons was not effective either, most likely, the authors think, because a trigger is recognised by patterns “among multiple neurons instead of single neurons”. Sharkey and colleagues’ 2025 review of open problems in the field cites Casper’s related finding that feature-synthesis methods struggle to identify trojans.

The clearest success comes from Anthropic, a vendor, in 2025. Marks and colleagues trained a language model with a hidden objective. Three of four blind teams found it, two with the help of interpretability tools. All three had the training data, where “the small number of synthetic documents” gave the game away. The team with neither weights nor data failed. The authors say training data that states the key information “may be important for many of our techniques”, interpretability included.

Where definitions disagree

NIST cites Casper and colleagues for “mechanistic interpretability”. The paper itself never uses the word; it benchmarks “interpretability tools”, 16 feature attribution tools and 9 feature synthesis methods. Sharkey and colleagues count feature synthesis as a mechanistic method, so NIST’s filing is defensible, but a reader following the citation will not find the term there.

Questions and answers

Can mechanistic interpretability find a backdoor in a model today?

Not dependably, and the published evidence says so. Casper and colleagues implanted 12 trojans, their word for backdoors, in image classifiers and asked people to find the triggers with the help of nine feature synthesis tools. Even the best of them "fell short of helping humans succeed 50% of the time on 8-option multiple-choice questions", and none worked on style-transfer trojans. The authors add that implanted trojans may be "a much easier challenge than real-world debugging". NIST lists the approach as proposed, not as available.

Is mechanistic interpretability the same as detecting a trigger at inference time?

No. NIST lists them as two separate proposed approaches to poisoned models: interpretability to "identify backdoor features" in the model, and "detecting and counteracting triggers when they are seen at inference time". The first looks for the backdoor inside the model; the second watches for its trigger in the inputs. In a different section, on defending against direct prompting attacks, NIST adds: "Early work suggests that interpretability-based techniques can also be used to detect anomalous input".

Does this mean weights can be audited like source code after all?

Not yet. NIST's statement that organizations and researchers "may not be able to audit and identify vulnerabilities encoded into a model's weights in the same way it is often possible to audit open-source software" still stands; mechanistic interpretability is the research aimed at changing it. Until it does, controls on where a model and its data came from, such as dataset provenance and an AI bill of materials, carry the load.

Sources

  1. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
  2. Mechanistic Interpretability for AI Safety: A ReviewarXiv (Bereska and Gavves, University of Amsterdam), 22 Apr 2024
  3. Red Teaming Deep Neural Networks with Feature Synthesis ToolsarXiv (Casper et al., MIT CSAIL), 8 Feb 2023
  4. Open Problems in Mechanistic InterpretabilityarXiv (Sharkey et al.), 27 Jan 2025
  5. Auditing Language Models for Hidden ObjectivesarXiv (Marks et al., Anthropic), 14 Mar 2025

Guides that use this term