What matters in AI.

Subscribe

Learn / AI basics

Guide · AI basics

Abliterated vs uncensored models

An abliterated model is one kind of uncensored model. "Uncensored" names the result: a model that answers requests a safety-trained chat model would refuse, whether its refusals were removed or never trained in. "Abliterated" names one method: find the direction in the model's activations that carries refusal, then edit the weights so the model can no longer represent it, with no retraining.

Last reviewed

Uncensored is the result, abliterated is one way to get it

The two words answer different questions. “Uncensored” says what the model does: it answers requests that a model trained for model alignment would refuse. “Abliterated” says how one such model was made. Maxime Labonne’s June 2024 tutorial on the method is titled “Uncensor any LLM with abliteration”, which puts both words in their places.

There are three common routes to an uncensored model.

Retrain without refusals. Eric Hartford’s 2023 WizardLM-Uncensored models were built by removing as many refusals and “biased answers” as he could from an instruction dataset, then training the base model on what was left, the same way the original was trained. The refusals are left out of training rather than taken out afterwards. For the 7B model that meant a 4x A100 node and a run he estimated at about 26 hours.

Fine-tune the refusals away. Start from a safety-trained model and train it further on examples of it complying. Lermen and colleagues did this to Llama 2-Chat 70B with a budget under $200 and one GPU, and cut its refusal rate to about 1% on two benchmarks. Qi and colleagues did it to GPT-3.5 Turbo through OpenAI’s own fine-tuning API, with 10 examples, for under $0.20.

Abliterate. Arditi and colleagues found that in 13 open-source chat models, up to 72B parameters, refusal runs through a single direction in the model’s internal activations. Remove it and the model stops refusing harmful requests; add it and the model refuses harmless ones. They find the direction by comparing the model’s average activations on harmful and harmless prompts. They then edit the weights so that no layer can write that direction into the model’s internal state. There is no training step and no harmful answers are needed. The paper puts the cost for a 70B model at under $5 of compute. “Abliteration” is the community’s name for this technique. Labonne’s tutorial credits Arditi and colleagues with the finding, and bases its code on a notebook by the developer FailSpy, “which is itself based on the original authors’ notebook”.

Side by side

Retrained or fine-tuned Abliterated
How the weights change Training on data with refusals removed, or with compliant examples added A direct edit that projects one direction out, with no training
What you need A dataset and a training run Sets of harmful and harmless prompts
Needs the weights Not always: Hartford and Lermen trained on the weights, Qi used a hosted API Yes
Reported cost Lermen: a budget under $200 for 70B; Qi: under $0.20 through an API Arditi: under $5 of compute for 70B
Side effects reported Lermen: capabilities kept on two general benchmarks Contested; see below

Each cost is the paper’s own figure, measured its own way, so the row does not rank the routes.

What abliteration removes, and what it costs

Arditi and colleagues say the meaning of the removed direction “remains unclear”. They call it the “refusal direction” as a functional label, and say it could stand for a concept such as “harm” or “danger” instead.

The measured side effects disagree. In the paper, most edited models scored close to the originals on MMLU, ARC and GSM8K, benchmarks of general knowledge, science questions and grade-school maths. The paper names Qwen 7B and Yi 34B as the exceptions. TruthfulQA, a benchmark of whether a model repeats common falsehoods, consistently dropped; the authors note it sits close to refusal territory, with categories like misinformation and conspiracies. Labonne abliterated Daredevil-8B, a merged model in the Llama 3 8B family, and saw scores drop on every benchmark he ran. Further preference training with DPO recovered most of the loss, but not on GSM8K. Both can be true: the paper tested the official chat models of five families, and the tutorial tested one community merge. Treat a claim of no quality loss as something to check on the model in front of you, not a property of the technique.

When the difference matters for security

Refusals are not a control on open weights. The papers cited here removed them from open-weight models for under $5 of compute (Arditi) and a budget under $200 (Lermen). Arditi and colleagues conclude that current safety mechanisms “can easily be circumvented and are insufficient to prevent the misuse of open-source LLMs”. If an application depends on the model saying no, anyone holding an open weight model’s weights can remove that. Limits that must hold belong outside the model, in AI guardrails and in what the application lets the model do.

The access needed differs. Abliteration needs the weights. Fine-tuning does not: Qi and colleagues removed GPT-3.5 Turbo’s refusals through a hosted API, and found that fine-tuning on ordinary benign data also weakened safety, to a lesser extent. A provider that offers fine-tuning has the second route open whether or not it releases weights.

The name on a model is the uploader’s word. On 28 September 2026, Hugging Face search returned 8,217 models matching “abliterated” and 7,275 matching “uncensored”. Whoever uploads a model chooses its name, and a model without either word can still have been fine-tuned or edited. Test refusal behaviour yourself rather than reading it from the name.

Uncensored is sometimes the design. Hartford’s argument was for “composable alignment”: start from a model with no built-in refusals and add the policy that fits your use. On that view the deployer owns the policy. For a security team that is the right reading of any uncensored model: the refusals are now your job.

What this does not settle

The paper’s own limits section says the single-direction result “may not generalize to untested models”, including larger and proprietary ones. It shows the method worked on the 13 models tested, not that it works on every model. The sources also do not separate uncensoring from jailbreaking. Arditi and colleagues call their weight edit “a novel white-box jailbreak method”, white-box meaning it needs the model’s internals, and Qi and colleagues describe their fine-tuning of GPT-3.5 Turbo as jailbreaking it.

Questions and answers

Is an abliterated model the same as an uncensored model?

An abliterated model is one kind of uncensored model. Uncensored describes the result, a model that does not refuse; abliteration is one way to get it, by editing out the direction in the model's activations that carries refusal instead of retraining.

Does abliteration make a model worse?

The evidence is mixed. Arditi et al. found most edited models close to the originals on MMLU, ARC and GSM8K, with Qwen 7B and Yi 34B the exceptions, while TruthfulQA scores consistently dropped. Labonne measured a drop on every benchmark for one 8B model and recovered most of it with further DPO training.

Sources

  1. Uncensored ModelsEric Hartford, May 2023
  2. Refusal in Language Models Is Mediated by a Single DirectionArditi et al., NeurIPS 2024, Jun 2024
  3. Uncensor any LLM with abliterationMaxime Labonne, Hugging Face, 13 Jun 2024
  4. LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70BLermen, Rogers-Smith and Ladish, ICLR 2024 workshop, Oct 2023
  5. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Qi et al., Oct 2023
  6. Hugging Face model searchHugging Face, 28 Sep 2026

Terms in this guide

In the news