heretic
View on GitHubFully automatic censorship removal for language models
CLI tool that fully automatically strips refusal behavior ("safety alignment") from transformer LLMs using directional ablation plus Optuna-based parameter search, minimizing refusals while preserving the original model's capabilities.
Use Cases
Automatically remove refusal behavior from open-weight LLMsProduce abliterated/uncensored model variants without post-trainingMinimize KL divergence while suppressing refusals via TPE parameter searchDecensor dense, MoE, and multimodal transformer modelsQuantize models with bitsandbytes to fit limited VRAMUpload decensored models to Hugging FaceRun MMLU/GSM8K benchmarks on modified modelsInterpretability research on residual vectors and refusal directionsChat-test decensored models interactivelyReproduce manual expert abliteration quality automatically
Built With
- Language
- Python
- Frameworks
- PyTorch · Hugging Face Transformers · Optuna · PEFT · bitsandbytes · accelerate · lm-eval · Hugging Face Hub · PaCMAP · NumPy · scikit-learn · Pydantic
Tags
abliteration · llm · transformer · model-editing · decensoring · safety-alignment · interpretability · quantization · gpu · cli · fine-tuning-alternative · optuna · peft · huggingface · research