Skip to content

All pages

Heretic

Fully automatic censorship removal for language models

Other21d ago
Category
Other
Website
github.com
Language
English
Listed
21d ago
Forks
3,576

p-e-w/heretic ↗, Python, AGPL-3.0, Last commit 1d ago

Heretic strips safety alignment out of transformer language models without any post-training. It combines directional ablation, the technique usually called abliteration, with a TPE parameter optimizer built on Optuna, and searches for settings that minimise refusals and KL divergence from the original weights at the same time. Because both objectives move together, the result keeps most of the base model's capability instead of trading it away, and the whole search runs without human tuning. Python, AGPL-3.0.

Like this product?