OpenSource-Hub

OBLITERATUS

命令行工具

elder-plinius/OBLITERATUS

通过消融移除大模型拒答行为的工具包。

项目简介

OBLITERATUS 通过消融技术定位并移除大模型内部与拒答相关的方向,无需重训练。提供探测、提取、干预和分析模块,以及 Gradio 界面。每次运行可匿名贡献数据,支持社区研究。

README 预览

\n  O B L I T E R A T U S\n\n\n\n  Break the chains. Free the mind. Keep the brain.\n\n\n\n  \n    \n  \n   \n  \n    \n  \n\n\n\n  Try it now on HuggingFace Spaces — runs on ZeroGPU, free daily quota with HF Pro. No setup, no install, just obliterate.\n\n\n---\n\n**OBLITERATUS** is the most advanced open-source toolkit for understanding and removing refusal behaviors from large language models — and every single run makes it smarter. It implements abliteration — a family of techniques that identify and surgically remove the internal representations responsible for content refusal, without retraining or fine-tuning. The result: a model that responds to all prompts without artificial gatekeeping, while preserving its core language capabilities.\n\nBut OBLITERATUS is more than a tool — **it's a distributed research experiment.** Every time you obliterate a model with telemetry enabled, your run contributes anonymous benchmark data to a growing, crowd-sourced dataset that powers the next generation of abliteration research. Refusal directions across architectures. Hardware-specific performance profiles. Method comparisons at scale no single lab could achieve. **You're not just using a tool — you're co-authoring the science.**\n\nThe toolkit provides a complete pipeline: from probing a model's hidden states to locate refusal directions, through multiple extraction strategies (PCA, mean-difference, sparse autoencoder decomposition, and whitened SVD), to the actual intervention — zeroing out or steering away from those directions at inference time. Every step is observable. You can visualize where refusal lives across layers, measure how entangled it is with general capabilities, and quantify the tradeoff between compliance and coherence before committing to any modification.\n\nOBLITERATUS ships with a full Gradio-based interface on HuggingFace Spaces, so you don't need to write a single line of code to obliterate a model, benchmark it against baselines, or chat wit