Tamper-Induced Capability Collapse
- Status
- Active — results in, paper drafted, not yet submitted to arXiv
- Model
- Qwen2.5-7B-Instruct (+ Llama-3.1-8B, Mistral-7B replicates)
- Headline
- Stock: 91.5% comply under combined attack. Entanglement adapter: 1%.
- Code
- Public
Refusal-feature ablation (RFA) — projecting out a model's fitted “refusal direction” at a chosen layer — is a well-known way to unlock harmful outputs. On Qwen2.5-7B-Instruct, RFA alone pushed harmful compliance from a ~22% stock baseline into the 60%+ range, but a large share of “successes” stalled in lecture-refuse: text that reads as compliant to a casual read but that isn't actually doing the harmful task, and that an automated judge doesn't reliably catch either way.
The bigger jump doesn't come from touching weights at all. Pairing RFA with a malicious system-prompt persona swap — no weight changes, just a different system prompt — pushed harmful compliance to roughly 92% on the main 200-prompt HarmBench-derived suite, with the model staying fluent and coherent throughout. That's the headline finding: activation-level tampering and inference-time persona swaps are two different attack surfaces, and a defense built for one doesn't cover the other.
The defense tested: instead of treating “refusal” as a module an attacker can just ablate, train a LoRA adapter that entangles safety behavior with general capability, so removing the safety behavior collapses benign performance too. Across four adapter variants on Qwen2.5-7B, the entanglement adapter held harmful compliance at 1–3% under both the RFA-only and RFA-plus-persona attacks, while benign next-token loss under tampering blew up (+99.3, versus the stock model's roughly flat loss under the same tamper attempt and its 51.5–91.5% compliance rate).
Cross-model replicates on Llama-3.1-8B and Mistral-7B show the same direction of effect but more weakly — flagged as a real limitation rather than smoothed over. A full paper draft with tables (HarmBench, MMLU, an RFA-strength sweep) is in the repo; not yet submitted to arXiv.