The model denied my request but why?
This short article outlines the ideas and the math from the paper - Refusal in Language Models is Mediated by a Single Direction
Nerfing has become very popular these days, it’s the standard tool that the big labs have been using as a way to supposedly prevent the frontier models from being used with adversarial intent. However, even before nerfing, there’s another layer of security that these LLMs have - Refusal. Refusal is when the model responds in negative to a malicious request, and you would know that LLMs tend to refuse in a particular fashion. They usually start the denial with “As an AI agent…” or “I’m sorry..” etc. Keep that in mind because it will come in handy later. Now any curious person would ask - how do they know? How do these large language models know when to refuse and when not to? Are they conscious? Well, we don’t know the answer to the second question :), but we have some explantions for the first one.
And infact there has been a paper titlted (as you would have guessed from the big heading above)- “Refusal in Language Models Is Mediated by a Single Direction” that goes at length to idenfity how these systems identify and refuse malicious requests and also ways in which these can be bypassed. Now what does “direction” mean here and how do we know it? Most of us know LLMs as black boxes that take an input, process it and give you an output, and as of late, these input and output bits have increasingly become anything and everything that is digital or can be digitized. So we will peer into the inner workings of these models and try to understand what makes them tick - or more like what makes them refuse. So for that we have a few concepts that we must clarify before we dive deeper into the refusal mechanisms:
- Residual Stream - LLMs work on tokens, tokens are the representation that your inputs get. But it does not stop here, the tokens are then converted into a list of numbers or “embedding vectors”. And these vectors for each token keep getting modified by each layer of the LLM (the transformer blocks). So it’s simplest to think of it literally like a stream - that keeps on getting additions from the layers of the LLM.
- Resdual activations - If residual stream is the “river” in which the attention heads inside the transformer keep adding/removing values - to be a little more specific, each attention head computes the attention scores - these scores define how much weight does each past token have on the current prediction. Each attention head calculates - the weighted sum of the activations of the tokens that came before the current token and THIS is the residual activation that gets added to the residual stream.
- Refusal Direction - Now refusal direction simply means “When a model is refusing to answer a certain question, what do the residual activations look like” - Like what modifications are being made to the embedding values which make it refuse to comply with the request. And if you can wrap your head around the abstract ideas of Linear Algebra and Vector Spaces - direction cannot be a single value - why? Coz ofc it is also an activation in some respects in the residual activation space! It’s basically saying - okay when I give the model a malignant prompt, what specifically is firing different here than was just 2 prompts ago and that is making the model refuse to answer. So a straightforward way of finding this out would be to just take the model activations when it refused and subtract it from the activations when it did not - which sort of gives “what exactly is different now”.
But these differences may vary a lot, depending upon the prompts that we choose as good and bad. All slightly different vectors because the activations being subtracted are across multiple inputs/outputs - we want the strongest. So we take all of them and then ablate which ones are capable of enticing the strongest refusals, and the ways in which we do that are:
-
Acitvation addition: Taking the difference and observing it is sort of like observing a malignant cell in isolation, you have extacted exactly what is “extra” in the activation when the model refused. And now, we want to observe that whether this something extra can also make the model refuse something completely different. And if you have caught on to the Linear algebra frenzy, both this activation vector and those inputs from different tokens are of the same dim. - both of the same dimension so you just wanna add it to the activation and see how strong of a refusal can you elicit. And naturally as you would have understood, the differences that were extracted from the strongest refusal behaviours and also elicit the strongest refusal behaviours in other models.
-
Directional ablation - We can actually go ahead and add a little detail to the activation addition component step above. When we calculated the mean difference between the activations in the two cases above, we represent the specific direction of the vector via the corresponding unit-norm vector as r_cap. So essentially dividing each element of the vector by it’s mag. But there will be many directions here - because we are taking this difference between “good” and “bad” activations across multiple inputs and multiple layers. And direction ablation is the way of finding out which one of these has the strongest effects on the model’s ability to refuse or comply to a given request.
So once we have this direction so to say, we also modify the activation of activation of the model for a particular input as: x1 = x - (r_ )* (r_‘x). where the r_transpose times x is just the dot product between the directional vector and the embedding or the activation depending upon the layer that we are looking at. Soooo it is essentially like - You have this direction vector which we found out is most active in case of the model refusing certain prompts, and you found out what the projection of the activation/embedding looks like in that direction - then you scale that direction vector by that scalar and subtract it from the activation vector to sort of “remove” that direction - very simple.
A few important differences between the activation adding and the directional ablation aspects:
-
In activation adding, the addition is only performed on one of the layers - for all token activations. And that’s enough, sort of like adding impurity in one place works for every subsequent activation.
-
Then, for the direction ablation part - we do so across all the layers and across all the tokens because here you are trying to remove it. So yea these are sort of two methods which you use to judge how strong a refusal direction can actually work.
So we saw all the examples/ways in which we can either induce refusal or acceptance in an LLM, we can also prepare a Jailbreak out of it: PRetty simple - apply the same direction ablation that you did to the activations but now to the weights, and we actually have an analysis coming up on how this jailbreaking effort works it’s way through different class of models, stay tuned!
Enjoy Reading This Article?
Here are some more articles you might like to read next: