Algoverse research | NeurIPS 2026 workshop | Lead author
Interpretability on vision-language models
More companies are putting AI models into real workflows, and that depends on being able to check why a model gave an answer. This paper tests one of the standard methods for doing that, and finds it can report a feature that is not there.
The field
The research area is called mechanistic interpretability: looking inside a model to find which parts of its internal state produce an answer. If you can find the part that encodes "red", you can check what the model is using instead of guessing.
The standard method
A common way to find one of these parts goes like this. Show the model many pairs of inputs that differ in one thing, like color. Average the difference in its internal state. That gives a "direction". Then add the direction back in and see if the answer changes. If it does, and a random direction of the same size does not, people report that they found the feature.
We showed a vision-language model two pictures and asked a question like "Which picture has the purple square? A or B." We fitted a color direction the standard way and added it in.
If the direction really encoded color, the model should mix up the two pictures and get almost everything wrong. Instead, it answered "A" every time, whatever it saw, which scores exactly 50% on a balanced test.
It fixed the model's answer in 43 of 96 test conditions; a random direction of the same size did it in none. The result also held on a second set of pictures. Under the usual standard, it would be reported as a color feature.
The direction was mostly the model's preference for answer letter A over B. Its cosine with the answer-letter direction is 0.99, close to identical. Added to a question about shape instead of color, it did the same thing.
The averaging explains it. Every pair shares the same pull toward one letter, but the picture details differ from pair to pair, so they cancel out. What survives the average is the letter.
We also tested the controls. A "sham" that adds the same vector to both pictures should do nothing. It changed 203 answers, all toward one letter. The random-direction control barely moves the model at all, so it has nothing to compare with and cannot catch this.
Why it matters
An interpretability result can pass the standard checks and still be wrong about what it found, so the controls need testing too. The code we will release has a test for each control that confirms the control can fail.