← Patrick Taylor
Algoverse research  |  NeurIPS 2026 workshop  |  Lead author

Interpretability on vision-language models

More companies are putting AI models into real workflows, and that depends on being able to check why a model gave an answer. This paper tests one of the standard methods for doing that, and finds it can report a feature that is not there.

The field

The research area is called mechanistic interpretability: looking inside a model to find which parts of its internal state produce an answer. If you can find the part that encodes "red", you can check what the model is using instead of guessing.

The standard method

A common way to find one of these parts goes like this. Show the model many pairs of inputs that differ in one thing, like color. Average the difference in its internal state. That gives a "direction". Then add the direction back in and see if the answer changes. If it does, and a random direction of the same size does not, people report that they found the feature.

Our test

The fitted "color" direction makes the model answer "A" every time
Share of answers that were "A", same question with three different additions inside the model
Normal, nothing added0.50
Add the fitted direction1.00
Add a random direction, same size0.50
Note: The test set is balanced, so answering "A" every time scores 0.50, chance. A direction that really swapped the two pictures would push accuracy toward 0. Accuracy was 1.00, 0.50 and 1.00 in the three runs.
Source: Taylor et al., NeurIPS 2026 InterpScience workshop. Qwen2.5-VL-3B-Instruct unless noted.
Patrick Taylor
Click to enlarge

We showed a vision-language model two pictures and asked a question like "Which picture has the purple square? A or B." We fitted a color direction the standard way and added it in.

If the direction really encoded color, the model should mix up the two pictures and get almost everything wrong. Instead, it answered "A" every time, whatever it saw, which scores exactly 50% on a balanced test.

It passes the usual check

The fitted direction passes the standard check; a random one never does
Test conditions, out of 96, where the model gave one fixed answer
Fitted direction43 of 96
Random direction, same size0 of 96
Note: Six cells at 16 non-zero doses. Beating a same-size random direction is the usual evidence that a direction is a real feature. The onset dose also held on a second, separate set of 60 image pairs.
Source: Taylor et al., NeurIPS 2026 InterpScience workshop. Qwen2.5-VL-3B-Instruct unless noted.
Patrick Taylor
Click to enlarge

It fixed the model's answer in 43 of 96 test conditions; a random direction of the same size did it in none. The result also held on a second set of pictures. Under the usual standard, it would be reported as a color feature.

What it really found

The directions point almost exactly along the answer-letter direction
Cosine similarity with the answer-letter direction at layer 29 (1.00 means the same direction)
Color direction0.99
Object direction0.99
Note: The color direction, applied unchanged to a question about shape, gave the same results (r = 0.9999 on prediction rate). A direction that only knows the answer letter would do that.
Source: Taylor et al., NeurIPS 2026 InterpScience workshop. Qwen2.5-VL-3B-Instruct unless noted.
Patrick Taylor
Click to enlarge

The direction was mostly the model's preference for answer letter A over B. Its cosine with the answer-letter direction is 0.99, close to identical. Added to a question about shape instead of color, it did the same thing.

The averaging explains it. Every pair shares the same pull toward one letter, but the picture details differ from pair to pair, so they cancel out. What survives the average is the letter.

The controls can miss it

The "do nothing" control changes answers too, and the random control cannot see it
Answer changes over 864 rows per arm, two ways of labelling the pictures
Sham, normal labels203
Sham, swapped labels49
Random, normal labels0
Random, swapped labels5
Note: The sham adds the same vector to both pictures, so it swaps nothing and is usually read as a null. Its answer changes land on one letter at rates of 1.00 and 0.88. The random control barely moves the model, so it has nothing to compare against.
Source: Taylor et al., NeurIPS 2026 InterpScience workshop. Qwen2-VL-7B-Instruct.
Patrick Taylor
Click to enlarge

We also tested the controls. A "sham" that adds the same vector to both pictures should do nothing. It changed 203 answers, all toward one letter. The random-direction control barely moves the model at all, so it has nothing to compare with and cannot catch this.

Why it matters

An interpretability result can pass the standard checks and still be wrong about what it found, so the controls need testing too. The code we will release has a test for each control that confirms the control can fail.

Authors
Patrick Taylor (lead), Preeyam Shah, David Wu, Joel Gurivireddy, Margaret Capetz
Venue
InterpScience workshop, NeurIPS 2026
Status
Accepted, September 2026
Code
Not public yet
Back to everything else