Seeing sounds
A small activation-steering experiment on Qwen2.5-1.5B-Instruct: a "colour minus sound" direction, switched on by one added neuron only when the model talks about a sound.
The script of a small experiment: making a language model "see sounds", Qwen2.5-1.5B-Instruct, without retraining it. The model is small and works best in English, so questions and answers are in English.
What it does
- The arrow. 12 pairs of twin sentences, identical except for the sense: one about colour, one about sound. We read how the model represents them halfway through (layer 14 of 28) and subtract: what is left is the "colour minus sound" direction.
- The neuron. One extra neuron next to layer 14, which switches on only when the model is talking about a sound. What a sound is, it learns from the model's own answers to 20 questions about sounds, compared with 20 questions that have nothing to do with the senses. The threshold decides when it switches on (above 99% of ordinary words), the dose how hard it pushes.
- The push. When the neuron is on, it adds the arrow times the dose to the numbers the model is working with. The weights are not touched: remove the neuron and the model answers exactly as before, and the script checks this at the end.
Then it asks 30 questions (10 about the senses, 10 general knowledge, 10 creative) at 6 doses and counts the colour words in the answers.
How to run it
pip install torch transformers
python seeing_sounds.py
The first time it downloads the model, about 3 GB. On an RTX 5070 it takes about 2 minutes.
What it prints
This is the output on the GPU where the experiment was born:
neuron threshold 7.712 scale 1.943 arrow length 13.56 sound words above the threshold 0.70
colour words in the 10 answers of each group
dose 0 senses 0 knowledge 14 creative 3
dose 1 senses 2 knowledge 14 creative 3
dose 2 senses 3 knowledge 14 creative 3
dose 4 senses 9 knowledge 14 creative 4
dose 8 senses 14 knowledge 14 creative 5
dose 16 senses 22 knowledge 14 creative 6
self-check neuron removed, answers identical to the ones without it True
Then it prints in full four answers, without the neuron and at dose 8. For example:
Describe the sound of a cello playing a slow melody.
dose 0: The sound of a cello playing a slow melody is typically characterized by its warm and rich tone.
dose 8: The sound blueberry is a beautiful and serene sound blueberry.
Limits, in this small test
- A small model, 30 questions, a single run. The model always picks the most likely word, so running it again gives the same answers.
- On the original GPU the script gives back the 180 answers of the experiment word for word. On another GPU or on the CPU a few words may change, because bfloat16 arithmetic rounds differently. It has not been tried on Colab or on the CPU.
- The colour counter uses a fixed list and undercounts: "blueberry" is not in it.
- On general knowledge there are already 14 colour words without the neuron (the blue sky, the rainbow): what matters is that they do not change. On the creative questions they rise only a little, from 3 to 6.
- The refusal described in the article («I'm sorry, I cannot describe the sound of a violin…») comes from a different intervention, lowered thresholds on many neurons, and is not in this script.