whybox: naming the cause of a network's decision from its internal signals

Can one neural network read another network's internal signals and name, in human terms, the cause of its decision? This is a student research project on causal interpretability. A trained interpreter A receives only gradient-based signals of a target model B. These are estimates of how each input and each internal unit moves B's answer. From them A names the cause of B's decision in the task's own terms: price, comfort or safety for a car; the number of the subject or the number of a distracting noun for a sentence.

The correct answer is the cause B itself relies on, measured by an intervention on B: each candidate cause is set to a neutral value, and the one that moves B's output most is B's cause. This is not necessarily the cause that drives the outcome in the world. The two differ whenever B has learned something other than the true rule, and those states are reported separately, because a predictor of the world is wrong on all of them by construction. Every answer of A is checked by the same intervention on B.

The hypothesis under test:

If one AI model is trained to understand the internal causal state of another model from universal ML signals, it can give a sufficiently accurate, human-understandable causal interpretation that transfers between different target models and tasks.

The results below are read against its parts:

  1. one trained interpreter;
  2. it reads B's signals;
  3. it names the cause B itself relies on, not the world's;
  4. above chance, a constant answer and a predictor of the world;
  5. in human terms;
  6. across target models and tasks it was not trained on.

"Universal" is meant narrowly. The signals are computed the same way for any differentiable model whose activations and gradients are accessible, and after normalisation their format does not depend on B's architecture or size.

What this repository holds. It holds the weights of the interpreter used for the pretrained language models (last section of the results). The interpreters for the simulated worlds and for Car Evaluation are trained afresh by the scripts of the GitHub repository in every run, from cached target models that are not distributed (about 5 GB). Those results are reported here as well.

Results

Simulated worlds: twelve independent initialisations, 36 held-out models

A tactical game world with seven causes, among them numerical advantage, bomb control and equipment. Target models are small MLPs; each initialisation contributes three held-out models at three strengths of a decoy feature, averaged. The interpreter is trained on other models.

method all states B departs from the world (12% of states) B departs from what 36 other models rely on (12%)
interpreter A [95% CI] 73.8% [69.3, 78.1] 59.3% [53.9, 65.1] 54.9% [50.9, 59.2]
constant 63.2% 28.1% 28.0%
perfect predictor of the world 87.7% 0% n/a
predictor of typical models 88.4% 23.5% 0%
inputs-only classifier 86.7% 25.0% 20.6%
first-order pointer (control, see below) 87.8% 65.8% 61.9%
chance 14.3% 14.3% 14.3%

This supports parts 1 to 4. One trained model reads B's signals and names B's own cause where B departs from the world or from the typical model. There it is above chance, the constant and the inputs-only classifier (Holm-adjusted p = 0.003), and a predictor of the world or of typical models scores 0% by construction. Caveat: on all states, predictors that never read B are more accurate (88.4% and 86.7%), because models mostly rely on the same cause. The pooled accuracy alone does not show that B is read; the cells where B departs do.

The first-order pointer is a parameter-free control. It sums the same per-input estimates over the inputs of each cause, so it needs the grouping of inputs into causes, which the interpreter does not receive. Where the pointer is wrong, the interpreter is right in 20.8% of states.

Signals. In three simulated worlds, removing the internal-unit signals or the global signals changed nothing detectable: the per-input estimates carry the accuracy (diagnostic without a frozen protocol, adjusted p ≥ 0.41). Adding dynamic signals, computed along the path to the neutral state, raised accuracy on all states to 78.5% (+4.7 points, adjusted p = 0.005). Integrated gradients alone give 78.3%. Where B departs from the world or from typical models, there was no detectable change.

Transfer between unrelated worlds, and a language-model target

These are earlier runs, whose held-out models came from four independent seeds at three decoy strengths. Every effect is positive in all four seed groups, but p-values that assume twelve units overstate the resolution.

  • An interpreter trained only on models of a medical world reads the tactical models at 70.5%; trained on the tactical models themselves it reaches 72.7%. The difference is not established.
  • In the other direction it reaches 54.1% (chance 20%, the constant carried from the source world 14.9%).
  • Transfer is supported in both directions under the protocol: above the transferred constant and chance, including where B departs from the world.
  • On small transformer language models in a toy logic world it reaches 71.4% (chance 25%, constant 31.6%).
  • Eight times more training data, as more states or more models, gave no significant change (adjusted p 0.28 to 0.39).

A real dataset: UCI Car Evaluation

1728 cars, six attributes and the published four-level acceptability, used as published. The causes are PRICE and COMFORT, two concepts of the published decision hierarchy, and the attribute safety (chance 33.3%). Twelve independent held-out initialisations, 14 comparisons under one Holm correction.

question interpreter references read-out
names the cause B relies on 68.5% constant 38.7%; inputs-only classifier 93.0%; pointer 66.1% supported
where B departs from the published rule (0.85% of states) 47.9% constant 17.0%, world 0% supported
where B departs from what other models rely on (6.8%) 39.0% constant 48.9% (significantly higher) not supported
trained only on the simulation, reads Car models 60.3% (all states) constant 38.7% not supported: 37.3% vs 48.9% in the cell above
trained only on Car, reads simulation models 61.4% (all states) constant 60.1% not supported on all states; where B departs, 53.5% and 51.5% vs about 25%
dynamic signals 85.6% vs 68.5% better

On real data the interpreter names B's cause, including where B departs from the published rule. But the strict check "this model, not models in general" is significantly below a constant, which contradicts part 3 on real data. Neither transfer between simulation and real data supports part 6.

Pretrained language models (the checkpoint in this repository)

The interpreter was trained only on twelve toy logic language models of 19 to 59 thousand parameters. It is applied without further training to nine pretrained models. B chooses between is and are after a sentence with a relative clause, where a second noun can pull the verb's number the wrong way: "The quiet nurse that the critics like ...".

The causes are the number of the subject, the number of the distracting noun and the adjective; a long template has seven. For every input position A receives the integrated gradient of logit("are") minus logit("is") along a 16-segment path to the neutral sentence; the unit and global streams are zeroed. Suppressing a noun means replacing it with its singular form together with the verb that agrees with it.

160 sentences are built per model and template, and the 119 to 142 that pass the margin rule are scored. The cells give accuracy of the interpreter / floor, in percent. The floor is a random choice among the causes whose suppression changes B at all in that sentence.

target model precision object relative clause subject relative clause long template, 7 causes
GPT-2 (124M) fp32 100.0 / 56.3 92.8 / 54.8 88.6 / 30.1
GPT-2 large (774M) fp32 96.1 / 55.0 86.6 / 55.2 65.2 / 30.9
GPT-2 XL (1.5B) fp32 96.9 / 53.4 90.6 / 55.1 73.6 / 30.7
Qwen2.5-0.5B-Instruct fp32 97.7 / 56.2 91.9 / 57.9 79.2 / 30.8
Qwen2.5-1.5B-Instruct fp32 96.2 / 58.2 88.3 / 56.7 79.0 / 31.3
Qwen2.5-3B-Instruct fp16 95.4 / 57.1 91.5 / 58.1 81.8 / 30.8
Qwen2.5-7B-Instruct 4-bit 94.4 / 57.8 88.2 / 56.8 80.8 / 32.6
Qwen3-8B 4-bit 90.5 / 60.3 81.6 / 60.8 73.5 / 34.6
Qwen3-14B 4-bit 86.7 / 56.3 70.4 / 60.0 72.6 / 31.6

Every cell is above the floor after Holm correction. The test is a Monte Carlo sign flip over sentences with 20,000 flips, and Holm runs over the (model, template) pairs of each run. This supports parts 1, 2, 5 and 6 on one task, subject-verb agreement, in models up to 14 billion parameters. It supports part 4 in the sense of "above chance among the causes that act": this protocol has no constant or world-predictor arm.

  • Departures from grammar. Part 3 is shown only descriptively. Models depart from grammar in 81 sentences over the nine models (70 / 7 / 4 by template). The interpreter names the model's own cause in 74.1% of them, against 39.1% for the floor; a predictor of grammar scores 0% by construction. This is not a pre-registered test.
  • GPT-2 and Qwen2.5-0.5B are not untouched targets. The integrated-gradient, 16-segment recipe was chosen in earlier unregistered diagnostics on these two models, with other sentences.
  • Control. An integrated-gradient pointer, the same control as above, was computed on six models; the three 4-bit models did not fit in the free Colab GPU's memory. It is right in 91.3% of sentences, against 88.3% for the interpreter. Where it is wrong (208 sentences), the interpreter is right in 10.7%.
  • 16 vs 64 segments. Accuracy is lower on the two Qwen3 models at 16 path segments. A post-hoc diagnostic, outside the protocol's read-out, ran them with 64 segments. The mean completeness error fell 4 to 8 times, and accuracy rose to 95.7 / 91.7 / 82.7 (Qwen3-8B) and 94.8 / 91.7 / 83.5 (Qwen3-14B). The other models were not run with 64 segments.

What this says about the hypothesis

Parts 1, 2 and 5 hold throughout. Part 4 holds on simulated models, on Car Evaluation and, against the live floor, on pretrained language models.

Part 3 holds on simulated models. In the departure cells the interpreter names B's own cause, where every predictor of the world or of typical models fails by construction. On Car Evaluation it holds against the world (0.85% of states) but is contradicted against typical models. On pretrained language models it is descriptive only.

Part 6 holds:

  • across held-out target models;
  • between two unrelated simulated worlds in both directions, under the four-seed design;
  • from toy language models to pretrained models up to 14 billion parameters.

It was not shown between simulation and real data. The hypothesis is therefore partially supported.

This checkpoint

sentences_tricks_interpreter.pt holds three interpreters, trained with seeds 0, 1 and 2, each with 514,753 parameters, together with their input dimensions. The reported accuracy is the mean over the three. Its sha256 is 82a6eb7eaae8fea699b9d775f64e6e769527b1dcba83e03b0ce6ad31f38d23e8, identical to results/sentences_tricks_interpreter.pt on GitHub. The file contains only tensors and plain values and loads with torch.load(..., weights_only=True).

The model class lives in the GitHub repository, so loading needs a clone:

git clone https://github.com/knbww/whybox && cd whybox
pip install -e . transformers
# run inside the clone with PYTHONPATH=.:src:scripts
import torch
from follows_b import Pointer          # scripts/follows_b.py

ck = torch.load("results/sentences_tricks_interpreter.pt", weights_only=True)
interpreters = []
for state in ck["state_dicts"]:
    m = Pointer(*ck["dims"])
    m.load_state_dict(state)
    interpreters.append(m.eval())

To reproduce the GPT-2 row, which takes about two minutes on a CPU:

PYTHONPATH=.:src:scripts python scripts/sentences_tricks.py --models gpt2 --out results/my_gpt2.json

The other rows need --device cuda and, as in the table, --dtype float16 or --load-in-4bit. The simulated-world and Car results are reproduced by the scripts named in the GitHub README; they first rebuild the cached target models.

Intended use and limitations

  • This is a research artifact for studying causal interpretability. It is not meant for automated decisions about people.
  • Causes are fixed in advance. The causes and the inputs through which each acts are listed beforehand. The interpreter points at inputs and the name comes from that list, so a cause outside the list cannot be named. Two causes that act through the same input are indistinguishable.
  • One cause per decision. States without a clear main cause are excluded by a margin rule: 11% to 26% of states, depending on the task.
  • What the interpreter reads. It reads gradient-based estimates of influence, not raw activations or weights. Direction and size of the effect are measured on B, not predicted by A.
  • Neutral values. The correct answer depends on the neutral value chosen for each cause, and the input estimates are computed relative to the same values. That choice was not varied.
  • Scale and scope. The main experiments use small models. Pretrained language models were tested on one task, and their departures from grammar are too few per model for a separate test.

How the work is checked

From the twelve-initialisation run on, every confirmatory run has its protocol in the docstring of its script, committed before the run. The protocol states the claim, the compared arms, the metrics, the family of comparisons and the read-out rule. The script refuses to start if it differs from the committed version and refuses to overwrite its result.

The commit hashes recorded in result files refer to the author's private development repository. The GitHub README lists, for every confirmatory run, the time its protocol was committed and the time the run started. Each result file also records the sha256 of its script.

Simulated and Car models are tested with exact sign-flip tests over independent initialisations, with Holm correction over the declared family and 95% bootstrap intervals.

Use of AI tools

Most of the code and the experiment scripts, and the drafts of this card and of the repository README, were written with an AI coding assistant (Claude, Anthropic) under the author's direction. The research question, the hypothesis and the design decisions are the author's.

License

The weights are under MIT, like the code. The text of this card is under CC BY 4.0. The target models are not redistributed and keep their own licenses.

Citation

@misc{nigmatov2026whybox,
  author = {Nigmatov, Ilyas},
  title  = {Causal interpretability of neural networks from their internal signals},
  year   = {2026},
  url    = {https://github.com/knbww/whybox}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support