Research · Brainscope · Video walkthrough
Inside Gemma 3 12B.
Watch an answer become readable. Then test which states can bring it back.
Three minutes inside the model.
Open full-size video ↗Read the text alternative ↓All research
The measured result
Change the subject. Restore a state.
We ask: “In which city is the Eiffel Tower? Answer with only the city name.” Replacing Eiffel with Tokyo changes one token in the same 24-token template. The preferred answer flips from Paris to Tokyo.
Each intervention puts one clean internal state into the counterfactual run. Restoring the residual state at layer 38, at the final prompt position, raises the probability of Paris to 0.97069.
The scan covers 48 layers, two prompt positions and three hook types: residual, MLP output and attention output. All 192 component-output patches round to zero recovery at the saved precision. Restoring every clean position at layer 0 recovers the clean probability with a reported error of 0.0.
About the recording
Inspect, then intervene.
The walkthrough shows Gemma 3 12B instruction-tuned running locally in Brainscope. It follows layer readouts, inspects 3,840 residual channels and compares controlled restoration patches. The film uses animated banners and has no audio track.
This is the completed recording, with edited pacing and focused crops. Its text alternative below repeats the on-screen explanations. The 1× raster label refers to the source interface before video scaling.
Text alternative
The walkthrough, in words.
The silent film’s on-screen explanations, in sequence. The result and its limitations are also written above.
Read all 18 scenes with timestamps
- Inside Gemma 3 12B
A Brainscope walkthrough: inspect the states, then test the answer. 12B · gemma 3 · it
- Where does Paris become readable?
Follow the answer across the model’s text decoder. 48 · layers
- Then test what changes the answer.
Layer readouts → residual channels → controlled restoration. 3 · views into the model
- Start with a familiar fact.
“The Eiffel Tower is located in the city of…” T = 0 · deterministic generation
- Gemma answers: Paris.
Real local execution. The recorded answer is Paris, France. Paris · observed output
- The answer is only the beginning.
Signed hooks collect the states in a replay after generation. 48 · layer readouts
- Follow the Paris column through depth.
Each row projects one layer’s residual state into vocabulary. LENS · a diagnostic readout
- Readable does not mean causal.
Intermediate predictions help us choose a test. ? · a hypothesis to test
- Pin one token. Inspect its evidence.
Paris: output probability, alternatives, and residual norms. 50429 · paris token id
- See all 3,840 residual channels.
Color shows deviation from each channel’s mean over the sequence. 3,840 · channels per layer
- Zoom to one pixel per channel.
Scale and saturation stay visible. Float32 preserves the large values. 1× · exact channel scale
- Change one token: Eiffel → Tokyo.
Same 24-token template. Different subject. Ask for only the city. 1 / 24 · prompt tokens changed
- The preferred city flips to Tokyo.
p(Paris): 1.000 → 0.000, at the saved precision. 1 → 0 · paris probability
- Restore one clean state at a time.
48 layers × 2 positions × 3 hook types. Other positions are unmeasured. 288 · measured patches
- A residual patch brings Paris back.
Layer 38, final prompt position: p(Paris) rises to 0.97069. 97.1% · paris after restoration
- The component patches do not recover it.
MLP and attention outputs: all 192 tested cells round to zero. 192 · component patches tested
- The full-restore control passes.
Restore all clean positions at layer 0. The evidence ledger verifies intact. 0.0 · reported probability error
- One example. A testable result.
Selected residual patches are sufficient for this prompt pair. Evidence retained. 33 · tests passed · 1 skipped
