Sekoskeys 1 · 2 · 3

Research · Brainscope · Video walkthrough

Inside Gemma 3 12B.

Watch an answer become readable. Then test which states can bring it back.

· Gemma 3 12B instruction-tuned · Local Apple MPS

Three minutes inside the model.

3:00 · 1080p · Silent, with animated banners. Recorded Brainscope UI; edited pacing and focused crops. Activation playback follows generation.

The measured result

Change the subject. Restore a state.

We ask: “In which city is the Eiffel Tower? Answer with only the city name.” Replacing Eiffel with Tokyo changes one token in the same 24-token template. The preferred answer flips from Paris to Tokyo.

Each intervention puts one clean internal state into the counterfactual run. Restoring the residual state at layer 38, at the final prompt position, raises the probability of Paris to 0.97069.

48Text decoder layers
288Measured interventions
97.1%Paris after the layer 38 residual patch

The scan covers 48 layers, two prompt positions and three hook types: residual, MLP output and attention output. All 192 component-output patches round to zero recovery at the saved precision. Restoring every clean position at layer 0 recovers the clean probability with a reported error of 0.0.

What this supports. Selected residual patches are sufficient for this prompt pair. This does not establish necessity, a unique fact circuit, or generalization. The other 22 prompt positions were unmeasured. A null component patch does not show that attention or MLPs are unimportant. Values shown as zero or one reflect saved numerical precision.

About the recording

Inspect, then intervene.

The walkthrough shows Gemma 3 12B instruction-tuned running locally in Brainscope. It follows layer readouts, inspects 3,840 residual channels and compares controlled restoration patches. The film uses animated banners and has no audio track.

This is the completed recording, with edited pacing and focused crops. Its text alternative below repeats the on-screen explanations. The 1× raster label refers to the source interface before video scaling.

Text alternative

The walkthrough, in words.

The silent film’s on-screen explanations, in sequence. The result and its limitations are also written above.

Read all 18 scenes with timestamps
  1. Inside Gemma 3 12B

    A Brainscope walkthrough: inspect the states, then test the answer. 12B · gemma 3 · it

  2. Where does Paris become readable?

    Follow the answer across the model’s text decoder. 48 · layers

  3. Then test what changes the answer.

    Layer readouts → residual channels → controlled restoration. 3 · views into the model

  4. Start with a familiar fact.

    “The Eiffel Tower is located in the city of…” T = 0 · deterministic generation

  5. Gemma answers: Paris.

    Real local execution. The recorded answer is Paris, France. Paris · observed output

  6. The answer is only the beginning.

    Signed hooks collect the states in a replay after generation. 48 · layer readouts

  7. Follow the Paris column through depth.

    Each row projects one layer’s residual state into vocabulary. LENS · a diagnostic readout

  8. Readable does not mean causal.

    Intermediate predictions help us choose a test. ? · a hypothesis to test

  9. Pin one token. Inspect its evidence.

    Paris: output probability, alternatives, and residual norms. 50429 · paris token id

  10. See all 3,840 residual channels.

    Color shows deviation from each channel’s mean over the sequence. 3,840 · channels per layer

  11. Zoom to one pixel per channel.

    Scale and saturation stay visible. Float32 preserves the large values. 1× · exact channel scale

  12. Change one token: Eiffel → Tokyo.

    Same 24-token template. Different subject. Ask for only the city. 1 / 24 · prompt tokens changed

  13. The preferred city flips to Tokyo.

    p(Paris): 1.000 → 0.000, at the saved precision. 1 → 0 · paris probability

  14. Restore one clean state at a time.

    48 layers × 2 positions × 3 hook types. Other positions are unmeasured. 288 · measured patches

  15. A residual patch brings Paris back.

    Layer 38, final prompt position: p(Paris) rises to 0.97069. 97.1% · paris after restoration

  16. The component patches do not recover it.

    MLP and attention outputs: all 192 tested cells round to zero. 192 · component patches tested

  17. The full-restore control passes.

    Restore all clean positions at layer 0. The evidence ledger verifies intact. 0.0 · reported probability error

  18. One example. A testable result.

    Selected residual patches are sufficient for this prompt pair. Evidence retained. 33 · tests passed · 1 skipped