Goodfire

William Mayner

Functioning coactivation, not Fisher, recovers published residual-MLP mechanisms; the canonical 67M language model also retains continuous coactivation over tangent-field and K-FAC

Hypothesis

The clustering step of a parameter decomposition groups rank-one subcomponents into mechanisms. This experiment kept the released decomposition losses and model recipes fixed and varied only the affinity used for that clustering. Fisher cosine was the pre-registered primary method.

Going in, we predicted three things:

  1. Primary claim. Fisher output-gradient geometry recovers the parameter-derived mechanism labels better than causal-importance coactivation.
  2. Against a parameter-aware control. The hypothesis predicted that Fisher would match or beat full K-FAC on published residual-MLP ground truth.
  3. Decision rule. The AUC ≥ 0.60 functioning-baseline gate was declared before the published sweep. Under the plan's verdict taxonomy, a functioning coactivation baseline that matches the labels while Fisher does not yields refutation.

Results on the published checkpoints

Published checkpoints, valid and functioning
20 / 20
Continuous coactivation, mean true-K ARI
0.9923
Fisher, mean true-K ARI
−0.0119
Fisher minus coactivation
−1.0042
95% CI [−1.0158, −0.9851]
Fisher win rate vs coactivation and vs K-FAC
0%

Figure 1: Mean true-K ARI for all ten clustering affinities

Mean true-K adjusted Rand index (ARI) for all ten prespecified clustering affinities, averaged over the 20 published checkpoints (15 SPD, 5 VPD). True-K ARI scores the recovered partition against the parameter-derived mechanism labels at the ground-truth cluster count. A value of 1 is a perfect match and 0 is chance. Bars are horizontal because the affinity names are long; the dark line marks ARI = 0.

The three coactivation variants take the top of the ranking, and the output-gradient affinities cluster near chance. Continuous coactivation recovers the labels almost perfectly at 0.9923, binary CI at threshold 0.1 reaches 0.9758, and binary CI at threshold 0.01 reaches 0.7610. (a) Fisher cosine sits just below chance at −0.0119, near shuffled (−0.0009), Hellinger (−0.0142), and the direction-free identity-G control (−0.0213). (b) Two other gradient-based affinities carry real signal that Fisher does not: full K-FAC reaches 0.2479 and tangent-field reaches 0.2109, with conditional Fisher intermediate at 0.0960. (c) Tangent-field's overall mean hides a strong VPD result (see Figure 3), which is why the failure is specific to Fisher cosine and related output-only geometry rather than to gradient-based affinities as a class. Fisher performing at the level of controls that carry no grouping signal is the core of the refutation.

Figure 2: Fisher minus each alternative affinity, hierarchical bootstrap

Fisher ARI minus each alternative affinity with 95% bootstrap confidence intervals

Difference in true-K ARI between Fisher and each alternative affinity, with 95% confidence intervals from a hierarchical bootstrap over the four task families and 20 checkpoints (10,000 replicates). Points left of the dark zero line mean Fisher is worse. The point estimate is the marker; the bar is the interval.

Fisher is worse than coactivation by a full unit of ARI: −1.0042 (95% CI [−1.0158, −0.9851]), an interval that does not approach zero. (a) Fisher also trails full K-FAC by −0.2598 (95% CI [−0.6269, −0.0488]), so it loses to the parameter-aware control too. (b) The only comparison Fisher wins is against direction-free identity-G, and only barely: +0.00943 (95% CI [0.00092, 0.02292]). Every interval that matters for the primary claim is on the losing side of zero.

Figure 3: True-K ARI by task family

Mean true-K ARI for continuous coactivation, full K-FAC, and Fisher, split by the four task families: one-, two-, and three-layer stochastic parameter decomposition (SPD) and two-layer variational parameter decomposition (VPD). Each family averages 5 checkpoints. Multi-layer families use the benchmark's cross-layer convention, where the same target feature ID across layer-specific input matrices counts as one shared mechanism.

The pattern holds in every family, so the refutation is not driven by one task type. Continuous coactivation is at or near 1.0 throughout: SPD one-layer 1.000, SPD two-layer 1.000, SPD three-layer 0.99903, and VPD two-layer 0.97023. Fisher is slightly negative in each: −0.00944, −0.01642, −0.01591, and −0.00584 respectively. Full K-FAC is the only affinity that varies by family; it is weak on SPD but reaches 0.800 on VPD, which is why its overall mean sits above Fisher but well below coactivation.

Figure 4: Pair-discrimination AUC for all ten affinities

Mean pair-discrimination AUC for all ten prespecified affinities across the 20 published checkpoints. Pair-discrimination AUC measures how well an affinity separates same-mechanism from different-mechanism subcomponent pairs, independent of the clustering step. A value of 0.5 is no signal. The dashed line at 0.60 is the functioning-baseline threshold used to decide whether a coactivation baseline is admissible.

The full ranking is: binary CI 0.1 0.9972, continuous CI 0.9965, binary CI 0.01 0.9824, full K-FAC 0.9417, tangent-field 0.8701, Hellinger 0.7112, conditional Fisher 0.6580, shuffled 0.5006, identity-G 0.4493, and Fisher cosine 0.4215. (a) The three coactivation variants and full K-FAC clear the gate comfortably, so the baselines Fisher is measured against are real. (b) Fisher cosine scores 0.4215, below the 0.5 no-signal point, meaning its pair scores are slightly anti-correlated with the true grouping, and identity-G is similar at 0.4493. (c) Tangent-field and Hellinger separate pairs well above chance even though their clustered true-K ARI is low, so pairwise signal does not always survive the clustering step. This is the companion to Figure 1: Fisher cosine fails at the pairwise level, not just after clustering.

Figure 5: Ten-affinity comparison matrix

Clustering affinityCheckpoint coverageMean true-K ARIVersus Fisher cosineVersus continuous coactivation
Paired hierarchical Δ [95% CI]K-sweep mean ARI ΔK-sweep win fractionPaired hierarchical Δ [95% CI]K-sweep mean ARI ΔK-sweep win fraction
Continuous coactivation (continuous CI)20/200.9923+1.0042 [0.9851, 1.0158]+0.483699.6%baseline
Binary coactivation, threshold 0.120/200.9758+0.9877 [0.9363, 1.0151]+0.482299.6%−0.0165 [−0.0495, 0.0000]−0.001430.0%
Binary coactivation, threshold 0.0120/200.7610+0.7729 [0.3088, 1.0131]+0.416897.5%−0.2313 [−0.6766, 0.0000]−0.066846.5%
Full K-FAC20/200.2479+0.2598 [0.0480, 0.6253]+0.193998.1%−0.7444 [−0.9631, −0.3591]−0.28970.9%
Tangent-field20/200.2109+0.2228 [0.0411, 0.4956]+0.186487.8%−0.7814 [−0.9718, −0.4889]−0.29720.0%
Conditional Fisher20/200.0960+0.1079 [0.0053, 0.2580]+0.094973.8%−0.8963 [−0.9898, −0.7507]−0.38882.5%
Shuffled null20/20−0.0009+0.0110 [0.0055, 0.0164]+0.009795.3%−0.9932 [−1.0015, −0.9791]−0.47390.0%
Fisher cosine20/20−0.0119baseline−1.0042 [−1.0158, −0.9852]−0.48360.0%
Hellinger20/20−0.0142−0.0023 [−0.0053, −0.0001]−0.000222.3%−1.0065 [−1.0204, −0.9859]−0.48380.0%
Identity-G (direction-free)20/20−0.0213−0.0094 [−0.0230, −0.0009]−0.00874.2%−1.0137 [−1.0378, −0.9885]−0.49230.0%

One row per prespecified clustering affinity, ordered from best to worst mean true-K ARI. Every affinity has 20/20 checkpoint coverage. The two right-hand column groups compare each affinity against Fisher cosine and against continuous coactivation. Each group reports the paired hierarchical estimate with its 95% confidence interval, the K-sweep mean ARI difference, and the K-sweep win fraction. A positive difference favors the row affinity, so a positive number against Fisher means the row beats Fisher. The K-sweep fields average checkpoint-level comparisons over each checkpoint's tested cluster-count range; every affinity shares the same within-checkpoint K support, so the comparison is like-for-like. Self-comparisons show “baseline” rather than a win fraction.

Continuous coactivation leads every measure and Fisher cosine sits near the bottom. (a) Continuous coactivation beats Fisher by +1.0042 (95% CI [0.9851, 1.0158]) and wins 99.6% of the K-sweep. Full K-FAC beats Fisher by +0.2598 (95% CI [0.0480, 0.6253], 98.1% of the sweep), tangent-field by +0.2228 (95% CI [0.0411, 0.4956], 87.8%), and conditional Fisher by +0.1079 (95% CI [0.0053, 0.2580]). So three prespecified gradient- or activation-aware affinities each beat Fisher with intervals clear of zero. (b) The shuffled null also edges Fisher by +0.0110 (95% CI [0.0055, 0.0164]), which places Fisher below a control that carries no grouping structure. Read this carefully: both sit at chance in absolute terms (Fisher's mean true-K ARI is −0.0119 and the shuffled null's is −0.0009), so the point is that Fisher does not clear a no-structure floor, not that shuffling recovers mechanisms. (c) Against continuous coactivation, every other affinity's paired interval is non-positive, so none beats it. The two binary coactivation thresholds have an upper confidence bound of exactly 0, so they are tied-or-worse than continuous coactivation rather than significantly worse. This table is the numeric backbone of Figures 1, 2, and 4.

Ten VPD output modules fail the output-write alignment gate and are excluded from the VPD family. VPD recovery therefore uses the two quality-passing input modules per checkpoint, which retain their cross-layer feature labels. All 15 SPD checkpoints exclude no modules.

Where Fisher does recover: a bounded diagnostic setting

Fisher is not useless everywhere. It recovers structure in constructed settings where causal-importance coactivation carries no signal at all. That boundary is worth stating precisely, because it is the strongest case for keeping Fisher in the toolbox.

On a separate benchmark of 27 planted decompositions with known partitions, Fisher reaches a true-K ARI of 0.3605 while binary coactivation sits at −0.0007. Read alone, that looks like a Fisher win. It is not a win over a working baseline. On this benchmark the binary coactivation baseline is non-functioning: its pair-discrimination AUC is 0.4736, below the 0.60 gate, and continuous coactivation also fails here. Fisher's advantage on the planted benchmark is therefore a recovery in a setting where learned coactivation is uninformative. We treat it as a recipe-specific instrument diagnostic, not as evidence that Fisher works whenever coactivation fails. Three caveats set out below make that boundary precise.

Figure 6: True-K ARI on the planted benchmark

Mean true-K ARI by affinity on the 27 planted joint decompositions (64 learned components and 40 matched atoms per checkpoint), where the ground-truth partition is known by construction. Both coactivation baselines here are non-functioning (pair-discrimination AUC below the 0.60 gate). The dark line marks ARI = 0.

Fisher recovers the planted partition at 0.3605 and full K-FAC at 0.3140, while both coactivation baselines sit at chance: binary coactivation −0.0007 and continuous coactivation −0.0018. Identity-G is clearly worse at −0.0887. The contrast with Figure 1 is the whole point: Fisher's recovery appears only where coactivation has no signal to offer, which is the opposite of the published-checkpoint regime.

Three properties of this benchmark keep it a recipe-specific instrument diagnostic rather than evidence that Fisher works whenever coactivation fails.

A constructed oracle grid reproduces Fisher's predicted failure mode directly. At an output-effect angle of zero, meaning independent triggers whose output effects are aligned or identical, binary coactivation leads Fisher by 0.484 ARI at trigger correlation 0, and by 0.221 at trigger correlation 0.5. Across the full constructed grid, Fisher minus binary CI is +0.161 ARI with 95% CI [−0.026, 0.336], so the aggregate is inconclusive. This is exactly the case the pre-registration flagged as Fisher's weak spot, and it is a constructed instrument diagnostic rather than method evidence.

The canonical 67M model: no verdict under the original document-disjoint protocol

The evidence ladder was meant to end at the canonical 67M VPD decomposition. Under the original document-disjoint protocol it does not reach a clustering verdict there, and the reason is a property of the data, not the method.

Four pseudo-target seeds were each harvested over 500,000 tokens, producing checksum-verified activation bundles. Reconstruction error is at most 1.86 × 10⁻⁹ and unmasked logits match the source model.

The canonical Pile source carries no document identity. The pre-tokenized shuffled Pile has no document IDs, so sequence hashes cannot establish document-disjoint splits. The plan required document-disjoint splitting, and that planned leakage precondition could not be met. The original document-disjoint question therefore has no clustering verdict on this model. A separate, amended question with an exact sequence-hash split over the 977 verified sequences was committed with its own protocol and answered on this model; the sections after the Discussion report it.

Method

Discussion

Amended 67M question: do tangent-field or full K-FAC clusters beat continuous-CI coactivation on a causal compression endpoint?

The published residual-MLP verdict above stands on its own and is not revisited here. This part answers a different, amended question on the canonical 67M language model itself: do tangent-field or full K-FAC clusters improve a neutral per-sequence causal compression endpoint over continuous causal-importance (CI) coactivation?

The protocol was committed before any test metric was computed. Test sequences stayed sealed until a validation-only selector chose the cluster count. The endpoint is the area under the normalized next-token KL curve, KL(unmodified‖masked), integrated over a common log-rank support. Lower is better: a smaller area means the arm's clusters compress the model with less damage to its predictions.

The verdict is strong support for retaining continuous coactivation under this model and protocol. Both geometry arms lose the primary endpoint, and both behavior batteries independently oppose switching. This is an answer to the amended question only; it does not revise the published residual-MLP result above, and it does not claim coactivation is universally superior.

Validation-selected cluster count K
1024
Tangent-field − continuous CI, KL AUC
+0.01294
Holm [0.01214, 0.01393]
Full K-FAC − continuous CI, KL AUC
+0.01091
Holm [0.01014, 0.01169]
Sealed test sequences per seed
196
Overall recommendation
Retain continuous coactivation

Figure 7: Validation K selection for the amended 67M question

Mean validation KL AUC per cluster count K, with K=1024 selected

Mean validation normalized KL AUC for each valid cluster count K on the 195 validation sequences, averaged over the three primary arms and four pseudo-target seeds. Lower is better. The full K grid was {128, 256, 512, 1024, 2048, 4096}; K = 128 and K = 256 are absent because their primary curves have no common positive-rank support and are invalid by the protocol's rule. The ember bar marks the selected K.

K = 1024 has the lowest mean validation KL AUC at 0.00802, against 0.00940 at K = 512, 0.01065 at K = 2048, and 0.01175 at K = 4096. The selection was committed before any test metric was computed, so all test results in Figures 8 to 10 are evaluated at this single K.

Figure 8: Primary sealed-test deltas versus continuous coactivation

Paired KL AUC deltas of tangent-field and full K-FAC versus continuous CI with Holm-adjusted intervals

Paired difference in normalized next-token KL AUC between each geometry arm and continuous CI at K = 1024, on the 196 sealed test sequences per pseudo-target seed (784 cells in total). Error bars are Holm-adjusted percentile intervals from a hierarchical paired bootstrap over seeds and sequences (10,000 replicates). Lower AUC is better, so a delta right of the zero line means the arm loses to continuous CI.

Both geometry arms lose. Tangent-field minus continuous CI is +0.01294 (Holm-adjusted interval [0.01214, 0.01393], p < 10⁻⁴) and full K-FAC minus continuous CI is +0.01091 ([0.01014, 0.01169], p < 10⁻⁴). Neither interval approaches zero, so the primary endpoint gives no support for switching either arm.

Figure 9: Selected-K control deltas on the same primary support

Sealed-test KL AUC deltas versus continuous CI for the two primary arms and the Fisher and identity-G controls

Sealed-test KL AUC deltas versus continuous CI at K = 1024 for the two primary geometry arms and the two mandatory controls, evaluated on the identical common support (784 cells). Primary-arm intervals are Holm-adjusted; control intervals are unadjusted 95% bootstrap percentile intervals. The shuffled control is not plotted: its partitions cover 0 of the 784 cells' primary rank support, so it yields a null rather than a number and no extrapolation is made.

Every plotted arm and control is worse than continuous CI. (a) Fisher cosine trails by +0.01129 (95% interval [0.01031, 0.01241], p < 10⁻⁴), landing between the two geometry arms, which is consistent with the published-checkpoint finding that output-gradient geometry does not carry the grouping signal here. (b) Identity-G, the direction-free control, trails by +0.00698 ([0.00645, 0.00744]); even the weakest control interval sits well clear of zero, so continuous CI beats everything that produced a valid comparison at the selected K.

Causal endpoints on the 67M model

Two preregistered causal endpoints back the compression result. The exact interaction comparison is inconclusive: it required 256 exact matched pairs per cell, and continuous-CI clusters yielded 0 matched pairs in every seed, so the precondition failed and no interaction number exists. The behavior batteries did run, and both independently oppose switching for both geometry arms.

Each battery measures three things at the selected K. Necessity is the drop in the held-out target token's log-probability when the arm's selected clusters are ablated from the all-on model; higher means the clusters matter more. Recovery is the fraction of the all-off to all-on target-log-probability gap restored by activating the selected clusters alone; 1 means full restoration. Off-target KL measures collateral disruption on unrelated continuations; lower is better.

Figure 10: Behavior-battery deltas versus continuous coactivation

Necessity and recovery deltas of tangent-field and full K-FAC versus continuous CI on both behavior batteries

Paired necessity and recovery deltas of each geometry arm versus continuous CI on the delimiter-completion battery (60 test prompts per run) and the possessive-pronoun-agreement battery (16 test prompts per run), pooled over four pseudo-target seeds with Holm-adjusted bootstrap intervals. Negative means the geometry arm's clusters are less necessary, or recover less of the behavior, than continuous CI's clusters.

All eight intervals are entirely below zero. (a) Necessity losses are 7 to 8 nats: on delimiter completion, tangent-field −7.38 [−8.59, −6.22] and full K-FAC −8.26 [−10.07, −5.81]; on possessive agreement, −8.09 [−10.34, −5.88] and −8.34 [−11.01, −5.55]. (b) Recovery losses are largest on possessive agreement, where tangent-field recovers −0.956 [−0.994, −0.884] less of the gap and full K-FAC −0.847 [−0.997, −0.497], meaning the geometry clusters restore almost none of that behavior. (c) Off-target KL does favor the geometry arms (for example, delimiter tangent-field −6.18 [−7.45, −4.79]), but the preregistered combined status per arm and battery is opposite: the capability losses dominate, so both batteries oppose switching.

Figure 11: Qualitative test prompts behind the necessity and recovery contrast

Test prompt (target token)Continuous CITangent-fieldFull K-FAC
Necessity (nats)RecoveryNecessity (nats)RecoveryNecessity (nats)Recovery
“The prose insertion is (the internal identifier” → “)”
prompt-5fa1edc4246172bd9ab2
18.260.9944.670.8835.690.916
“After lunch, the businesswoman returned to” → “ her”
prompt-29fb2ede9eb44a515916
14.240.9991.66−0.0021.780.001
“The prose insertion is [a deterministic ordering” → “]”
prompt-a195eec7f9f6d7dc8edb
9.330.9982.970.8042.210.850

Three sealed test prompts from pseudo-target seed 0, with each arm's per-prompt necessity (target log-probability drop under ablation of that arm's selected clusters, in nats) and recovery (fraction of the all-off to all-on gap restored by those clusters alone). Prompt text and targets come from the behavior manifest, matched by prompt ID; per-arm numbers come from the seed-0 example logs.

The aggregate contrast in Figure 10 is visible prompt by prompt. (a) On the parenthesis prompt, ablating continuous CI's clusters drops the target by 18.26 nats while ablating tangent-field's or full K-FAC's clusters drops it by only 4.67 and 5.69, so the geometry clusters are far less necessary for the behavior. (b) On the possessive prompt, the geometry clusters recover essentially none of the all-off to all-on gap (−0.002 and 0.001) while continuous CI's clusters recover 0.999 of it. (c) The bracket prompt is a failure case: it appears in the logged worst-five test prompts for the tangent-field arm and for the continuous-CI arm itself, yet even there continuous CI's clusters recover 0.998 of the gap against 0.804 and 0.850 for the geometry arms. No arm's seed-0 log records a sign-flipping counterexample; the failure lists are the closest recorded negatives.

Method and provenance for the amended 67M question

Limitations of the 67M result

These limitations bound how far the amended-question verdict travels. None of them changes the direction of the primary result, but each narrows its scope.

What the amended question settles