When a measurement is wrong, it doesn't necessarily crash. It produces a number. The number goes into a table. Sometimes it gets an interval. Then I make a decision with it and stop asking the question.

That was the expensive part of this project: not the runs that failed visibly, but the ones that appeared to settle something.

Here are eight places where the machinery or its interpretation needed repair. Some came out during the research; others, noted below, became clear when I revisited the series. The fixes are more useful than a general promise to be more careful.

1. A Search Result Outlived Its Training Setup

The normalization history has two distinct mistakes. February's stability patch normalized the input to the whole Transformer. I called it Pre-LN in the blog, but Pre-LN changes the operation order inside each layer. The archived patch left those layers at post-norm. Part 1 now says what the code did.

Later experiments genuinely tried both internal norm orderings. Four-layer post-norm models trained at batch size 512, with slightly better recorded scores than their pre-norm counterparts. This comment wasn't invented:

norm_first=False,  # Hardcoded based on research

But the batch size subsequently rose to 8,192 and the training setup changed. Three- and six-layer runs failed under that later recipe. Their failure was not a clean measurement of what deeper models could learn.

Restoring the option, using Pre-LN with a lower learning rate and warmup, and letting the health check wait until warmup ended made the ladder trainable. It wasn't one forgotten flag, and it wasn't the first time a deeper model had worked. A choice measured under one regime had been frozen across changes to that regime.

2. A Row Count Wearing a False Moustache

My notes described a Transformer with 22 million parameters.

The Transformer I was measuring had 1,074,305 parameters. Its training set contained 22,196,565 matches. The resemblance explains the likely transcription error, not why I let it survive being quoted repeatedly.

The correction changed the capacity question. I had been discussing the limits of a substantially larger model than the one I'd actually built. Counting the parameters from the artifact was cheaper than every conversation based on the wrong count.

3. Grouping by Size Instead of Situation

I added a ranking loss intended to compare picks within similar draft contexts. The first version grouped by the number of heroes present. Two entirely different lineups could therefore become the same "situation."

Turning up that objective hurt the ranking diagnostic. It looked like a tidy negative result about ranking losses; it was a result about this particular misgrouped loss.

The repaired version grouped using context information. At weight 0.3, its counter-interaction partial rank correlation rose from 0.6889 to 0.7116. The earlier, count-grouped version had moved from 0.6250 to 0.5994.

Those are changes from different controls. Normalization, warmup, learning rate, and the training recipe had changed between the experiments. It was wrong to write "same everything but the grouping." The comparison supports a useful response to the corrected objective within its own recipe, not a numerical attribution of the entire cross-experiment gain to one fix.

Paired repetitions across three training seeds kept the corrected improvement positive, averaging about +0.016, with gains from about +0.006 to +0.023. That was more informative than the first attractive +0.023 alone.

4. Random Heroes In, Fixed Slots Out

The deeper evaluation hid a random subset of heroes, then tried to reconstruct the visible context by reading only the first few slots.

A hero visible in slot four could be absent from the grouping key. At the worst one-hero-per-side case, 98.4% of comparisons landed in one bucket, the bucket labelled as if no heroes were visible.

A bar chart with eight pairs of bars, one pair for each number of heroes on the board from two to nine. For each pair the red bar shows the broken key and the green bar shows the corrected one. At two heroes the red bar reaches 91% and the green 22%; at three, 83% against 11%; at four, 69% against 8%; and both fall away to near zero by eight and nine heroes.

A second version of the same mistake inserted the candidate into a fixed slot, which could already contain a revealed hero. The supposed next pick sometimes evicted a teammate. The repair reads every visible slot and reveals the candidate in its own previously hidden slot.

The corrected evaluation is what Part 9 now reports. One primary criterion flipped to pass; the full acceptance test still failed, because the model remained worse than the table at several shallow states.

A cleaner-looking result after a fix deserves checking too. Improvement is not an exemption from verifying which computation changed.

5. The Optimizer Was Moving the Experimental Knob

One model had a learned knob controlling how strongly it mixed direct pair estimates with a factored alternative. Runs started far apart and converged toward the same floor. I read that as identifying a preferred mixture.

Weight decay was moving the knob even without evidence from the data. A no-learning calculation showed that the optimizer's shrinkage could account for its trip toward the floor.

Agreement between starting points had not isolated a data preference. I needed to separate the intended learning signal from the regularization applied to the parameter being interpreted. After changing the optimizer treatment, the mixture behaved differently; that still didn't make its fitted value uniquely identified.

6. A Query and Its Scoring Input Shared a Name

The editorial audit found another error in the draft-state story. A user with one teammate and no visible enemy asks at 1/0. The evaluator adds a candidate and scores 2/0.

My diagram put 2/0 on a grid of pre-pick requests, found no observations there, and declared the one-teammate question irrelevant. The actual 1/0 request accounts for 3.87% of the serialized decision points in the trace sample. Only the two-teammate, no-enemy request is absent under that convention.

This distinction also explains why a complete board can matter to a draft advisor: it's the scoring input after adding a candidate for the final pick. Training on request states alone can omit scoring inputs the model needs.

Naming the coordinate system is not paperwork. Without it, two individually correct arrays can produce a wrong conclusion when compared.

7. The Supposed Ceiling Was an Estimator

I selected the best-looking heroes using empirical rates and scored them on those same rates. Unsurprisingly, that made the ranking look excellent.

Fitting on one subset and scoring on another removed some of the selection optimism. At sparse contexts the resulting lookup could even lose to the baseline. That does not mean the true opportunity is negative. It means a noisy estimator is not a measurement of the best any method could achieve.

I could compare estimators. I could not turn the strongest-looking row in that comparison into a universal ceiling, then report a fraction of it as "signal captured."

8. Seed Variation Was Not Measurement Error

The scheduler wasn't exposing the training seed, so repeated queued runs were not exploring different initializations. Once I exposed it, three repetitions of a Transformer recipe spanned about 0.076 on the counter-ranking correlation, while validation log loss varied by about 0.00004.

Those are different trained models. The spread is not the error bar for grading one fixed model, and a three-run range is not a universal noise floor. The structured model varied much less in its repetitions, which also argued against blaming everything on an intrinsically unreliable metric.

This cost Part 9 its "bigger is worse" headline. The size-ladder endpoints differed by about 0.085, not less than 0.076. But one run per size and a seed check at one size do not establish a monotonic capacity effect. Nor may I paste that one seed range onto every point as if each had been independently repeated.

Checks I Can Actually Run

The response is not another paragraph promising vigilance. It is a few concrete requirements:

  • Controls with a stated expected result. A constant correlation can be undefined; a random ordering should approach 0.5 concordance. "Should score zero" is not a universal specification.
  • Behavioral tests. Hide a hero outside the first slots. Add a candidate without overwriting anyone. Compare a grouping's actual membership, not just a comment describing it.
  • One coordinate convention at each boundary. State whether a count is before or after candidate insertion, then test the conversion.
  • Artifact-derived metadata. Record parameter counts, fitting populations, configurations, and seeds with the result rather than copying old prose.
  • Separate uncertainty questions. Resampling evaluation data and retraining with another seed answer different questions. Neither two runs nor an overlapping error bar settles every comparison.

Several failures made promising ideas look unhelpful. That is a reason to revisit negative results when their instruments change, not evidence that measurement errors generally favour false negatives.

The next post assembles the two estimators into a switching rule. That rule needs its own check too; neither a good table nor a good model certifies their combination.