The lookup table had been sitting in my results for a while. I had put it in a row called ceiling, which is a wonderful way to stop yourself competing with something.
It counted how often each hero won against each enemy hero, then ranked the candidates. No attention mechanism. No GPU. Just the sort of thing a database is very pleased to do for you.
I was treating its score as the limit to approach, rather than asking whether it was already the better estimator.
Give the Table a Fair Race
Part 7 compared the neural model with a baseline that ranked heroes by their overall win rates. Beating that says more than beating random choices, but it doesn't establish that a model is needed. A pair lookup can use the enemy hero too.
An initial cross-fitted check - fit the lookup on one half, score it on the other - made it clear that the row deserved a proper comparison. The subsequent matched experiment used the newer 27.7-million-match snapshot, with both estimators fitted from its 22.2 million training matches and pair outcomes counted on validation.
For the question "they have one hero; what should I pick?", the recorded top-five historical-rate scores were:
| Ranking method | Score |
|---|---|
| Overall hero strength | 54.05% |
| Mask-trained Transformer, job 176 | 54.78% |
| Pair lookup | 55.17% |
| Pair lookup with empirical-Bayes shrinkage | 55.17% |
This is the Part 7 statistic: average the observed matchup win rates of the five selected heroes, then average equally across eligible enemy contexts. It is not the win rate of people assigned those recommendations.
The lookup was ahead by about 0.39 percentage points on that proxy. Calling it a ceiling had hidden an ordinary engineering obligation: compare the complicated thing with the strongest simple thing I can actually build.
A lookup is an estimator. It has fitted quantities even if nobody ran gradient descent to obtain them. Shrinkage pulls noisy pair effects toward a simpler estimate, rather than letting a rare pairing's spectacular result dictate the recommendation. That would become an important baseline as the questions got more specific.
What the Pair Counts Still Couldn't Tell Me
The result did not remove the selection problem.
Ancient Apparition's observed record against Necrophos belongs to the players who chose that matchup. Some may be specialists. A comparison across rank brackets doesn't eliminate that possibility: specialists can exist in every bracket.
I had checked whether pair effects looked similar across ranks and read that as ruling out skill. It did not. The snapshot has no player identifiers with which to separate hero familiarity from the matchup itself. The surviving claim is about associations in the collected matches, not what a particular player would gain by following the list.
Nor was the empirical best-looking ranking an upper bound on all possible methods. Selecting and scoring on the same noisy rates is optimistic; splitting the data exposes that optimism, but a split-fitted lookup still has estimation error. Another estimator can pool information differently and beat it.
So I had a competitor, not a ceiling, and an evaluation proxy, not a causal trial. That was enough to run a useful race, provided I stopped renaming those things into stronger claims.
Which Board Is the Question About?
I also revisited the masking experiment. Reweighting training toward states observed in draft traces had not delivered the hoped-for improvement on the early-state tests. Before allocating another run, I needed to say what a "state" meant.
There are three different objects:
- A state a recorded draft passes through.
- The information entered when somebody asks for a pick.
- The board the model scores after inserting a candidate.
My 5,121 validated traces supplied 51,210 serialized pick events. Counting the state immediately before each event, from the picking team's perspective, gave 14 distinct request-count pairs. That is a frozen evaluation convention derived from traces, not a measurement of actual website usage. It also doesn't prove that every intermediate recorded prefix was visible to players during a simultaneous picking phase.
The distinction matters even before those limitations. When I ask what to pick against one enemy, the request is 0/1: zero of ours, one of theirs. Scoring a candidate produces 1/1. Those coordinates cannot be put on the same grid without moving the point.
On revisiting the series, I found that my diagram had done exactly that. It put the oracle's scoring coordinates on the request grid and declared two teammate questions nonexistent. The corrected plot marks requests before candidate insertion, with the two coordinate frames labelled separately:
The corresponding request-to-score mapping is:
| Question | Request: ours/theirs | Board scored | Share of recorded requests |
|---|---|---|---|
| One enemy, no teammate | 0/1 | 1/1 | 6.13% |
| One teammate, no enemy | 1/0 | 2/0 | 3.87% |
| Two teammates, no enemy | 2/0 | 3/0 | 0% |
The one-teammate question does occur under this convention. Its disappointing validation result in Part 7 cannot be explained away by calling it irrelevant. Only the two-teammate request is absent from these traces.
Likewise, a completed draft is not automatically wasted training material. A last-pick request has nine heroes; adding each candidate produces a complete board. A model that ranks the last pick has to score those boards.
The Next Comparison
The question was no longer whether a neural model could beat a static tier list. It was whether a fitted model could improve on a strong counting estimator across the situations the evaluation was meant to represent.
Both approaches can compose pair effects to score a lineup absent from training. The table doesn't literally need a stored copy of the entire draft. Sparse full-lineup counts instead make evaluation difficult: once several heroes are fixed, there may be too few matching games to estimate each candidate's conditional win rate directly.
That meant the race needed two things: a clearly defined population of questions, and another way to compare estimators at depth. If a model helped only in one region, I wanted to know where. If counting remained better, I wanted a product that could use the table without being embarrassed about it.