Juggernaut and Lifestealer are each other's closest neighbours in my model. So are Shadow Demon and Oracle. Chen and Io, two supports with rather unusual job descriptions, have found each other too.
Those are relationships a Dota player can make sense of. The model hadn't been given their names, abilities, or roles. It had been trained on hero IDs and team membership, with the match outcome as its target.
After all the time I'd spent watching accuracy move by fractions of a point, this was something I wanted to look at properly.
The Model Behind the Map
The scheduler from Part 3 let me queue experiments rather than attend to each one personally. The sweep ran through late February and early March: learning rates, weight decay, embedding sizes, layer counts, and the size of the final prediction network.
The winning configuration was not especially grand: two Transformer layers, 128-dimensional hero embeddings, and a 512-unit decoder. It used a maximum learning rate of 0.00145, weight decay of 0.05, and a thirty-epoch cosine schedule that gradually lowered the learning rate.
Job 116 reported 58.05% test accuracy on the historical lower-rank, complete-draft dataset. That number needs its original label: I was looking at test results while choosing configurations. Checkpoints were selected using validation loss, but the subsequent comparison of their test scores made that holdout part of development. It wasn't an untouched final exam, and I didn't preserve the dataset and split identities needed to reproduce it exactly.
There was no single architectural breakthrough that explains crossing 58%. The move from the earlier 57.99% run changed both layer count and learning rate. Some larger configurations stopped learning under the settings I tried; that isn't evidence that larger models are intrinsically worse.
Job 117 repeated the winning settings and reported the same accuracy. I left both rows on the dashboard because deleting one felt like throwing away the witness. But this was a repeat under the fixed-seed setup, not a test of robustness to different initializations.
One detail matters later: the preserved model uses post-norm inside its Transformer layers. Earlier jobs had tested both norm orderings, and four-layer post-norm models did train at batch size 512. This wasn't a reversal of the February input-normalization fix described in Part 1; those are different operations. Freezing the internal norm choice and later changing the training setup becomes a problem in Part 10.
A Map Without Hero Descriptions
A hero embedding is a learned list of numbers - 128 in this model - used to represent that hero. I extracted the vectors for the 127 named heroes in the roster file, not every allocated row of the larger embedding table.
To look at them, I used t-SNE, which arranges high-dimensional points on a two-dimensional page. It's useful for exploring neighbourhoods, but it distorts distances and can make groups look more definite than they are.
For the figure, I kept the 98 heroes tagged Carry or Support, but not both. Carries typically need resources to become major damage dealers; supports help the team without taking as many of those resources. I used those tags to colour the points after training. They weren't inputs to the win predictor.
The carry and support regions were the first thing I noticed. Then came the smaller groups: ranged carries together, illusion heroes near one another, and supports with familiar jobs sitting side by side.
Three pairs particularly pleased me:
- Juggernaut and Lifestealer: melee carries with temporary debuff immunity.
- Shadow Demon and Oracle: supports with ways to save a teammate who appears thoroughly committed to dying.
- Chen and Io: specialist supports with unusual ways of helping their team.
None of those descriptions had been supplied to the model. But I had supplied them to the story. A plot full of heroes I know gives me plenty of opportunities to find something that looks sensible.
Checking the Neighbours
The plot came first. These numerical checks were added later using the preserved checkpoint, in the original 128-dimensional space rather than on the picture.
For each of the 98 role-tagged heroes, I found its five nearest neighbours within that same set. I used cosine distance, which compares the directions of their vectors. 80.4% of those neighbour relationships joined heroes with the same role tag. Shuffling the labels ten thousand times on the fixed neighbour graph gave an average of 52.5%. That comparison keeps the unequal class sizes - 61 carries and 37 supports - instead of assuming chance means 50%.
The named pairs survived a broader search too. Among all 127 roster entries, Juggernaut and Lifestealer were mutual nearest neighbours; so were Shadow Demon and Oracle, and Chen and Io.
That confirms their proximity isn't just an artifact of squeezing the vectors onto a page. It is not an independent discovery sample: I selected those pairs after looking at the projection. And not every grouping I narrated was as tidy. Void Spirit's nearest neighbour was Puck, not one of the other Spirits I'd grouped him with.
The result I can keep is specific: this outcome-trained representation carries information related to the roster's role tags. It doesn't establish that the network understands abilities, that attention was necessary to learn the structure, or that my interpretation of every cluster is right. Co-occurrence, individual hero strength, and interactions could all contribute.
The map was worth preserving. It gave me something concrete to investigate inside a model whose job, up to this point, had been to output one number.