I've been playing Dota 2 on and off for years. Two teams of five players choose heroes before the match, and those choices matter. Ancient Apparition counters Alchemist's healing. Silencer can ruin Enigma's day. Five melee heroes against an Earthshaker is asking for trouble.
When it's my turn to pick, though, I don't calmly weigh every counter and every hole in our lineup. I usually just slam Invoker down mid and hope for the best.
I wanted a second opinion: enter the heroes already picked, then get a ranked list of candidates for my team. Building a neural network for casual games with friends is wildly over-engineered. That was part of the appeal.
This series follows the attempt to build that advisor. My early data mixed public game modes; the later research focuses on Ranked All Pick, where players choose their own heroes. Results from those games don't establish that the tool works for organized Captains Mode, the captain-led pick-and-ban format that gave the project its original working title.
Predict the Outcome, Then Rank the Picks
OpenDota's API supplied completed matches: the heroes on each side and who won. I rewrote my prototype collector to handle duplicates, resume interrupted collection, and batch database inserts in groups of 100. The first dataset reached 226,520 matches.
The model's job was to estimate:
P(Radiant wins | the heroes shown)
Radiant and Dire are the two sides. To recommend a pick, I'd insert each available candidate into my team's next empty slot, score the resulting lineup, and sort by my team's estimated win probability.
This is a prediction built from historical choices, not a measurement of what would happen if a player followed the advice. Skill, familiarity with a hero, and everything that happens after the draft also affect the result. A higher model score is a preference to investigate, not a promised increase in wins.
I chose a Transformer because attention lets each hero's representation depend on its teammates and opponents. That seemed a useful fit for synergies and counters. It gives the model a way to represent those relationships; whether it learns them better than a simpler model is another question.
The input has ten slots, five per side. Hero IDs become learned vectors, and position embeddings distinguish the slots. Zero means empty, with an attention mask preventing empty slots from being used as if they contained real heroes. I remembered the masking technique from re-implementing BERT during an internship. Here it let the same network accept an unfinished lineup.
Ten Views Are Still One Match
Accepting a partial draft is not the same as knowing how to score it. My data contained final lineups, not the sequence of decisions that produced them.
I generated ten views of each match, revealing one hero at a time in an assumed alternating order. Each view inherited the match's win/loss label.
That label is a noisy observation, not a claim that the visible heroes determined the outcome. With suitable sampling, these examples can teach a conditional outcome predictor: among matches containing this information, how often did Radiant win? The hidden picks and the players still matter.
There are two catches. First, ten views of one game are not ten independent games. The revised loader kept all views of a match in the same training, validation, or test partition, then shuffled training examples for batching. Shuffling alone would not prevent leakage across those partitions.
Second, my assumed reveal order wasn't recovered from actual drafts. It created synthetic questions with a frequency I chose, not necessarily the questions a player would ask. The masking made those inputs representable; it did not validate them as recommendation requests. That distinction becomes important throughout the series.
Getting Through an Epoch
The first obstacle was less philosophical:
CUDA error: device-side assert triggered
I'd allowed hero IDs through 150. The data contained 155. Hero IDs have gaps: the size of the game's roster, the number of distinct heroes observed in a sample, and the largest ID are three different counts. An embedding indexed by raw ID needs room for the last one.
I expanded it to IDs through 200, plus the empty-slot row. That buys space, not knowledge. An unused hero's vector starts random; allocating it doesn't make recommendations involving that hero trustworthy. I also added a utility to expand the table while preserving existing weights, so future training needn't start from scratch just because the vocabulary grows.
Then training produced NaN losses. The stabilization patch added gradient
clipping, gentler initialization, and normalization of the input to the
Transformer. Training resumed, but several changes had landed together; that
wasn't a diagnosis of the original failure or proof of a single cure.
I originally described that normalization as Pre-LN. The archived code does something different: it normalizes the input once, while each Transformer layer still uses post-norm internally. The distinction matters when deeper models return in Part 10.
After the first epoch, training accuracy was about 53%. Radiant had won roughly 53% of the matches. My neural network had become competitive with a constant. Machine learning is truly majestic.
It could now finish an epoch, accept incomplete lineups, and produce scores. Whether those scores were useful was still open. Before trusting its advice, I needed to let it train and find out what much simpler models could do.