I let the model from Part 1 keep training. Forty-three epochs, early stopping, a saved checkpoint. Everything looked extremely professional.

The reported test accuracy was 54.77%. Always predicting a Radiant win scored about 53.2%. A million parameters had bought roughly a point and a half over ignoring the heroes entirely.

By then the collection had grown to about 417,000 matches. I had enough results to start comparing approaches, though not the frozen datasets and recorded split identities I'd introduce later. The numbers here are historical reports, not benchmarks I can now reproduce exactly.

The Small Model Earns Its Seat

Before tuning the Transformer further, I tried two simpler alternatives.

The first was a hero win-rate lookup: average the five Radiant heroes' training win rates, do the same for Dire, and choose the side with the higher average. It reported 53.90% accuracy. Cheap, but not much above the side-rate baseline.

The second was a small neural network, an MLP. It summed the hero embeddings on each team, joined those two vectors, and passed them through a few dense layers. No attention, and no dependence on the order of heroes within a team. About 37,000 parameters rather than the Transformer's roughly 1.1 million.

It reported 54.88%, close to the Transformer's 54.77%. The larger model hadn't earned a convincing advantage in these runs. That didn't establish whether the limitation was the architecture, its training, or the data. It did make attention something I had to justify.

I also tried dropping the ten masked views per match and training on complete lineups alone:

Model With the original partial-draft setup Complete-draft setup
Small MLP 54.88% 55.74%
Transformer 54.77% 55.90%

Those higher scores convinced me to proceed with complete drafts. They aren't, by themselves, a clean ablation of augmentation: the surviving February script also expands validation and test matches into masked views. A complete-draft test supplies more information. Without the original evaluation records, I can't establish that these runs differed only in their training examples.

Nor had ten views meant ten times as many independent games. I had changed how the same matches were presented, not collected more evidence. And a model trained only on complete lineups still needed a separate test before I could trust its advice halfway through a draft.

Surely Better Players Would Help

My next idea was to filter for high-ranked games. Good players understand counters, build more deliberate lineups, and generally know what their teammates are trying to do. I expected their match outcomes to offer a clearer lesson about drafting.

Filtering to Divine and above left 32,080 matches. That model reported 53.77%, below the complete-draft model's 55.90%. Not the jump I'd expected.

But the filter had changed both the amount of training data and the population being tested. Different brackets can also have different Radiant base rates. This was a failed attempt to obtain a higher headline accuracy, not a controlled measurement of how much draft matters at different skill levels.

I then fitted separate models by bracket. The archived chart data makes the sample-size difference visible:

Two aligned panels for Herald through Divine. Dataset sizes range from 32,080 to 89,322 matches; reported accuracy ranges from 52.41% to 55.94%. The sample sizes differ substantially across brackets.
These were different-sized, mixed-mode datasets, not a controlled experiment on the effect of player rank. The plotted values are the historical reports.

This was a separate sweep, not a repetition of the 53.77% run. Its lower-bracket models scored better on their own test populations. The training populations were plainly unequal, though, and the collection mixed game modes. Matching architecture settings didn't make those other differences disappear.

I had a plausible explanation: conspicuously awkward lineups might be more common lower on the ladder, leaving more recognizable patterns in hero-only data. A team without a support can have a very predictable problem even when none of its players is predictable individually.

That is a hypothesis, not something these tests established. I hadn't measured how often high-ranked players drafted badly. The model didn't even receive intended roles or player histories. It couldn't distinguish a poor lineup from one that was sensible but badly executed.

A Working Population, Not the Right Population

I settled on Herald-through-Archon games: known ranks below 50, with the retained training policy also restricting duration to 20-60 minutes. The resulting working dataset was about 200,000 matches, and the Transformer reported 56.27% test accuracy.

That became the starting point for the next round of experiments. It was not proof that I'd found "clean signal in clean data," nor a measured 0.37-point benefit from filtering on a shared test set. I'd chosen a narrower problem and scored the model within it.

There was a practical reason to spend time there. I'm not an Immortal player. Neither are my friends. When we queue for a pub game, someone may still type "pos 5 pudge" and lock in Pudge. Lower-bracket performance was relevant to the tool I wanted to use, even if my grand explanation for it remained unproven.

The useful change in my thinking was to treat the dataset boundary as a modeling decision. "Which matches belong in training?" needed an experiment, not an answer borrowed from my opinions about good Dota.

The filter wasn't permanent. Part 6 returns with a single-mode corpus, recorded chronological splits, and linear baselines fitted within each bracket. Those still don't equalize training-set size, but they show that higher brackets needn't be written off. A rank-conditioned Transformer then lets me keep the data and tell the model which bracket it is looking at.

For the early experiments, though, I had a working population and a baseline to improve. What I didn't have was the patience to submit every training job by hand.