I typed in a single enemy hero - Necrophos - and asked the model what I should pick.
It said Ancient Apparition.
If you play Dota, you just nodded. Ancient Apparition's ultimate prevents healing, and Necrophos wins fights by out-sustaining people. It is the textbook answer.
My model had never read the textbook. Its inputs were hero IDs and team and rank information, not names, abilities, or role descriptions. I sat there feeling extremely clever.
Then I got suspicious. I'd asked it one question I already knew the answer to, and graded it myself.
Any recommendation list puts something at the top. If I recognize a counter, I feel validated. If I don't, I have a rich vocabulary available: "interesting," "maybe it knows something I don't," "that's a flex pick." I am not a neutral instrument.
The useful question was harder: without labelled examples of good advice, what could I actually measure?
First, Make Missing Heroes Part of Training
The current baseline had been trained on complete drafts. February's prototype had tried synthetic partial views, but the later sweeps returned to complete lineups. Accepting zeros for missing heroes didn't establish that this newer model could score unfinished drafts sensibly.
Predicting the next hero picked would teach what players choose, which isn't necessarily what helps them win.
Instead, I kept the match outcome and randomly hid heroes during training. For each match presentation, I sampled how many heroes to retain on each side, from zero to five independently, then chose the visible heroes at random. This differs from February's fixed alternating reveal order.
The outcome is a noisy label for the information left visible. It teaches associations under this masking scheme, not observed pick chronology.
I also wanted to replace the order-sensitive slot decoder with a permutation-invariant set encoder, which pools hero representations. The slots were not actual pick chronology: the snapshot sorted hero IDs within each team. A set seemed a better fit, but changing the architecture and masking together would leave me unable to tell which change helped.
So I ran the four neural combinations: each architecture, trained either on complete drafts or with random masking. They used the same snapshot and 24-epoch recipe, and were evaluated on identical masked versions of the 2,671,633 validation matches. A mask-trained linear baseline supplied another comparison. Evaluation used one count pattern per total - 1/0, 1/1, 2/1, and so on - rather than every possible partial board.
The metric was mean log loss, in natural-log units per match: it penalizes confidently wrong probabilities, and lower is better. The dotted reference is the log loss of always predicting the roughly 53.15% Radiant win rate.

With one hero visible on each side, the complete-draft Transformer scored 0.70749, worse than the constant forecast's 0.69116. Training it with masking brought that down to 0.68861.
Pooling alone did not fix the problem. The set encoder trained on complete drafts scored 0.85092 at that same state. Once both designs were trained with masking, their log losses differed by at most about 0.00016 across the tested non-empty states.
In this recipe, mask training was the important change. I kept the Transformer. The architecture-only arm had earned its training time by stopping me crediting the wrong change.
An Answer Key Made of Other People's Choices
For recommendation testing I used a rerun of the mask-trained Transformer recipe, job 155. I called its evaluator an "oracle," which was a grand name for a table of historical pair outcomes.
Fix Necrophos as the enemy. For each candidate hero, count the matches where our team had that candidate and the opposing team had Necrophos, then compute our team's win rate. Repeat for each enemy hero.
These are outcomes for players who chose those heroes, not outcomes from players assigned my recommendations. Specialists, roles, teammates, and selection effects remain in the rates. The snapshot has no player identifiers with which to adjust for individual familiarity. This is an observational proxy for evaluating rankings, not a measurement of advice-following win-rate uplift.
My top-five score worked like this:
- Insert each candidate into the model input and rank the candidates. With one enemy and no teammate already selected, that means scoring one hero on each side after insertion, not an input containing only the enemy.
- Take the arithmetic mean of the five selected candidates' observed matchup win rates.
- Average those means equally across enemy-hero contexts.
The evaluator retained candidate pairs with at least 300 matching games and contexts with at least 20 eligible candidates. It pooled both sides of the map. Model scores likewise averaged the two side assignments and the rank mix of the population being counted.
The baseline ignored the enemy and ranked eligible heroes by their overall training-set win rates. On the 21,373,140 training matches, the results were:
| Ranking method | Top-five score |
|---|---|
| Overall hero win rates, ignoring the enemy | 54.08% |
| Complete-draft Transformer | 53.39% |
| Mask-trained Transformer | 54.97% |
The mask-trained model beat that baseline on this statistic. The complete-draft model did not. But these rates came from the same matches that trained the models; they were a first diagnostic, not a generalization result.
The Control With No Opinions
A second diagnostic asked whether the ranking reflected more than overall hero strength. A model can look good simply by putting generally successful heroes near the top, regardless of the enemy.
My first attempt subtracted a hero-strength estimate from both model scores and observed pair rates, then correlated the remainders. On an early teammate-pair check, a constant predictor scored +0.0343. A model with no preferences had acquired some through my metric.
In this construction, subtracting the same term created shared variation instead of isolating matchup information. I replaced it with partial Spearman correlation: a unitless score from -1 to 1 that compares rankings after linearly adjusting both sets of ranks for the baseline's ranks.
Under the corrected diagnostic, a constant predictor returned undefined - it has no ranking variation - and random scores landed near zero. Those controls then travelled with the real model's results.
That adjustment controls for the chosen hero-strength baseline. It does not remove specialist-player effects or turn correlation into causation. Its useful contribution was more immediate: a check that could reject a flattering number before I built an explanation around it.
Did the Gain Leave the Training Set?
I recounted pair outcomes on the 2,671,633 validation matches, which were held out from weight updates. This was still the validation set used during model development, not a new untouched test set.
The checkpoint stayed fixed, and the baseline continued using hero win rates from training. The same top-five statistic could be applied to a teammate question: with one teammate already selected, insert a candidate and score the two-hero team with no enemies revealed.
| Question | Advantage over the baseline on training rates | On validation rates |
|---|---|---|
| One enemy already selected | +0.894 percentage points | +0.824 percentage points |
| One teammate already selected | +0.244 percentage points | -0.165 percentage points |
These are differences in the defined historical-rate score, not extra wins caused by the recommendations. The enemy-context test covered 127 enemy heroes; the teammate test covered 127 locked teammates. Eligibility was checked on the counted split, so validation had fewer qualifying candidate pairs than training.
The teammate gain did not carry over in this evaluation. That was enough to withdraw the stronger claim I'd been making, but not to diagnose exactly what the network had memorized: time, pair coverage, and the rank mix also changed across the split.
The countering result was more promising. It justified further testing of one-enemy rankings against stronger baselines, not declaring that I'd built a validated counter-picker for real players.
Back to Necrophos
I could now return to the original recommendation as an illustration rather than an exam I'd marked myself.

The figure selects the nine largest promotions and nine largest demotions for this one enemy. Its annotations are training-set matchup win rate minus the hero's overall win rate, in percentage points, not causal effects.
Ancient Apparition moves from 71st to 1st. Its observed win rate against Necrophos is about 2.35 points above its overall rate. All nine displayed demotions have negative observed differences, but two promotions go the wrong way: Outworld Destroyer and Viper.
The list responds to the enemy, rather than merely repeating a tier list. One selected matchup does not establish that the model generally knows traps better than opportunities, or that every move it makes is sensible. Ancient Apparition was a satisfying example. It still wasn't an answer key.
What to Ask Next
The cheaper training from Part 6 also funded a short search and longer runs. These used a refreshed 27.7-million-match snapshot, not the 26.7-million cohort behind the results above. Five startup trials ran before I stopped the search; one was aborted after a constant-prediction first epoch. This was not an exhaustive search or a demonstration of TPE's adaptive advantage.
On that newer cohort, a 48-epoch schedule improved validation log loss for complete drafts, but not the one-enemy top-five score counted from training pairs. The schedules were separate runs, not one continued checkpoint. Nor are complete inputs useless to a recommender: scoring candidates for the last pick produces a complete draft. The question is which improvements help the decisions people actually make.
The refreshed snapshot also contained 5,121 validated draft traces. Their prefixes occupied 20 of the 36 possible team-count combinations. My uniform mask assigned 16/36, or 44.4%, of its sampling probability to combinations absent from those traces. That is not a measurement of wasted gradient or a proof that every unobserved state is impossible.
It gave me the next comparison to run: change the masking distribution, and check which states the advisor should actually be asked about. States a draft passes through, requests a player makes, and boards scored after inserting a candidate are related, but they are not interchangeable.