Anti-Mage finished second on his team in net worth in 23.8% of 13,759 recorded games. My automatic role table took that as evidence that he belonged in position 2.

That's not what the statistic says. In the usual role convention Anti-Mage is a carry, and another player finishing richer doesn't make him the mid. The measurement recorded the ordering of farm at the end, not the players' intended jobs at the start.

A tighter interval around that percentage wouldn't repair the substitution.

A Small Matching Problem With a Large Assumption

I wanted the advisor to flag candidates that left no plausible assignment of the five heroes to the five conventional positions.

Given allowed positions for each hero, that is a small matching problem: can each hero occupy a different slot? The algorithm can answer it exactly relative to those allowances. What it cannot establish is whether the allowances are right. Dota's roles are conventions and strategic choices, not five mandatory character classes enforced by the game.

I tried inferring them from 97,091 matches with player-level statistics. Rank each team by net worth, gold per minute, or last hits, then count where each hero tends to finish. Some results were convincing: supports tended to finish low in farm, carries high. Enough familiar examples worked to make the exceptions look like threshold problems.

They weren't only threshold problems.

Farm Outcome Is Not Draft Intent

Net worth reflects kills and other sources of income as well as creep farming. Last hits offer a different view, but they still measure what happened during the game. A carry who was shut down can finish below the mid in both statistics.

My original explanation of Invoker's result was wrong: last hits include creep kills made with spells, not just right-click attacks. His farm-rank distribution cannot be explained by saying the metric ignores spell farming. I hadn't isolated why the two summaries disagreed for him.

Combining the signals brought back some plausible role assignments, but did not turn them into labels of intent. At a permissive cutoff almost every hero acquired almost every role. A stricter cutoff constrained the suggestions, while still retaining mistakes such as interpreting Anti-Mage's second-place farm as mid-lane evidence.

I had useful evidence for reviewing a role table, not an automatic source of role truth.

So I Typed It In

I built an editor with one row per hero, seeded it from the farm summaries, and reviewed all 127 entries. The result is a hand-authored role table informed by data, not a measured classifier of how every hero must be played.

That distinction determines the behaviour. A candidate that doesn't fit the remaining assignments is flagged, not removed or demoted. If my role judgment is wrong, the player sees a note they can disagree with rather than losing an option from the ranking.

A screenshot of the draft advisor at dota2.scottliu.com/draft. The left side shows a draft board: four heroes on the player's team - Io, Nature's Prophet, Crystal Maiden and Puck - with one empty slot, and five enemy heroes: Batrider, Alchemist, Rubick, Enigma and Templar Assassin. Below them a row of ten banned heroes, greyed out and crossed through. Below that a search box and a grid of hero portraits with position filters. The right panel is headed 'Recommended next pick', notes that with nine heroes on the board the answer comes from the trained model, and lists Spectre as the top pick clear of the field, followed by Legion Commander at minus 0.0327, Slark at minus 0.0475, Phoenix at minus 0.0524 and Drow Ranger at minus 0.0560. Each hero shows its playable positions. A panel underneath reads 'positions still open' with 1 and 3 highlighted, and one slot left.
The released interface combines rankings, score-gap hints, and hand-authored role notes. The uncertainty hints have the limitations described in Part 13; neither they nor the role table certify a best pick. Role flags do not change the ranking.

The export uses a counting estimator below five already-selected heroes and a trained pairwise model from five onward. One selected training seed supplies the model scores; two additional seeds supply a limited variability estimate. It is not the Transformer, and it is not a three-model prediction ensemble. The inspected compressed NumPy artifact is about 2.47 MB.

The Release Did Not Pass Its Gate

Part 11 had evaluated the combined research hybrid directly. The later E41 checks addressed a different boundary: the exported arrays, their serving arithmetic, bracket handling, and the switch around them. It was wrong to say nobody had ever evaluated a combined estimator before this point.

The first export fitted its counting tables on training data alone while its acceptance baseline had training plus validation. Refitting the export's tables on the same 24.97 million matches removed that asymmetry.

The E41b report passed both primary conditions on its evaluation adapter: +0.00081 weighted concordance over the measured 80% of request mass, and +0.00195 on the designated deep states. Those are ordering-statistic margins, not extra wins. As the later check below makes clear, the adapter was not a complete end-to-end test of the public ranking call.

One guard still failed. For the one-enemy choice, labelled 1/1 after candidate insertion in the legacy guard, the historical top-one outcome score was about 0.057 percentage points below the reference estimator. A separate concordance result near zero did not cancel that failure: the two statistics ask different questions.

A diagnostic partially attributed the difference but did not close the guard. The rule had said that failure blocked launch. I chose to release the tool as a personal experiment despite it. That was explicit risk acceptance, not continued compliance with the original acceptance rule and not a passing test.

The consequences are concrete: no claim of acceptance-test passage, no claim of measured additional wins from following advice, and no demonstrated performance on a genuinely new patch. Revising the estimator to win on the same repeatedly inspected split would not supply that missing evidence.

What I Am Signing Off On

The useful distinction at the end is between three things: what was measured, what I chose by judgment, and what remains unresolved.

The role table belongs to judgment. The offline comparisons belong to specific estimators, datasets, and instruments. The release is an experimental tool with known limits, not the point where those limits disappear.

That is a smaller ending than "the neural network learned to draft." It is also a system I can explain, inspect, and revise without pretending that every number beside a hero answers the same question.


Editorial Check: A Boundary Bug Found and Fixed

A September 10 source check found an additional integration mismatch. The public ranking method correctly used the number of heroes before adding a candidate. The old acceptance adapter counted the board after insertion. At four existing heroes and threshold five, the ranking method used the table while the adapter used the model. A small local check reproduced that difference using the exported artifact; no training or production change was made.

I corrected the adapter by making it share an explicit scoring contract with the public ranker: distinguish a bare context from a board containing a candidate, then select the arm using the original request's count. The threshold remains five. Regression tests cover mixed batches, rank settings, and both sides of the boundary; the switch check now also compares with public ranking calls.

This is a code-and-test fix, not a new deployment result or a rerun of the research gate. The old E41b numbers still belong to its old adapter, its failed guard remains failed, and the fresh-data requirement has not gone away.