I went to top up my training data, like a responsible adult. I came back with 26.8 million match records and considerably more questions than I'd started with.
These were records retrieved from OpenDota's public_matches table, not a
census of every Dota game played. Walking a provider's table doesn't establish
that the provider has the whole population. But it was still a rather large
change in what I could afford to investigate.
A Different Door
My old collector used /publicMatches, which returned at most a hundred
sampled matches per request. Under the allowance I was working with - three
thousand requests a day - collecting 26.8 million rows would consume roughly
three months of quota even at perfect yield, before duplicates or other API use.
More paging would not turn a sampled endpoint into a guaranteed complete feed.
I'd been improving the machinery around that limit: resumable collection,
backoff, quota tracking. Then I looked at /explorer.
It accepts read-only SQL against OpenDota's database. A bounded query returned 50,000 rows in one request, spending the same one-call allowance as a hundred-row page. Not every query returned that many, but the economics were suddenly different.
I rewrote collection as a backward walk through match-ID windows. The windows could be resumed, inserts tolerated duplicates, and a window hitting the row limit could be split rather than silently treated as complete. I requested Ranked All Pick records from normal or ranked lobbies, without filtering rank.
The first bulk attempt stopped partway through. The resumed run finished the walk, bringing the combined total to 26,797,863 records. About 100 minutes elapsed between the first start and the resumed completion, including the gap. The resumable machinery still mattered; it just had a much better source to work with.
The growing archive and the research dataset are different things. The first frozen snapshot contained 26,716,438 eligible matches, after applying the date window, complete-lineup and outcome checks, and a ten-minute minimum duration. The archive continued receiving top-ups; the snapshot stayed fixed, with recorded chronological train/validation/test splits and a content fingerprint.
That's the 26.8-to-26.7 million change: raw collection versus an explicitly selected research population, not two counts of the same object.
Bristleback Depends on the Room
The early experiments had been shaped by the data I happened to have. Now I could compare the same heroes across much larger rank cohorts. Before training anything, I wanted to look.
The first question was simple: how does a hero's observed win rate vary with the match's rank bracket?
For this comparison I used the snapshot's matches with known average rank, from Herald at the lower end through Divine at the higher end. The project's June-July cohort was labelled 7.41d using a time-based patch mapping, not a patch field in the bulk records. Read the rates as a photograph of this selected window, not permanent properties of the heroes.

Bristleback wins 53.4% of his appearances in Herald-bracket matches and 45.5% in Divine-bracket matches. His observed win rate declines at every intervening bracket too. Enigma goes the other way: 47.9% to 54.6%.
Better opponents might handle Bristleback more effectively; a team that follows up Enigma's Black Hole might get more from the pick. Those are plausible Dota stories, not mechanisms this comparison measured. The rates mix different players, roles, teammates, and opponents; they don't track what happens to one player as that player climbs the ladder.
Nor does the spread of individual hero win rates tell me how much predictive information the whole draft contains. This doesn't settle the explanation for Part 2's bracket results. It does give me a reason to test a model that knows the bracket, rather than pooling these populations and hoping one Bristleback estimate suits everyone.
I now had a dataset I could name and revisit, and enough coverage within it to ask more specific questions. How do these rates change over time? Which combinations matter beyond individual hero strength? Can knowing the bracket improve predictions? Final lineups were now plentiful; the actual order of draft picks still needed separate collection.
And one familiar face was doing very well in the popularity department. Pudge appeared in 7.54 million matches - about 28% of the snapshot. His win rate ran from 51.6% in Herald to 50.9% in Divine: a much less dramatic change than Bristleback's.
There is a lesson in there somewhere, and I intend to keep not learning it.