My Magic: The Gathering cube analyzer was finally done. It recommends cards for my cube - a hand-picked collection my friends and I use to build decks for game night. I'd just built a scoring engine that could either use pre-computed values or recalculate them live, and put it at mtg.scottliu.com/cube/recommend. Finished, I thought.

And then the elephants showed up. Or rather, they didn't.

The card was "Call of the Herd," a green spell that makes elephants. For some reason, I found myself staring at its art - two elephants in a forest - for an uncomfortably long time. Their eyes started to look... creepy. Sentient, even.

The Magic card Call of the Herd: a green sorcery whose art shows two elephants walking through a forest, facing the viewer.

The Crime Scene

The "Card Lookup" tool gave Call of the Herd a very respectable 0.890. Higher scores mean a better fit for my cube: popular cards, manageable rules text, and a price I can live with.

Then I went to "Recommended Adds" and filtered by green. The recommendations ran from 0.925 down to 0.813. At 0.890, Call of the Herd should have been sitting comfortably in the middle - unless some filter had ruled it out.

But it wasn't there.

The Card Lookup panel open over the Recommended Adds grid. The panel shows Call of the Herd with a score of 0.890. Behind it, the recommendation cards visible on the same screen show Tamiyo's Safekeeping at 0.891 and Lotus Cobra at 0.888.
Lookup says 0.890; the cards it appears to belong between say 0.891 and 0.888. Either something excluded it, or those scores aren't comparable.

I put on my detective hat, poured a cup of coffee, and began the investigation.

The Investigation

Suspect #1: A Simple Rule-Breaker? I had a long list of house rules hard-coded into the analyzer: no vintage frames, no multi-faced cards, no cards with certain banned keywords. Was Call of the Herd secretly one of the offenders?

I wrote a diagnostic script to check it against the exclusion rules. Among the checks:

  • Not already in the cube (is_in_cube = 0)? Pass.
  • Modern frame (frame = '2015')? Pass.
  • Only one face (num_faces = 1)? Pass.

The checks came back clean. The card had an alibi.

Suspect #2: The Heuristic Dragnet? My next suspect was my own optimization. The precomputed "Adds" path takes a shortlist of the top 500 candidates before applying the slow text filters. Could Call of the Herd be just below that cutoff? I cranked CANDIDATE_POOL_SIZE up to 5,000 and refreshed.

Still no elephants. Casting a net ten times wider hadn't changed the result.

Suspect #3: A Case of Mistaken Identity? Maybe the database didn't agree that the card was green. I opened the SQLite console and checked its vitals: color_identity: 'G'. Exactly who it said it was.

A Clue in the Logs

The exclusion checks weren't explaining it. I added logging to print the final SQL query, rather than keep reasoning about what I thought the application was asking the database.

While looking for that output, I opened the systemd service logs for the first time in a while. There was a warning I'd overlooked:

(!) WARNING: Formula has changed. Using slower 'Live Tuning Mode'.

The call was coming from inside the house.

The engine stores a hash - a fingerprint of the scoring settings - alongside its precomputed scores. If the current settings have a different fingerprint, those scores are stale and the app needs to calculate fresh ones.

Three things had happened:

  1. The Motive: I'd made the scoring settings stricter, but hadn't rerun the bootstrap script to rebuild the stored scores.
  2. The Live Reality: The "Adds" list checked the hash, noticed the change, and calculated scores using the new settings.
  3. The Blind Spot: Card Lookup didn't check. I'd written it as a simple database fetch, so it displayed the old, stored 0.890.

I wasn't comparing one score to a list of scores. I was comparing a score from one reality to a list of scores from another.

That also explained why widening the shortlist did nothing: live tuning bypasses the 500-card cutoff entirely. I had been adjusting a knob on a path the request wasn't taking.

I reran the bootstrap. Call of the Herd's score under the current settings was 0.810, below the displayed recommendations. The list had been right about where it belonged; the lookup had been giving it a glowing reference from its previous employer.

The Card Lookup panel after re-running the bootstrap. Call of the Herd now scores 0.810, and its Card Strength bar has turned amber where it was previously green.
The same lookup after re-running the bootstrap. 0.810, and the Card Strength bar has gone amber - the score the list had been using all along.

The Resolution

Rebuilding the scores cleared the symptom. It did not fix the lookup: the next scoring adjustment would put the two screens out of step again.

The code fix was to move the hash check into one function that both paths call, and make Card Lookup recalculate a card's scores when that check says the stored values are stale. For the same card data and scoring settings, lookup and recommendation must agree on the score. Fetching a stored number is only valid after checking which settings produced it.

A regression test should exercise the transition, not just the freshly built database: store a card's score, change a scoring weight without rebuilding, then ask both paths for that card. Both should return its newly calculated score, not merely agree with each other about the old one. Rebuild with the new settings and check again through the cached path. That's the case a test of the scoring formula alone would miss.

I had spent three hours checking why the card was excluded, when the wrong number was in the tool I was using to check the list. I was checking the wrong thing carefully.

The warning had also been sitting in my service logs the whole time, which is its own lesson about logs you've stopped reading.

And as for the elephants? Their eyes don't look so creepy anymore. Now I can finally get back to the actual work of sleeving cards.