Steam Neural Recommender: evaluation card
Temporal top-K recommendation on the Kaggle Steam recommendations archive (v28, CC0).
Every number below is regenerated from retained run manifests for split 70edc4434b998994;
the test split is scored once per final seed and never tuned on.
Two-tower test NDCG@10
0.0279
3 seeds
Two-tower test Recall@10
0.0554
held-out next positive
CPU retrieval p50
15.9 ms
AMD64 Family 25 Model 97 Stepping 2, AuthenticAMD
Recent-popularity test NDCG@10
0.0312
strongest zero-parameter baseline
Preregistered hypotheses
- H1: held. two-tower beats MF on test NDCG@10 with a paired bootstrap interval above zero for every seed.
- H2: failed. both learned models beat recent popularity on test NDCG@10 and Recall@10.
- H3: held. on held-out items released after the cutoff, the metadata-aware two-tower scores above zero Recall@10 and MF does not exceed the popularity fallback.
- H4: failed. the hybrid policy (recent popularity below a validation-chosen history threshold, two-tower otherwise) beats recent popularity on test NDCG@10 with a paired bootstrap interval above zero for every seed and on mean Recall@10, and keeps Coverage@10 above 0.05.
- operational: held. single-query CPU retrieval p50 under 50 ms and item index under 20 MB.
Test split: NDCG@10 by model (mean over seeds, whisker = ±1 std)
Test split: all headline metrics
| Model | Seeds | NDCG@10 | Recall@10 | Coverage@10 | Popularity bias@10 | CPU p50 ms | Params |
|---|---|---|---|---|---|---|---|
| Global popularity | 1 | 0.0175 | 0.0373 | 0.0006 | 0.9995 | 0.4289 | 0 |
| Recent popularity (90 d) | 1 | 0.0312 | 0.0635 | 0.0003 | 0.9978 | 0.4096 | 0 |
| Matrix factorisation (BPR) | 3 | 0.0188 ± 0.0005 | 0.0413 ± 0.0014 | 0.0020 ± 0.0002 | 0.9994 ± 0.0000 | 0.7633 ± 0.0385 | 831,292,024 |
| Two-tower retriever | 3 | 0.0279 ± 0.0006 | 0.0554 ± 0.0009 | 0.1653 ± 0.0026 | 0.9611 ± 0.0008 | 15.9381 ± 0.1267 | 6,557,184 |
| Hybrid (recent popularity below threshold, else two-tower) | 3 | 0.0279 ± 0.0006 | 0.0554 ± 0.0009 | 0.1653 ± 0.0026 | 0.9611 ± 0.0008 | n/a | 6,557,184 |
Test split: slices
Cold start (Recall@10)
| Model | cold_start | has_history |
|---|---|---|
| Global popularity | 0.0938 n=5000 | 0.0232 n=20000 |
| Recent popularity (90 d) | 0.0938 n=5000 | 0.0559 n=20000 |
| Matrix factorisation (BPR) | 0.0938 ± 0.0000 n=5000 | 0.0282 ± 0.0017 n=20000 |
| Two-tower retriever | 0.0938 ± 0.0000 n=5000 | 0.0459 ± 0.0011 n=20000 |
| Hybrid (recent popularity below threshold, else two-tower) | 0.0938 ± 0.0000 n=5000 | 0.0459 ± 0.0011 n=20000 |
History length (Recall@10)
| Model | long_10_plus | medium_3_9 | short_1_2 |
|---|---|---|---|
| Global popularity | 0.0102 n=3935 | 0.0219 n=7794 | 0.0544 n=13271 |
| Recent popularity (90 d) | 0.0384 n=3935 | 0.0531 n=7794 | 0.0771 n=13271 |
| Matrix factorisation (BPR) | 0.0152 ± 0.0017 n=3935 | 0.0249 ± 0.0007 n=7794 | 0.0587 ± 0.0017 n=13271 |
| Two-tower retriever | 0.0423 ± 0.0011 n=3935 | 0.0500 ± 0.0019 n=7794 | 0.0625 ± 0.0003 n=13271 |
| Hybrid (recent popularity below threshold, else two-tower) | 0.0423 ± 0.0011 n=3935 | 0.0500 ± 0.0019 n=7794 | 0.0625 ± 0.0003 n=13271 |
History polarity (Recall@10)
| Model | has_negatives | positive_only |
|---|---|---|
| Global popularity | 0.0189 n=6557 | 0.0439 n=18443 |
| Recent popularity (90 d) | 0.0433 n=6557 | 0.0707 n=18443 |
| Matrix factorisation (BPR) | 0.0225 ± 0.0010 n=6557 | 0.0480 ± 0.0015 n=18443 |
| Two-tower retriever | 0.0389 ± 0.0012 n=6557 | 0.0613 ± 0.0008 n=18443 |
| Hybrid (recent popularity below threshold, else two-tower) | 0.0389 ± 0.0012 n=6557 | 0.0613 ± 0.0008 n=18443 |
Held-out item popularity (Recall@10)
| Model | head_top10pct | tail |
|---|---|---|
| Global popularity | 0.0482 n=19350 | 0.0000 n=5650 |
| Recent popularity (90 d) | 0.0821 n=19350 | 0.0000 n=5650 |
| Matrix factorisation (BPR) | 0.0534 ± 0.0018 n=19350 | 0.0000 ± 0.0000 n=5650 |
| Two-tower retriever | 0.0694 ± 0.0012 n=19350 | 0.0076 ± 0.0005 n=5650 |
| Hybrid (recent popularity below threshold, else two-tower) | 0.0694 ± 0.0012 n=19350 | 0.0076 ± 0.0005 n=5650 |
Held-out item release cohort (Recall@10)
| Model | catalogue_before_cutoff | released_after_cutoff |
|---|---|---|
| Global popularity | 0.0477 n=19560 | 0.0000 n=5440 |
| Recent popularity (90 d) | 0.0786 n=19560 | 0.0094 n=5440 |
| Matrix factorisation (BPR) | 0.0528 ± 0.0018 n=19560 | 0.0001 ± 0.0001 n=5440 |
| Two-tower retriever | 0.0700 ± 0.0012 n=19560 | 0.0031 ± 0.0006 n=5440 |
| Hybrid (recent popularity below threshold, else two-tower) | 0.0700 ± 0.0012 n=19560 | 0.0031 ± 0.0006 n=5440 |
Validation split (hyperparameter selection only)
| Model | Seeds | NDCG@10 | Recall@10 | Coverage@10 | Popularity bias@10 | CPU p50 ms | Params |
|---|---|---|---|---|---|---|---|
| Global popularity | 1 | 0.0193 | 0.0408 | 0.0007 | 0.9995 | 0.4289 | 0 |
| Recent popularity (90 d) | 1 | 0.0253 | 0.0573 | 0.0004 | 0.9978 | 0.4096 | 0 |
| Matrix factorisation (BPR) | 3 | 0.0213 ± 0.0004 | 0.0460 ± 0.0011 | 0.0026 ± 0.0001 | 0.9994 ± 0.0000 | 0.7633 ± 0.0385 | 831,292,024 |
| Two-tower retriever | 3 | 0.0367 ± 0.0007 | 0.0681 ± 0.0011 | 0.1641 ± 0.0025 | 0.9584 ± 0.0008 | 15.9381 ± 0.1267 | 6,557,184 |
| Hybrid (recent popularity below threshold, else two-tower) | 3 | 0.0367 ± 0.0007 | 0.0681 ± 0.0011 | 0.1641 ± 0.0025 | 0.9584 ± 0.0008 | n/a | 6,557,184 |