NeuroGolf Solver — What We Know
Scoring Formula (neurogolf_utils.py, May 14)
score = max(1.0, 25.0 - math.log(max(1.0, memory + params)))
memory= sum of intermediate tensor bytes from ORT Profiler. Excludes 'input' and 'output'.params= sum of initializer element counts + Constant node element counts.- Initializer BYTES are NOT memory. Only their element counts go to params.
input→Conv→outputhas memory=0. Cost = W.size + B.size = 910 for ks=3. Score = 18.19.- NO MACs in formula (removed May 4).
Submissions
| File | LB Score | Source |
|---|---|---|
| submission-6043.zip | 6043 | External. Not from our solver. |
| submission-5743.zip | 5743 | External. Partially reverse-engineered (wave22). |
| submission-v90.zip | 5322 | Our solver. |
| V3 (merged output) | 5560 | merge_best_of_both.py + stage1_trivial_optimizer.py |
6043 Model Classification (400 models)
| Pattern | Count | Score | Example Tasks |
|---|---|---|---|
| Direct Conv(W,B)→output | 22 | 18.19 | 15, 73, 95, 127, 171, 283... |
| Gather (color perm) | 12 | 22.7 | 16, 276, 309, 337... |
| GridSample | 9 | ~16.4 | 83, 108, 142, 307... |
| Conv+ReLU (two-layer) | 6 | varies | 4, 37, 90, 93... |
| fp16 various | ~80 | varies | scattered |
| Complex (10-14000 nodes) | ~270 | 12-15 | most tasks |
What Works
| Technique | Where | Result |
|---|---|---|
| Direct Conv→output (0 intermediates) | 6043 has 22 tasks | 18.19 pts each |
| Gather for color permutation | stage1_trivial_optimizer.py | 22.7 pts |
| Transpose→output | stage1_trivial_optimizer.py | 25.0 pts |
| Reverse-engineering into DSL | wave22.py (from 5743) | 143 tasks |
| lstsq Conv fitting | conv.py | 35 tasks |
| Merge lowest-cost per task | merge_best_of_both.py | +200 pts |
What Does NOT Work
| Technique | Why |
|---|---|
| fp16 weight conversion | Precision loss → fails arc-gen validation |
| SVD factorization | Approximation error → fails validation |
| Post-hoc ONNX optimization | Marginal gains, wrong approach |
| Merging with V90 | V90 worse than 5743 for 375/400 tasks |
| sign_corrected_conv.py fitting | Only works 3/18 ks=3 tasks, fails the rest |
Static Score Estimate vs Kaggle Reality
Static shape inference gives ~6168 for 5743, but Kaggle gives 5743. Gap of ~400 pts because ORT Profiler takes MAX tensor size across all examples. Static estimates are optimistic. Only use for relative comparison.
Key Mistakes (don't repeat)
- Trying to "optimize" ONNX models without understanding the underlying task logic
- fp16 conversion (always breaks arc-gen)
- Using intermediate count as a proxy for cost (wrong — use measure_cost with shape inference)
- Writing optimizer scripts instead of solver code
- Adding version numbers to filenames (use git)