Self-distillation NNUE: the student/teacher transfer ratio drops every generation and plateaus after three

Discussion of chess software programming and technical issues.

Moderator: Ras

Spunc595
Posts: 35
Joined: Mon Jul 06, 2026 12:15 am
Full name: Daniele Marpino

Self-distillation NNUE: the student/teacher transfer ratio drops every generation and plateaus after three

Post by Spunc595 »

Hi everyone,

I've been bootstrapping an NNUE for my engine with a strict self-imposed rule: every position and label must come from my engine's own search. No external evaluator anywhere in the pipeline. I've now hit a wall I can measure but can't really explain, and I'd rather ask here than keep losing nights over it.

Setup per generation: self-play at 3k nodes, ~3M unique positions, labels (eval + best move) from my search at 20k nodes using the previous generation's net as teacher. Net architecture is 768x4 king buckets -> 1024, SCReLU, akimbo-style. Training saturates very quickly and early-stops.

The ladder flattens fast:

Code: Select all

           rho vs Stockfish      step
        (static, same 2,000 positions)

  gen1        0.587
  gen2        0.679             +0.092
  gen3        0.700             +0.021
It looks like the student is capturing less of its teacher with each step. Teacher measured in search at 20k nodes, student measured static:

Code: Select all

  gen2:   0.679 / 0.825  =  0.823
  gen3:   0.700 / 0.876  =  0.800
So the teacher gains +0.051, but the student only gets +0.021. My guess is that as the teacher gets stronger, more of its strength comes from search/tree expansion, which a static net just can't represent — making a growing part of the label look like pure noise to the student. But that's just a guess.

Two things I already tested and ruled out:

Code: Select all

  Capacity:     1024 -> 2048 hidden = 0.695 (no gain, actually slightly
                worse) and 26% lower nps. Same dataset, only L1 changed.

  Data volume:  val loss bottoms out at epoch 3 with 3.1M positions,
                and was epoch 4 with 2.97M. More data just meant faster
                saturation.
And the cost in Elo is heavy (measured on the same binary, only changing the net file, 500 games per pairing at 20+0.2):

gen3 is -759 ± 58 against the akimbo net, and -277 ± 36 against my own old net that was trained on external labels. gen3 did not win a single one of the 500 games against akimbo.

A few questions for the forum:

1. Is this declining student/teacher transfer ratio something you've seen before? Is it just how self-distillation behaves, or a sign that I'm doing something wrong?

2. What scale does this usually need (positions per iteration, total iterations)? I'm starting to think 3M per generation is off by an order of magnitude or two.

3. Quiet-position filtering: I currently only discard positions in check. Around 19.7% of my positions have a capture as best move. Is filtering those out standard practice, and does it actually make a big difference?

4. Is it better to label the static eval of the PV leaf instead of the root search score? That would address the "unlearnable search noise" problem, but maybe I'm just labeling the wrong node.

Everything is open source, including the raw measurements and a couple of predictions I registered in advance and got completely wrong: https://github.com/Spunc595/Luna-CE-NNUE

Thanks,

Daniele
Aleks Peshkov
Posts: 1012
Joined: Sun Nov 19, 2006 9:16 pm
Location: Russia
Full name: Aleks Peshkov

Re: Self-distillation NNUE: the student/teacher transfer ratio drops every generation and plateaus after three

Post by Aleks Peshkov »

For a simplest 1 hidden layer net it is commonly suggested to have 1M positions in dataset per neuron. King input buckets need magnitude more.
Spunc595
Posts: 35
Joined: Mon Jul 06, 2026 12:15 am
Full name: Daniele Marpino

Re: Self-distillation NNUE: the student/teacher transfer ratio drops every generation and plateaus after three

Post by Spunc595 »

Thanks — that rule of thumb turned out to be the piece I was missing.
I'd trained three generations of my own net for Luna and none came close to the akimbo net the engine currently ships (MIT, with attribution). The dataset had 11M positions for a 1024-neuron net — about 1% of the ~1B your rule implies. That alone was enough to explain it.
Petrel was the other half of the answer. The recipe is documented in net/petrel.rs, which made it directly usable rather than something to guess at, and it gave me an existence proof that the architecture wasn't the limit.
So I switched to bullet with Linrock's filtered T80 data. Where it stands, measured as Spearman rank correlation against 435k Stockfish 18 depth-12 evals:

previous best (own net) 0.8253
37M samples (13 min) 0.8687
250M 0.8870
1G 0.8952
akimbo net (currently used) 0.9036

An 8-epoch run over the 1B set is going now.
To be clear about what that is: static correlation with Stockfish, not playing strength. None of these nets has played a game yet. The SPRT comes after.
One question, if you have a view: if that curve flattens short of the target, my guess is it's capacity rather than data — Petrel and my net are plain 768 inputs, while the akimbo net uses 768×4 king buckets, so four times the input parameters. Is that where you'd look first, or is 1B simply still too small to tell?
Aleks Peshkov
Posts: 1012
Joined: Sun Nov 19, 2006 9:16 pm
Location: Russia
Full name: Aleks Peshkov

Re: Self-distillation NNUE: the student/teacher transfer ratio drops every generation and plateaus after three

Post by Aleks Peshkov »

Training or validation loss nor correlation do not work, you need to switch to direct metric -- whole engine playing strength. I suggest to build the simplest reasonably small net: plain dual perspective with 32/64 neurons to make sure that further complexities pay off in practice.
Spunc595
Posts: 35
Joined: Mon Jul 06, 2026 12:15 am
Full name: Daniele Marpino

Re: Self-distillation NNUE: the student/teacher transfer ratio drops every generation and plateaus after three

Post by Spunc595 »

Thanks, both for the earlier rule of thumb and for this. The first shaped
how much data I collected; the second matches what I found the hard way
over the last two weeks. Some numbers, in case they're useful, and the
result of the match I promised.

You're right that correlation stops working. I used Spearman rank
correlation against Stockfish labels as my gate, and it carried me a long
way: from 0.8253 to 0.9043 on my eval set. But above roughly 0.90 it goes
almost flat while playing strength keeps moving. One of my nets sat 0.0011
below akimbo's network on that statistic (0.9026 vs 0.9036) and lost by
about 50 Elo in play. Two of my own nets, 0.9026 and 0.9043, differ by
0.0017 in Spearman and were about 37 Elo apart in an SPRT. Rank correlation
is invariant to any monotone transformation of the output, so it is
structurally blind to a whole class of differences that decide games. I
now treat it as an entry filter ("is this net broken?") and nothing more.

One thing I'd add to your advice, because it nearly invalidated my
measurements. Before you can compare nets by game strength, you have to
check each net's output scale against the engine's fixed centipawn
constants. My training recipe (a WDL ramp) inflated the output by a
factor of about 1.116 relative to akimbo's network: 1.1156 and 1.1164 on
two different architectures trained with the same recipe, so I attribute
it to the recipe, not to the net. That factor is invisible to rank
correlation. Correcting a single constant (SCALE 400 -> 358) was worth
+15.9 Elo, SPRT accepted, LOS 99.5%. I haven't isolated which search
constants account for it, but without that correction a ladder of nets
compared by game strength measures calibration rather than architecture.

On the rule of thumb: I tested a corollary of it and it predicted the
wrong thing, though the extension was mine, not yours. I went from 768
inputs to 768x4 king buckets on the same 1B distinct positions. By the
million-per-neuron logic each bucket then sees only ~250M positions, so I
expected it to need more data before it could pay. It won by +36.9 +/- 20.5
Elo (SPRT H1 accepted in 529 games).

Result: a direct fixed-length match against akimbo's network running in my
engine, 2,000 games at 10+0.1: +492 =1018 -490, +0.3 +/- 10.7 Elo (95% CI),
zero losses on time. So a net trained on my own 1B distinct positions, with
768x4 king buckets and the output scale corrected, is statistically level
with it. Two caveats: the interval runs from roughly -10 to +11, so "level"
is the honest claim and not "better", and it is a single training run.

Where I'd hesitate on the 32/64-neuron net: for me the bottleneck was never
the cost of measuring. My SPRTs resolved in between 1.5 and 6 hours on 4
Arm cores, which is what makes direct strength measurement affordable at
full size. A minimal net would be cheaper to train, but it would answer
questions about a regime the engine doesn't operate in, and it wouldn't
have surfaced the scale issue above.

Next I'll try more distinct positions and GPU training. Thanks again for
the pointers.
Spunc595
Posts: 35
Joined: Mon Jul 06, 2026 12:15 am
Full name: Daniele Marpino

Re: Self-distillation NNUE: the student/teacher transfer ratio drops every generation and plateaus after three

Post by Spunc595 »

Correction: the 1B positions are Linrock's public Leela-derived data, not my own
Aleks Peshkov
Posts: 1012
Joined: Sun Nov 19, 2006 9:16 pm
Location: Russia
Full name: Aleks Peshkov

Re: Self-distillation NNUE: the student/teacher transfer ratio drops every generation and plateaus after three

Post by Aleks Peshkov »

Linrock dataset has 12B augmented (reshuffled) positions from 1B of unique positions. 12B > 1B, even if 12 are not perfectly unique.
Spunc595
Posts: 35
Joined: Mon Jul 06, 2026 12:15 am
Full name: Daniele Marpino

Re: Self-distillation NNUE: the student/teacher transfer ratio drops every generation and plateaus after three

Post by Spunc595 »

Thanks, both for the earlier rule of thumb and for this. The first shaped
how much data I used; the second matches what I found the hard way over the
last two weeks. Some numbers, in case they're useful, and the result of the
match I promised.

You're right that correlation stops working. I used Spearman rank
correlation against Stockfish labels as my gate, and it carried me a long
way: from 0.8253 to 0.9043 on my eval set. But above roughly 0.90 it goes
almost flat while playing strength keeps moving. One of my nets sat 0.0011
below akimbo's network on that statistic (0.9026 vs 0.9036) and lost by
about 50 Elo in play. Two of my own nets, 0.9026 and 0.9043, differ by
0.0017 in Spearman and were about 37 Elo apart in an SPRT. Rank correlation
is invariant to any monotone transformation of the output, so it is
structurally blind to a whole class of differences that decide games. I
now treat it as an entry filter ("is this net broken?") and nothing more.

One thing I'd add to your advice, because it nearly invalidated my
measurements. Before you can compare nets by game strength, you have to
check each net's output scale against the engine's fixed centipawn
constants. My training recipe (a WDL ramp) inflated the output by a
factor of about 1.116 relative to akimbo's network: 1.1156 and 1.1164 on
two different architectures trained with the same recipe, so I attribute
it to the recipe, not to the net. That factor is invisible to rank
correlation. Correcting a single constant (SCALE 400 -> 358) was worth
+15.9 Elo, SPRT accepted, LOS 99.5%. I haven't isolated which search
constants account for it, but without that correction a ladder of nets
compared by game strength measures calibration rather than architecture.

On the rule of thumb: I tested a corollary of it and it predicted the
wrong thing, though the extension was mine, not yours. I went from 768
inputs to 768x4 king buckets on the same 1B positions. By the
million-per-neuron logic each bucket then sees only ~250M positions, so I
expected it to need more data before it could pay. It won by +36.9 +/- 20.5
Elo (SPRT H1 accepted in 529 games).

Result: a direct fixed-length match against akimbo's network running in my
engine, 2,000 games at 10+0.1: +492 =1018 -490, +0.3 +/- 10.7 Elo (95% CI),
zero losses on time. So a net trained by me on 1B positions per epoch from
Linrock's public Leela-derived bullet data, with 768x4 king buckets and the
output scale corrected, is statistically level with it. Two caveats: the
interval runs from roughly -10 to +11, so "level" is the honest claim and
not "better", and it is a single training run.

On your point about Linrock's dataset: I counted. Of the 1B positions I fed
to the trainer, 914,416,689 are unique by exact board + side-to-move match
(91.4%); 910,466,531 (91.0%) if a position and its mirror image are treated
as the same. That counts exact matches only, so it says nothing about
near-duplicates from the same games; if that's what you meant by augmented,
I'd be glad to hear how you'd measure it. I've dropped "distinct" from my
own wording.

Where I'd hesitate on the 32/64-neuron net: for me the bottleneck was never
the cost of measuring. My SPRTs resolved in between 1.5 and 6 hours on 4
Arm cores, which is what makes direct strength measurement affordable at
full size. A minimal net would be cheaper to train, but it would answer
questions about a regime the engine doesn't operate in, and it wouldn't
have surfaced the scale issue above.

Next I'll try more unique positions and GPU training. Thanks again for the
pointers.