Mindchess: raw, aggregate and hashed perft on Ryzen 9 9950X3D

Discussion of chess software programming and technical issues.

Moderator: Ras

vtlmks
Posts: 8
Joined: Fri Aug 07, 2026 4:59 am
Full name: Peter Fors

Mindchess: raw, aggregate and hashed perft on Ryzen 9 9950X3D

Post by vtlmks »

Mindchess: raw, aggregate and hashed perft on Ryzen 9 9950X3D

I think this will be the last Mindchess update for a while, so I wanted to collect the current results in one place.

Mindchess was never intended to become a chess engine. I don't have a particularly strong interest in chess itself; I started this because perft looked like a good optimization target. The original goal was simply to see how far I could push raw move generation, with Chessbit as the obvious program to compare against.

I reached that goal, but then stayed around long enough to optimize the other two variants as well: depth-2 aggregate counting and transposition-table perft.

Raw single-threaded perft

This is still what I consider the primary result. There is no transposition table or subtree reuse here; every legal leaf is actually enumerated.

The results are medians of five runs on my Ryzen 9 9950X3D:

Code: Select all

Position       Depth   Mindchess MNPS   Chessbit MNPS   Lead
--------------------------------------------------------------
Start             7          3680.26          2560.29   +43.744%
Kiwipete          6          4662.31          3455.52   +34.924%
Position 3        8          3015.91          2068.48   +45.803%
Position 4        6          3821.15          2956.15   +29.261%
Position 5        6          3882.35          2912.85   +33.284%
Position 6        6          5241.79          4173.73   +25.590%

Geometric mean                                          +35.241%
Rates are raw MNPS.

Raw multithreaded perft

The multithreaded version runs the same raw enumeration across workers, without a TT.

Start position, depth 9, 16 workers:

Code: Select all

Nodes:       2,439,530,234,167
Time:        43.758 s
Rate:        55,749.87 raw MNPS
Workers:     16
Split depth: 3
TT:          none
That corresponds to a 14.55x speedup and 90.93% scaling efficiency across 16 physical cores.

Depth-2 aggregate counting

After finishing the raw version I also experimented with avoiding work at depth 2 by proving groups of replies without constructing every child position.

For this comparison both Mindchess and Chessbit were built with ICX 2026.0.0, without PGO, and Chessbit's transposition table was disabled. The values are medians of five alternating runs:

Code: Select all

Position       Depth   Mindchess MNPS   Chessbit MNPS   Lead
--------------------------------------------------------------
Start             7          8641.51          7912.69    +9.211%
Kiwipete          6          7016.32          6896.50    +1.737%
Position 3        8          3276.70          3115.89    +5.161%
Position 4        6          4793.57          4474.56    +7.129%
Position 5        6          6441.44          5959.36    +8.089%
Position 6        6          9119.73          8835.76    +3.214%

Geometric mean                                           +5.724%
Again, these are MNPS, but unlike the raw numbers this version avoids some work at depth 2, so the rates should not be compared directly with raw MNPS.

Hashed / transposition-table perft

The final experiment was a multithreaded transposition-table version. This is a different benchmark again: transposed subtrees are reused, so the reported rate is logical nodes divided by elapsed time rather than physically enumerated leaves.

Start position depth 10, 69,352,859,712,417 logical nodes, 32 workers and an 8 GiB TT:

Code: Select all

Engine                 Build       Time       Logical MNPS
----------------------------------------------------------

Chessbit / ICX          non-PGO    12.020 s    5,769,788.66
Mindchess / Clang       non-PGO    10.204 s    6,796,816.11
Mindchess / Clang       PGO         9.579 s    7,240,438.70
The directly comparable non-PGO Mindchess result uses 15.11% less elapsed time than Chessbit and has 17.80% greater logical throughput.

Thomas deserves credit here. Chessbit provided the target that made this interesting in the first place, and he pointed me toward the depth-2 approach. Studying Chessbit also gave me useful ideas around recursive batching and cache-line bucket organization.

At this point I think I have done what I came for. I started out wanting to beat the raw implementation, achieved that, and then ended up staying around long enough to beat the aggregate and hashed versions as well.

There are certainly more things that could be optimized, but for now I feel done with the experiment.

I've cleaned up the repository so it now contains the three implementations separately, along with the build scripts, reproduction commands and transposition-table documentation:

https://github.com/vtlmks/mindchess

The two earlier threads contain the development history:

Raw perft:
https://talkchess.com/viewtopic.php?t=86611

Hashed perft:
https://talkchess.com/viewtopic.php?t=86623
chessbit
Posts: 50
Joined: Fri Dec 29, 2023 4:47 pm
Location: Belgium
Full name: thomas huijbregts

Re: Mindchess: raw, aggregate and hashed perft on Ryzen 9 9950X3D

Post by chessbit »

I took some time to compare the builds, but I was not able to get the increase in performance from your version. I used the same build options as I did with my engine (with PGO).
Today, I also improved by adding large pages in my TT (I found the reason why it was slower when I did in the past, and I had discarded it) and improved my logic a bit, which surprisingly gave me a significant ~6.5% performance boost with PGO (by adding a single condition for my pins, which was branchless before).

Anyway, I digress. I compared the numbers, and while your version (aggregate one) had somewhat similar numbers, now chessbit is clearly faster on my machine:

Mindchess:

Code: Select all

Perft Start   d7 :     3195901860       409 ms   7817.25 MNodes/s  OK
Perft Kiwi    d6 :     8031647685      1360 ms   5907.22 MNodes/s  OK
Perft Midgame d6 :     6923051137       860 ms   8049.96 MNodes/s  OK
Perft Endgame d7 :    24958831314      4307 ms   5795.47 MNodes/s  OK

Perft aggregate: 43109431996  6935 ms  6216.14 MNodes/s
Chessbit:

Code: Select all

Perft Start 1: 20 0ms 10 MNodes/s
Perft Start 2: 400 0ms 18.1818 MNodes/s
Perft Start 3: 8902 0ms 240.595 MNodes/s
Perft Start 4: 197281 0ms 2293.97 MNodes/s
Perft Start 5: 4865609 0ms 6683.53 MNodes/s
Perft Start 6: 119060324 16ms 7425.49 MNodes/s
Perft Start 7: 3195901860 399ms 7999.53 MNodes/s
OK

Perft Kiwi 1: 48 0ms 16 MNodes/s
Perft Kiwi 2: 2039 0ms 679.667 MNodes/s
Perft Kiwi 3: 97862 0ms 3624.52 MNodes/s
Perft Kiwi 4: 4085603 0ms 6209.12 MNodes/s
Perft Kiwi 5: 193690690 23ms 8131.09 MNodes/s
Perft Kiwi 6: 8031647685 1185ms 6774.94 MNodes/s
OK

Perft Midgame 1: 46 0ms inf MNodes/s
Perft Midgame 2: 2079 0ms 1039.5 MNodes/s
Perft Midgame 3: 89890 0ms 4731.05 MNodes/s
Perft Midgame 4: 3894594 0ms 8304.04 MNodes/s
Perft Midgame 5: 164075551 19ms 8598 MNodes/s
Perft Midgame 6: 6923051137 823ms 8403.12 MNodes/s
OK

Perft Endgame 1: 38 0ms inf MNodes/s
Perft Endgame 2: 1129 0ms 1129 MNodes/s
Perft Endgame 3: 37035 0ms 193.901 MNodes/s
Perft Endgame 4: 1023977 0ms 5851.3 MNodes/s
Perft Endgame 5: 31265700 4ms 6478.6 MNodes/s
Perft Endgame 6: 849167880 135ms 6278.13 MNodes/s
Perft Endgame 7: 24958831314 3896ms 6404.75 MNodes/s
OK

Perft aggregate: 43109431996 6517ms 6614.3 MNodes/s
I'm not sure how to get a definitive idea of the fastest version. It seems the quality of the executable is very dependent on the compiler (+ options) and the CPU used, making the comparison difficult.

PS: I lost access to my old github. I have added the latest code on my new account (see signature)
https://github.com/thomashuijbregts

chessbit - fast CPU perft engine
Crownd - strong UCI engine with NNUE evaluation
vtlmks
Posts: 8
Joined: Fri Aug 07, 2026 4:59 am
Full name: Peter Fors

Re: Mindchess: raw, aggregate and hashed perft on Ryzen 9 9950X3D

Post by vtlmks »

Excellent, this will be fun to investigate. I’ll compare your latest version with the one I tested and see whether I can reproduce the difference. The compiler, PGO workload, and CPU clearly matter quite a bit here. Thanks for publishing the updated code!