Are the bots fixed strength or do they learn and improve with experience?

Public forum discussion.

#1 · tiyusufaly

5 Jun 2024

Maybe it’s my imagination, but I think even the weaker bots have gotten stronger in the few weeks since they were introduced. Are they self-learning and dynamic, then? Or are the weights retrained and fixed?

#2 · tiyusufaly

5 Jun 2024

“Pretrained and fixed”*

#3 · hexanna

6 Jun 2024

It seems like even the weaker bots play very strong in the opening, often stronger than top human players. They tend to fall completely apart and blunder in the middlegame or endgame though. There might be a bug with KaTrain when used for hex.

#4 · Arek Kulczycki

7 Jun 2024

I think it's not objectively clear what it means to play strong in opening. I suppose what you mean is that weaker bots choose the same initial moves as the stronger bots. In fact the same is true about low rated and high rates human players.

#5 · tiyusufaly

7 Jun 2024

Is that really true, though? I know that in playing hexanna, I've definitely learned a lot about some principles of 'smart' opening play that I certainly didn't know before. For example:
- Don't play near your own edge for the first 30-40 moves unless you have a very good reason.
- Don't rush to get the fastest possible connection, but space your moves out and mini-max.
- Defend at a distance, make each move as 'efficient' as possible, with multiple threats both offensive and defensive.
- Go for the 4-4 obtuse corners and 5-5 acute corners.

#6 · add3993

7 Jun 2024

I think it's not objectively clear what it means to play strong in opening. I suppose what you mean is that weaker bots choose the same initial moves as the stronger bots.

This isn't all they mean or all we should infer. Suppose we have the hypothesis H0 that, while KataHex is very strong overall, its first ~10 moves are just educated guesses and not much better than some other strategy (provided by a human or simple bot). ...then, there is a natural "hybrid" strategy in which the other strategy is played for the first 10 moves, then KH takes over.

This hybrid can be empirically tested against pure KH, a matchup which also tests the hypothesis H0. I think pure KH will do better. The reason is that I think (1) the style of play in the opening is important, (2) it is very hard to determine best opening play by pure case-analysis and logical deduction, or any well-explained heuristic known to the community, but (3) good opening play is characterized by general patterns of the sort that a deep convolutional neural network can learn.

Based on the above, I treat KH's opening style as a provisional best-guess about good form. It is still useful for humans to learn and describe this style, and even to ask questions like "but why not play this other way instead?".

In fact the same is true about low rated and high rates human players.

I don't think this is true, not when one looks closely at many human games (esp. with AI assistance). Bearing in mind that KataHex not only has a score-evaluation function, but can also play against itself from any given position to self-assess the accuracy of this scoring.

#7 · hexanna

7 Jun 2024

Yes, "play very strong" means if you analyze the moves with full-strength KataHex, even the weaker bots tend to play moves with win percentages quite close to what KataHex thinks is the best move. I am specifically referring to 19x19. (I don't think this is true for the weakest 3 or 4 bots which do play noticeably worse than strong humans in the opening.)

Go for the 4-4 obtuse corners and 5-5 acute corners.

Close, but to be precise, you should play the 5-4 acute corner slightly closer to your opponent's edge, like d5 for Red/Black and e4 for Blue/White. https://www.hexwiki.net/index.php/Corner_move

#8 · Arek Kulczycki

8 Jun 2024

I'm not quite convinced that what KataHex does in opening is the strongest play, talking about 19x19 in particular. The opening play in general terms is very simple - one just occupies impactful cells like those that extended end up connecting to your edge (either bridge-by-bridge or cell-by-cell along the diagonal). The difficult part is to understand how subtleties from opening impact the following part of the game. It's something that a human (me at least) tries to generalize by creating variety of principles. An engine on the other hand is super strong about it because good play gets encoded into the network based on factual self-play statistics.

If we nonetheless use the engine to evaluate the opening moves I think we find lower rated players do equally well to higher rated ones. To use an example, let's take current championship and games between the top 2 and the bottom 2 players. Intuitively I don't see a difference in quality between those and KataHex might even find the lower rated game more precise because it follows a more conventional approach.

@tiyusufaly
I don't know if your initial question got answered: The bots are fixed as they are until someone manually uploads different ones. It's possible that someone changes a model underneath but keeps the bot name, in such case it would look as if the bot improved. However they don't learn from LG games.

#9 · tiyusufaly

9 Jun 2024

Thanks Arek. I’m guessing that’s what’s happening, because the bots do seem to be getting stronger.

#10 · tiyusufaly

9 Jun 2024

FWIW, I’ve started to notice a consistent pattern where I’m in a tight midgame / endgame, and a bot will be playing really well, minimaxing, responding in must play zones for a really long time. Then suddenly, around move 130 ~ 150, it will randomly play one ‘brain dead’ move, where it plays in a dead or captured region of the board even though there is a clear immediate threat to be taken care of. And that one brain dead move usually is a sharp transition from “the bot is probably winning” to “the bot has definitely lost”.

It’s interesting that it always seems to happen around the same point in the game…

#11 · Arek Kulczycki

9 Jun 2024

I have no insight to how the "weakness" is implemented. One option is that it's not implemented very well, but another option is that what seems to you "bot is maybe winning" for the bot is already lost and therefore it doesn't distinguish a "brain dead move" from any other move if all options are recognized as 0% chance. Bot comprehension is very different to human.

#12 · hexanna

9 Jun 2024

I have looked at many of the games from tiyusufaly and in every case where the bot loses in move >100, the bot was completely winning out of the opening. Furthermore, its moves were consistently close to the top move according to full-strength KataHex.

A good example is https://www.littlegolem.net/jsp/game/game.jsp?gid=2451814&nmove=24. The bot played nearly perfectly in the first 24 moves and ended up with a 99.5% win percentage according to the full-strength evaluation. That is, all of moves 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23 were either the top move according to full-strength KataHex, or within 1% of the top move. It's hard to evaluate the moves after that because the win rate is already >99%. The bot maintained >99% until its blunder 95. a5, which brought it from 99% to 1%. The same general pattern occurs for basically all of the other games (except the weakest bots where the human wins in moves).

#13 · Arek Kulczycki

11 Jun 2024

Wow then it's a real bad implementation. Do the bots play reasonable first move now?