They are using Monte Carlo Methods for looking around, more like a tiny sampling, it's an incomprehensibly large space. They have both value and policy (deep) neural nets. Train one to get the sense of good/bad individual board states (value) and one to get the sense of good/bad trajectories (policy, evaluates chains of moves).
Chess has a fairly straight forward ranking of pieces, fairly well established ranking of piece power. A knight is more or less always a knight.
Go, you kinda make up your "pieces" from scratch and there isn't any single common ranking. This makes the problem of even evaluating the board difficult (value), which is a sub step of making your higher plans (policy).
The output of the two dNN help cut off, or prevent the sampling of the bad game states and encourage looking into the fruitful areas.
Chess has a fairly straight forward ranking of pieces, fairly well established ranking of piece power. A knight is more or less always a knight.
Go, you kinda make up your "pieces" from scratch and there isn't any single common ranking. This makes the problem of even evaluating the board difficult (value), which is a sub step of making your higher plans (policy).
The output of the two dNN help cut off, or prevent the sampling of the bad game states and encourage looking into the fruitful areas.
Here is a (edit: not old) new paper, https://gogameguru.com/i/2016/03/deepmind-mastering-go.pdf