28 января 2016 · Комментарий

Без заголовка

Ничего абсолютного, конечно, нет. С одной стороны, это явно не подход GOFAI, а с другой стороны, не end-to-end learning. Вот более точно (игра сама с собой там была -- playing thousands of games between its neural networks): We trained the neural networks on 30 million moves from games played by human experts, until it could predict the human move 57 percent of the time (the previous record before AlphaGo was 44 percent). But our goal is to beat the best human players, not just mimic them. To do this, AlphaGo learned to discover new strategies for itself, by playing thousands of games between its neural networks, and adjusting the connections using a trial-and-error process known as reinforcement learning. Of course, all of this requires a huge amount of computing power, so we made extensive use of Google Cloud Platform. Это из https://googleblog.blogspot.co.uk/2016/01/alphago-machine-learning-game-go.html (я добавил сейчас ещё и эту ссылку). Мы абсолютно разное значение придаём этому моменту, а он становится стандартным сегодня. Я предпочитаю трактовать это как вариант transfer learning: сначала учим на известном, а затем на базе выученного играем сами с собой, достигая заоблачных высот. И да, transfer learning это некоторое искусство сейчас (нужно и старое не забыть, и новое нарастить -- это трудно при нынешних алгоритмах). Вы делаете акцент на первичном обучении (на базе человечьих партий), а я на вторичном (когда сетки тренируются в игре между собой).

К записи · К обсуждению