ailev.ru

4 ноября 2016 · Комментарий

Без заголовка

У меня ж это не первый и не второй текст, где я об этом пишу. Первый раз как-то развёрнуто об этом я писал, ссылаясь вот тут: http://ailev.livejournal.com/1240509.html
Интереснейшие замечания про word embeddings ровно вдоль этой линии рассуждений (сетка мощна, всё будет в ней) дал приглашённый Нандо де Фрейтасом в чат Ed Grefenstette (http://egrefen.com/news.html, он активно работает с лингвистикой -- вот тут его свежие результаты: http://arxiv.org/find/cs/1/au:+Grefenstette_E/0/1/0/all/0/1). Тут нужно заметить, что в ответе он главным образом адресует "заранее выученные" word embeddings, которыми завалена Сеть: -- самые разные подходы к word embeddings по сути are effectively equivalent in performance and representational power, up to the correct choice of hyperparameters. Вот тут многочисленные ссылки по embeddings, подтверждающие эту одинаковость, включая дополнительное замечание, что для их построения никакой глубокой сетки не нужно, хватает мелкой сетки: http://gavagai.se/blog/2015/09/30/a-brief-history-of-word-embeddings/. И ещё https://levyomer.wordpress.com/2015/03/30/improving-distributional-similarity-with-lessons-learned-from-word-embeddings/ и https://levyomer.wordpress.com/2014/09/10/neural-word-embeddings-as-implicit-matrix-factorization/ (и там ещё много -- https://levyomer.wordpress.com/category/word-embeddings/). Хе-хе, не могу не заметить, что matrix factorization любимая тема для compressive sensing -- https://sites.google.com/site/igorcarron2/matrixfactorizations и оттуда прямой ход на вычислительную оптику! Хотя в этом пункте про word embeddings это явно оффтоп. -- word embeddings ... in no way a good general representation of semantics, but rather just one very successful example of an application transfer learning between contextual prediction (word given context, or context given word) and other domains with very different objectives (sentiment analysis, language modelling, question answering), either by serving as representations in their own right, or as initial settings to aid training. -- эмпирическое возражение: если данных достаточно, то ничего заранее выученного не нужно. Более того, если вам нужно будет развести значения "тачка" и "автомобиль" для сленга и официальной речи, то прихват заранее выученных word embeddings будет даже лишним, и лучше бы для этого учить модель языка заново. -- концептуальное возражение: для RNN embedding matrix это часть самой сетки! ... embeddings are just weights of a linear transform from the one-hot input into vectors used by the network's internal dynamics. Meaning and interpretation, if there is such things, are present in the state of the network, rather than solely in the embeddings, and it makes as much sense to seek to interpret the weights that constitute embeddings as it does to seek to interpret any other weight in the network. Pre-training embeddings and using them in another network, under this view, is even more explicitly just a form of transfer learning, in that we are initialising the weights of part of a task-specific network, and perhaps freezing them, with information obtained from another task. It's not a bad strategy, but I think people focus too much on this very specific form of transfer learning rather than, more generally, on other options there are out there (or yet to be discovered) to help us deal with data paucity, and to best share information across similar tasks. Вот! Meaning and interpretation -- Эд различает "значение и смысл" и напирает на то, что нужно работать и со значением/знанием (переносимым из других ситуаций опытом), и со смыслом (transactional information -- имеющей смысл только в контексте конкретного действия). Для меня это различение всегда было важно, и теперь понятно, какая группа в deep learning с этим работает -- купленные в состав DeepMind люди из http://www.darkbluelabs.com/ -- они там все из Оксфорда.
Но потом в разговорах с avlasov я начал все такие рассуждения называть работой с priors -- всё одно байесовская терминология и идеи начали глубоко проникать в эту предметную область, и они оказываются очень удобными для описания.

К записи · К обсуждению