ailev.ru

Обсуждение

В архиве: 10 комментариев.

Читать и комментировать в ЖЖ ↗

Имя не сохранено · 27 марта 2015

Комментарий

Ну вот оно... Дожили!!!

Имя не сохранено · 27 марта 2015

Комментарий

Враньё это всё, точнее, профанация -- анализ от того самого Andrej Karpathy показывает, что ещё как минимум раза в 2 надо ошибку уменьшить для паритета с человеком, потому что компьютер совершает такие ошибки, которые человек не совершил бы.

Анатолий Левенчук · 27 марта 2015

Комментарий

Я думаю, в данной ситуации всё уже не так однозначно. Ибо человек тоже совершает такие ошибки, которые компьютер бы не совершил. Компьютерну не нужно моделировать профиль человеческих ошибок, ему нужно иметь число ошибок меньше, чем у человека.

Ответ на комментарий

Имя не сохранено · 27 марта 2015

Комментарий

Думаю "неоднозначно" самое правильное слово. Человек может ошибиться (грубо) двумя способами: 1. Он правда не смог разобрать, что там на картинке 2. "чисто по человечески": Тупо опечатался, так как он человек. У него начали глаза слипаться на сотой картинке теста и качество распознавания снизилось. Для производства это разделение не важно: есть пост контроля и выбор, кого на него сажать, трех-четырех операторов в сутки или одну железяку, а вот с научной-соревновательной точки зрения комп хорошо бы сравнивать только против ошибок первого рода. Не читал подробно описания, как получали процент для человека, может там и отброшен второй вариант ошибок

Ответ на комментарий

Имя не сохранено · 27 марта 2015

Комментарий

А можно просто взять и почитать: http://karpathy.github.io/2014/09/02/what-i-learned-from-competing-against-a-convnet-on-imagenet/ Types of error that both GoogLeNet and human are susceptible to: Multiple objects. Both GoogLeNet and humans struggle with images that contain multiple ILSVRC classes (usually many more than five), with little indication of which object is the focus of the image. This error is only present in the Classification setting, since every image is constrained to have exactly one correct label. In total, we attribute 24 (24%) of GoogLeNet errors and 12 (16%) of human errors to this category. It is worth noting that humans can have a slight advantage in this error type, since it can sometimes be easy to identify the most salient object in the image. Incorrect annotations. We found that approximately 5 out of 1500 images (0.3%) were incorrectly annotated in the ground truth. This introduces an approximately equal number of errors for both humans and GoogLeNet. Types of error that GoogLeNet is more susceptible to than human: Object small or thin. GoogLeNet struggles with recognizing objects that are very small or thin in the image, even if that object is the only object present. Examples of this include an image of a standing person wearing sunglasses, a person holding a quill in their hand, or a small ant on a stem of a flower. We estimate that approximately 22 (21%) of GoogLeNet errors fall into this category, while none of the human errors do. In other words, in our sample of images, no image was mislabeled by a human because they were unable to identify a very small or thin object. This discrepancy can be attributed to the fact that a human can very effectively leverage context and affordances to accurately infer the identity of small objects (for example, a few barely visible feathers near person's hand as very likely belonging to a mostly occluded quill). Image filters. Many people enhance their photos with filters that distort the contrast and color distributions of the image. We found that 13 (13%) of the images that GoogLeNet incorrectly classified contained a filter. Thus, we posit that GoogLeNet is not very robust to these distortions. In comparison, only one image among the human errors contained a filter, but we do not attribute the source of the error to the filter. Abstract representations. We found that GoogLeNet struggles with images that depict objects of interest in an abstract form, such as 3D-rendered images, paintings, sketches, plush toys, or statues. An example is the abstract shape of a bow drawn with a light source in night photography, a 3D-rendered robotic scorpion, or a shadow on the ground, of a child on a swing. We attribute approximately 6 (6%) of GoogLeNet errors to this type of error and believe that humans are significantly more robust, with no such errors seen in our sample. Miscellaneous sources. Additional sources of error that occur relatively infrequently include extreme closeups of parts of an object, unconventional viewpoints such as a rotated image, images that can significantly benefit from the ability to read text (e.g. a featureless container identifying itself as "face powder"), objects with heavy occlusions, and images that depict a collage of multiple images. In general, we found that humans are more robust to all of these types of error.

Ответ на комментарий

Имя не сохранено · 30 марта 2015

Комментарий

Одна оценка (пессимистичная) следующая: пока что вычислительная мощность ИНС-компьютеров сравнима с мозгом лягушки, менее 1% от мозга человека. Вычислительная мощность быстро не прибавляется, но, возможно, дальнейшими изменениями алгоритма можно ещё немного увеличить его точность. Но проблема отсутствия избытычности остаётся (мозг распознаёт изображения, имея запас нейронов и доучиваясь на ходу). Вторая (оптимистичная) -- что точность уже и так неплохая, ну, пусть даже 99% против человеческих 99.99% (для известных классов -- мы никогда не перепутаем, скажем, человека и собаку -- хотя бы исходя из их роста!), но для видео-трекинга объектов точность будет значительно больше ("усредним ошибки распознавания по времени" или смахнём пыль с алгоритма "Predator" ), и есть существенный резерв увеличения количества иерархически распознаваемых объектов, в то время, как у человека опыт весьма ограничен из-за его низкой скорости обучения. На крайний случай, зальют всё в железо -- но, вероятнее всего, через пару годиков. А за полгода вряд ли что-то значительно изменится с картинками -- просто фокус начнёт смещаться на видео.

Ответ на комментарий