10 февраля 2023 · Комментарий

Без заголовка

Archive of chatGPT failures
https://docs.google.com/spreadsheets/d/1kDSERnROv5FgHbVN8z_bXH9gak2IXRtoqz0nwhrviCw/htmlview#gid=1302320625

A Categorical Archive of ChatGPT Failures

Large language models have been demonstrated to be valuable in differentfields. ChatGPT, developed by OpenAI, has been trained using massive amounts ofdata and simulates human conversation by comprehending context and generatingappropriate responses. It has garnered significant attention due to its abilityto effectively answer a broad range of human inquiries, with fluent andcomprehensive answers surpassing prior public chatbots in both security andusefulness. However, a comprehensive analysis of ChatGPT's failures is lacking,which is the focus of this study. Ten categories of failures, includingreasoning, factual errors, math, coding, and bias, are presented and discussed.The risks, limitations, and societal implications of ChatGPT are alsohighlighted. The goal of this study is to assist researchers and developers inenhancing future language models and chatbots.

arxiv.org

A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity

This paper proposes a framework for quantitatively evaluating interactiveLLMs such as ChatGPT using publicly available data sets. We carry out anextensive technical evaluation of ChatGPT using 21 data sets covering 8different common NLP application tasks. We evaluate the multitask, multilingualand multi-modal aspects of ChatGPT based on these data sets and a newlydesigned multimodal dataset. We find that ChatGPT outperforms LLMs withzero-shot learning on most tasks and even outperforms fine-tuned models on sometasks. We find that it is better at understanding non-Latin script languagesthan generating them. It is able to generate multimodal content from textualprompts, via an intermediate code generation step. Moreover, we find thatChatGPT is 64.33% accurate on average in 10 different reasoning categoriesunder logical reasoning, non-textual reasoning, and commonsense reasoning,hence making it an unreliable reasoner. It is, for example, better at deductivethan inductive reasoning. ChatGPT suffers from hallucination problems likeother LLMs and it generates more extrinsic hallucinations from its parametricmemory as it does not have access to an external knowledge base. Finally, theinteractive feature of ChatGPT enables human collaboration with the underlyingLLM to improve its performance, i.e, 8% ROUGE-1 on summarization and 2% ChrF++on machine translation, in a multi-turn "prompt engineering" fashion.

arxiv.org

And I don't even start to put links to Gary Marcus' critics or the last revelations of Le Cun himself.

К записи · К обсуждению