Awesome Deep Learning
Section: Datasets · Stanford released ~100,000 English QA pairs and ~50,000 unanswerable questions
Entry
Appears in 4 awesome lists
Stanford Question Answering Dataset (SQuAD) is a reading comprehension dataset, consisting of questions posed by crowdworkers on a set of Wikipedia articles, where the answer to every question is a segment of text, or span, from the corresponding reading passage, or the question might be…
Section: Datasets · Stanford released ~100,000 English QA pairs and ~50,000 unanswerable questions
Section: Some Datasets · Question answering dataset that can be explored online, and a list of models performing well on that dataset.
Section: Question Answering and Reading Comprehension · extractive reading comprehension.
Section: Datasets · Stanford Question Answering Dataset (SQuAD) is a reading comprehension dataset, consisting of questions posed by crowdworkers on a set of Wikipedia articles, where the answer to every question is a segment of text, or span, from the corresponding reading passage, or the question might be…
Reading Wikipedia to answer open-domain questions.
(2023) - retrieval, generation, and self-critique.
TriviaQA is a reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaQA includes 95K question-answer pairs authored by trivia enthusiasts and independently gathered evidence documents, six per question on average, that provide high quality distant supervision…
general AI assistant benchmark including multi-step QA.
Huge free English speech dataset with balanced genders and speakers, that seems to be of high quality.
MNIST like fashion product dataset consisting of a training set of 60,000 examples and a test set of 10,000 examples. Each example is a 28x28 grayscale image, associated with a label from 10 classes.
This could be used for a chatbot.
and FiD - retrieve-then-read; the standard pre-LLM open-domain QA pipeline.