Awesome Information Retrieval
A curated list of awesome information retrieval resources
Information Retrievalinvolves finding relevant information for user queries, ranging from simple domain of database search to complicated…
Introduction to Information RetrievalC.D. Manning, P. Raghavan, H. Schütze. Cambridge UP, 2008. (First book for getting started with Information Retrieval).
Search Engines: Information Retrieval in PracticeBruce Croft, Don Metzler, and Trevor Strohman. 2009. (Great book for readers interested in knowing how Search Engines…
Modern Information RetrievalR. Baeza-Yates, B. Ribeiro-Neto. Addison-Wesley, 1999.
Information Retrieval in PracticeB. Croft, D. Metzler, T. Strohman. Pearson Education, 2009.
Mining the Web: Analysis of Hypertext and Semi Structured DataS. Chakrabarti. Morgan Kaufmann, 2002.
Language Modeling for Information RetrievalW.B. Croft, J. Lafferty. Springer, 2003. (Handles Language Modeling aspect of Information Retrieval. It also…
Information Retrieval: A SurveyEd Greengrass, 2000. (Comprehensive survey of Conventional Information Retrieval, before Deep Learning era).
Introduction to Modern Information RetrievalG.G. Chowdhury. Neal-Schuman, 2003. (Intended for students of library and information studies).
Text Information Retrieval SystemsC.T. Meadow, B.R. Boyce, D.H. Kraft, C.L. Barry. Academic Press, 2007 (library/information science perspective).
INF384H / CS395T / INF350E: Concepts of Information Retrieval (and Web Search)Matthew Lease (University of Texas at Austin).
CS 276 / LING 286: Information Retrieval and Web SearchChris Manning and Pandu Nayak (Stanford University).
CS 371R: Information Retrieval and Web SearchRaymond J. Mooney (University of Texas at Austin).
CS 172: Introduction to Information RetrievalVagelis Hristidis (University of California - Riverside).
SIMS 240: Principles of Information RetrievalRay R. Larson (UC berkeley).
11-442 / 11-642: Search EnginesJamie Callan (CMU).
600.466: Information Retrieval and Web AgentsDavid Yarowsky (John Hopkins University).
CS 435: Information Retrieval, Discovery, and DeliveryAndrea LaPaugh (Princeton University).
Information Retrieval and Data MiningDr. Jilles Vreeken , Prof. Dr. Gerhard Weikum (MPI).
Coursera - Text Retrieval and Search EnginesProf. ChengXiang Zhai (University of Illinois at Urbana-Champaign).
Apache LuceneOpen Source Search Engine that can be used to test Information Retrieval Algorithm. Twitter uses this core for its…
The Lemur ProjectThe Lemur Project develops search engines, browser toolbars, text analysis tools, and data resources that support…
Indri Search EngineAnother Open Source Search Engine competitor of Apache Lucene.
Lemur ToolkitOpen Source Toolkit for research in Language Modeling, filtering and categorization.
DBPediaLinked data web.
Cranfield CollectionsThis is one of the first collections in IR domain, however the dataset is too small for any statistical significance…
TREC CollectionsTREC is the benchmark dataset used by most IR and Web search algorithms. It has several tracks, each of which consists…
BlogExplore information seeking behavior in the blogosphere.
Chemical IRAddress challenges in building large chemical testbeds for chemical IR.
Clinical Decision SupportInvestigate techniques to link medical cases to information relevant for patient care.
ConfusionStudy Known Item Searching problem.
Contextual SuggestionInvestigate search techniques for complex information needs (context and user interests based).
CrowdsourcingExplore crowdsourcing methods for performing and evaluating search.
EnterpriseStudy search over the organization data.
EntityPerform entity-related search (find entities and their properties) on Web data.
FilteringBinarily decide retrieval of new incoming documents given a stable information need.
Federated Web SearchStudy merge performance for results from various search services.
GenomicsStudy retrieval efficiency of genomics data and corresponding documentation.
HARDObtain High Accuracy Retrieval from Documents by leveraging searcher's context.
Interactive TrackStudy user interaction with text retrieval systems.
Knowledge base accelerationStudy algorithms that improve efficiency of human Knowledge Base.
Legal TrackStudy retrieval systems that have high recall for legal documents use case.
Medical TrackExplore unstructured search performance over patients record data.
Microblog TrackExamine satisfaction of real-time information need for microblogging sites.
Million Query TrackExplore ad-hoc retrieval over large set of queries.
Novelty TrackInvestigate systems' abilities to locate new (non-redundant) information.
Question Answering TrackTest systems that scale beyond document retrieval, to retrieve answers to factoid, list and definition type questions.
Relevance Feedback TrackFor deep evaluation of relevance feedback processes.
Robust TrackStudy individual topic's effectiveness.
Session TrackDevelop methods for measuring multiple-query sessions where information needs drift.
SPAM TrackBenchmark spam filtering approaches.
Tasks TrackTest if systems can induce possible tasks, users might be trying to accomplish for the query.
Temporal Summarization TrackDevelop systems that allow users to efficiently monitor the information associated with an event over time.
Terabyte TrackTest scalability of IR systems to large scale collection.
Web TrackExplore information seeking behaviors common in general web search.
GOV2 Test CollectionThis is one of the largest Web collection of documents obtained from crawl of government websites by Charlie Clarke…
NTCIR Test CollectionThis is collection of wide variety of dataset ranging from Ad-hoc collection, Chinese IR collection, mobile…
CLIR Test CollectionsThis dataset can be used for cross lingual IR between CJKE (Chinese-Japanese-Korean-English) languages. It is suitable…
Cross Language Q&A (CLQA) dataset collectionIt supports following bi-lingua and mono-lingua:; Bi-lingua; Japanese to English.; Chinese to English.; English to…
Advanced Cross Linugal Information Retrieval and Question Answering (ACLIA)The dataset is used for the task of cross-lingual question answering but the complexity of the task is higher than…
Conference and Labs of the Evaluation Forum (CLEF) datasetIt contains a multi-lingual document collection. The test suite includes:; AdHoc - News Test suite.; Domain Specific…
Reuters CorporaThe corpora is now available through NIST. The corpora includes following:; RCV1 (Reuter's Corpus Volume 1) - Consists…
20 Newsgroup datasetThis data set consists of 20000 newsgroup messages.posts taken from 20 newsgroup topics.
English Gigaword Fifth EditionThis data set is a comprehensive archive of English newswire text data including headlines, datelines and articles.
Document Understanding Conference (DUC) datasetsPast newswire/paper datasets (DUC 2001 - DUC 2007) are available upon request.
CMU ListStanford ListUniversity of Tennesse KnoxvilleExtreme Classification: A New Paradigm for Ranking & RecommendationManik Verma (Microsoft Research)
The next webTim Berners-Lee (Ted Talk) [Tim Berners-Lee invented the World Wide Web. He leads the World Wide Web Consortium (W3C),…
Is Pivot a turning point for web exploration?Gary Flake, Technical Fellow at Microsoft (TED Talks).
Challenges in Building Large-Scale Information Retrieval SystemsJeff Dean (WSDM Conference, 2009).
Knowledge-based Information Retrieval with WikipediaDavid Wilne (The University of Waikato, 2008).
Music Information Retrieval Using Locality Sensitive HashingSteve Tjoa (RackSpace Developers) [This talk shows that IR is not just text and images].
The Functional Web -- The Future of Apps and the WebLiron Shapira (Box Tech Talk).
Information Experience - Solution to Information Overload on WebDoug Imbruce (Techcrunch Disrupt)[Doug Imbruce is the Founder of Qwiki, Inc, a technology startup in New York, NY,…
Internet PrivacyDr. Alma Whitten (Google Brussels Tech Talk).
The moral bias behind your search resultsAndreas Ekström (Swedish Author & Journalist, TED Talk).
Beware online "filter bubbles"Eli Pariser (Author of the Filter Bubble, TED Talk).
Think your email's private? Think againAndy Yen (CERN, TED Talk) [This talk talks about privacy, which Search Engines intrude into, and how can people…
Do we have the right to be forgotten?Michael Douglas [TEDx SouthBank].
The case for anonymity onlineChristopher "moot" Poole" (Ted Talks) [Christopher "moot" Poole is founder of 4chan, an online imageboard whose…
Information Retrieval and the WebGoogle Research.
IR ThoughtsDr. Edel Garcia.
Deep Neural Network Learns to Judge Books by Their CoversInformation Extraction.
Can Deep Learning help solve Deep LearningInformation Retrieval from Lip Reading.
To reduce biases in machine learning start with openly discussing the problemBias in Relevance.
Whoa, Google’s AI Is Really Good at PictionarySketch-based search.
Neural Network Learns to Identify Criminals by Their FacesInformation Extraction.