CLARIN-D web tools
Tools for Analysing Research Data
A curated list of anything remotely related to linguistics
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
Tools for Analysing Research Data
Software for corpus linguists and text/data mining enthusiasts. The CorpusExplorer combines over 50 interactive visualizations under a user-friendly interface.
Early linguistical analysis and natural language processing library for Haxe.
The most complete platform for building Python programs to work with human language data.
Snowball is a language in which stemming algorithms can be easily represented.
, webservice via WebLicht
Easy-to-use text annotation tool for teams with most comprehensive auto-annotation features. Supports NER, relations and document classification as well as OCR annotation for invoice labeling.
Nice alternative for spacy (see above).
A utility for finding Typo-Bridges.
An open source Python library for processing morphologically rich and, for the most part, endangered Uralic languages. It can do morphological analysis, generation, lemmatization, disambiguation and lexical lookup for a great many Uralic languages.
Various stemming algorithms from snowball.
The ‘official’ home page for distribution of the Porter Stemming Algorithm, written and maintained by its author, Martin Porter.
JSON formatted Pan-Romance word lists.
German corpus based on CommonCrawl
(CommonCrawl)
German Reference Corpus
sampled sentences in different languages.
big german internet corpus
A list of resources for conservation, development, and documentation of low resource (human) languages.
Language Science Press is a born-digital scholar-led open access publisher in linguistics.
(NER exploration)
(from UKPLab) - Widely used encoders computing dense vector representations for sentences, paragraphs, and images.
is a machine learning algorithm that is used solved calssification problems. It's based on applying Bayes' theorem with strong independence assumptions between the features.
Lectures for University of Maryland class on computational linguistics.
CC-licensed educational videos interconnected with Marburg University's e-learning platform of the same name.
An introductory book (2nd edition).
The book from the NLTK package.
This book serves as an introduction of text mining using the tidytext package and other tidy tools in R. Authors: Julia Silge and David Robinson.
A curated list of datasets for natural language processing (NLP) tasks.
Repository to track the progress in Natural Language Processing (NLP), including the datasets and the current state-of-the-art for the most common NLP tasks.
Articles on natural language processing.
A ranked list of awesome Python libraries for natural language processing (NLP).
curated list of resources for Danish language technology.
Information retrieval resources
curated list of open-access, open-source, and off-the-shelf resources and tools developed with a focus on German.
Linguistic Resources for doing NLP & CL on Spanish
hesreallyhim/awesome-claude-code
A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…
VoltAgent/awesome-agent-skills
A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.
josephmisiti/awesome-machine-learning
A curated list of awesome Machine Learning frameworks, libraries and software.
EthicalML/awesome-production-machine-learning
A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning
academic/awesome-datascience
:memo: An awesome Data Science repository to learn and apply for real world problems.
analysis-tools-dev/static-analysis
⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…