awesome-huggingface
Section: Official Libraries · Fast state-of-the-Art tokenizers optimized for research and production.
Entry
Appears in 6 awesome lists
Hugging Face's tokenizers for modern NLP pipelines (original implementation) with bindings for Python.
Section: Official Libraries · Fast state-of-the-Art tokenizers optimized for research and production.
Section: Large Model Serving · 💥 Fast State-of-the-Art Tokenizers optimized for Research and Production
Section: Rust · Fast State-of-the-Art Tokenizers optimized for Research and Production
Section: Libraries · Tokenizers optimized for Research and Production.
Section: 1. Core Frameworks & Libraries · Fast state-of-the-art tokenizers for training and inference.
Section: Artificial Intelligence · Hugging Face's tokenizers for modern NLP pipelines (original implementation) with bindings for Python.
(formerly known as pytorch-transformers and pytorch-pretrained-bert) provides state-of-the-art general-purpose architectures (BERT, GPT-2, RoBERTa, XLM, DistilBert, XLNet, CTRL...) for Natural Language Understanding (NLU) and Natural Language Generation (NLG) with over 32+ pretrained models in…
State-of-the-art serving engine with PagedAttention and continuous batching. Currently the fastest production-grade LLM server.
Flowise simplifies the creation of applications leveraging large language models (LLMs) by providing a drag-and-drop interface for customizing AI workflows, offering easy installation, Docker support, development tools, and documentation for integrating various functionalities such as…
Python-free Rust inference server with OpenAI API compatibility. Supports GGUF and SafeTensors formats with hot model swap, auto-discovery, and single binary deployment for zero-dependency inference. Apache 2.0 licensed.
Ports for inferencing LLaMA in C/C++ running on CPUs, supports alpaca, gpt4all, etc.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference…
Accelerate abstracts exactly and only the boilerplate code related to multi-GPU/TPU/mixed-precision and leaves the rest of your code unchanged.
The largest hub of ready-to-use NLP datasets for ML models with fast, easy-to-use and efficient data manipulation tools.