Skip to content
87

Awesome LLM Resources

🧑‍🚀 全世界最好的LLM资料总结(多模态生成、Agent、辅助编程、AI审稿、数据处理、模型训练、模型推理、o1 模型、MCP、小语言模型、视觉语言模型) | Summary of the world's best LLM resources.

9k stars1,007 forks983 entriesLast push Sep 21, 2026 (9 days ago)License Apache-2.0

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

推荐 Suggestion

谷歌AI的14年、Gemini翻身之战,与视觉理解模型:专访DeepMind前核心科学家Andrew Dai|Neolabs特辑

140. 对姚顺宇的4小时访谈:请允许我小疯一下!在Anthropic和Gemini训模型、技术预测、英雄主义已过去

张驰: A Year Inside ByteDance's AI Lab

Luo Fuli: OpenClaw, Agent Frameworks — The AI Paradigm Has Already Changed Dramatically!

A 7-hour marathon interview with Saining Xie: World Models, AMI Labs, Yann LeCun, Fei-Fei Li, and 42

翁家翌:OpenAI,GPT,强化学习,Infra,后训练,天授,tuixue,开源,CMU,清华|WhynotTV Podcast

Lovart 创始人陈冕×罗永浩!且让我大闹一场,然后悄然离去

MiniMax 创始人闫俊杰×罗永浩!大山并非无法翻越

影视飓风TIM×罗永浩!用影像打开世界的梦想家

129. 全球大模型第一股的上市访谈,和智谱CEO张鹏聊:敢问路在何方?

128. Manus决定出售前最后的访谈:啊,这奇幻的2025年漂流啊…

122. 朱啸虎现实主义故事的第三次连载:人工智能的盛筵与泡泡

119. Kimi Linear、Minimax M2?和杨松琳考古算法变种史,并预演未来架构改进方案

118. 对李想的第二次3小时访谈:CEO大模型、MoE、梁文锋、VLA、能量、记忆、对抗人性、亲密关系、人类的智慧

115. 对OpenAI姚顺雨3小时访谈:6年Agent研究、人与系统、吞噬的边界、既单极又多元的世界

113. 和杨植麟时隔1年的对话:K2、Agentic LLM、缸中之脑和“站在无限的开端”

数据 Data

AotoLabel

Label, clean and enrich text datasets with LLMs.

In 2 lists

LabelLLM

The Open-Source Data Annotation Platform.

data-juicer

A one-stop data processing system to make data higher-quality, juicier, and more digestible for LLMs!

OmniParser

a native Golang ETL streaming parser and transform library for CSV, JSON, XML, EDI, text, etc.

In 3 lists

MinerU (🔥)

MinerU is a one-stop, open-source, high-quality data extraction tool, supports PDF/webpage/e-book extraction.

In 5 listsDetails

PDF-Extract-Kit

A Comprehensive Toolkit for High-Quality PDF Content Extraction.

In 2 lists

Parsera

Lightweight library for scraping web-sites with LLMs.

Sparrow

Sparrow is an innovative open-source solution for efficient data extraction and processing from various documents and images.

In 2 lists

Docling

Get your documents ready for gen AI.

GOT-OCR2.0

OCR Model.

LLM Decontaminator

Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.

DataTrove

DataTrove is a library to process, filter and deduplicate text data at a very large scale.

In 3 lists

llm-swarm

Generate large synthetic datasets like Cosmopedia.

Distilabel

Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.

In 3 lists

Common-Crawl-Pipeline-Creator

The Common Crawl Pipeline Creator.

Tabled

Detect and extract tables to markdown and csv.

Zerox

Zero shot pdf OCR with gpt-4o-mini.

In 2 lists

DocLayout-YOLO

Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception.

TensorZero

make LLMs improve through experience.

In 4 listsDetails

Promptwright

Generate large synthetic data using a local LLM.

pdf-extract-api

Document (PDF) extraction and parse API using state of the art modern OCRs + Ollama supported models.

pdf2htmlEX

Convert PDF to HTML without losing text or format.

Extractous

Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.

MegaParse

File Parser optimised for LLM Ingestion with no loss.

In 2 lists

MarkItDown

Python tool for converting files and office documents to Markdown.

In 9 listsDetails

datasketch

datasketch gives you probabilistic data structures that can process and search very large amount of data super fast, with little loss of accuracy.

In 2 lists

semhash

lightweight and flexible tool for deduplicating datasets using semantic similarity.

ReaderLM-v2

a 1.5B parameter language model that converts raw HTML into beautifully formatted markdown or JSON.

Bespoke Curator

Data Curation for Post-Training & Structured Data Extraction.

In 2 lists

LangKit

An open-source toolkit for monitoring Large Language Models (LLMs). Extracts signals from prompts & responses, ensuring safety & security.

In 4 listsDetails

olmOCR

A toolkit for training language models to work with PDF documents in the wild.

In 5 listsDetails

Easy Dataset (🔥)

A powerful tool for creating fine-tuning datasets for LLM.

In 2 lists

BabelDOC

PDF scientific paper translation and bilingual comparison library.

Dolphin

Document Image Parsing via Heterogeneous Anchor Prompting.

In 3 lists

EasyDistill

Easy Knowledge Distillation for Large Language Models.

ContextGem

a free, open-source LLM framework that makes it radically easier to extract structured data and insights from documents.

In 2 lists

OCRFlux

a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex layout handling, complicated table parsing and cross-page content merging.

In 2 lists

DataFlow

Easy Data Preparation with latest LLMs-based Operators and Pipelines.

In 3 lists

DatasetLoom (multimodal)

一个面向多模态大模型训练的智能数据集构建与评估平台.

Logics-Parsing

DeepSeek-OCR

PaddleOCR-VL

Chandra

a highly accurate OCR model that converts images and PDFs into structured HTML/Markdown/JSON while preserving layout information.

In 2 lists

HunyuanOCR

a leading end-to-end OCR expert VLM powered by Hunyuan's native multimodal architecture.

DeepSeek-OCR-2

Visual Causal Flow.

PaddleOCR-VL-1.5 (🔥)

Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing.

GLM-OCR

a multimodal OCR model for complex document understanding, built on the GLM-V encoder–decoder architecture.

微调 Fine-Tuning

LLaMA-Factory (🔥)

Unify Efficient Fine-Tuning of 100+ LLMs.

In 3 lists

360-LLaMA-Factory

Unify Efficient Fine-Tuning of 100+ LLMs. (add Sequence Parallelism for supporting long context training)

unsloth (🔥)

2-5X faster 80% less memory LLM finetuning.

In 8 listsDetails

TRL

Transformer Reinforcement Learning.

Firefly

Firefly: 大模型训练工具,支持训练数十种大模型

Xtuner

An efficient, flexible and full-featured toolkit for fine-tuning large models.

In 2 lists

torchtune

A Native-PyTorch Library for LLM Fine-tuning.

In 2 lists

Swift

Use PEFT or Full-parameter to finetune 200+ LLMs or 15+ MLLMs.

AutoTrain

A new way to automatically train, evaluate and deploy state-of-the-art Machine Learning models.

OpenRLHF

An Easy-to-use, Scalable and High-performance RLHF Framework (Support 70B+ full tuning & LoRA & Mixtral & KTO).

Ludwig

Low-code framework for building custom LLMs, neural networks, and other AI models.

In 6 listsDetails

mistral-finetune

A light-weight codebase that enables memory-efficient and performant finetuning of Mistral's models.

aikit

Fine-tune, build, and deploy open-source LLMs easily!

In 2 lists

H2O-LLMStudio

H2O LLM Studio - a framework and no-code GUI for fine-tuning LLMs.

In 3 lists

LitGPT

Pretrain, finetune, deploy 20+ LLMs on your own data. Uses state-of-the-art techniques: flash attention, FSDP, 4-bit, LoRA, and more.

In 4 listsDetails

LLMBox

A comprehensive library for implementing LLMs, including a unified training pipeline and comprehensive model evaluation.

In 2 lists

PaddleNLP

Easy-to-use and powerful NLP and LLM library.

In 4 listsDetails

workbench-llamafactory

This is an NVIDIA AI Workbench example project that demonstrates an end-to-end model development workflow using Llamafactory.

TinyLLaVA Factory

A Framework of Small-scale Large Multimodal Models.

LLM-Foundry

LLM training code for Databricks foundation models.

lmms-finetune

A unified codebase for finetuning (full, lora) large multimodal models, supporting llava-1.5, qwen-vl, llava-interleave, llava-next-video, phi3-v etc.

Simplifine

Simplifine lets you invoke LLM finetuning with just one line of code using any Hugging Face dataset or model.

Transformer Lab

Open Source Application for Advanced LLM Engineering: interact, train, fine-tune, and evaluate large language models on your own computer.

In 2 lists

Liger-Kernel

Efficient Triton Kernels for LLM Training.

In 3 lists

ChatLearn

A flexible and efficient training framework for large-scale alignment.

In 2 lists

nanotron

Minimalistic large language model 3D-parallelism training.

In 3 lists

Proxy Tuning

Tuning Language Models by Proxy.

Effective LLM Alignment

Effective LLM Alignment Toolkit.

Autotrain-advanced

AutoTrain Advanced is a no-code solution that allows you to train machine learning models in just a few clicks.

In 4 listsDetails

Meta Lingua

a lean, efficient, and easy-to-hack codebase to research LLMs.

Vision-LLM Alignemnt

This repository contains the code for SFT, RLHF, and DPO, designed for vision-based LLMs, including the LLaVA models and the LLaMA-3.2-vision models.

finetune-Qwen2-VL

Quick Start for Fine-tuning or continue pre-train Qwen2-VL Model.

Online-RLHF

A recipe for online RLHF and online iterative DPO.

InternEvo

an open-sourced lightweight training framework aims to support model pre-training without the need for extensive dependencies.

veRL (🔥)

Volcano Engine Reinforcement Learning for LLM.

In 3 lists

Axolotl

Axolotl is designed to work with YAML config files that contain everything you need to preprocess a dataset, train or fine-tune a model, run model inference or evaluation, and much more.

Oumi

Everything you need to build state-of-the-art foundation models, end-to-end.

In 3 lists

Kiln

The easiest tool for fine-tuning LLM models, synthetic data generation, and collaborating on datasets.

In 5 listsDetails

DeepSeek-671B-SFT-Guide

An open-source solution for full parameter fine-tuning of DeepSeek-V3/R1 671B, including complete code and scripts from training to inference, as well as some practical experiences and conclusions.

MLX-VLM

MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.

In 3 lists

RL-Factory

Train your Agent model via our easy and efficient framework.

In 2 lists

RM-Gallery

A One-Stop Reward Model Platform.

ART

rain multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training.

In 3 lists

LMMs-Engine

A simple, any-to-any modality framework for pretraining and finetuning. Lean, flexible, and built for research.

dLLM

a library that unifies the training and evaluation of diffusion language models, bringing transparency and reproducibility to the entire development pipeline. diffusion

Miles

an enterprise-facing reinforcement learning framework for large-scale MoE post-training and production workloads.

In 2 lists

Skills

a collection of pipelines to improve "skills" of large language models (LLMs).

Twinkle

a lightweight, client-server training framework engineered with modular, high-cohesion interfaces.

NeMo AutoModel

Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support.

In 2 lists

VeOmni

Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo.

In 2 lists

Soup

One-config CLI for LLM post-training (SFT/DPO/GRPO/KTO/ORPO). Layer streaming trains an 8B model on a 4 GB laptop GPU by streaming the frozen base from host RAM one decoder layer at a time.

In 2 lists

Agentic RL

veRL (🔥)

Volcano Engine Reinforcement Learning for LLM.

In 3 lists

AReaL:

AntGroup/Tsinghua

In 2 lists

slime (🔥):

LLM post-training framework for RL Scaling from THUDM. Supports SFT and RL training with multi-turn compilation feedback, powering projects like TritonForge for automated GPU kernel generation. Apache 2.0 licensed.

In 5 listsDetails

Agent Lightning:

The absolute trainer to light up AI agents.

In 3 lists

Molt:

NVIDIA (NeMo Labs)

In 2 lists

prime-rl:

Agentic RL Training at Scale from Prime Intellect. Framework for large-scale reinforcement learning capable of scaling to 1000+ GPUs with fully asynchronous RL, FSDP2 training, and vLLM inference. Apache 2.0 licensed.

In 3 lists

推理 Inference

ollama (🔥)

Get up and running with Llama 3, Mistral, Gemma, and other large language models.

In 12 listsDetails

Open WebUI

User-friendly WebUI for LLMs (Formerly Ollama WebUI).

In 11 listsDetails

Text Generation WebUI

A Gradio web UI for Large Language Models. Supports transformers, GPTQ, AWQ, EXL2, llama.cpp (GGUF), Llama models.

In 7 listsDetails

Xinference

A powerful and versatile library designed to serve language, speech recognition, and multimodal models.

In 5 listsDetails

LangChain

Build context-aware reasoning applications.

In 20 listsDetails

LlamaIndex

A data framework for your LLM applications.

In 14 listsDetails

lobe-chat

an open-source, modern-design LLMs/AI chat framework. Supports Multi AI Providers, Multi-Modals (Vision/TTS) and plugin system.

In 8 listsDetails

TensorRT-LLM

TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs.

In 7 listsDetails

vllm (🔥)

A high-throughput and memory-efficient inference and serving engine for LLMs.

In 11 listsDetails

LlamaChat

Chat with your favourite LLaMA models in a native macOS app.

In 2 lists

NVIDIA ChatRTX

ChatRTX is a demo app that lets you personalize a GPT large language model (LLM) connected to your own content—docs, notes, or other data.

LM Studio (🔥)

Discover, download, and run local LLMs.

In 11 listsDetails

chat-with-mlx

Chat with your data natively on Apple Silicon using MLX Framework.

LLM Pricing

Quickly Find the Perfect Large Language Models (LLM) API for Your Budget! Use Our Free Tool for Instant Access to the Latest Prices from Top Providers.

Open Interpreter

A natural language interface for computers.

In 9 listsDetails

Chat-ollama

An open source chatbot based on LLMs. It supports a wide range of language models, and knowledge base management.

chat-ui

Open source codebase powering the HuggingChat app.

In 4 listsDetails

MemGPT

Create LLM agents with long-term memory and custom tools.

In 2 lists

koboldcpp

A simple one-file way to run various GGML and GGUF models with KoboldAI's UI.

In 5 listsDetails

LLMFarm

llama and other large language models on iOS and MacOS offline using GGML library.

enchanted

Enchanted is iOS and macOS app for chatting with private self hosted language models such as Llama2, Mistral or Vicuna using Ollama.

Flowise

Drag & drop UI to build your customized LLM flow.

In 12 listsDetails

Jan

Jan is an open source alternative to ChatGPT that runs 100% offline on your computer. Multiple engine support (llama.cpp, TensorRT-LLM).

In 8 listsDetails

LMDeploy

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

In 3 lists

RouteLLM

A framework for serving and evaluating LLM routers - save LLM costs without compromising quality!

MInference

About To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.

Mem0

The memory layer for Personalized AI.

In 13 listsDetails

SGLang (🔥)

SGLang is yet another fast serving framework for large language models and vision language models.

In 9 listsDetails

AirLLM

AirLLM optimizes inference memory usage, allowing 70B large language models to run inference on a single 4GB GPU card without quantization, distillation and pruning. And you can run 405B Llama3.1 on 8GB vram now.

In 3 lists

LLMHub

LLMHub is a lightweight management platform designed to streamline the operation and interaction with various language models (LLMs).

YuanChat

LiteLLM (🔥)

Call all LLM APIs using the OpenAI format [Bedrock, Huggingface, VertexAI, TogetherAI, Azure, OpenAI, Groq etc.]

In 16 listsDetails

GuideLLM

GuideLLM is a powerful tool for evaluating and optimizing the deployment of large language models (LLMs).

LLM-Engines

A unified inference engine for large language models (LLMs) including open-source models (VLLM, SGLang, Together) and commercial models (OpenAI, Mistral, Claude).

OARC

ollama_agent_roll_cage (OARC) is a local python agent fusing ollama llm's with Coqui-TTS speech models, Keras classifiers, Llava vision, Whisper recognition, and more to create a unified chatbot agent for local, custom automation.

g1

Using Llama-3.1 70b on Groq to create o1-like reasoning chains.

MemoryScope

MemoryScope provides LLM chatbots with powerful and flexible long-term memory capabilities, offering a framework for building such abilities.

OpenLLM

Run any open-source LLMs, such as Llama 3.1, Gemma, as OpenAI compatible API endpoint in the cloud.

In 10 listsDetails

Infinity

The AI-native database built for LLM applications, providing incredibly fast hybrid search of dense embedding, sparse embedding, tensor and full-text.

In 9 listsDetails

optillm

an OpenAI API compatible optimizing inference proxy which implements several state-of-the-art techniques that can improve the accuracy and performance of LLMs.

In 2 lists

LLaMA Box

LLM inference server implementation based on llama.cpp.

ZhiLight

A highly optimized inference acceleration engine for Llama and its variants.

DashInfer

DashInfer is a native LLM inference engine aiming to deliver industry-leading performance atop various hardware architectures.

LocalAI

The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required.

In 14 listsDetails

ktransformers

A Flexible Framework for Experiencing Cutting-edge LLM Inference Optimizations.

In 4 listsDetails

SkyPilot

Run AI and batch jobs on any infra (Kubernetes or 14+ clouds). Get unified execution, cost savings, and high GPU availability via a simple interface.

In 4 listsDetails

Chitu

High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.

TokenSwift

From Hours to Minutes: Lossless Acceleration of Ultra Long Sequence Generation.

Cherry Studio (🔥)

a desktop client that supports for multiple LLM providers, available on Windows, Mac and Linux.

In 4 listsDetails

Shimmy

Python-free Rust inference server — OpenAI-API compatible. GGUF + SafeTensors, hot model swap, auto-discovery, single binary.

In 7 listsDetails

LlamaBarn

Run local LLMs on your Mac with a simple menu bar app.

In 2 lists

Parallax

a distributed model serving framework that lets you build your own AI cluster anywhere.

xLLM

A high-performance inference engine for LLMs, optimized for diverse AI accelerators.

In 2 lists

Rapid-MLX

OpenAI-compatible local LLM inference server for Apple Silicon, 2-4x faster than Ollama.

In 4 listsDetails

TokenSpeed

a speed-of-light LLM inference engine designed for agentic workloads, with TensorRT-LLM-level performance and vLLM-level usability. Our goal is to be the most performant inference engine for production agentic workloads.

In 2 lists

FreeToken

Unlock datacenter-class intelligence on the hardware you already own.

In 2 lists

评估 Evaluation

lm-evaluation-harness

A framework for few-shot evaluation of language models.

In 7 listsDetails

opencompass (🔥)

OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.

In 4 listsDetails

llm-comparator

LLM Comparator is an interactive data visualization tool for evaluating and analyzing LLM responses side-by-side, developed.

In 2 lists

EvalScope (🔥)

Streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking. One-stop evaluation solution with 80+ benchmarks. Apache 2.0 licensed.

In 3 lists

Weave

A lightweight toolkit for tracking and evaluating LLM applications.

MixEval

Deriving Wisdom of the Crowd from LLM Benchmark Mixtures.

Evaluation guidebook

If you've ever wondered how to make sure an LLM performs well on your specific task, this guide is for you!

Ollama Benchmark

LLM Benchmark for Throughput via Ollama (Local LLMs).

VLMEvalKit

Open-source evaluation toolkit of large vision-language models (LVLMs), support ~100 VLMs, 40+ benchmarks.

In 4 listsDetails

AGI-Eval

DeepEval

a simple-to-use, open-source LLM evaluation framework, for evaluating and testing large-language model systems.

In 9 listsDetails

Lighteval

Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends.

In 3 lists

QwQ/eval

QwQ is the reasoning model series developed by Qwen team, Alibaba Cloud.

Evalchemy

A unified and easy-to-use toolkit for evaluating post-trained language models.

In 3 lists

MathArena

Evaluation of LLMs on latest math competitions.

YourBench

A Dynamic Benchmark Generation Framework.

MedEvalKit

A Unified Medical Evaluation Framework.

OpenJudge

A Unified Framework for Holistic Evaluation and Quality Rewards.

体验 Usage

LM Arena

Design Arena

Vals AI

evaluation

蚂蚁数科

知识库 RAG

AnythingLLM

The all-in-one AI app for any LLM with full RAG and AI Agent capabilites.

In 8 listsDetails

MaxKB

基于 LLM 大语言模型的知识库问答系统。开箱即用,支持快速嵌入到第三方业务系统

In 4 listsDetails

RAGFlow

An open-source RAG (Retrieval-Augmented Generation) engine based on deep document understanding.

In 7 listsDetails

Dify

An open-source LLM app development platform. Dify's intuitive interface combines AI workflow, RAG pipeline, agent capabilities, model management, observability features and more, letting you quickly go from prototype to production.

In 14 listsDetails

FastGPT

A knowledge-based platform built on the LLM, offers out-of-the-box data processing and model invocation capabilities, allows for workflow orchestration through Flow visualization.

In 4 listsDetails

Langchain-Chatchat

基于 Langchain 与 ChatGLM 等不同大语言模型的本地知识库问答

In 3 lists

QAnything

Question and Answer based on Anything.

In 2 lists

Quivr

A personal productivity assistant (RAG) ⚡️🤖 Chat with your docs (PDF, CSV, ...) & apps using Langchain, GPT 3.5 / 4 turbo, Private, Anthropic, VertexAI, Ollama, LLMs, Groq that you can share with users ! Local & Private alternative to OpenAI GPTs & ChatGPT powered by retrieval-augmented generation.

In 5 listsDetails

RAG-GPT

RAG-GPT, leveraging LLM and RAG technology, learns from user-customized knowledge bases to provide contextually relevant answers for a wide range of queries, ensuring rapid and accurate information retrieval.

Verba

Retrieval Augmented Generation (RAG) chatbot powered by Weaviate.

In 2 lists

FlashRAG

A Python Toolkit for Efficient RAG Research.

In 2 lists

GraphRAG

A modular graph-based Retrieval-Augmented Generation (RAG) system.

In 5 listsDetails

LightRAG

LightRAG helps developers with both building and optimizing Retriever-Agent-Generator pipelines.

GraphRAG-Ollama-UI

GraphRAG using Ollama with Gradio UI and Extra Features.

nano-GraphRAG

A simple, easy-to-hack GraphRAG implementation.

RAG Techniques

This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. RAG systems combine information retrieval with generative models to provide accurate and contextually rich responses.

In 5 listsDetails

ragas

Evaluation framework for your Retrieval Augmented Generation (RAG) pipelines.

In 3 lists

kotaemon

An open-source clean & customizable RAG UI for chatting with your documents. Built with both end users and developers in mind.

In 3 lists

RAGapp

The easiest way to use Agentic RAG in any enterprise.

In 2 lists

TurboRAG

Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text.

LightRAG

Simple and Fast Retrieval-Augmented Generation.

In 5 listsDetails

TEN

the Next-Gen AI-Agent Framework, the world's first truly real-time multimodal AI agent framework.

AutoRAG

RAG AutoML tool for automatically finding an optimal RAG pipeline for your data.

In 5 listsDetails

KAG

KAG is a knowledge-enhanced generation framework based on OpenSPG engine, which is used to build knowledge-enhanced rigorous decision-making and information retrieval knowledge services.

Fast-GraphRAG

RAG that intelligently adapts to your use case, data, and queries.

Tiny-GraphRAG

DB-GPT GraphRAG

DB-GPT GraphRAG integrates both triplet-based knowledge graphs and document structure graphs while leveraging community and document retrieval mechanisms to enhance RAG capabilities, achieving comparable performance while consuming only 50% of the tokens required by Microsoft's GraphRAG. Refer to…

In 7 listsDetails

Chonkie

The no-nonsense RAG chunking library that's lightweight, lightning-fast, and ready to CHONK your texts.

RAGLite

RAGLite is a Python toolkit for Retrieval-Augmented Generation (RAG) with PostgreSQL or SQLite.

In 2 lists

CAG

CAG leverages the extended context windows of modern large language models (LLMs) by preloading all relevant resources into the model’s context and caching its runtime parameters.

MiniRAG

an extremely simple retrieval-augmented generation framework that enables small models to achieve good RAG performance through heterogeneous graph indexing and lightweight topology-enhanced retrieval.

XRAG

a benchmarking framework designed to evaluate the foundational components of advanced Retrieval-Augmented Generation (RAG) systems.

Rankify

A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation.

In 2 lists

RAG-Anything

All-in-One RAG System.

In 3 lists

智能体 Agents

AutoGen

AutoGen is a framework that enables the development of LLM applications using multiple agents that can converse with each other to solve tasks. AutoGen AIStudio

In 14 listsDetails

CrewAI

Framework for orchestrating role-playing, autonomous AI agents. By fostering collaborative intelligence, CrewAI empowers agents to work together seamlessly, tackling complex tasks.

In 4 listsDetails

Coze

ByteDance agent builder. Visual workflow. Plugin marketplace.

In 2 lists

AgentGPT

Assemble, configure, and deploy autonomous AI Agents in your browser.

In 9 listsDetails

XAgent

An Autonomous LLM Agent for Complex Task Solving.

In 5 listsDetails

MobileAgent

The Powerful Mobile Device Operation Assistant Family.

In 4 listsDetails

Lagent

A lightweight framework for building LLM-based agents.

In 3 lists

Qwen-Agent

Agent framework and applications built upon Qwen2, featuring Function Calling, Code Interpreter, RAG, and Chrome extension.

In 4 listsDetails

LinkAI

一站式 AI 智能体搭建平台

Baidu APPBuilder

agentUniverse

agentUniverse is a LLM multi-agent framework that allows developers to easily build multi-agent applications. Furthermore, through the community, they can exchange and share practices of patterns across different domains.

LazyLLM

低代码构建多Agent大模型应用的开发工具

AgentScope

Start building LLM-empowered multi-agent applications in an easier way.

In 4 listsDetails

AgentField

Open-source control plane for building and operating AI agents like APIs at scale, with routing, memory, observability, identity, auth, and policy controls.

In 3 lists

MoA

Mixture of Agents (MoA) is a novel approach that leverages the collective strengths of multiple LLMs to enhance performance, achieving state-of-the-art results.

In 2 lists

Agently

AI Agent Application Development Framework.

OmAgent

A multimodal agent framework for solving complex tasks.

In 3 lists

Tribe

No code tool to rapidly build and coordinate multi-agent teams.

CAMEL

First LLM multi-agent framework and an open-source community dedicated to finding the scaling law of agents.

In 7 listsDetails

PraisonAI

PraisonAI application combines AutoGen and CrewAI or similar frameworks into a low-code solution for building and managing multi-agent LLM systems, focusing on simplicity, customisation, and efficient human-agent collaboration.

In 9 listsDetails

IoA

An open-source framework for collaborative AI agents, enabling diverse, distributed agents to team up and tackle complex tasks through internet-like connectivity.

In 2 lists

llama-agentic-system

Agentic components of the Llama Stack APIs.

In 3 lists

Agent Zero

Agent Zero is not a predefined agentic framework. It is designed to be dynamic, organically growing, and learning as you use it.

Agents

An Open-source Framework for Data-centric, Self-evolving Autonomous Language Agents.

In 3 lists

FastAgency

The fastest way to bring multi-agent workflows to production.

In 3 lists

Swarm

Framework for building, orchestrating and deploying multi-agent systems. Managed by OpenAI Solutions team. Experimental framework.

In 6 listsDetails

Agent-S

an open agentic framework that uses computers like a human.

In 6 listsDetails

PydanticAI

Agent Framework / shim to use Pydantic with LLMs.

In 12 listsDetails

Agentarium

open-source framework for creating and managing simulations populated with AI-powered agents.

In 2 lists

smolagents

a barebones library for agents. Agents write python code to call tools and orchestrate other agents.

In 6 listsDetails

Cooragent

Cooragent is an AI agent collaboration community.

Agno

Agno is a lightweight library for building Agents with memory, knowledge, tools and reasoning.

In 7 listsDetails

Suna

Open Source Generalist AI Agent.

rowboat

Let AI build multi-agent workflows for you in minutes.

In 4 listsDetails

EvoAgentX

Building a Self-Evolving Ecosystem of AI Agents.

In 2 lists

ii-agent

a new open-source framework to build and deploy intelligent agents.

In 2 lists

OWL

Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation.

In 3 lists

OpenManus

No fortress, purely open ground. OpenManus is Coming.

In 3 lists

JoyAgent-JDGenie

业界首个开源高完成度轻量化通用多智能体产品.

coze-studio

An AI agent development platform with all-in-one visual tools, simplifying agent creation, debugging, and deployment like never before.

OxyGent

An advanced Python framework that empowers developers to quickly build production-ready intelligent systems.

LazyCraft

LazyCraft 是一个基于 LazyLLM 构建的 AI Agent 应用开发与管理平台,旨在协助开发者以 低门槛、低成本 快速构建和发布大模型应用。

OpenAgents

AI Agent Networks for Open Collaboration.

In 3 lists

SandBox

All-in-One Sandbox for AI Agents that combines Browser, Shell, File, MCP and VSCode Server in a single Docker container.

In 3 lists

DeepAnalyze

First agentic LLM for autonomous data science, supporting specific data tasks (data preparation, analysis, modeling, visualization, and insight) and data-oriented deep research (produce analyst-grade research reports).

In 8 listsDetails

Astron Agent

Enterprise-grade, commercial-friendly agentic workflow platform for building next-generation SuperAgents.

In 2 lists

Youtu-Agent

A simple yet powerful agent framework that delivers with open-source models.

In 4 listsDetails

MiroThinker

an open-source search agent model, built for tool-augmented reasoning and real-world information seeking, aiming to match the deep research experience of OpenAI Deep Research and Gemini Deep Research.

Nexent

A zero-code platform for auto-generating agents — no orchestration, no complex drag-and-drop required, using pure language to develop any agent you want.

In 2 lists

Yunjue-Agent

A Fully Reproducible, Zero-Start In-Situ Self-Evolving Agent System for Open-Ended Tasks.

Hindsight

State-of-the-art long-term memory for AI agents by Vectorize. Open-source, self-hostable, with integrations for LangChain, CrewAI, LlamaIndex, MCP, and more.

In 4 listsDetails

AgentsMesh

The AI Agent Workforce Platform. Self-hostable multi-agent orchestration with remote AI workstations (AgentPods), PTY sandbox + git worktree isolation, channels-based agent collaboration, built-in Kanban, and per-pod MCP server. Supports Claude Code, Codex CLI, Gemini CLI, Aider, OpenCode.

In 4 listsDetails

BitFun

Open-source agentic development environment with a Rust/Tauri desktop app and CLI for coding, research, office work, browser and desktop automation, extensible through MCP, Skills, and custom agents.

In 4 listsDetails

pi

AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI.

In 4 listsDetails

DeepSeek Harness

Everything is a Plugin.

In 5 listsDetails

OpenSquilla

a token-efficient, microkernel AI agent.

In 2 lists

PenguinHarness

Your Automated Agent Builder, Right on Your Desktop / Server.

FrontierAgent

an open-source agent runtime, terminal product, and evaluation suite for long-horizon research and file-based work.

In 2 lists

研究 Research

PaperDebugger:

Paper Debugger is the best overleaf companion

In 2 lists

XtraGPT served as refiner:

Chat Overleaf:

文智云助手:

LiteWrite:

Prism:

claude-prism:

Offline-first scientific writing workspace powered by Claude, integrating LaTeX, Python, and 100+ scientific skills with local execution, Zotero integration, and privacy-focused design (2026)

In 2 lists

PaperReview:

aiXiv:

OpenJudge Review:

PPTAgent:

Beyond text-to-slides generation with PPTEval multi-dimensional evaluation (EMNLP 2025)

In 2 lists

Paper PPT Agent:

PPT Master:

AI turns documents or topics into real, native PowerPoint decks—with native shapes, transitions and animations, data-backed charts and tables on demand, audio narration from speaker notes, and support for your own .pptx templates. · by Hugo He

In 3 lists

LandPPT:

Kami:

Kami是一款专注于将Markdown内容转化为高质量PDF排版工具的开源项目,它精准解决了创作者在使用常规编辑器导出文档时遭遇的格式错乱、字体生硬及页眉页脚配置繁琐等痛点。该项目在同类工具中展现出三大核心优势:内置经过精细调校的学术级排版引擎,能够自动处理行距、段首缩进与代码块样式,彻底免除了手动调整的冗余;提供可视化主题编辑界面,用户无需编写CSS即可实时预览并切换封面、页码及版权信息,大幅降低了定制门槛;深度集成Pandoc底层转换能力却剥离了复杂的命令行依赖,通过轻量级Web服务实现一键导出,在兼容性与易用性之间取得了卓越平衡。从技术架构来看,Kami本质上扮演了“数字排版工厂”的角…

In 3 lists

beautiful-html-templates:

guizang-ppt-skill:

GordenSuperPPTSkills:

dashiAI-ppt-skill:

Paper2Video:

First benchmark for automatic video generation from scientific papers (NeurIPS 2025)

In 2 lists

Paper2Poster:

Multi-agent system with Parser-Planner-Painter architecture converting paper.pdf to editable poster.pptx, outperforms GPT-4o with 87% fewer tokens

In 2 lists

AutoPR:

Fix issues with AI-generated pull requests, powered by ChatGPT

In 3 lists

Auto-Slides:

EvoPresent:

Paper2All:

AI-powered pipeline converting papers into interactive websites, posters, and multimedia presentations with "Let's Make Your Paper Alive!" philosophy

In 2 lists

AutoPage:

pdf2video:

Idea2Paper:

PaperX:

figures4papers:

PaperBanana:

Automated academic illustration generation for AI scientists, converting research papers into publication-ready figures using VLMs and diffusion models with iterative refinement (PKU & Google Research, 6.2K+ stars, 2026)

In 2 lists

PaperBanana-Pro:

AutoFigure:

FigureWeave:

EditDeck:

AutoFigure-Edit:

Kahneman4Review:

Academic Figure Generator:

PaperFit:

EvoScientist:

Self-evolving AI scientist with 6 specialized sub-agents (plan/research/code/debug/analyze/write) and persistent memory, #1 on DeepResearch Bench II and AstaBench, supporting multi-provider LLMs and multi-channel deployment (Apache 2.0, 2026)

In 2 lists

Auto-claude-code-research-in-sleep:

ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works with Claude Code, Codex, OpenClaw, or any LLM agent.

In 7 listsDetails

ArgusBot:

Station:

Dr.Claw:

Open-source research workspace with sequential idea-to-paper pipelines and integrated autoresearch tool packs.

In 2 lists

Redigg:

AutoResearchClaw:

April 2026 open-source human-in-the-loop system with six intervention modes (full-auto, gate-only, checkpoint, step-by-step, co-pilot, custom), SmartPause confidence-driven dynamic suspension, and Intervention Learning from human corrections. The cost-guardrail system — aborting runs that exceed…

In 5 listsDetails

NanoResearch:

End-to-end autonomous AI research engine that turns an idea into a complete LaTeX paper by dispatching real computational experiments to local GPUs or SLURM clusters, collecting actual results, generating figures/tables, and writing a data-grounded manuscript rather than LLM hallucinations…

In 3 lists

ScienceClaw:

EurekaClaw:

Claude-scholar:

Semi-automated research assistant for academic research and software development, supporting Claude Code, Codex CLI, Kimi Code CLI, and OpenCode across ideation, coding, experiments, writing, and publication (Galaxy-Dawn, 4.5K+ stars, MIT License, 2026)

In 3 lists

claude-scientific-skills:

by K-Dense - "A set of ready-to-use Agent Skills for research, science, engineering, analysis, finance and writing." That's their description - modest, simple. That's how you can tell this is really one of the best skills repos on GitHub. If you've ever thought about getting a PhD... just read all…

In 5 listsDetails

K-Dense BYOK:

Free, open-source desktop AI research assistant that runs locally and turns natural-language requests into real data analysis, literature search, figure generation, and manuscript review; ships with 149 scientific skills, 326 workflow templates, and 229 databases across genomics, proteomics, drug…

In 2 lists

latex-paper-skills:

NeuriCo:

AutoResearch :

Andrej Karpathy's autonomous LLM research framework: AI agent runs overnight experiments on a real training setup, auto-editing code→5min training→evaluation in a loop, ~100 experiments per night on a single GPU

In 5 listsDetails

RD-Agent :

Open-source LLM-powered R&D agent framework automating data-driven AI solution building through automated research, development, and evolution; achieves top open-source performance on MLE-Bench with dual Researcher-Developer agents and supports research copilot, data mining, Kaggle, and quant R&D…

In 5 listsDetails

DeepScientist :

First system progressively surpassing human SOTA on frontier AI tasks (183.7%, 1.9%, 7.9% improvements), month-long autonomous discovery with 20,000+ GPU hours

In 2 lists

Deep Researcher Agent:

academic-research-skills:

Comprehensive Claude Code skill suite covering the full academic pipeline from deep research and paper writing to multi-perspective peer review, revision, and finalization; features multi-agent teams, PRISMA systematic review, style calibration, claim-level citation audits, integrity gates, and…

In 4 listsDetails

Supervisor-Skills:

代码 Coding

Cloi CLI

Local debugging agent that runs in your terminal.

Devin

Autonomous AI software engineer by Cognition with its own IDE, shell, browser, and cloud sandbox.

In 5 listsDetails

v0

Prompt-driven UI generation for React and Next.js, creating production-ready components.

In 9 listsDetails

Blot.new

Build, edit, and deploy full-stack web apps in the browser using natural language with one-click deployment.

In 9 listsDetails

cursor

AI Code Editor with Cloud Agents, JetBrains integration, and 30+ plugins from partners like Atlassian, Datadog, and GitLab.

In 12 listsDetails

Windsurf

agentic IDE, "where the work of developers and AI truly flow together, allowing for a coding experience that feels like literal magic"

In 5 listsDetails

cline

autonomous coding agent right in your IDE, capable of creating/editing files, executing commands, using the browser, and more with your permission every step of the way

In 7 listsDetails

Trae

Free AI IDE by ByteDance with Builder Mode, free access to GPT-4o, Claude Sonnet, and DeepSeek R1.

In 4 listsDetails

MGX

Roo Code

An AI-powered autonomous coding agent integrated directly into VS Code. #opensource

In 3 lists

Kilo Code

Open Source AI coding assistant for planning, building, and fixing code. We're a superset of Roo, Cline, and our own features. Follow us: kilocode.ai/social

In 4 listsDetails

AugmentCode

AI coding platform with deep cross-repo codebase understanding via its Context Engine, built for large enterprise codebases.

In 4 listsDetails

Claude Code (🔥)

Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.

In 4 listsDetails

Gemini CLI

The official open-source AI agent that brings the power of Gemini directly into your terminal. Features context-aware coding assistance, file manipulation, and command execution capabilities.

In 9 listsDetails

Serena

Powerful MCP toolkit for coding agents providing semantic retrieval and editing capabilities. Integrates language servers for IDE-level code understanding. MIT licensed.

In 4 listsDetails

Claudia

OpenCode

Open-source terminal AI agent (95K+ GitHub stars) supporting 75+ providers. Free, privacy-first, with LSP integration.

In 4 listsDetails

Kiro

Spec-driven development. Write specs → auto-generate tasks → implement. DevOps automation.

In 4 listsDetails

CodeBuddy

In 2 lists

CodeX (🔥)

OpenAI's official autonomous coding agent CLI — the open-source reference implementation of the Codex harness with sandboxed tool execution, multi-file editing, and a streaming agent loop. Worth studying because it is the most widely adopted terminal-native coding agent harness and exposes the…

In 9 listsDetails

Kimi-CLI

[Archived] Legacy Python Kimi CLI, no longer maintained. Please use Kimi Code CLI: https://github.com/MoonshotAI/kimi-code (archived)

In 3 lists

opencode

Open-source terminal-native AI coding agent with 131K+ stars and 2.5M+ monthly active developers. Provider-agnostic architecture supports 75+ LLM providers plus native LSP auto-configuration, multi-session parallel agents, and MCP extensibility. The build/plan agent split and client/server…

In 5 listsDetails

Multica

Managed agents platform where you assign tasks, track progress, and let agents compound skills between runs.

In 3 lists

Atomic Agent

Local-first coding agent that runs open-weight models entirely on your machine via a llama.cpp fork. 56 built-in tools (browser, filesystem, git, memory, vision), MCP support, and a five-layer memory system. No account or API key required.

In 2 lists

DeepSeek Harness

Everything is a Plugin.

In 5 listsDetails

视频 Video

HunyuanVideo

HunyuanVideo is an open-source video foundation model that demonstrates performance in video generation comparable to, or even surpassing, leading closed-source models. It encompasses a comprehensive framework integrating data curation, advanced architectural design, progressive model scaling and…

In 4 listsDetails

CogVideo

Tsinghua/Zhipu open-source, multiple sizes.See also: Software Reference → AI Video Generation Software

In 5 listsDetails

Wan2.1

Wan: Open and Advanced Large-Scale Video Generative Models.

In 4 listsDetails

Open-Sora

Open-Sora: Democratizing Efficient Video Production for All

In 4 listsDetails

Open-Sora-Plan

This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.

In 4 listsDetails

LTX-Video

LTX-Video is the first DiT-based video generation model that can generate high-quality videos in real-time. It can generate 24 FPS videos at 768x512 resolution, faster than it takes to watch them.

In 5 listsDetails

Step-Video-T2V

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

In 2 lists

Step1X-Edit

Editing

Wan2.1-VACE

Editing

ICEdit

Editing

mochi-1-preview

Wan2.1-Fun

Wan2.1-FLF2V

首尾帧

MAGI-1

自回归模型

SkyReels-V2

FramePack

Lets make video diffusion practical!

In 2 lists

Pusa-VidGen

Wan2.2

Wan: Open and Advanced Large-Scale Video Generative Models.

In 3 lists

MoGA

长视频

LongCat-Video

HunyuanVideo-1.5

HunyuanVideo-1.5: A leading lightweight video generation model.

In 2 lists

LTX-2

Training

Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.

In 3 lists

daVinci-MagiHuman

LongLive

LongLive: Real-time Interactive Long Video Generation.

In 2 lists

JoyAI-Echo

NAVA

LTX-2.3 (🔥)

LingBot-Video

MiniMax-H3 (🔥)

MAGI-2-preview

LTX-2.5 (🔥)

Ditto:

Bernini (🔥):

JoyAI-Video-Edit:

https://github.com/hao-ai-lab/FastVideo

https://github.com/tdrussell/diffusion-pipe

https://github.com/VideoVerses/VideoTuna

(🔥)

High-Resolution Editable Toon Shading via Diffusion Models.

In 3 lists

https://github.com/huggingface/diffusers

A library that provides pre-trained diffusion models for generating and editing images, audio, and video.

In 5 listsDetails

https://github.com/kohya-ss/musubi-tuner

https://github.com/spacepxl/HunyuanVideo-Training

https://github.com/Tele-AI/TeleTron

https://github.com/Yaofang-Liu/Mochi-Full-Finetuner

https://github.com/bghira/SimpleTuner

https://github.com/X-GenGroup/Flow-Factory

https://github.com/shengshu-ai/minWM

world model

https://github.com/ModelTC/LightX2V

https://github.com/thu-ml/TurboDiffusion

PySceneDetect

Python and OpenCV-based scene cut/transition detection program & library.

In 3 lists

DOVER

Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives.

ArtiMuse

Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding.

图片 Image

awesome-nano-banana

Awesome-Nano-Banana-images

https://huggingface.co/inclusionAI/TwinFlow

ERNIE-Image:

ERNIE-Image-Turbo:

Nucleus-Image:

HiDream-O1-Image:

Ideogram 4:

Boogu-Image:

Krea-2-Raw:

SenseNova-U1:

Mage-Flow:

Qwen-Image-2.1:

a unified text-to-image generation and image editing model in the Qwen family

In 2 lists

ChronoEdit-14B:

Eigen-Banana-Qwen-Image-Edit:

Qwen-Image-Edit-2509:

Upscale:

Multiple-angles:

Multi-Angle-Lighting:

LongCat-Image-Edit:

Qwen-Image-Edit-2511:

Qwen-Image-Edit-2511-Upscale2K:

Qwen-Image-Edit-2511-Multiple-Angles-LoRA:

FireRed-Image-Edit:

JoyAI-Image-Edit:

GLM-Image:

an image generation model

In 2 lists

https://huggingface.co/black-forest-labs/FLUX.2-klein-4B

https://huggingface.co/black-forest-labs/FLUX.2-klein-9B

DreamLite:

SenseNova-U1:

https://github.com/kohya-ss/musubi-tuner

https://github.com/bghira/SimpleTuner

MS Training:

Finetune HunyuanImage-3.0:

OneTrainer:

One-stop solution for all your Diffusion training needs. Supports FLUX, Stable Diffusion 1.5/2.x/3.x/SDXL, Würstchen, PixArt, Hunyuan Video and more. Features full fine-tuning, LoRA, embeddings, masked training, automatic backups, and TensorBoard integration. GPL-3.0 licensed.

In 2 lists

Finetune LongCat-Image and Edit:

https://github.com/X-GenGroup/Flow-Factory

UniRL:

TypemovieInfer:

OpenSearch GPT

SearchGPT / Perplexity clone, but personalised for you.

MindSearch

An LLM-based Multi-agent Framework of Web Search Engine (like Perplexity.ai Pro and SearchGPT).

In 3 lists

nanoPerplexityAI

The simplest open-source implementation of perplexity.ai.

curiosity

Try to build a Perplexity-like user experience.

MiniPerplx

A minimalistic AI-powered search engine that helps you find information on the internet.

In 2 lists

语音 Speech

SpeechGPT-2.0-preview:

kokoro:

https://hf.co/hexgrad/Kokoro-82M

In 2 lists

Higgs Audio V2:

【Training】

In 2 lists

KittenTTS:

Kitten TTS is an open-source realistic text-to-speech model with just 15 million parameters, designed for lightweight deployment and high-quality voice synthesis.

In 4 listsDetails

ZipVoice:

VyvoTTS:

VibeVoice:

VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking.

In 5 listsDetails

Index-TTS-2:

FireRedTTS2:

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot.

In 2 lists

VoxCPM:

Open-sourced tokenizer-free multilingual speech synthesis model with high-quality TTS and style transfer workflows.

In 4 listsDetails

Neutts-Air:

Maya1:

VibeVoice:

GLM-TTS:

Fun-CosyVoice3:

Ming-Omni-TTS:

VoxCPM2:

OmniVoice:

High-Quality Voice Cloning TTS for 600+ Languages

In 2 lists

MOSS-TTS-v1.5:

Higgs Audio v3 TTS:

Confucius4-TTS:

Dia-1.6B:

FireRedTTS3:

Kyutai:

Whisper:

Whisper is a general-purpose speech recognition model that can be run locally offline. It can transcribe audio from and to multiple languages.

In 11 listsDetails

Audio Flamingo 3:

Voxtral:

Step-Audio2:

Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation.

In 2 lists

SoulX-Podcast:

Omnilingual ASR:

Fun-ASR:

FunASR:

industrial-grade ASR toolkit; 170× realtime on GPU, 50+ languages, built-in VAD, punctuation, speaker diarization, and emotion detection. Includes non-autoregressive SenseVoice and LLM-based Fun-ASR-Nano models.

In 9 listsDetails

SenseVoice:

SenseVoice is a speech foundation model with multiple speech understanding capabilities, including automatic speech recognition (ASR), spoken language identification (LID), speech emotion recognition (SER), and audio event detection (AED).

In 2 lists

VibeVoice-ASR:

Qwen3-ASR:

Mega-ASR:

MOSS-Transcribe-Diarize:

Fun-Audio-Chat:

Chroma 1.0:

AudioInteraction:

世界模型 World Models

ShadowDancer:

PhiZero:

Wonder:

open-dreamer:

Cosmos3-Edge:

WorldWander:

Matrix-Game-3.5:

ABot-World:

LingBot-World-V2:

AlayaWorld:

MoWorld:

Warp-as-History:

MIRA:

Multiplayer

ActWorld:

Sana-WM:

Cosmos-3:

Cosmos is a world model development platform that consists of world foundation models, tokenizers and video processing pipeline to accelerate the development of Physical AI at Robotics & AV labs.

In 3 lists

DreamX-World-5B:

DreamX-World-5B-Cam:

Astra:

Yume-5B:

LingBot-World:

HY-WorldPlay:

Matrix-Game-3.0:

Matrix-Game-2.0:

Waypoint-1.5-1B:

Hunyuan-GameCraft:

nano-world-model:

Minimalist, batteries-included repository for training video world models with diffusion-forcing, supporting long-horizon rollouts, 3D point-cloud generation, and model-predictive control with pretrained checkpoints (Simchowitz Lab, 700+ stars, MIT License, 2026)

In 2 lists

https://github.com/shengshu-ai/minWM

world model

stable-worldmodel:

OpenWorldLib:

A unified codebase for world models, providing a standardized pipeline interface over existing open-source models (Matrix-Game-2, Hunyuan-GameCraft, FlashWorld, Cosmos-Predict-2.5, and others) across video generation, 3D scene generation, and reasoning. Apache-2.0.

In 2 lists

BiWM:

龙虾 OpenClaw

MultiUserClaw:

ClawManager:

Qclaw:

NEXU:

The simplest desktop client for OpenClaw 🦞 — bridge your Agent to WeChat, Feishu, Slack & Discord in one click. Works with Claude Code, Codex & any LLM. BYOK, Oauth, local-first, chat from your phone 24/7.

In 2 lists

OpenHanako:

统一模型 Unified Model

MetaMorph:

Ullava:

ILLUME:

Transfusion:

JanusFlow:

UniUGG:

3D

LightBagel:

DreamLLM:

X-Omni:

Ming-flash-omni-Preview:

Omni-View:

NExT-OMNI:

Uni-MoE-2.0-Omni:

LongCat-Flash-Omni:

ShapeLLM-Omni:

UniGen-1.5:

Jodi:

UniModel:

TUNA:

HBridge:

EMMA:

OpenOmni:

Ming-Flash-Omni:

STAR:

InternVL-U:

LongCat-Next:

SenseNova-U1:

TUNA-2:

Lance:

JoyAI-VL-Interaction:

Interaction

MOSS-VL-Realtime:

Interaction

SenseNova-Vision:

Mage-VL:

Interaction

书籍 Book

《大规模语言模型:从理论到实践》

《大语言模型》

《动手学大模型Dive into LLMs》

《动手做AI Agent》

《Build a Large Language Model (From Scratch)》

Implementing a ChatGPT-like LLM from scratch, step by step

In 5 listsDetails

《多模态大模型》

《Generative AI Handbook: A Roadmap for Learning Resources》

《Understanding Deep Learning》

website with the book draft and Google Colabs of the book by Simon J.D. Prince

In 5 listsDetails

《Illustrated book to learn about Transformers & LLMs》

《Building LLMs for Production: Enhancing LLM Abilities and Reliability with Prompting, Fine-Tuning, and RAG》

《大型语言模型实战指南:应用实践与场景落地》

《Hands-On Large Language Models》

Covers LLM fundamentals, prompt engineering, and fine-tuning.

In 2 lists

《自然语言处理:大模型理论与实践》

《动手学强化学习》

《面向开发者的LLM入门教程》

《大模型基础》

Taming LLMs: A Practical Guide to LLM Pitfalls with Open Source Software

Foundations of Large Language Models

by Tong Xiao and Jingbo Zhu

In 2 lists

Textbook on reinforcement learning from human feedback

《大模型算法:强化学习、微调与对齐》

《从零开始构建智能体》——从零开始的智能体原理与实践教程

📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程

In 3 lists

《Hands-On Modern RL》

课程 Course

LLM Resources Hub

斯坦福 CS224N: Natural Language Processing with Deep Learning

This course provides a comprehensive insight into Deep Learning for NLP using PyTorch, emphasizing end-to-end neural models, eliminating the need for task-specific feature engineering, and equipping students with the skills to craft their own neural network solutions.

In 3 lists

吴恩达: Generative AI for Everyone

吴恩达: LLM series of courses

Focused courses on current generative AI engineering techniques.

In 2 lists

ACL 2023 Tutorial: Retrieval-based Language Models and Applications

llm-course: Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.

Free LLM course with roadmap-style progression and practical Colab notebooks.

In 4 listsDetails

微软: Generative AI for Beginners

21 lessons covering generative AI fundamentals, prompt engineering, RAG applications, fine-tuning, and LLM app deployment with practical exercises.

In 5 listsDetails

微软: State of GPT

HuggingFace NLP Course

清华 NLP 刘知远团队大模型公开课

斯坦福 CS25: Transformers United V4

This course delves into the transformative role of Transformers in deep learning, particularly their impact on the advancement of language models like ChatGPT and GPT-4.

In 5 listsDetails

斯坦福 CS324: Large Language Models

普林斯顿 COS 597G (Fall 2022): Understanding Large Language Models

An advanced exploration into the transformative realm of LLMs, discussing state-of-the-art models, their profound capabilities, and associated challenges, with an emphasis on in-depth research, ethical considerations, and hands-on project experience, tailored for seasoned students versed in…

In 2 lists

约翰霍普金斯 CS 601.471/671 NLP: Self-supervised Models

李宏毅 GenAI课程

openai-cookbook

Examples and guides for using the OpenAI API.

In 9 listsDetails

Hands on llms

Learn about LLM, LLMOps, and vector DBS for free by designing, training, and deploying a real-time financial advisor LLM system.

In 3 lists

滑铁卢大学 CS 886: Recent Advances on Foundation Models

Mistral: Getting Started with Mistral

Coursera: Chatgpt 应用提示工程

LangGPT

Empowering everyone to become a prompt expert!

In 3 lists

mistralai-cookbook

Introduction to Generative AI 2024 Spring

build nanoGPT

Video+code lecture on building nanoGPT from scratch.

LLM101n

Let's build a Storyteller.

Knowledge Graphs for RAG

LLMs From Scratch (Datawhale Version)

OpenRAG

通往AGI之路

Andrej Karpathy - Neural Networks: Zero to Hero

Build neural networks and language models from first principles.

In 3 lists

Interactive visualization of Transformer

Interactive visualization of how transformer-based LLMs work, running a live GPT-2 model in the browser. #opensource

In 2 lists

andysingal/llm-course

LM-class

Google Advanced: Generative AI for Developers Learning Path

Anthropics:Prompt Engineering Interactive Tutorial

Prompt Engineering Interactive Tutorial by Anthropic

In 4 listsDetails

《大型语言模型实战指南:应用实践与场景落地》

Large Language Model Agents

In 2 lists

Cohere LLM University

free course on LLMs, embeddings, semantic search, and NLP applications.

In 3 lists

LLMs and Transformers

Smol Vision

Recipes for shrinking, optimizing, customizing cutting edge vision models.

Multimodal RAG: Chat with Videos

LLMs Interview Note

RAG++ : From POC to production

Advanced RAG course.

Weights & Biases AI Academy

Finetuning, building with LLMs, Structured outputs and more LLM courses.

Prompt Engineering & AI tutorials & Resources

Learn RAG From Scratch – Python AI Tutorial from a LangChain Engineer

LLM Evaluation: A Complete Course

Learn to build modern software with LLMs using the newest tools and techniques in the field.

In 2 lists

HuggingFace Learn

Free hands-on courses using only open models.

In 3 lists

Andrej Karpathy: Deep Dive into LLMs like ChatGPT

LLM技术科普

RAG Techniques

This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. RAG systems combine information retrieval with generative models to provide accurate and contextually rich responses.

In 5 listsDetails

100+ LLM & RL Algorithm Maps | 原创 LLM / RL 100+原理图

Reinforcement Learning of Large Language Models

NanoChat

The best ChatGPT that $100 can buy.

In 4 listsDetails

斯坦福CS146S: The Modern Software Developer

教程 Tutorial

动手学大模型应用开发

AI开发者频道

B站:五里墩茶社

B站:木羽Cheney

YTB:AI Anytime

B站:漆妮妮

Prompt Engineering Guide

This guide introduces Prompt Engineering, a discipline that optimizes interactions with Large Language Models, offering extensive resources, research, and tools.

In 5 listsDetails

YTB: AI超元域

B站:TechBeat人工智能社区

B站:黄益贺

B站:深度学习自然语言处理

LLM Visualization

In 2 lists

知乎: 原石人类

B站:小黑黑讲AI

B站:面壁的车辆工程师

B站:AI老兵文哲

Large Language Models (LLMs) with Colab notebooks

YTB:IBM Technology

YTB: Unify Reading Paper Group

Chip Huyen

ML Engineering, MLOps, and the use of ML in startups

In 2 lists

How Much VRAM

Blog: 科学空间(苏剑林)

In 2 lists

YTB: Hyung Won Chung

Blog: Tejaswi kashyap

Blog: 小昇的博客

知乎: ybq

W&B articles

Huggingface Blog

Blog: GbyAI

LLM-Action

Blog: Lil’Log (OponAI)

In 2 lists

B站: 毛玉仁

AI-Guide-and-Demos

cnblog: 第七子

Implementation of all RAG techniques in a simpler way.

Implementation of all RAG techniques in a simpler way

In 2 lists

Theoretical Machine Learning: A Handbook for Everyone

鱼皮的 Vibe Coding 零基础教程

程序员鱼皮的 AI 资源大全 + Vibe Coding 零基础教程,分享 OpenClaw 保姆级教程、大模型玩法(DeepSeek / GPT / Gemini / Claude / GLM)、最新 AI 资讯、Prompt 提示词大全、AI 知识百科(Agent Skills / RAG / MCP / A2A)、AI 编程教程(Harness Engineering)、AI 工具用法(Cursor / Claude Code / TRAE / Codex / Copilot)、AI 开发框架教程(Spring AI / LangChain)、AI 产品变现指南,帮你快速掌握 AI…

In 5 listsDetails

论文 Paper

Hermes-3-Technical-Report

The Llama 3 Herd of Models

(Meta, 2024-2025) - widely adopted open-weight family; default base for fine-tuning across NLP tasks.

In 2 lists

Qwen Technical Report

Qwen2 Technical Report

Qwen2-vl Technical Report

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Baichuan 2: Open Large-scale Language Models

DataComp-LM: In search of the next generation of training sets for language models

OLMo: Accelerating the Science of Language Models

MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series

Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Jamba-1.5: Hybrid Transformer-Mamba Models at Scale

Jamba: A Hybrid Transformer-Mamba Language Model

(2024) - hybrid Mamba-Transformer-MoE architecture.

In 2 lists

Textbooks Are All You Need

Preprint

In 2 lists

Unleashing the Power of Data Tsunami: A Comprehensive Survey on Data Assessment and Selection for Instruction Tuning…

data

OLMoE: Open Mixture-of-Experts Language Models

Model Merging Paper

Baichuan-Omni Technical Report

1.5-Pints Technical Report: Pretraining in Days, Not Months – Your Language Model Thrives on Quality Data

Baichuan Alignment Technical Report

Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models

TÜLU 3: Pushing Frontiers in Open Language Model Post-Training

(AI2, 2024) - fully open post-training recipe with state-of-the-art results among open models.

In 2 lists

Phi-4 Technical Report

(Microsoft, 2024) - small models trained on curated data, competitive with much larger ones on NLP benchmarks.

In 2 lists

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Qwen2.5 Technical Report

YuLan-Mini: An Open Data-efficient Language Model

An Introduction to Vision-Language Modeling

DeepSeek V3 Technical Report

2 OLMo 2 Furious

(AI2, 2025) - fully open: weights, training data, code; reproducibility benchmark.

In 2 lists

Yi-Lightning Technical Report

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning.

In 2 lists

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

Qwen2.5-VL Technical Report

Baichuan-M1: Pushing the Medical Capability of Large Language Models

Predictable Scale: Part I -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

SkyLadder: Better and Faster Pretraining via Context Window Scheduling

Qwen2.5-Omni technical report

Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs

Gemma 3 Technical Report

(Google, 2025) - 1B-27B open models with high local-to-global attention ratio to keep KV-cache tractable at 128K context.

In 2 lists

Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources

Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs

MiMo: Unlocking the Reasoning Potential of Language Model – From Pretraining to Posttraining

Llama-Nemotron: Efficient Reasoning Models

Qwen3 Technical Report

Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.

In 5 listsDetails

MiMo-VL Technical Report

Kwai Keye-VL Technical Report

Kimi K2 Technical Report

Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters.

In 3 lists

KAT-V1: Kwai-AutoThink Technical Report

Step3

SAIL-VL2 Technical Report

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

85M-Midtraining Data 22M Instruct Data

Olmo3

Charting a path through the model flow to lead open-source AI. Website

OpenMMReasoner

Qwen3-VL Technical Report

社区 Community

魔乐社区

HuggingFace

Popular open platform for sharing ML models, datasets, and collaborating on NLP and generative AI projects.

In 7 listsDetails

ModelScope

WiseModel

OpenCSG

模型上下文协议 MCP

MCP是啥?技术原理是什么?一个视频搞懂MCP的一切。Windows系统配置MCP,Cursor,Cline 使用MCP

MCP是什么?为啥是下一代AI标准?MCP原理+开发实战!在Cursor、Claude、Cline中使用MCP,让AI真正自动化!

从零编写MCP并发布上线,超简单!手把手教程

smithery.ai

mcp.so

Platform for MCP server resources and community.

In 2 lists

modelcontextprotocol/servers

Anthropic's official reference MCP server implementations (GitHub, Slack, Postgres, Puppeteer, etc.). The authoritative source for understanding correct MCP server structure before building your own.

In 6 listsDetails

mcp.ad

pulsemcp.com

awesome-mcp-servers

Curated community list of MCP servers.

In 4 listsDetails

glama.ai

mcp.composio.dev

Connect Cursor, Windsurf, and Claude to 100+ fully managed MCP Servers with built-in auth; These servers are built by the community and are hosted by Composio

In 2 lists

awesome-mcp-list

mcpo

FastMCP

The fast, Pythonic way to build Model Context Protocol servers 🚀

In 2 lists

sharemcp.cn

mcpstore.co

FastAPI-MCP

Expose your FastAPI endpoints as Model Context Protocol (MCP) tools, with Auth!

In 2 lists

modelscope/mcp

mcpm.sh

技能 Skills

Agent Skills (Claude Skills) 详细攻略,一期视频精通

OpenClaw

awesome-claude-skills

Curated list of Claude Skills, plugins, resources, and custom commands to extend terminal and API workflows. Apache 2.0 licensed.

In 4 listsDetails

Anthropics Skills

by Anthropic - Anthropic's official repository for Agent Skills — the SKILL.md format, a skill template, and example skills, the same format Claude Code loads natively.

In 9 listsDetails

Skillsmp

awesome-claude-skills

security-flavoured skill list with design-engineering crossover

In 2 lists

ClawHub

水产市场

Skills.Sh

Open ecosystem by Vercel for installing reusable AI agent skills with a single command across 18+ platforms.

In 2 lists

awesome-agent-skills

1000+ skills incl. design-md, enhance-prompt, react-components, shadcn-ui

In 3 listsDetails

llmbase

awesome-openclaw-skills

Details

claude-scientific-skills:

by K-Dense - "A set of ready-to-use Agent Skills for research, science, engineering, analysis, finance and writing." That's their description - modest, simple. That's how you can tell this is really one of the best skills repos on GitHub. If you've ever thought about getting a PhD... just read all…

In 5 listsDetails

LLMs-Universal-Life-Science-and-Clinical-Skills-

SkillHub

LabClaw

Skill operating layer for biomedical AI agents with 211 production-ready SKILL.md files across 7 domains (biology, pharmacology, medicine, data science, literature search), enabling modular dry-lab reasoning and protocol composition for Stanford LabOS-compatible agents

In 2 lists

Modelscope Skills

Agent Skill

mmx-cli

cocoloop hub

推理 Open o1

https://github.com/atfortes/Awesome-LLM-Reasoning

https://github.com/hijkzzz/Awesome-LLM-Strawberry

https://github.com/wjn1996/Awesome-LLM-Reasoning-Openai-o1-Survey

https://github.com/srush/awesome-o1

https://github.com/open-thought/system-2-research

https://github.com/ninehills/blog/issues/121

https://github.com/OpenSource-O1/Open-O1

https://github.com/GAIR-NLP/O1-Journey

https://github.com/marlaman/show-me

A visual and transparent alternative to open-source ChatGPT O1

In 2 lists

g1

Using Llama-3.1 70b on Groq to create o1-like reasoning chains.

https://github.com/Jaimboh/Llamaberry-Chain-of-Thought-Reasoning-in-AI

https://github.com/pseudotensor/open-strawberry

https://huggingface.co/collections/peakji/steiner-preview-6712c6987110ce932a44e9a6

https://github.com/SimpleBerry/LLaMA-O1

https://huggingface.co/collections/Skywork/skywork-o1-open-67453df58e12f6c3934738d0

https://huggingface.co/collections/Qwen/qwq-674762b79b75eac01735070a

https://github.com/SkyworkAI/skywork-o1-prm-inference

https://github.com/RifleZhang/LLaVA-Reasoner-DPO

https://github.com/ADaM-BJTU

https://github.com/ADaM-BJTU/OpenRFT

https://github.com/RUCAIBox/Slow_Thinking_with_LLMs

https://github.com/richards199999/Thinking-Claude

https://huggingface.co/AGI-0/Art-v0-3B

https://huggingface.co/deepseek-ai/DeepSeek-R1

https://huggingface.co/deepseek-ai/DeepSeek-R1-Zero

https://github.com/huggingface/open-r1

Fully open reproduction of DeepSeek-R1

In 3 lists

https://github.com/hkust-nlp/simpleRL-reason

https://github.com/Jiayi-Pan/TinyZero

Minimal reproduction of DeepSeek R1-Zero

In 2 lists

https://github.com/baichuan-inc/Baichuan-M1-14B

https://github.com/EvolvingLMMs-Lab/open-r1-multimodal

https://github.com/open-thoughts/open-thoughts

Mini-R1:

LLaMA-Berry:

MCTS-DPO:

OpenR:

https://arxiv.org/abs/2410.02725

LLaVA-o1:

code model

In 2 lists

Marco-o1:

[code] [model]

In 2 lists

OpenAI o1 report:

DRT-o1:

https://arxiv.org/abs/2412.09413

https://arxiv.org/abs/2501.02497

https://arxiv.org/abs/2501.18585

https://github.com/simplescaling/s1

s1: Simple test-time scaling.

In 2 lists

https://github.com/Deep-Agent/R1-V

https://github.com/StarRing2022/R1-Nature

https://github.com/Unakar/Logic-RL

https://github.com/datawhalechina/unlock-deepseek

https://github.com/GAIR-NLP/LIMO

https://github.com/Zeyi-Lin/easy-r1

https://github.com/jackfsuia/nanoRLHF/tree/main/examples/r1-v0

https://github.com/FanqingM/R1-Multimodal-Journey

https://github.com/dhcode-cpp/X-R1

https://github.com/agentica-project/deepscaler

https://github.com/ZihanWang314/RAGEN

https://github.com/sail-sg/oat-zero

https://github.com/TideDra/lmm-r1

https://github.com/FlagAI-Open/OpenSeek

https://github.com/SwanHubX/ascend_r1_turtorial

https://github.com/om-ai-lab/VLM-R1

https://github.com/wizardlancet/diagnosis_zero

https://github.com/lsdefine/simple_GRPO

https://github.com/brendanhogan/DeepSeekRL-Extended

https://github.com/Wang-Xiaodong1899/Open-R1-Video

https://github.com/Open-Reasoner-Zero/Open-Reasoner-Zero

https://github.com/lucasjinreal/Namo-R1

https://github.com/hiyouga/EasyR1

Efficient, scalable, multi-modality RL training framework based on veRL. Extends veRL to support vision-language models with GRPO algorithm for efficient RL training. Apache 2.0 licensed.

In 3 lists

https://github.com/Fancy-MLLM/R1-Onevision

https://github.com/tulerfeng/Video-R1

https://huggingface.co/qihoo360/TinyR1-32B-Preview

https://github.com/facebookresearch/swe-rl

Meta/UIUC/CMU

In 2 lists

https://github.com/turningpoint-ai/VisualThinker-R1-Zero

https://github.com/yuyq96/R1-Vision

https://github.com/sungatetop/deepseek-r1-vision

https://huggingface.co/qihoo360/Light-R1-32B

https://github.com/Liuziyu77/Visual-RFT

Shanghai AI Lab / SJTU

In 2 lists

https://github.com/Mohammadjafari80/GSM8K-RLVR

https://github.com/ModalMinds/MM-EUREKA

https://github.com/joey00072/nanoGRPO

https://github.com/PeterGriffinJin/Search-R1

UIUC/Google

In 2 lists

https://openi.pcl.ac.cn/PCL-Reasoner/GRPO-Training-Suite

https://github.com/dvlab-research/Seg-Zero

https://github.com/HumanMLLM/R1-Omni

https://github.com/OpenManus/OpenManus-RL

UIUC/MetaGPT

In 2 lists

https://arxiv.org/pdf/2503.07536

https://github.com/Osilly/Vision-R1

https://github.com/LengSicong/MMR1

https://github.com/phonism/CP-Zero

https://github.com/SkyworkAI/Skywork-R1V

https://arxiv.org/abs/2503.13939v1

https://github.com/0russwest0/Agent-R1

USTC

In 2 lists

https://github.com/MetabrainAGI/Awaker2.5-R1

https://github.com/LG-AI-EXAONE/EXAONE-Deep

https://github.com/qiufengqijun/open-r1-reprod

https://github.com/SUFE-AIFLM-Lab/Fin-R1

https://github.com/sail-sg/understand-r1-zero

https://github.com/baibizhe/Efficient-R1-VLLM

https://arxiv.org/abs/2502.19655

https://arxiv.org/abs/2503.21620v1

https://arxiv.org/abs/2503.16081

https://github.com/ShadeCloak/ADORA

https://github.com/appletea233/Temporal-R1

AReaL:

AntGroup/Tsinghua

In 2 lists

https://github.com/lzhxmu/CPPO

https://arxiv.org/abs/2503.23829

https://github.com/TencentARC/SEED-Bench-R1

https://github.com/McGill-NLP/nano-aha-moment

https://github.com/VLM-RL/Ocean-R1

https://github.com/OpenGVLab/VideoChat-R1

https://github.com/ByteDance-Seed/Seed-Thinking-v1.5

https://github.com/SkyworkAI/Skywork-OR1

Skywork AI

In 2 lists

https://github.com/MoonshotAI/Kimi-VL

https://arxiv.org/abs/2504.08600

https://github.com/ZhangXJ199/TinyLLaVA-Video-R1

https://arxiv.org/abs/2504.11914

https://github.com/policy-gradient/GRPO-Zero

https://github.com/linkangheng/PR1

https://github.com/jiangxinke/Agentic-RAG-R1

PKU

In 2 lists

https://github.com/shangshang-wang/Tina

https://github.com/aliyun/qwen-dianjin

https://github.com/RAGEN-AI/RAGEN

RAGEN-AI

In 2 lists

MiMo: Unlocking the Reasoning Potential of Language Model – From Pretraining to Posttraining

https://github.com/yuanzhoulvpi2017/nano_rl

https://huggingface.co/a-m-team/AM-Thinking-v1

https://huggingface.co/Intelligent-Internet/II-Medical-8B

https://github.com/CSfufu/Revisual-R1

推理 Open o3

Mini-o3:

Simple-o3:

Thyme:

Open o3 Video:

小语言模型 Small Language Model

https://github.com/jiahe7ay/MINI_LLM

https://github.com/jingyaogong/minimind

Train a 64M-parameter LLM from scratch in just 2 hours for $3. Complete from-scratch implementation covering MoE, data cleaning, pretraining, SFT, LoRA, RLHF (DPO/PPO/GRPO), tool use, and model distillation. All core algorithms implemented in pure PyTorch without high-level abstractions.…

In 4 listsDetails

https://github.com/DLLXW/baby-llama2-chinese

https://github.com/charent/ChatLM-mini-Chinese

https://github.com/wdndev/tiny-llm-zh

https://github.com/Tongjilibo/build_MiniLLM_from_scratch

https://github.com/jzhang38/TinyLlama

https://github.com/AI-Study-Han/Zero-Chatgpt

https://github.com/loubnabnl/nanotron-smol-cluster

(使用Cosmopedia训练cosmo-1b)

https://github.com/charent/Phi2-mini-Chinese

https://github.com/allenai/OLMo

Open Language Model

In 2 lists

https://github.com/keeeeenw/MicroLlama

https://github.com/Chinese-Tiny-LLM/Chinese-Tiny-LLM

https://github.com/leeguandong/MiniLLaMA3

https://github.com/Pints-AI/1.5-Pints

https://github.com/zhanshijinwat/Steel-LLM

https://github.com/RUC-GSAI/YuLan-Mini

https://github.com/Om-Alve/smolGPT

https://github.com/skyzh/tiny-llm

learn LLM inference system on Apple Silicon for systems engineers: build a tiny vLLM + Qwen

In 2 lists

https://github.com/qibin0506/Cortex

https://github.com/huggingface/picotron

https://github.com/Alic-Li/Mini_RWKV_7

https://huggingface.co/Nanbeige/Nanbeige4-3B-Thinking-2511

23T tokens预训练模型

https://github.com/stepfun-ai/SteptronOss

https://github.com/huggingface/nanowhale

https://github.com/OpenBMB/ForgeTrain

https://github.com/vukrosic/glm-5.3-flash-from-scratch

小多模态模型 Small Vision Language Model

https://github.com/jingyaogong/minimind-v

🚀 「大模型」3小时从0训练27M参数的视觉多模态VLM!🌏 Train a 27M-parameter VLM from scratch in just 3 hours!

In 2 lists

https://github.com/yuanzhoulvpi2017/zero_nlp/tree/main/train_llava

https://github.com/AI-Study-Han/Zero-Qwen-VL

https://github.com/Coobiw/MPP-LLaVA

https://github.com/qnguyen3/nanoLLaVA

TinyLLaVA Factory

A Framework of Small-scale Large Multimodal Models.

https://github.com/ZhangXJ199/TinyLLaVA-Video

https://github.com/Emericen/tiny-qwen

Smol Vision

Recipes for shrinking, optimizing, customizing cutting edge vision models.

https://github.com/huggingface/nanoVLM

https://github.com/GeeeekExplorer/nano-vllm

Minimalist vLLM implementation in ~1,200 lines of Python. Educational yet performant with prefix caching, tensor parallelism, and CUDA graph acceleration. Comparable inference speeds to full vLLM. MIT licensed.

In 5 listsDetails

https://github.com/ritabratamaiti/AnyModal

https://github.com/yujunhuics/Reyes

https://github.com/Victorwz/Open-Qwen2VL

https://github.com/jingyaogong/minimind-o

技巧 Tips

What We Learned from a Year of Building with LLMs (Part I)

What We Learned from a Year of Building with LLMs (Part II)

What We Learned from a Year of Building with LLMs (Part III): Strategy

轻松入门大语言模型(LLM)

LLMs for Text Classification: A Guide to Supervised Learning

Unsupervised Text Classification: Categorize Natural Language With LLMs

Text Classification With LLMs: A Roundup of the Best Methods

LLM Pricing

Uncensor any LLM with abliteration

Tiny LLM Universe

https://github.com/AI-Study-Han/Zero-Chatgpt

https://github.com/AI-Study-Han/Zero-Qwen-VL

finetune-Qwen2-VL

Quick Start for Fine-tuning or continue pre-train Qwen2-VL Model.

https://github.com/Coobiw/MPP-LLaVA

https://github.com/Tongjilibo/build_MiniLLM_from_scratch

https://github.com/wdndev/tiny-llm-zh

https://github.com/jingyaogong/minimind

Train a 64M-parameter LLM from scratch in just 2 hours for $3. Complete from-scratch implementation covering MoE, data cleaning, pretraining, SFT, LoRA, RLHF (DPO/PPO/GRPO), tool use, and model distillation. All core algorithms implemented in pure PyTorch without high-level abstractions.…

In 4 listsDetails

LLM-Travel

致力于深入理解、探讨以及实现与大模型相关的各种技术、原理和应用

Knowledge distillation: Teaching LLM's with synthetic data

Part 1: Methods for adapting large language models

Part 2: To fine-tune or not to fine-tune

Part 3: How to fine-tune: Focus on effective datasets

Reader-LM: Small Language Models for Cleaning and Converting HTML to Markdown

LLMs应用构建一年之心得

LLM训练-pretrain

pytorch-llama

LLaMA 2 implemented from scratch in PyTorch.

Preference Optimization for Vision Language Models with TRL

【support model】

Fine-tuning visual language models using SFTTrainer

【docs】

A Visual Guide to Mixture of Experts (MoE)

Role-Playing in Large Language Models like ChatGPT

Distributed Training Guide

Best practices & guides on how to write distributed pytorch training code.

Chat Templates

Top 20+ RAG Interview Questions

LLM-Dojo 开源大模型学习场所,使用简洁且易阅读的代码构建模型训练框架

o1 isn’t a chat model (and that’s the point)

Beam Search快速理解及代码解析

基于 transformers 的 generate() 方法实现多样化文本生成:参数含义和算法原理解读

The Ultra-Scale Playbook: Training LLMs on GPU Clusters

In 2 lists
See category
94

Table of Contents

hesreallyhim/awesome-claude-code

A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…

Fresh★ 55k202 entriesPushed today
94

Awesome Agent Skills

VoltAgent/awesome-agent-skills

A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.

Fresh★ 35k839 entriesPushed today
93

Awesome Machine Learning

josephmisiti/awesome-machine-learning

A curated list of awesome Machine Learning frameworks, libraries and software.

Fresh★ 74k1188 entriesPushed 7 days ago
92

Awesome Production Machine Learning

EthicalML/awesome-production-machine-learning

A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning

Fresh★ 21k519 entriesPushed 3 days ago
92

AWESOME DATA SCIENCE

academic/awesome-datascience

:memo: An awesome Data Science repository to learn and apply for real world problems.

Fresh★ 30k881 entriesPushed today
91

Static Analysis

analysis-tools-dev/static-analysis

⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…

Fresh★ 15k528 entriesPushed 8 days ago