Skip to content
87

Awesome local LLM

A curated list of awesome platforms, tools, practices and resources that helps run LLMs locally

2.9k stars412 forks309 entriesLast push Sep 28, 2026 (yesterday)License MIT

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

Inference platforms

LM Studio

discover, download and run local LLMs

In 11 listsDetails

unsloth

unified web UI for training and running open models like Qwen, DeepSeek, and Gemma locally

In 8 listsDetails

LocalAI

the free, open-source alternative to OpenAI, Claude and others

In 14 listsDetails

jan

an open source alternative to ChatGPT that runs 100% offline on your computer

In 2 lists

ChatBox

user-friendly desktop client app for AI models/LLMs

In 3 lists

lemonade

a local LLM server with GPU and NPU Acceleration

In 2 lists

Inference engines

ollama

get up and running with LLMs

In 12 listsDetails

llama.cpp

LLM inference in C/C++

In 10 listsDetails

vllm

a high-throughput and memory-efficient inference and serving engine for LLMs

In 11 listsDetails

exo

run your own AI cluster at home with everyday devices

In 4 listsDetails

BitNet

official inference framework for 1-bit LLMs

In 5 listsDetails

sglang

a fast serving framework for large language models and vision language models

In 9 listsDetails

omlx

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

In 3 lists

Nano-vLLM

a lightweight vLLM implementation built from scratch

In 5 listsDetails

TensorRT-LLM

provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs

In 7 listsDetails

koboldcpp

run GGUF models easily with a KoboldAI UI

In 5 listsDetails

dynamo

a datacenter scale distributed inference serving framework

In 3 lists

mlx-lm

generate text and fine-tune large language models on Apple silicon with MLX

In 3 lists

mistral.rs

fast, flexible LLM inference

In 3 lists

flashinfer

kernel library for LLM serving

In 2 lists

LiteRT-LM

Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices

In 4 listsDetails

gpustack

simple, scalable AI model deployment on GPU clusters

In 5 listsDetails

mlx-vlm

a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX

In 3 lists

mini-sglang

a lightweight yet high-performance inference framework for Large Language Models

In 3 lists

executorch

on-device AI across mobile, embedded and edge for PyTorch

In 2 lists

LiteRT

Google's on-device framework for high-performance ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and optimization

In 4 listsDetails

distributed-llama

connect home devices into a powerful cluster to accelerate LLM inference

In 3 lists

ik_llama.cpp

llama.cpp fork with additional SOTA quants and improved performance

In 3 lists

tokenspeed

a speed-of-light LLM inference engine

In 2 lists

sonar

large-scale LLM inference engine based on vLLM

FastFlowLM

run LLMs on AMD Ryzen™ AI NPUs

krasis

a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware

llm-scaler

run LLMs on Intel Arc™ Pro B60 and B70 GPUs

vllm-gfx906

vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60

User Interfaces

Open WebUI

User-friendly AI Interface (Supports Ollama, OpenAI API, ...)

In 11 listsDetails

Lobe Chat

an open-source, modern design AI chat framework

In 8 listsDetails

Text generation web UI

LLM UI with advanced features, easy setup, and multiple backend support

In 7 listsDetails

SillyTavern

LLM Frontend for Power Users

In 4 listsDetails

Page Assist

Use your locally running AI models to assist you in your web browsing

Large Language Models >Explorers, Benchmarks, Leaderboards

Arena

benchmark & compare the best AI models

In 3 lists

AI Models & API Providers Analysis

understand the AI landscape to choose the best model and provider for your use case

In 3 lists

SWE-rebench

a continuously evolving and decontaminated benchmark for software engineering LLMs

BullshitBench

measure whether AI models challenge nonsensical prompts instead of confidently answering them

LLM Explorer

explore list of the open-source LLM models

Dubesor LLM Benchmark table

small-scale manual performance comparison benchmark

oobabooga benchmark

a list sorted by size (on disk) for each score

CyberGym

evaluating AI agents' real-world cybersecurity capabilities at scale

vakra

a benchmark for evaluating multi-hop, multi-source tool-calling in AI agents

swe-serve

a benchmark for evaluating multi-hop, multi-source tool-calling in AI agents

Large Language Models >Model providers

Qwen

powered by Alibaba Cloud

Mistral AI

a pioneering French artificial intelligence startup

Tencent

a profile of a Chinese multinational technology conglomerate and holding company

Unsloth AI

focusing on making AI more accessible to everyone (GGUFs etc.)

bartowski

providing GGUF versions of popular LLMs

Beijing Academy of Artificial Intelligence

a private non-profit organization engaged in AI research and development

Open Thoughts

a team of researchers and engineers curating the best open reasoning datasets

Large Language Models >Specific models

DeepSeek-V4

a collection of the DeepSeek V4 LLMs

Qwen3.8

a collection of the latest generation Qwen LLMs

NVIDIA Nemotron v3

a family of open models from NVIDIA with open weights, training data and recipes, delivering leading efficiency and accuracy for building specialized AI agents

Gemma 4

a family of open models built by Google DeepMind, that are multimodal, handling text and image input (with audio supported on small models) and generating text output

Mistral Medium 3.5

The first flaship models from Mistral AI handling instruction-following, reasoning, and coding in a single set of opened-weights

gpt-oss

a collection of open-weight models from OpenAI, designed for powerful reasoning, agentic tasks, and versatile developer use cases

gpt-oss-puzzle-88B

a deployment-optimized large language model developed by NVIDIA, derived from OpenAI's gpt-oss-120b

Hunyuan

a collection of Tencent's open-source efficient LLMs designed for versatile deployment across diverse computational environments

Phi-4

a family of small language, multi-modal and reasoning models from Microsoft

OpenReasoning-Nemotron

a collection of models from NVIDIA, trained on 5M reasoning traces for math, code and science

Kimi K2.5

a collection of open-source, native multimodal agentic models from Moonshot AI that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration

GLM-5.3

a Z.ai's flagship model for long-horizon tasks

Ling-3.0-flash

a native hybrid reasoning model from inclusionAI, operating with 124B total and 5.1B active parameters

Granite 4.1

efficient language models from IBM for multilingual generation, coding, RAG, and AI assistant workflows

Ornith-1.5

a collection of open-source models for agentic tasks and coding

EXAONE-4.5

LG's First Open-Weight Vision-Language Model for Industrial Intelligence

Step-3.7-Flash

Mixture-of-Experts vision-language model for developers who need to scale agentic workflows that combine perception, search, and reasoning

Occamy-1.0

a compact agentic model purpose-built for real-world co-work

MiniCPM5

a collection of SOTA on-device LLMs, small yet powerful

Qwen3-Coder-Next

a collection of Qwen's open-weight language models designed specifically for coding agents and local development

Devstral 2

a couple of agentic LLMs for software engineering tasks, excelling at using tools to explore codebases, edit multiple files, and power SWE Agents

Mellum 2

an assistant model trained by JetBrain

MiniMax-M3

a native multimodal model with 1M context

MiniMax-M2

a collection of SOTA models for real-world dev & agents

Laguna-S-2.1

a 118B total parameter Mixture-of-Experts model with 8B activated parameters per token, designed for agentic coding and long-horizon work

SWE-FastContext

a family of code-search models from Microsoft powering the Explore subagent for coding agents

OmniCoder-9B

a 9-billion parameter coding agent model built by Tesslate, fine-tuned on top of Qwen3.5-9B's hybrid architecture

NousCoder-14B

a competitive programming model post-trained on Qwen3-14B via reinforcement learning

MusaCoder-27B

a code model developed by Moore Threads for PyTorch-to-CUDA/MUSA native kernel generation

Qwen3-Omni

a collection of the natively end-to-end multilingual omni-modal foundation models from Qwen

GLM-4.6V

a collection of open source multimodal models with native tool use from Zhipu AI

Qwen-Image-2.1

a unified text-to-image generation and image editing model in the Qwen family

In 2 lists

Qwen3-VL

a collection of the most powerful vision-language models in the Qwen series to date

GLM-Image

an image generation model

In 2 lists

Ming-Image-0.1-Design

a 6B text-to-image model for UI, infographics, posters, and other text-rich visual designs

Granite Vision

multimodal models from IBM built for visual document analysis and image understanding

HunyuanImage

a collection of image generation models from Tencent

HunyuanVideo

a collection of video generation models from Tencent

Vidi

a collection of models for multimodal video understanding and creation

FastVLM

a collection of VLMs with efficient vision encoding from Apple

MiniCPM-o & MiniCPM-V

multimodal models with leading performance

LFM2.5-VL

a collection of vision-language models, designed for on-device deployment

ClipTagger-12b

a vision-language model (VLM) designed for video understanding at massive scale

whisper-large-v3

a state-of-the-art model for automatic speech recognition (ASR) and speech translation from OpenAI

Nemotron Speech

a collection of open, state-of-the-art, production‑ready enterprise speech models from NVIDIA for ASR, TTS, Speaker Diarization and S2SOpenAI

NVIDIA NemotronLabs VoiceChat 11B

a 11B end-to-end, real-time speech full duplex (FD) model from NVIDIA for conversational AI that jointly performs streaming speech understanding and speech generation

Qwen3-ASR

a collection of models that support language identification and ASR for 52 languages and dialects

Qwen3-TTS

a collection of TTS models that cover 10 major languages as well as multiple dialectal voice profiles to meet global application needs

Granite Speech

a collection of compact and efficient speech-language models from IBM, specifically designed for multilingual automatic speech recognition (ASR) and bidirectional automatic speech translation (AST)

Voxtral-Small-24B-2507

an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance

Voxtral-Mini-4B-Realtime-2602

a multilingual, realtime speech-transcription model and among the first open-source solutions to achieve accuracy comparable to offline systems with a delay of <500ms

Voxtral-4B-TTS-2603

frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents

chatterbox

first production-grade open-source TTS model

VibeVoice

a collection of frontier text-to-speech models from Microsoft

Kitten TTS

a collection of open-source realistic text-to-speech models designed for lightweight deployment and high-quality voice synthesis

Streaming Sortformer Diarizer 4spk v2.1

a streaming version of a novel end-to-end neural model for speaker diarization from NVIDIA

Nemotron RAG

a set of tools to build retrieval-augmented generation (RAG) systems, improve search and ranking accuracy, and extract structured data from complex docs

Qwen3-Embedding

a collection of the latest proprietary Qwen models, specifically designed for text embedding and ranking tasks

Qwen3-VL-Embedding

an addition to the Qwen embedding models, specifically designed for multimodal information retrieval and cross-modal understanding

Qwen3-Reranker

a collection of the latest proprietary Qwen models, engineered to refine embedding results

Shieldstral 1.0 3B

a compact 3B-parameter, policy-adaptive multimodal safety classifier

Granite Guardian

a collection of safety models from IBM for detecting risks, toxicity, and hallucinations in LLM workflows

Qwen3Guard

a collection of safety moderation models built upon Qwen3

NemoGuard

a collection of models from NVIDIA for content safety, topic-following and security guardrails

Nemotron-3.5-Content-Safety

a small language model (SLM) that uses Google's Gemma-3-4B-it as the base and is fine-tuned by NVIDIA on multimodal, multilingual, and reasoning-oriented content-safety datasets

Privasis

a collection of lightweight text-sanitization models from NVIDIA designed to remove or abstract sensitive information from text according to a user-provided sanitization instruction

SingGuard

a collection of policy-adaptive multimodal LLM Guardrails with dynamic reasoning

HARC

a family of safety-aligned instruction models from Microsoft trained with HARC

gpt-oss-safeguard

a collection of safety reasoning models built-upon gpt-oss from OpenAI

privacy-filter

a bidirectional token-classification model from OpenAI for personally identifiable information (PII) detection and masking in text

AprielGuard

a safeguard model designed to detect and mitigate both safety risks and security threats in LLM interactions

Intern-S2

a collection of multimodal foundation models for scientific intelligence and long-horizon agents

Holo4

a collection of Visual Language Models for computer/mobile use, tool calls, and code

Marco-MoE

a suit of multilingual MoE models with highly-sparse architectures

OUI-1

a diffusion model built for generative UI

Nemotron-Orchestrator-8B

a state-of-the-art 8B orchestration model designed to solve complex, multi-turn agentic tasks by coordinating a diverse set of expert models and tools

Arch-Router-1.5B

the fastest LLM router model that aligns to subjective usage preferences

Waypoint

a collection of real-time interactive video world models

Hunyuan3D

a collection of everything related (models, datasets etc.) to 3D assets generation from Tencent

Hunyuan-GameCraft-1.0

a novel framework for high-dynamic interactive video generation in game environments

void-model

a model from Netflix that removes objects from videos along with all interactions they induce on the scene — not just secondary effects like shadows and reflections, but physical interactions like objects falling when a person is removed

Tools >Models

llmfit

hundreds of models & providers, one command to find what runs on your hardware

In 4 listsDetails

outlines

structured outputs for LLMs

In 6 listsDetails

llama-swap

reliable model swapping for any local OpenAI compatible server - llama.cpp, vllm, etc.

In 5 listsDetails

llguidance

super-fast structured outputs

Tools >Agent Frameworks

AutoGPT

a powerful platform that allows you to create, deploy, and manage continuous AI agents that automate complex workflows

In 10 listsDetails

langflow

a powerful tool for building and deploying AI-powered agents and workflows

In 9 listsDetails

langchain

build context-aware reasoning applications

In 20 listsDetails

pi

AI agent toolkit: coding agent CLI, unified LLM API, TUI & web UI libraries, Slack bot, vLLM pods

In 4 listsDetails

anything-llm

the all-in-one Desktop & Docker AI application with built-in RAG, AI agents, No-code agent builder, MCP compatibility, and more

In 8 listsDetails

autogen

a programming framework for agentic AI

In 14 listsDetails

crewAI

a framework for orchestrating role-playing, autonomous AI agents

In 10 listsDetails

Flowise

build AI agents, visually

In 12 listsDetails

llama_index

the leading framework for building LLM-powered agents over your data

In 14 listsDetails

agno

a full-stack framework for building Multi-Agent Systems with memory, knowledge and reasoning

In 7 listsDetails

sim

open-source platform to build and deploy AI agent workflows

In 4 listsDetails

openai-agents-python

a lightweight, powerful framework for multi-agent workflows

In 9 listsDetails

NemoClaw

run OpenClaw more securely inside NVIDIA OpenShell with managed inference

In 2 lists

pydantic-ai

a Python agent framework designed to help you quickly, confidently, and painlessly build production grade applications and workflows with Generative AI

In 12 listsDetails

SuperAGI

an open-source framework to build, manage and run useful Autonomous AI Agents

In 5 listsDetails

camel

the first and the best multi-agent framework

In 7 listsDetails

agent-framework

a framework for building, orchestrating and deploying AI agents and multi-agent workflows with support for Python and .NET

In 4 listsDetails

txtai

all-in-one open-source AI framework for semantic search, LLM orchestration and language model workflows

In 9 listsDetails

archgw

a high-performance proxy server that handles the low-level work in building agents: like applying guardrails, routing prompts to the right agent, and unifying access to LLMs, etc.

In 2 lists

genkit

open-source framework for building AI-powered apps in JavaScript, Go, and Python, built and used in production by Google

In 4 listsDetails

ClaraVerse

privacy-first, fully local AI workspace with Ollama LLM chat, tool calling, agent builder, Stable Diffusion, and embedded n8n-style automation

NeMo-Agent-Toolkit

an open-source library for efficiently connecting and optimizing teams of AI agents

In 2 lists

ragbits

building blocks for rapid development of GenAI applications

Tools >Model Context Protocol

mindsdb

federated query engine for AI - the only MCP Server you'll ever need

In 11 listsDetails

playwright-mcp

Playwright MCP server

In 5 listsDetails

github-mcp-server

GitHub's official MCP Server

In 4 listsDetails

chrome-devtools-mcp

Chrome DevTools for coding agents

In 3 lists

n8n-mcp

a MCP for Claude Desktop / Claude Code / Windsurf / Cursor to build n8n workflows for you

awslabs/mcp

AWS MCP Servers — helping you get the most out of AWS, wherever you use MCP

mcp-atlassian

MCP server for Atlassian tools (Confluence, Jira)

In 3 lists

dbhub

zero-dependency, token-efficient database MCP server for Postgres, MySQL, SQL Server, MariaDB, SQLite

In 2 lists

Tools >Retrieval-Augmented Generation

pathway

Python ETL framework for stream processing, real-time analytics, LLM pipelines and RAG

In 8 listsDetails

LightRAG

simple and fast RAG

In 5 listsDetails

graphrag

a modular graph-based RAG system

In 5 listsDetails

onyx

the AI platform connected to your company's docs, apps, and people

In 5 listsDetails

graphiti

build real-time knowledge graphs for AI Agents

In 3 lists

haystack

AI orchestration framework to build customizable, production-ready LLM applications, best suited for building RAG, question answering, semantic search or conversational agent chatbots

In 13 listsDetails

vanna

an open-source Python RAG framework for SQL generation and related functionality

In 4 listsDetails

claude-context

make entire codebase the context for any coding agent

In 2 lists

pipeshub-ai

a fully extensible and explainable workplace AI platform for enterprise search and workflow automation

In 2 lists

Tools >Coding Agents

opencode

a AI coding agent built for the terminal

zed

a next-generation code editor designed for high-performance collaboration with humans and AI

In 6 listsDetails

OpenHands

a platform for software development agents powered by AI

In 7 listsDetails

cline

autonomous coding agent right in your IDE, capable of creating/editing files, executing commands, using the browser, and more with your permission every step of the way

In 7 listsDetails

goose

an open-source, extensible AI agent that goes beyond code suggestions

In 3 lists

aider

AI pair programming in your terminal

In 5 listsDetails

continue

create, share, and use custom AI code assistants with our open-source IDE extensions and hub of models, rules, prompts, docs, and other building blocks

In 10 listsDetails

tabby

an open-source GitHub Copilot alternative, set up your own LLM-powered code completion server

In 7 listsDetails

void

an open-source Cursor alternative, use AI agents on your codebase, checkpoint and visualize changes, and bring any model or host locally

In 2 lists

crush

the glamourous AI coding agent for your favourite terminal

In 4 listsDetails

kilocode

open source AI coding assistant for planning, building, and fixing code

In 4 listsDetails

Roo-Code

a whole dev team of AI agents in your code editor

In 3 lists

humanlayer

the best way to get AI coding agents to solve hard problems in complex codebases

In 2 lists

openchamber

Agentic Development Environment based on OpenCode AI agent

In 2 lists

99

neovim AI agent done right

ProxyAI

the leading open-source AI copilot for JetBrains

In 2 lists

Tools >Computer Use

open-interpreter

a natural language interface for computers

In 9 listsDetails

cua

the Docker Container for Computer-Use AI Agents

In 3 lists

OmniParser

a simple screen parsing tool towards pure vision based GUI agent

In 2 lists

openwork

an open-source alternative to Claude Cowork, powered by OpenCode

In 4 listsDetails

Agent-S

an open agentic framework that uses computers like a human

In 6 listsDetails

self-operating-computer

a framework to enable multimodal models to operate a computer

In 6 listsDetails

OpenRoom

a browser-based desktop where AI Agent operates every app through natural language, from MiniMaxAI

Tools >Browser Automation

firecrawl

turn entire websites into LLM-ready markdown or structured data

In 2 lists

browser-use

make websites accessible for AI agents

In 10 listsDetails

puppeteer

a JavaScript API for Chrome and Firefox

In 5 listsDetails

playwright

a framework for Web Testing and Automation

In 10 listsDetails

stagehand

the AI Browser Automation Framework

In 2 lists

nanobrowser

open-source Chrome extension for AI-powered web automation

In 3 lists

Tools >Memory Management

mem0

universal memory layer for AI Agents

In 13 listsDetails

mempalace

the highest-scoring AI memory system ever benchmarked

supermemory

memory engine and app that is extremely fast, scalable

In 3 lists

cognee

memory for AI Agents in 5 lines of code

In 6 listsDetails

letta

the stateful agents framework with memory, reasoning, and context management

In 4 listsDetails

LMCache

supercharge your LLM with the fastest KV Cache Layer

In 4 listsDetails

memU

an open-source memory framework for AI companions

In 2 lists

reasoning-bank

a memory mechanism for agents that learns from both successful and failed trajectories, with reasoning stored as memory content

Tools >Testing, Evaluation and Observability

langfuse

an open-source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Integrates with OpenTelemetry, Langchain, OpenAI SDK, LiteLLM, and more

In 10 listsDetails

opik

debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards

In 16 listsDetails

openllmetry

an open-source observability for your LLM application, based on OpenTelemetry

In 7 listsDetails

giskard

an open-source evaluation & testing for AI & LLM systems

In 2 lists

agenta

an open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place

In 6 listsDetails

Evaluator

open-source library for scalable, reproducible evaluation of AI models and benchmarks

Tools >Research

open-notebook

an open-source implementation of Notebook LM with more flexibility and features

In 2 lists

Perplexica

an open-source alternative to Perplexity AI, the AI-powered search engine

In 4 listsDetails

gpt-researcher

an LLM based autonomous agent that conducts deep local and web research on any topic and generates a long report with citations

In 9 listsDetails

SurfSense

an open-source alternative to NotebookLM / Perplexity / Glean

In 5 listsDetails

RD-Agent

automate the most critical and valuable aspects of the industrial R&D process

In 5 listsDetails

local-deep-researcher

fully local web research and report writing assistant

In 2 lists

local-deep-research

an AI-powered research assistant for deep, iterative research

In 7 listsDetails

maestro

an AI-powered research application designed to streamline complex research tasks

Tools >Training and Fine-tuning

heretic

fully automatic censorship removal for language models

In 3 lists

sentence-transformers

a Python library for using and training embedding and reranker models for applications like retrieval augmented generation, semantic search, and more

In 3 lists

trl

train transformer language models with reinforcement learning

In 8 listsDetails

OpenRLHF

an easy-to-use, high-performance open-source RLHF framework built on Ray, vLLM, ZeRO-3 and HuggingFace Transformers, designed to make RLHF training simple and accessible

In 4 listsDetails

slime

an LLM post-training framework for RL Scaling

In 5 listsDetails

Kiln

the easiest tool for fine-tuning LLM models, synthetic data generation, and collaborating on datasets

In 5 listsDetails

miles

an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime

In 2 lists

OpenEnv

an interface library for RL post training with environments

In 2 lists

RL

scalable toolkit for efficient model reinforcement

In 3 lists

augmentoolkit

train an open-source LLM on new facts

Gym

evaluate and improve models and agents using environments

In 2 lists

SpecForge

train speculative decoding models effortlessly and port them smoothly to SGLang serving

Tools >Security and Sandboxing

garak

the LLM vulnerability scanner from NVIDIA

In 4 listsDetails

OpenShell

the safe, private runtime for autonomous AI agents from NVIDIA

In 4 listsDetails

Guardrails

an open-source toolkit from NVIDIA for easily adding programmable guardrails to LLM-based conversational systems

In 6 listsDetails

CubeSandbox

instant, concurrent, secure & lightweight sandbox for AI agents

In 4 listsDetails

Tools >Miscellaneous

context7

up-to-date code documentation for LLMs and AI code editors

In 6 listsDetails

deepwiki-open

open source DeepWiki: AI-powered wiki generator for GitHub/Gitlab/Bitbucket repositories

In 3 lists

cai

Cybersecurity AI (CAI), the framework for AI Security

In 4 listsDetails

speakr

a personal, self-hosted web application designed for transcribing audio recordings

presenton

an open-source AI presentation generator and API

OmniGen2

exploration to advanced multimodal generation

In 2 lists

4o-ghibli-at-home

a powerful, self-hosted AI photo stylizer built for performance and privacy

Observer

local open-source micro-agents that observe, log and react, all while keeping your data private and secure

mobile-use

a powerful, open-source AI agent that controls your Android or IOS device using natural language

gabber

build AI applications that can see, hear, and speak using your screens, microphones, and cameras as inputs

promptcat

a zero-dependency prompt manager/catalog/library in a single HTML file

Hardware

Alex Ziskind

tests of pcs, laptops, gpus etc. capable of running LLMs

Digital Spaceport

reviews of various builds designed for LLM inference

Donato Capitella

practical and insightful tutorials on running LLMs locally

JetsonHacks

information about developing on NVIDIA Jetson Development Kits

Miyconst

tests of various types of hardware capable of running LLMs

Kolosal - LLM Memory calculator

estimate the RAM requirements of any GGUF model instantly

LLM Inference VRAM & GPU Requirement Calculator

calculate how many GPUs you need to deploy LLMs

ZLUDA

CUDA on non-NVIDIA GPUs

In 3 lists

Strix Halo AI Toolboxes

toolboxes for GenAI on AMD Ryzen AI MAX+: containerized environments for LLMs, Image Generation, and Fine-tuning

Strix Halo Wiki

a website to gather important information and practical guides for systems powered by AMD Ryzen AI MAX and MAX+ processors

ai-notes

random AI notes for working with local models or playing around with random machine learning bits

Tutorials >Models

Let's reproduce GPT-2 (124M)

nanochat

a full-stack implementation of an LLM like ChatGPT in a single, clean, minimal, hackable, dependency-lite codebase, designed to run on a single 8XH100 node via scripts like speedrun.sh, that run the entire pipeline start to end

In 4 listsDetails

Knowledge Distillation: How LLMs train each other

gguf-docs

Docs for GGUF quantization (unofficial)

Embarrassingly Simple Self-Distillation Improves Code Generation

Tutorials >Prompt Engineering

Prompt Engineering Guide

guides, papers, lecture, notebooks and resources for prompt engineering

In 11 listsDetails

Prompt Engineering by NirDiamant

a comprehensive collection of tutorials and implementations for Prompt Engineering techniques, ranging from fundamental concepts to advanced strategies

In 5 listsDetails

Prompting guide 101

a quick-start handbook for effective prompts by Google

Prompt Engineering by Google

prompt engineering by Google

Prompt Engineering by Anthropic

prompt engineering by Anthropic

In 3 lists

Prompt Engineering Interactive Tutorial

Prompt Engineering Interactive Tutorial by Anthropic

In 4 listsDetails

system-prompts-and-models-of-ai-tools

a collection of system prompts extracted from AI tools

In 3 lists

system_prompts_leaks

a collection of extracted System Prompts from popular chatbots like ChatGPT, Claude & Gemini

In 4 listsDetails

Prompt from Codex

Prompt used to steer behavior of OpenAI's Codex

In 9 listsDetails

Tutorials >Context Engineering

Context-Engineering

a frontier, first-principles handbook inspired by Karpathy and 3Blue1Brown for moving beyond prompt engineering to the wider discipline of context design, orchestration, and optimization

In 3 lists

Awesome-Context-Engineering

a comprehensive survey on Context Engineering: from prompt engineering to production-grade AI systems

In 2 listsDetails

Tutorials >Inference

vLLM Production Stack

vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization

In 2 lists

Tutorials >Agents

superpowers

an agentic skills framework & software development methodology that works

In 7 listsDetails

GenAI Agents

tutorials and implementations for various Generative AI Agent techniques

In 6 listsDetails

500+ AI Agent Projects

a curated collection of AI agent use cases across various industries

In 3 lists

12-Factor Agents

principles for building reliable LLM applications

In 2 lists

Agents towards production

end-to-end, code-first tutorials covering every layer of production-grade GenAI agents, guiding you from spark to scale with proven patterns and reusable blueprints for real-world launches

In 2 lists

agents.md

a simple, open format for guiding coding agents

Agent Skills

a simple, open format for giving agents new capabilities and expertise

In 3 lists

skills

Hugging Face Skills are definitions for AI/ML tasks like dataset creation, model training and evaluation

In 2 lists

LLM Agents & Ecosystem Handbook

one-stop handbook for building, deploying, and understanding LLM agents with 60+ skeletons, tutorials, ecosystem guides, and evaluation tools

601 real-world gen AI use cases

601 real-world gen AI use cases from the world's leading organizations by Google

A practical guide to building agents

a practical guide to building agents by OpenAI

In 2 lists

Tutorials >Retrieval-Augmented Generation

Pathway AI Pipelines

ready-to-run cloud templates for RAG, AI pipelines, and enterprise search with live data

In 5 listsDetails

RAG Techniques

various advanced techniques for Retrieval-Augmented Generation (RAG) systems

In 5 listsDetails

Controllable RAG Agent

an advanced Retrieval-Augmented Generation (RAG) solution for complex question answering that uses sophisticated graph based algorithm to handle the tasks

In 2 lists

LangChain RAG Cookbook

a collection of modular RAG techniques, implemented in LangChain + Python

Tutorials >Miscellaneous

local-llm

everything jamesob knows about running LLMs locally

Self-hosted AI coding that just works

Communities

LocalLLaMA

Go-to subreddit for local/open-source LLM topics.

In 3 lists

LLMDevs

LocalLLM

LocalAIServers

GenAI monitor

monitoring updates & fresh releases related to LLMs, diffusion models and Generative AI

See category
94

Table of Contents

hesreallyhim/awesome-claude-code

A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…

Fresh★ 55k202 entriesPushed today
94

Awesome Agent Skills

VoltAgent/awesome-agent-skills

A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.

Fresh★ 35k839 entriesPushed today
93

Awesome Machine Learning

josephmisiti/awesome-machine-learning

A curated list of awesome Machine Learning frameworks, libraries and software.

Fresh★ 74k1188 entriesPushed 7 days ago
92

Awesome Production Machine Learning

EthicalML/awesome-production-machine-learning

A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning

Fresh★ 21k519 entriesPushed 3 days ago
92

AWESOME DATA SCIENCE

academic/awesome-datascience

:memo: An awesome Data Science repository to learn and apply for real world problems.

Fresh★ 30k881 entriesPushed today
91

Static Analysis

analysis-tools-dev/static-analysis

⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…

Fresh★ 15k528 entriesPushed 8 days ago