Skip to content

Entry

shimmy

Appears in 7 awesome lists

Python-free Rust inference server with OpenAI API compatibility. Supports GGUF and SafeTensors formats with hot model swap, auto-discovery, and single binary deployment for zero-dependency inference. Apache 2.0 licensed.

Open github.commichael-a-kuykendall/shimmy

Found in these lists

awesome-ChatGPT-repositories

Section: Langchain · ⚡ Python-free Rust inference server — OpenAI-API compatible. GGUF + SafeTensors, hot model swap, auto-discovery, single binary. FREE now, FREE forever.

FreshScore 87

Awesome LLM Resources

Section: 推理 Inference · Python-free Rust inference server — OpenAI-API compatible. GGUF + SafeTensors, hot model swap, auto-discovery, single binary.

FreshScore 87

Awesome LLMOps

Section: Large Model Serving · Python-free Rust inference server with OpenAI API compatibility and hot model swapping

ActiveScore 75

Awesome Machine Learning

Section: Rust · Python-free Rust inference server for NLP models with OpenAI API compatibility and hot model swapping.

FreshScore 93

Awesome Open Source AI

Section: 3. Inference Engines & Serving · Python-free Rust inference server with OpenAI API compatibility. Supports GGUF and SafeTensors formats with hot model swap, auto-discovery, and single binary deployment for zero-dependency inference. Apache 2.0 licensed.

FreshScore 89

Awesome Privacy

Section: Artificial Intelligence · Privacy-focused AI inference server with OpenAI API compatibility, zero cloud dependencies, and local model processing.

ActiveScore 84

Awesome Rust

Section: Artificial Intelligence · [shimmy] - Pure-Rust WebGPU inference engine with an OpenAI-compatible API and native GGUF support.

FreshScore 94

LangChain

Langchain integrates various providers like Anthropic, AWS, and OpenAI, and offers tools for components such as LLMs, chat models, and data analysis, supporting functionalities from Alpha Vantage to YouTube github | docs

In 20 listsDetails

LiteLLM

Unified proxy and SDK that routes to 100+ LLM providers behind a single OpenAI-compatible interface, with a Router handling retry/fallback across deployments, per-project cost and rate-limit tracking, and OTEL callback integrations. The right infrastructure layer when your harness needs provider…

In 16 listsDetails

LlamaIndex

(MIT) provides modules for structured outputs at different levels of abstraction, including output parsers for text completion endpoints, Pydantic programs for mapping prompts to structured outputs using function calling or output parsing, and pre-defined Pydantic programs for specific output types.

In 14 listsDetails

LocalAI

robot: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. Features: Generate Text, Audio, Video,…

In 14 listsDetails

Mem0

Mem0 is an intelligent memory layer for Large Language Models that enhances personalized AI experiences by retaining and utilizing contextual information across various applications. github | website | docs | discord | twitter | github profile | linkedin

In 13 listsDetails

Ollama

Ollama is a tool for running large language models locally, offering easy setup for macOS, Windows, Linux, and Docker, along with a library of models and quickstart guides for customization and integration github | github profile

In 12 listsDetails

vLLM

State-of-the-art serving engine with PagedAttention and continuous batching. Currently the fastest production-grade LLM server.

In 11 listsDetails

Flowise

Flowise simplifies the creation of applications leveraging large language models (LLMs) by providing a drag-and-drop interface for customizing AI workflows, offering easy installation, Docker support, development tools, and documentation for integrating various functionalities such as…

In 12 listsDetails