Awesome Cloud Native
Section: AI & Machine Learning Platforms · Standardized Serverless ML Inference Platform on Kubernetes.
Entry
Appears in 6 awesome lists
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Section: AI & Machine Learning Platforms · Standardized Serverless ML Inference Platform on Kubernetes.
Section: Tools · Standardized serverless inference platform for deploying and serving machine learning models on Kubernetes.
Section: ML Platforms · Standardized Serverless ML Inference Platform on Kubernetes
Section: 8. MLOps / LLMOps & Production · Kubernetes-based model serving.
Section: Deployment and Serving · KServe provides a Kubernetes Custom Resource Definition for serving predictive and generative ML.
Section: Other · Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
How to use the Hexagon Delegate to speed up model inference on mobile and edge devices. Also see blog post Accelerating TensorFlow Lite on Qualcomm Hexagon DSPs.
(label: good first issue) PyTorch is an open source machine learning library based on the Torch library, used for applications such as computer vision and natural language processing.
Unified proxy and SDK that routes to 100+ LLM providers behind a single OpenAI-compatible interface, with a Router handling retry/fallback across deployments, per-project cost and rate-limit tracking, and OTEL callback integrations. The right infrastructure layer when your harness needs provider…
Comet's open-source AI observability and evaluation platform: deep tracing of LLM calls, conversation logging, and agent activity, plus built-in eval metrics, prompt versioning, guardrails, and the Opik Agent Optimizer. Worth including because it unifies observability, verification, and…
(from Hpcaitech) - A Unified Deep Learning System for Large-Scale Parallel Training (1D, 2D, 2.5D, 3D and sequence parallelism, and ZeRO protocol).
Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search,…