Inference & Serving
Platforms and engines for running, serving, and routing requests to LLMs, from local inference to cloud-scale gateways.
-
OllamaThe default way to run open-weight LLMs locally: a single command pulls and runs a model like Llama, DeepSeek, or Qwen on your own machine, with an OpenAI-compatible API.
-
vLLMOpen-source, high-throughput inference engine for serving LLMs at scale, using PagedAttention to significantly cut memory waste and boost GPU utilization.
-
Together AICloud platform for running and fine-tuning open-source LLMs at scale, offering fast inference, dedicated GPU clusters, and a library of pre-hosted open models.
-
GroqInference platform running on custom LPU chips rather than GPUs, delivering notably fast token generation speeds for open models like Llama and DeepSeek.
-
ReplicatePlatform for running open-source machine learning models via a simple API, covering LLMs, image generation, and other model types with pay-per-use pricing and no infrastructure to manage.
-
OpenRouterUnified API and marketplace for LLMs, routing requests to the best available provider for a given model and handling failover automatically across dozens of providers.
-
LiteLLMOpen-source Python library and proxy server that calls 100+ LLM providers through one OpenAI-compatible interface, with built-in cost tracking and load balancing.
-
PortkeyAI gateway for production LLM apps, combining provider routing and failover with caching, guardrails, and observability in one managed layer.
-
Fireworks AI
Fast inference platform for open-source models with serverless and dedicated deployments, function calling and fine-tuning.
-
Modal
Serverless cloud for running Python code on GPUs, widely used to deploy inference endpoints, batch jobs and sandboxes for AI workloads.
-
llama.cpp
C/C++ LLM inference engine and GGUF format that runs quantized models efficiently on CPUs, Apple Silicon and GPUs, with an OpenAI-compatible server.
-
LM Studio
Desktop app for discovering, downloading and running local LLMs, with a built-in chat UI and an OpenAI-compatible local server.
-
SGLang
High-performance open-source serving framework for LLMs and multimodal models, with RadixAttention prefix caching and structured output support.
-
Cerebras Inference
Inference cloud running open models on Cerebras wafer-scale chips for very high tokens-per-second throughput, with an OpenAI-compatible API.
-
Baseten
Production inference platform for deploying open-source and custom models with autoscaling GPUs, model APIs and the open-source Truss packaging format.