LLM Inference
34 providers — compare pricing, models, speed, and capabilities. Click any provider to see full model listings and token costs.
GPT-4o, o3, and the world's most widely-used AI API
Claude — safety-focused frontier AI with exceptional coding ability
Gemini 2.5 — the largest context window at the lowest frontier price
European frontier AI — Mistral Large, Codestral, and open models
Enterprise NLP — Command R+ with retrieval-augmented generation
Chinese frontier lab — DeepSeek V3 and R1 at remarkably low prices
Grok 3 — Elon Musk's frontier AI with real-time web access
Qwen — frontier open-weight models with competitive pricing
Kimi — long-context frontier models from China's leading AI lab
Hunyuan LLMs from China's largest tech company
GLM frontier models with 1M context
Long-context frontier models with 1M token windows
Open-weight image generation with Stable Diffusion
FLUX — the new standard for image generation quality
Cinematic AI video generation from text and images
High-throughput frontier reasoning with Celeris-1
Open-source model hosting with competitive inference pricing
LPU-powered inference — the fastest tokens per second available
Production-grade open-source inference with fast cold starts
Sonar — search-augmented LLMs with real-time web grounding
Wafer-scale AI chips — 4,500 tokens/sec, the fastest inference on earth
Serverless LLM inference with a developer-first API
Open-source inference marketplace — Llama, DeepSeek R1, and more
Budget-friendly open-source inference with broad model selection
European cloud AI — affordable open-source inference from ex-Yandex team
The cheapest inference API for open-weight models — Llama, Mistral, and more
One API for every major LLM — route to the cheapest or fastest provider automatically
Run open-source AI models with a simple API — no infrastructure required
Fast serverless inference for image, video, and audio AI models
State-of-the-art embedding and reranking models
Hyper-realistic AI voice and speech synthesis
Llama 4 & Muse Spark — the world's most widely deployed open-weight models
AWS-native LLM access — Nova, Claude, Llama, and more via one API
French cloud provider with sovereign EU LLM inference