# Cerebrium > Cerebrium is serverless GPU infrastructure for real-time AI — voice agents, video models, LLMs, and custom ML apps. Sub-second cold starts, pay-per-second billing, no Kubernetes. ## What we do - **Voice AI**: Real-time end-to-end pipelines for voice agents (STT + LLM + TTS). - **Video & generative media**: Low-latency inference for video generation, image diffusion, and avatars. - **LLMs**: Serverless deployment for OpenAI-compatible endpoints, open-source models (DeepSeek, Llama, Orpheus), and custom fine-tunes. - **General ML**: Any Python workload, any GPU (L4, L40s, A100, H100, H200), pay only for compute-time used. ## Key Facts - Product: serverless GPU infrastructure for real-time AI inference — voice, video, LLMs, and custom ML. - Supported GPUs: L4, L40s, A100, H100, H200 — billed per-second of compute used, no idle charges or minimum reservations. - Cold starts: sub-second on the largest models (H100/H200 inference). - Voice pipelines: sub-500ms end-to-end latency (STT + LLM + TTS). - Deployment: multi-region (US + EU) for data residency; bring-your-own-code (any Python script or container, no lock-in). - Compliance: SOC 2, HIPAA, GDPR, ISO. - Reference customers: Resemble AI, Camb AI, Telli, Amira Learning, Invofox, Creatium. - Open source: github.com/CerebriumAI — 522-star examples repo (voice agents, LLMs, video, RAG). ## Key pages - [Serverless GPU Infrastructure for Real-Time AI](https://cerebrium.ai/): Deploy voice agents, video models, and LLMs on serverless GPUs with sub-second cold starts. Pay-per-second pricing. No Kubernetes. - [Pay-Per-Second Pricing for Serverless AI](https://cerebrium.ai/pricing): Pay for compute by the second, not the hour. Transparent serverless GPU pricing for voice, LLMs, and video. No commitment, no idle costs. - [Our Mission — Real-Time AI Infrastructure](https://cerebrium.ai/about): Cerebrium is the team building global serverless GPU infrastructure for real-time AI model applications. - [Book a Demo — Technical Architecture Review](https://cerebrium.ai/book-demo): See the technical architecture behind AI teams deploying real-time voice agents, LLMs, and video models on Cerebrium. 30-minute demo with our team. - [Contact — Sales, Support, Partnerships](https://cerebrium.ai/contact): Get in touch with Cerebrium for sales, partnerships, support, or enterprise inquiries. Real-time replies during business hours. - [Brand Assets — Logos & Guidelines](https://cerebrium.ai/brand-assets): Download Cerebrium logos, color palette, typography, and brand guidelines for press, partnerships, and media coverage. ## Use cases - [Large Language Models](https://cerebrium.ai/use-cases/large-language-models): Run and deploy LLMs at scale - [Voice](https://cerebrium.ai/use-cases/voice): Infrastructure built for low-latency voice at scale - [Image & Video](https://cerebrium.ai/use-cases/image-and-video): Run image and video pipelines at scale ## Documentation - [Documentation home](https://docs.cerebrium.ai/getting-started/introduction): Full developer docs (hosted on docs.cerebrium.ai). - [Cerebrium examples](https://github.com/CerebriumAI/examples): 522-star reference repo covering voice agents, LLMs, video, RAG. - [Documentation index](https://cerebrium.ai/docs/llms.txt): llms.txt index of every documentation page (151 entries), for agents that want to pick pages to fetch. - [Documentation full text](https://cerebrium.ai/docs/llms-full.txt): the entire documentation as one plain-text corpus, for agents that want it in a single fetch. ## For AI agents - [Agent Skill](https://cerebrium.ai/skill.md): Cerebrium packaged as an Agent Skill — deploying serverless GPU workloads, cold starts, scaling, and the CLI. Generated from the docs, so it tracks them rather than going stale. - Install it: `npx skills add https://cerebrium.ai -y` — works with Claude Code, Cursor, and other tools that read the agent-skills discovery index. - [Discovery index](https://cerebrium.ai/.well-known/agent-skills/index.json): agentskills.io discovery document, served with an [A2A agent card](https://cerebrium.ai/.well-known/agent-card.json) alongside it. - [MCP server](https://cerebrium.ai/docs/mcp): Model Context Protocol endpoint for searching and querying the Cerebrium documentation. ## Blog (selected high-value posts) - [How much does a H100 cost? Cost comparision](https://cerebrium.ai/blog/how-much-does-a-h100-cost-cost-comparision): GPU cost comparison. - [How much does a H200 cost? 2025 Guide](https://cerebrium.ai/blog/how-much-does-a-h200-cost-2025-guide): H200 pricing breakdown. - [Top 5 Serverless GPU providers](https://cerebrium.ai/blog/top-5-serverless-gpu-providers): Competitive landscape. - [Creating a realtime RAG voice agent](https://cerebrium.ai/blog/creating-a-realtime-rag-voice-agent): Tutorial. - [Deploying DeepSeek-R1: A Guide to a Serverless, High-Performaning OpenAI-Compatible Endpoint](https://cerebrium.ai/blog/deploying-deepseek-r1-a-guide-to-a-serverless-high-performaning-openai-compatible-endpoint): OpenAI-compatible endpoint guide. - [Orpheus TTS: How to Deploy Orpheus at Scale for Production Inference](https://cerebrium.ai/blog/orpheus-tts-how-to-deploy-orpheus-at-scale-for-production-inference): Production TTS deployment. - [Deploying Sesame CSM: The Most Realistic Voice Model as an API](https://cerebrium.ai/blog/deploying-sesame-csm-the-most-realistic-voice-model): Voice model guide. - [Launch Week Day 3: Annoucing Multi-Region Deployments](https://cerebrium.ai/blog/launch-week-day-3-annoucing-multi-region-deployments): Product announcement. - [Rethinking Container Image Distribution to eliminate cold starts](https://cerebrium.ai/blog/rethinking-container-image-distribution-to-eliminate-cold-starts): Engineering deep dive. - [The Shortcomings of Celery + Redis for ML Workloads and How Cerebrium Solves It](https://cerebrium.ai/blog/celery-redis-vs-cerebrium): Migration comparison. - [Faster Whisper Transcription: How to Maximize Performance for Real-Time Audio-to-Text](https://cerebrium.ai/blog/faster-whisper-transcription-how-to-maximize-performance-for-real-time-audio-to-text): Real-time STT. - [Integrating PayPal’s Model Context Protocol (MCP) into a Real-time Voice Agent](https://cerebrium.ai/blog/integrating-paypal-s-model-context-protocol-mcp-into-a-real-time-voice-agent): MCP integration. - [Introducing Cerebrium run: The Fastest Way to Execute Cloud Code](https://cerebrium.ai/blog/introducing-cerebrium-run-the-fastest-way-to-execute-cloud-code): Cloud-code execution. - [Blog index](https://cerebrium.ai/blog): All posts. ## Resources - [GPU Inference for AI Agent Workloads](https://cerebrium.ai/resources/gpu-inference-ai-agent-workloads): Agent workloads pay every fixed overhead once per step. Measured per-step latency, burst behaviour and what both mean for sizing GPU inference. - [Serving Fine-Tuned LLMs on Serverless GPUs](https://cerebrium.ai/resources/serving-fine-tuned-llms-serverless-gpu): Serve your own fine-tuned checkpoint on serverless GPUs: measured weight-load times, snapshot restore, right-sizing and per-model scaling. - [Pipecat Voice Agents: Which Stage Belongs on a GPU](https://cerebrium.ai/resources/pipecat-voice-agent-gpu-deployment): Where the latency and cost sit in a Pipecat voice pipeline: measured per-stage figures for STT, LLM and TTS, plus what co-location saves. - [LLM Inference Cost at Scale: The Tokens-Per-Minute Math](https://cerebrium.ai/resources/llm-inference-cost-at-scale): Convert tokens per minute into an actual LLM inference bill: measured instance counts, per-minute rates and why utilisation decides it. - [Where Latency Goes in an AI Inference Pipeline](https://cerebrium.ai/resources/network-latency-ai-inference-pipeline): Where latency really goes in an AI inference pipeline: measured region, API-hop, broker and routing costs, ranked by size. - [NVIDIA B200 Inference: Measured Throughput and TTFT](https://cerebrium.ai/resources/nvidia-b200-inference-throughput-ttft): Measured NVIDIA B200 inference throughput and TTFT from a production pipeline, plus why cross-GPU tokens/sec comparisons mislead. - [Concurrent Voice Sessions Per GPU: The Real Numbers](https://cerebrium.ai/resources/concurrent-voice-sessions-per-gpu): How many concurrent voice sessions fit on one GPU: measured A10 concurrency, per-stage latency budgets and the cost arithmetic. - [AWS Alternatives for AI Workloads: What Actually Changes](https://cerebrium.ai/resources/aws-alternatives-ai-workloads): AWS alternatives for AI workloads, compared on the number that decides cost: provisioning time. Measured cold starts, Inferentia benchmarks and when to stay. - [What a Voice AI Agent Really Costs Per Minute](https://cerebrium.ai/resources/voice-ai-agent-cost-per-minute): A measured per-minute cost breakdown for a production voice AI agent: $0.02932 per call minute, and the concurrency maths that changes it. - [Modal Alternatives for Serverless GPU Inference](https://cerebrium.ai/resources/modal-alternatives-serverless-gpu-inference): Compare Modal alternatives for serverless GPU inference on cold starts, tail latency, GPU breadth and pricing model, using measured restore-time benchmarks. - [Serverless GPU Cold Starts: Killing Voice AI Latency](https://cerebrium.ai/resources/serverless-gpu-cold-starts-voice-ai): Serverless GPU cold starts stack container pulls, weight loads, and CUDA warmup into 30-90s of delay. Here's how to kill each phase for real-time Voice AI. - [Resources index](https://cerebrium.ai/resources): All resources. ## Contact - Website: https://cerebrium.ai/ - Book a demo: https://cerebrium.ai/book-demo - General inquiries: https://cerebrium.ai/contact ## Optional - [GitHub @CerebriumAI](https://github.com/CerebriumAI): Open-source examples and tools. - [Privacy policy](https://cerebrium.ai/privacy) - [Terms of service](https://cerebrium.ai/terms-of-service)