> ## Agent Instructions > Cerebrium's documentation MCP server is available at https://cerebrium.ai/docs/mcp for searching and querying these docs directly. Install the Cerebrium agent skill with `npx skills add https://cerebrium.ai/docs`. Append .md to any docs page URL to fetch that page as plain Markdown. API keys and authentication tokens are created in the Cerebrium dashboard at https://dashboard.cerebrium.ai. # Create API Key Source: https://cerebrium.ai/docs/api-reference/api-keys/create-api-key https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json post /v2/projects/{project_id}/api-keys Create a new API key for a project. # Delete API Key Source: https://cerebrium.ai/docs/api-reference/api-keys/delete-api-key https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json delete /v2/projects/{project_id}/api-keys/{api_key_id} Delete an API key and permanently deny its token. Inference requests using the key start failing within a few seconds. This cannot be undone: the key is not recoverable and the token stays denied. The caller must be an owner on the project. Service account keys are rejected here — delete the service account instead. # List API Keys Source: https://cerebrium.ai/docs/api-reference/api-keys/list-api-keys https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/api-keys List all API keys for a project. # Create App Source: https://cerebrium.ai/docs/api-reference/apps/create-app https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json post /v2/projects/{project_id}/apps Create a new app for a specific project using manual upload. # Create GitHub App Source: https://cerebrium.ai/docs/api-reference/apps/create-github-app https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json post /v2/projects/{project_id}/apps/github Create a new app linked to a GitHub repository for automatic deployments. # Create Run App Source: https://cerebrium.ai/docs/api-reference/apps/create-run-app https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json post /v3/projects/{project_id}/apps/{app_id}/create-run-app Create an app if it does not already exist, used by the CLI before running code. # Delete App Source: https://cerebrium.ai/docs/api-reference/apps/delete-app https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json delete /v2/projects/{project_id}/apps/{app_id} Remove a specific app from a project. # Get Active Revision Source: https://cerebrium.ai/docs/api-reference/apps/get-active-revision https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/active-revision Retrieve the active revision for a specific app. # Get an app's request and cost breakdown over time Source: https://cerebrium.ai/docs/api-reference/apps/get-an-apps-request-and-cost-breakdown-over-time https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/dashboard-metrics How much traffic an app served and what it cost to run, over a chosen period. Requests are counted per time bucket and split by response status (2xx/4xx/5xx, normal and error WebSocket closes, and anything else); cost is split across GPU, CPU and memory. Every bucket in the period is returned, zero-filled where there was no activity, so the series can be charted without gap handling. # Get App Source: https://cerebrium.ai/docs/api-reference/apps/get-app https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id} Retrieve details for a specific app in a project. # Get App Cost Source: https://cerebrium.ai/docs/api-reference/apps/get-app-cost https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/cost Retrieve cost breakdown for a specific app. # Get App Logs Source: https://cerebrium.ai/docs/api-reference/apps/get-app-logs https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/logs Retrieve logs for a specific app. # Get App Resource Metrics Source: https://cerebrium.ai/docs/api-reference/apps/get-app-resource-metrics https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/resource-metrics Retrieve CPU, memory, and GPU utilization metrics for an app over a time period. # List Apps Source: https://cerebrium.ai/docs/api-reference/apps/list-apps https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps Retrieve a list of apps for a specific project. # Modify App Source: https://cerebrium.ai/docs/api-reference/apps/modify-app https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json patch /v2/projects/{project_id}/apps/{app_id} Update the configuration or metadata of a specific app. # Cancel Build Source: https://cerebrium.ai/docs/api-reference/builds/cancel-build https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json delete /v2/projects/{project_id}/apps/{app_id}/builds/{build_id} Cancel an ongoing build for an app. # Create Base Image Hash Source: https://cerebrium.ai/docs/api-reference/builds/create-base-image-hash https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json post /v3/projects/{project_id}/apps/{app_id}/base-image Generate a SHA256 hash from dependency lists to determine if a base image rebuild is needed. # Download Build Source: https://cerebrium.ai/docs/api-reference/builds/download-build https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/builds/{build_id}/download Download the build ZIP file for a specific build. # Get Build Source: https://cerebrium.ai/docs/api-reference/builds/get-build https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/builds/{build_id} Retrieve details for a specific build. # Get Build Zip Contents Source: https://cerebrium.ai/docs/api-reference/builds/get-build-zip-contents https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/builds/{build_id}/zip-contents List the files contained in a build's uploaded ZIP archive. # Health Check Source: https://cerebrium.ai/docs/api-reference/builds/health-check https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/build-service/status Check if the build service is operational. # List Build Logs Source: https://cerebrium.ai/docs/api-reference/builds/list-build-logs https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/builds/{build_id}/logs Retrieve logs for a specific build of an app. # List Builds Source: https://cerebrium.ai/docs/api-reference/builds/list-builds https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/builds Retrieve a list of builds for a specific app. # List Image Files Source: https://cerebrium.ai/docs/api-reference/builds/list-image-files https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/builds/{build_id}/image/files Browse files inside a built container image at a specified path. # Rebuild Source: https://cerebrium.ai/docs/api-reference/builds/rebuild https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json post /v2/projects/{project_id}/apps/{app_id}/builds/{build_id}/rebuild Trigger a new build using the same source as an existing build. # Get Container Source: https://cerebrium.ai/docs/api-reference/containers/get-container https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/containers/{container_id} Retrieve details for a specific container. # Get Container Events Source: https://cerebrium.ai/docs/api-reference/containers/get-container-events https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v3/projects/{project_id}/apps/{app_id}/containers/{container_id}/events Retrieve lifecycle events for a container such as eviction, OOM, spot interruption, and checkpoint save or restore. # Get Queue Depth Source: https://cerebrium.ai/docs/api-reference/containers/get-queue-depth https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v3/projects/{project_id}/apps/{app_id}/queue-depth Get real-time queue depth counts (proxyQueued, containerQueued, processing) for an app # List Active Containers Source: https://cerebrium.ai/docs/api-reference/containers/list-active-containers https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v3/projects/{project_id}/apps/{app_id}/containers/active Retrieve all active containers, including those in terminating state. # List Container Readiness Source: https://cerebrium.ai/docs/api-reference/containers/list-container-readiness https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v3/projects/{project_id}/apps/{app_id}/containers/status Retrieve the ready/not-ready status of each container. # List Container Resource Usage Source: https://cerebrium.ai/docs/api-reference/containers/list-container-resource-usage https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v3/projects/{project_id}/apps/{app_id}/containers/cpu-mem Retrieve current CPU and memory usage per container with totals. # List Recent Containers Source: https://cerebrium.ai/docs/api-reference/containers/list-recent-containers https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/containers Retrieve a list of recent containers for a specific app. # Search Containers Source: https://cerebrium.ai/docs/api-reference/containers/search-containers https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v3/projects/{project_id}/apps/{app_id}/containers/search Search containers by ID or status with pagination. # Stop Container Source: https://cerebrium.ai/docs/api-reference/containers/stop-container https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json delete /v2/projects/{project_id}/apps/{app_id}/containers/{container_id} Stop a specific container. # Get Execution Time Metrics Source: https://cerebrium.ai/docs/api-reference/metrics/get-execution-time-metrics https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/metrics/run-execution-time Retrieve run execution time percentiles for an app. # Get Response Time Metrics Source: https://cerebrium.ai/docs/api-reference/metrics/get-response-time-metrics https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/metrics/response-time Retrieve end-to-end response time percentiles for an app. # Get Startup Time Metrics Source: https://cerebrium.ai/docs/api-reference/metrics/get-startup-time-metrics https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/metrics/startup-time Retrieve cold start and container startup time metrics for an app. # Create Project Source: https://cerebrium.ai/docs/api-reference/projects/create-project https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json post /v2/projects Create a new project. # Delete Project Source: https://cerebrium.ai/docs/api-reference/projects/delete-project https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json delete /v2/projects/{project_id} Remove a specific project. # Get Project Source: https://cerebrium.ai/docs/api-reference/projects/get-project https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id} Retrieve details of a specific project by its ID. # Get Project Cost Source: https://cerebrium.ai/docs/api-reference/projects/get-project-cost https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/cost Retrieve current billing period cost breakdown for a project. # List Projects Source: https://cerebrium.ai/docs/api-reference/projects/list-projects https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects Retrieve a list of projects. # Modify Project Source: https://cerebrium.ai/docs/api-reference/projects/modify-project https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json patch /v2/projects/{project_id} Update the configuration or metadata of a specific project. # Count Queued Runs Source: https://cerebrium.ai/docs/api-reference/runs/count-queued-runs https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/runs/queueDepth Retrieve the number of queued runs for a specific app. # Get Run Source: https://cerebrium.ai/docs/api-reference/runs/get-run https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/runs/{run_id} Retrieve details for a specific run of an app. # Get Runs Chart Data Source: https://cerebrium.ai/docs/api-reference/runs/get-runs-chart-data https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/runs/charts Retrieve aggregated run data optimized for charting over long time ranges. # List Runs Source: https://cerebrium.ai/docs/api-reference/runs/list-runs https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/runs Retrieve a list of runs for a specific app. # Calculating compute cost Source: https://cerebrium.ai/docs/calculating-cost Understand how Cerebrium bills GPU, CPU, and memory per second, what counts toward build and runtime charges, and estimate monthly deployment costs. Deployment cost is based on the hardware selected and the execution time. Every time code runs or you configure a machine to stay running, compute is billed. GPU, CPU, and Memory usage are charged per second; persistent storage is charged per GB per month. View compute pricing on the [pricing page](https://www.cerebrium.ai/pricing). Listed rates apply to the default `interruptible` compute tier. Apps configured with the `protected` tier are billed at 2x these rates across GPU, CPU, and memory. See [Compute Tier](/docs/scaling/scaling-apps#compute-tier). Deploying a model incurs two billable processes: 1. **Build process** — sets up the app environment: a Python environment with the specified parameters, required apt packages, Conda and Python packages, and any model files. A build is only charged when the environment needs rebuilding, i.e., a `build` or `deploy` command runs with changed requirements, parameters, or code. Each build step is cached, so subsequent builds cost substantially less than the first. 2. **App runtime** — the time code runs from start to finish on each request. Three cost components apply: * Cold-start: The time to spin up server(s), load the environment, connect storage, etc. Cerebrium continuously optimizes cold-start latency. Cold-start time is not billed. * Model initialization: Code outside the request function that only runs on cold start (e.g., loading a model into GPU RAM, importing packages). This time is billed. * Function runtime: Code inside the request function, executed on every request. **Example cost calculation** A model deployment requires: * 24 GB VRAM (A10): \$0.000306 per second * 2 CPU cores: 2 \* \$0.00000655 per second * 20GB Memory: 20 \* \$0.00000222 per second Assume the app works on the first deployment, incurring a single 2-minute build. The app has 10 cold starts per day with an average initialization of 2 seconds and an average runtime (predict) of 2 seconds. The expected monthly volume is 100,000 inferences. ```python theme={null} # Your variables average_initialization_time = 2 cold_starts_per_month = 300 # 10 a day for 30 days average_inference_time = 2 # seconds number_of_inferences = 100000 # number of inferences per month GPU_cost = 0.000306 # per second CPU_cost = 0.00000655 # per second per core memory_cost = 0.00000222 # per second per GB num_of_cpu_cores = 2 gb_of_RAM = 20 build_seconds = 120 # 2 minutes # cost calculation compute_rate = GPU_cost + (CPU_cost * num_of_cpu_cores) + (memory_cost * gb_of_RAM) total_build_compute_cost = build_seconds * compute_rate total_initialization_time = average_initialization_time * cold_starts_per_month total_inference_time = average_inference_time * number_of_inferences initialization_compute_cost = total_initialization_time * compute_rate inference_compute_cost = total_inference_time * compute_rate storage_cost = gb_of_persistent_storage * persistent_storage_cost total_cost = inference_compute_cost + storage_cost + total_build_compute_cost + initialization_compute_cost print(f"Build Compute cost: ${total_build_compute_cost :.2f}/month", f"Initialization Compute cost: ${initialization_compute_cost :.2f}/month", f"Inference Compute cost: ${inference_compute_cost :.2f}/month", f"\nStorage cost: ${storage_cost :.2f}/month", f"\nTotal cost: ${total_cost :.2f}/month") ``` # Custom Dockerfiles Source: https://cerebrium.ai/docs/container-images/custom-dockerfiles Deploy containerized apps on Cerebrium with your own Dockerfile, from Python FastAPI servers to compiled Rust binaries, using the custom runtime config. Cerebrium supports deploying existing containerized apps — from standard Python apps to compiled Rust binaries — using a custom Dockerfile. This allows portable, locally reproducible deployment environments. ## Building Dockerized Python Apps A simple containerized FastAPI server: ```python theme={null} from fastapi import FastAPI app = FastAPI() @app.post("/hello") def hello(): return {"message": "Hello Cerebrium!"} @app.get("/health") def health(): return "OK" @app.get("/ready") def ready(): return "OK" ``` The corresponding Dockerfile: ```dockerfile theme={null} # Base image FROM python:3.12-bookworm RUN apt-get update && apt-get install dumb-init RUN update-ca-certificates # Source code COPY . . # Dependencies RUN pip install -r requirements.txt # Configuration EXPOSE 8192 CMD ["python", "-m", "uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8192"] ``` Dockerfiles for Cerebrium have three requirements: 1. Expose a port with the `EXPOSE` command. This port is referenced in `cerebrium.toml` 2. Include a `CMD` command to specify the container's startup process (typically the server) 3. Set the working directory with `WORKDIR` to ensure correct file paths (defaults to root if not specified) Update cerebrium.toml to include a custom runtime section with the `dockerfile_path` parameter: ```toml theme={null} [cerebrium.runtime.custom] port = 8192 healthcheck_endpoint = "/health" readycheck_endpoint = "/ready" dockerfile_path = "./Dockerfile" ``` The configuration requires four key parameters: * `port`: The port the server listens on. * `healthcheck_endpoint`: The endpoint used to confirm instance health. If unspecified, defaults to a TCP ping on the configured port. If the health check registers a non-200 response, Cerebrium considers the instance *unhealthy* and restarts it if it does not recover in time. * `readycheck_endpoint`: The endpoint used to confirm if the instance is ready to receive. If unspecified, defaults to a TCP ping on the configured port. If the ready check registers a non-200 response, Cerebrium does not route requests to the instance. * `dockerfile_path`: The relative path to the Dockerfile used to build the app. If the Dockerfile omits a `CMD` clause, specify the `entrypoint` parameter in `cerebrium.toml`: ```toml theme={null} [cerebrium.runtime.custom] entrypoint = ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8192"] ... ``` When specifying a `dockerfile_path`, all dependencies and necessary commands should be installed and executed within the Dockerfile. Dependencies listed under `cerebrium.dependencies.*`, as well as `cerebrium.deployment.shell_commands` and `cerebrium.deployment.pre_build_commands`, will be ignored. ## Building Generic Dockerized Apps Cerebrium supports non-Python apps as long as you provide a Dockerfile. The following example shows a Rust-based API server using the Axum framework: ```rust theme={null} use axum::{ routing::{get, post}, Json, Router, }; use serde_json::json; async fn hello() -> Json { Json(json!({ "message": "Hello Cerebrium!" })) } async fn health() -> &'static str { "OK" } async fn ready() -> &'static str { "OK" } #[tokio::main] async fn main() { let app = Router::new() .route("/hello", post(hello)) .route("/health", get(health)) .route("/ready", get(health)); tracing::info!("Listening on port 8192"); let listener = tokio::net::TcpListener::bind("0.0.0.0:8192").await.unwrap(); axum::serve(listener, app).await.unwrap(); } ``` A multi-stage Dockerfile separates the build step from the runtime, producing a smaller and more secure image: ```dockerfile theme={null} # Build stage FROM rust:bookworm as build RUN apt-get update && apt-get install dumb-init RUN update-ca-certificates # Project setup RUN USER=root cargo new --bin rs_server WORKDIR /rs_server # Dependencies COPY Cargo.lock ./Cargo.lock COPY Cargo.toml ./Cargo.toml # Cache dependencies RUN cargo build --release RUN rm src/*.rs # Source code COPY src/* src/ # Build RUN rm ./target/release/deps/rs_server* RUN cargo build --release # Runtime stage FROM gcr.io/distroless/base-debian12 WORKDIR / COPY --from=build /etc/ssl/certs/ca-certificates.crt /etc/ssl/certs/ COPY --from=build /lib/x86_64-linux-gnu/libgcc_s.so.1 /lib/x86_64-linux-gnu/libgcc_s.so.1 COPY --from=build /rs_server/target/release/rs_server /rs_server COPY --from=build /usr/bin/dumb-init /usr/bin/dumb-init EXPOSE 8192 CMD ["dumb-init", "--", "/rs_server"] ``` Configure the application in `cerebrium.toml` the same way as the FastAPI example: ```toml theme={null} [cerebrium.runtime.custom] port = 8192 healthcheck_endpoint = "/health" readycheck_endpoint = "/ready" dockerfile_path = "./Dockerfile" ``` # Custom Python Web Servers Source: https://cerebrium.ai/docs/container-images/custom-web-servers Run FastAPI and other ASGI or WSGI Python web servers on Cerebrium with a custom runtime by setting the entrypoint, port, and health check endpoints. Cerebrium's default runtime covers most app needs. For more control, use ASGI or WSGI servers through the custom runtime feature. This enables custom authentication, dynamic batching, frontend dashboards, public endpoints, and WebSocket connections. ## Setting Up Custom Servers A basic FastAPI server running as a custom server on Cerebrium: ```python theme={null} from fastapi import FastAPI app = FastAPI() @app.post("/hello") def hello(): return {"message": "Hello Cerebrium!"} @app.get("/health") def health(): return "OK" @app.get("/ready") def ready(): return "OK" ``` Configure this server in `cerebrium.toml` by adding a custom runtime section: ```toml theme={null} [cerebrium.runtime.custom] port = 5000 entrypoint = ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "5000"] healthcheck_endpoint = "/health" readycheck_endpoint = "/ready" [cerebrium.dependencies.pip] pydantic = "latest" numpy = "latest" loguru = "latest" fastapi = "latest" ``` The configuration requires four key parameters: * `entrypoint`: The command that starts your server * `port`: The port your server listens on * `healthcheck_endpoint`: The endpoint used to confirm instance health. If unspecified, defaults to a TCP ping on the configured port. If the health check registers a non-200 response, Cerebrium considers the instance *unhealthy* and restarts it if it does not recover in time. * `readycheck_endpoint`: The endpoint used to confirm if the instance is ready to receive. If unspecified, defaults to a TCP ping on the configured port. If the ready check registers a non-200 response, Cerebrium does not route requests to the instance. For ASGI applications like FastAPI, include the appropriate server package (like `uvicorn`) in your dependencies. After deployment, your endpoints become available at `https://api.cerebrium.ai/v4/p-xxxxxxxx/[app-name]/your/endpoint`. The [FastAPI Server Example](https://github.com/CerebriumAI/examples) provides a complete implementation. ## Request Headers Custom web servers receive the Cerebrium run ID in the `X-Request-Id` header on every request. This corresponds to the internal `run_id` and is useful for tracking and debugging. # Defining Container Images Source: https://cerebrium.ai/docs/container-images/defining-container-images Define your Cerebrium container image in cerebrium.toml, from Python, pip, apt, and conda dependencies to custom Docker base images and build commands. ## Introduction Cerebrium abstracts infrastructure management into configuration, so teams focus on app code. A single TOML file manages environment setup, deployments, and scaling: tasks that typically require dedicated teams. Unlike traditional Docker or Kubernetes setups with multiple configuration files and orchestration rules, Cerebrium uses a single `cerebrium.toml` file. The system handles container lifecycle, networking, and scaling automatically based on this configuration. ## Why TOML? Python decorators scatter infrastructure settings throughout code files, making changes risky and reviews difficult. TOML centralizes configuration in one place, making it easier to track changes and maintain consistency. Its hierarchical structure maps naturally to app requirements without the accidental complexity of code-based configuration. ### Getting Started Run `cerebrium init` to create a `cerebrium.toml` file in the project root. Edit it to match the app's requirements. You can initialize an existing project by adding a `cerebrium.toml` file to the root of your codebase. Define your entrypoint (`main.py` if using the default runtime, or add an entrypoint to the .toml file if using a custom runtime) and include the necessary files in the `deployment` section of your `cerebrium.toml` file. ## Hardware Configuration Configure GPU type and memory allocations in the hardware section: ```toml theme={null} [cerebrium.hardware] compute = "AMPERE_A10" # GPU selection memory = 16.0 # Memory allocation in GB cpu = 4 # Number of CPU cores gpu_count = 1 # Number of GPUs ``` For detailed hardware specifications see the [toml reference](/docs/toml-reference/toml-reference#hardware-configuration). ## Dependency Management ### Selecting a Python Version The Python runtime version forms the foundation of every Cerebrium app. Supported versions: 3.10 to 3.13. Specify the version in the deployment section: ```toml theme={null} [cerebrium.deployment] python_version = 3.11 ``` The Python version affects the entire dependency chain. For instance, some packages may not support newer Python versions immediately after release. To use a later Python version, please use a [Dockerfile](/docs/container-images/custom-dockerfiles). Changes to the Python version trigger a full rebuild since they affect both the base environment and all Python package installations. ### Adding Python Packages Manage Python dependencies directly in TOML or through requirement files: ```toml theme={null} [cerebrium.dependencies.pip] torch = "==2.0.0" transformers = "==4.30.0" numpy = "latest" ``` Or using an existing requirements file: ```toml theme={null} [cerebrium.dependencies.paths] pip = "requirements.txt" ``` For GitHub repositories, use shell commands instead of pip dependencies to ensure proper versioning. Cerebrium caches pip packages at the node level - including wheel files and compiled binaries - so subsequent builds only install new or updated packages. This significantly reduces build times. ### Adding APT Packages Declare system-level packages (image-processing libraries, audio codecs, etc.) under `[cerebrium.dependencies.apt]`: ```toml theme={null} [cerebrium.dependencies.apt] ffmpeg = "latest" libopenblas-base = "latest" libomp-dev = "latest" ``` Alternatively, reference a text file listing system dependencies: ```toml theme={null} [cerebrium.dependencies.paths] apt = "deps_folder/pkglist.txt" ``` Changes to APT packages trigger a full rebuild of the container image, so builds take longer than when modifying Python packages alone. ### Conda Packages Conda excels at managing complex system-level Python dependencies, particularly for GPU support and scientific computing: ```toml theme={null} [cerebrium.dependencies.conda] cuda = ">=11.7" cudatoolkit = "11.7" opencv = "latest" ``` Alternatively, reference a conda environment file: ```toml theme={null} [cerebrium.dependencies.paths] conda = "conda_pkglist.txt" ``` Like APT packages, Conda packages modify system-level components. Changes trigger a full rebuild. Batch Conda dependency updates together to minimize rebuild time. ## Build Commands The build process includes two command types that execute at different stages during container image creation. ### Pre-build Commands Pre-build commands execute at the start of the build process, before dependency installation. Use them to set up the build environment: ```toml theme={null} [cerebrium.deployment] pre_build_commands = [ # Add specialized build tools "curl -o /usr/local/bin/pget -L 'https://github.com/replicate/pget/releases/download/v0.6.2/pget_linux_x86_64'", "chmod +x /usr/local/bin/pget" ] ``` Common uses: installing build tools, configuring system settings, or preparing the environment for subsequent build steps. ### Shell Commands Shell commands execute after all dependencies install and the application code copies into the container. This later timing ensures access to the complete environment: ```toml theme={null} [cerebrium.deployment] shell_commands = [ # Initialize application resources "python -m download_models", "python -m compile_assets", "python -m init_app" ] ``` Use shell commands for tasks that require the fully configured environment, such as compiling code that depends on installed libraries or downloading resources. ## Custom Docker Base Images The base image determines the OS foundation for the container. The default Debian slim image works for most Python apps; other validated base images support specific requirements. ### Supported Base Images Supported base image categories include NVIDIA, Ubuntu, and Python images. ```toml theme={null} [cerebrium.deployment] docker_base_image_url = "debian:bookworm-slim" # Default minimal image #docker_base_image_url = "nvidia/cuda:12.0.1-runtime-ubuntu22.04" # CUDA-enabled images #docker_base_image_url = "ubuntu:22.04" # debian images ``` Starting with a minimal Debian or Ubuntu base image is recommended, as CUDA images include many pre-installed components that increase container size. While the relationship isn't strictly linear, larger container sizes generally lead to longer cold-starts and build times. Begin with a lean base image and add only essential components as needed. #### Public Docker Hub Images with Namespaces Public Docker Hub images with a namespace (e.g., `bob/infinity`, `huggingface/transformers`) require a local Docker Hub login, even though the image is public. Cerebrium reads `~/.docker/config.json` to authenticate image pulls. ```bash theme={null} # Login to Docker Hub with username (required for namespace/image format) docker login -u your-dockerhub-username # Enter your password or access token when prompted ``` After logging in, you can use the image in your configuration: ```toml theme={null} [cerebrium.deployment] docker_base_image_url = "bob/infinity:latest" ``` Official Docker Hub images without a namespace (like `python:3.11`, `debian:bookworm`, `ubuntu:22.04`) work without requiring a Docker login. Only images in the `namespace/image` format require authentication. Use `docker login -u username` instead of just `docker login`. The latter may use Docker's web-based OAuth flow which creates tokens that are incompatible with our build system. #### Public AWS ECR Images Public ECR images from the `public.ecr.aws` registry work without authentication: ```toml theme={null} [cerebrium.deployment] docker_base_image_url = "public.ecr.aws/lambda/python:3.11" ``` However, **private ECR images** require authentication. See [Using Private Docker Registries](/docs/container-images/private-docker-registry) for setup instructions. ## Custom Runtimes Cerebrium's default runtime covers most apps. Custom runtimes provide more control, enabling features like custom authentication, dynamic batching, public endpoints, or WebSocket connections. ### Basic Configuration Define a custom runtime by adding the `cerebrium.runtime.custom` section to the configuration: ```toml theme={null} [cerebrium.runtime.custom] entrypoint = ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8080"] port = 8080 healthcheck_endpoint = "" # Empty string uses TCP health check readycheck_endpoint = "" # Empty string uses TCP health check ``` Key parameters: * `entrypoint`: Command to start the app (string or string list) * `port`: Port the app listens on * `healthcheck_endpoint`: The endpoint used to confirm instance health. If unspecified, defaults to a TCP ping on the configured port. If the health check registers a non-200 response, it will be considered *unhealthy*, and be restarted should it not recover timely. * `readycheck_endpoint`: The endpoint used to confirm if the instance is ready to receive. If unspecified, defaults to a TCP ping on the configured port. If the ready check registers a non-200 response, it will not be a viable target for request routing. Check out [this example](https://github.com/CerebriumAI/examples/tree/master/11-python-apps/1-asgi-fastapi-server) for a detailed implementation of a FastAPI server that uses a custom runtime. ### Self-Contained Servers Custom runtimes also support apps with built-in servers. For example, deploying a VLLM server requires no Python code: ```toml theme={null} [cerebrium.runtime.custom] entrypoint = "vllm serve meta-llama/Meta-Llama-3-8B-Instruct --host 0.0.0.0 --port 8000 --device cuda" port = 8000 healthcheck_endpoint = "/health" healthcheck_endpoint = "/ready" [cerebrium.dependencies.pip] torch = "latest" vllm = "latest" ``` ### Important Notes * Code is mounted in `/cortex`. Adjust paths accordingly. * The port in your entrypoint must match the `port` parameter. * Install any required server packages (uvicorn, gunicorn, etc.) via pip dependencies. * All endpoints will be available at `https://api.cerebrium.ai/v4/p-xxxxxxxx/{app-name}/your/endpoint`. Deploy with `cerebrium deploy -y`. The system automatically detects custom runtime configuration. ## Deployment Process Deployment process The build process follows a sequence that transforms source code into a production-ready container image: ### Stage 1: App Upload Code is uploaded to Cerebrium, including all source files, configuration, and additional assets needed for the app. ### Stage 2: Image Creation The system creates a container image through the following sequential steps: 1. **Pre-build Commands Execute**: First, any pre-build commands run. These set up the build environment and compile necessary assets before the main installation steps begin. 2. **APT Dependencies Install**: System-level packages install next, establishing the foundation for all other dependencies. 3. **Conda Dependencies Install**: After APT packages are in place, Conda packages install. 4. **Pip Dependencies Install**: Python packages install last, ensuring they have access to all necessary system libraries and binaries. 5. **Python Code Copy**: The app's source code copies into the container, placing it in the correct directory structure. 6. **Shell Commands Execute**: Finally, any build-time shell commands run to complete the image setup. ### Stage 3: Production Image The result is a production-ready container image that contains everything needed to run the app. This image serves as a blueprint for creating individual containers when the app receives requests. # Using Private Docker Registries Source: https://cerebrium.ai/docs/container-images/private-docker-registry Authenticate with Docker Hub, AWS ECR, or other private registries and use private Docker images as base images for your Cerebrium app deployments. Cerebrium supports private Docker images as base images for deployments, including images from Docker Hub and AWS ECR. The build system pulls private images using Docker credentials stored in `~/.docker/config.json`, then builds the app on top of that image. Since these credentials are used during deployment, they should be team-accessible rather than tied to a single individual. Your Docker images MUST support `linux/amd64` architecture. This is required for Cerebrium's build environment. ## Step-by-Step Setup ### Step 1: Login to Your Registry Based on your registry, login using one of the following commands: **Docker Hub:** ```bash theme={null} docker login -u your-dockerhub-username # Enter your password or access token when prompted ``` Use `docker login -u username` instead of just `docker login`. The latter may use Docker's web-based OAuth flow which creates tokens that are incompatible with our build system. **AWS ECR:** ```bash theme={null} aws ecr get-login-password --region us-east-1 | \ docker login --username AWS --password-stdin \ 123456789.dkr.ecr.us-east-1.amazonaws.com ``` **Generic Registry:** ```bash theme={null} docker login -u your-username registry.company.com # Enter credentials when prompted ``` ### Step 2: Verify Login Check that your credentials are saved: ```bash theme={null} cat ~/.docker/config.json | jq '.auths | keys' ``` The output lists registered registry URLs. ### Step 3: Configure Your Project Set the `docker_base_image_url` in `cerebrium.toml` to the registry image URL: ```toml theme={null} [cerebrium.deployment] name = "my-app" python_version = "3.11" docker_base_image_url = "your-registry.com/your-org/your-image:tag" # Examples: # docker_base_image_url = "mycompany/ml-base:v2.1" # Docker Hub # docker_base_image_url = "123456.dkr.ecr.us-east-1.amazonaws.com/ml-base:latest" # ECR # docker_base_image_url = "gcr.io/project-id/ml-base:latest" # GCR ``` ### Step 4: Deploy Run: ```bash theme={null} cerebrium deploy ``` The application builds as normal. ## Security Notes Cerebrium applies the following security measures to protect registry credentials: 1. **In Transit:** Credentials are sent over HTTPS 2. **In Storage:** Stored in DynamoDB with automatic TTL-based deletion after build completion 3. **In Logs:** Never logged or displayed (using obfuscated.String type) 4. **In Build:** Only accessible during build, then discarded # CI/CD Pipelines Source: https://cerebrium.ai/docs/deployments/ci-cd Set up a CI/CD pipeline with GitHub Actions and Cerebrium service account keys to automatically deploy your app when a branch is pushed or merged. Configure a Continuous Integration or Continuous Deployment system (CI/CD) to automatically deploy a new version of an app to production/development when a branch is pushed or workflow triggered. This guide sets up a CI/CD pipeline using GitHub Actions and Cerebrium's Service Accounts. Maintaining **separate development and production apps** in separate projects is recommended, so that you can safely test changes before going live. ### 1. Authenticating to Cerebrium The GitHub Action workflow uses a **Service Account key** to authenticate to Cerebrium. Service Account keys are credentials tied to a specific service, not an individual. 1. Go to the [Cerebrium Dashboard](https://dashboard.cerebrium.ai/) and open the **API Keys** page. 2. Click **Create Service Account**, name it `"GitHub Actions CI/CD"`, choose an expiry date, and click **Create**. 3. **Copy the key** generated for the new service account. Cerebrium API Keys dashboard ### 2. Define Secrets in a GitHub Environment Store this key in a secret for use in GitHub Actions workflows. 1. Navigate to the **GitHub repository → Settings → Environments** 2. Create environments named `"dev"` and/or `"prod"` (Adjust to specific needs) 3. Inside each environment, add 2 secrets: * **Name**: `CEREBRIUM_SERVICE_ACCOUNT_TOKEN` * **Value**: the key copied in the previous step * **Name**: `CEREBRIUM_PROJECT_ID` * **Value**: your project ID > 🔒 Use *secrets*, not plain environment variables, to prevent leaking sensitive data in logs. GitHub ### 3. GitHub Actions Workflow GitHub Actions use workflow files in the `.github/workflows/` directory. Create one to define the deployment process. 1. Create a `.github/workflows` directory in the root of a new or existing GitHub repository 2. Within that directory, create a new file named `cerebrium-deploy.yml` 3. Paste the following YAML into the file: ```yaml theme={null} name: Cerebrium Deployment on: push: branches: - master workflow_dispatch: inputs: environment: description: "Environment" type: choice options: - "dev" - "prod" required: true default: "dev" pull_request: branches: - master - development jobs: deployment: runs-on: ubuntu-latest environment: ${{ github.event.inputs.environment || (github.ref == 'refs/heads/master' && 'prod') || 'dev' }} env: ENV: ${{ github.event.inputs.environment || (github.ref == 'refs/heads/master' && 'prod') || 'dev' }} CEREBRIUM_SERVICE_ACCOUNT_TOKEN: ${{ secrets.CEREBRIUM_SERVICE_ACCOUNT_TOKEN }} CEREBRIUM_PROJECT_ID: ${{ secrets.CEREBRIUM_PROJECT_ID }} steps: - uses: actions/checkout@v4.0 - uses: actions/setup-python@v5.0 with: python-version: "3.10" - name: Install Cerebrium run: pip install cerebrium - name: Set Cerebrium Project ID run: cerebrium project set ${{ secrets.CEREBRIUM_PROJECT_ID }} - name: Deploy App run: cerebrium deploy ``` Once committed, this workflow will automatically: * Deploy a production application on every push to master. * Deploy a development app on pull requests to master or development branches. Monitor workflow runs in the “Actions” tab of the repository on GitHub. # Gradual Roll-out Source: https://cerebrium.ai/docs/deployments/gradual-roll-out Use roll_out_duration_seconds in cerebrium.toml to gradually shift traffic between revisions after a deploy and minimize disruption in production. This feature is available from CLI version 1.38.2 The `roll_out_duration_seconds` parameter in the `[cerebrium.scaling]` section of your `cerebrium.toml` file controls how quickly traffic transitions between revisions after a successful build. ## Overview Each deployment creates a new revision. The `roll_out_duration_seconds` parameter determines how long traffic takes to transition from the old revision to the new one. Traffic shifts in 5 batches of 20% each over the specified duration, minimizing disruptions. ## Configuration Add the `roll_out_duration_seconds` parameter to the `[cerebrium.scaling]` section of your `cerebrium.toml` file: ```toml theme={null} [cerebrium.scaling] roll_out_duration_seconds = 0 # Default value ``` ## Parameters * **Valid range**: 0-600 seconds * **Default value**: 0 (immediate transition) ## Best Practices * **Development environments**: Keep the value at 0 during development for immediate transitions * **Production environments**: Use lower values to optimize cost and resources while ensuring smooth transitions * **High-traffic applications**: Consider using higher values for gradual transitions to minimize disruption # Multi-Region Deployment Source: https://cerebrium.ai/docs/deployments/multi-region-deployment Run a Cerebrium app globally across multiple regions for more GPU capacity and lower latency, or pin it to one region for data residency needs. Deploy an app once and run it in multiple regions. The `region` parameter in `cerebrium.toml` controls placement: run globally on whatever capacity is available (recommended), or pin the app to a specific region. The parameter is optional; when omitted, the platform chooses placement automatically based on the app's hardware requirements. Multi-region deployment is currently in **beta**. We will make rapid updates and improvements over the next few months to bring full functionality to life. Please reach out on our [Discord](https://discord.gg/ATj6USmeE2) about features/functionality you would like to see. ## Why Use Multi-Region Deployment * **More capacity, less queueing**: An app that can run anywhere draws on the GPU capacity of every region. Constraining placement shrinks the pool of available hardware. * **High availability**: Traffic shifts away from degraded or busy regions automatically. * **Reduced latency**: Requests are routed to healthy capacity, preferring lower-latency regions. * **Data residency**: Pin an app to a region to keep sensitive data within a specific geographic region and comply with regulations like GDPR and CCPA. * **No duplicate apps**: One app, one endpoint, and one configuration instead of a separate deployment per region. ## Run Globally Set `region = "global"` to let the platform place the app in any region with available capacity that matches its hardware requirements: ```toml theme={null} [cerebrium.hardware] region = "global" compute = "AMPERE_A10" cpu = 2 memory = 8.0 ``` When Cerebrium brings new capacity online, global apps can run there without a configuration change or redeploy. ## Run in a Geographical Zone Set `region = "us"` or `region = "eu"` to run the app in any region within that geographical zone: ```toml theme={null} [cerebrium.hardware] region = "eu" compute = "HOPPER_H100" cpu = 4 memory = 32.0 ``` * `"us"`: US regions, such as us-east-1 and us-central1 * `"eu"`: EU regions, such as eu-north-1 and eu-north1 This keeps the app within the geographical zone while drawing on the capacity of all of its regions, so queuing is less likely than when pinning a single region. ## Pin a Region Set `region` to a specific region to run the app only there: ```toml theme={null} [cerebrium.hardware] region = "us-east-1" compute = "AMPERE_A10" cpu = 2 memory = 8.0 ``` Pinning guarantees placement in that region but limits the app to that region's capacity, which increases the likelihood of request queuing when the region is busy. ### Available Regions #### United States * **us-east-1** (N. Virginia) * **us-central1** (Kansas City) #### Europe * **eu-north-1** (Stockholm) * **eu-north1** (Finland) #### Available on request Contact [support](mailto:support@cerebrium.ai) to enable access to these regions: * **us-west-2** (Oregon) * **eu-west-2** (United Kingdom) * **eu-central-1** (Frankfurt) * **ap-south-1** (Mumbai) * **ap-northeast-1** (Tokyo) * **sa-east-1** (São Paulo) * **ca-central-1** (Montreal) * **me-central-1** (Dubai) Some regions run on cloud providers other than AWS: when pinning `us-central1` or `eu-north1`, set `provider = "nebius"`. ## Hardware Preferences Placement respects both the region and the hardware requirements of an app. Specify acceptable GPU types in preference order with a `compute` list (see [GPU preference lists](/docs/hardware/using-gpus#gpu-preference-lists)): ```toml theme={null} [cerebrium.hardware] region = "global" compute = ["HOPPER_H100", "HOPPER_H200", "AMPERE_A100_80GB"] ``` The set of eligible regions is the intersection of both constraints: any region with several acceptable GPU types is the widest pool; one region with a single GPU type is the narrowest. ## Endpoints Apps keep a single endpoint no matter how many regions they run in: ```text theme={null} https://api.cerebrium.ai/v4/p-xxxxxxxx/{app-name}/{function-name} ``` Requests enter through the global router and are directed to a region where the app is running. Regioned hostnames continue to work but add a proxy hop. See [REST API](/docs/endpoints/inference-api). ## How Requests Are Routed * A request enters through the region closest to where it is made, then routes to a region with a warm instance before a cold start is triggered elsewhere. * Routing prefers lower-latency regions but does not guarantee the nearest region. * There is no session affinity. Consecutive requests can be served from different regions. ## Scaling Across Regions `min_replicas`, `max_replicas`, and `scaling_buffer` apply to the app as a whole, not per region. The platform distributes instances across eligible regions based on traffic and available capacity. View where containers are running: ```bash theme={null} cerebrium containers list my-app ``` The output includes the region of each container. ## Storage Persistent storage is managed per region: an app has an independent `/persistent-storage` volume in each region it runs in, however placement is configured. Files written in one region are not guaranteed to be available in other regions. Region-local caches, such as model weights downloaded on first load, fill independently per region. Apps deployed with `region = "global"` also mount `/global-persistent-storage`, a single volume shared across every region the app runs in. Files written there are visible from all regions, and reads are cached per region. Use the global volume for data that must be available everywhere and `/persistent-storage` for region-local data. Manage files on the global volume by passing `--region global` to the file commands. See [Managing Files](/docs/storage/managing-files#global-storage). Target a region with the file commands using `--region` (or its short form `-r`). This works for all file commands: `cp`, `ls`, `rm`, and `download`. ```bash theme={null} cerebrium ls --region eu-north-1 cerebrium cp model.bin -r eu-north-1 ``` Set a default region for CLI commands: ```bash theme={null} cerebrium region set us-east-1 ``` `--region` applies only to the command it's attached to. The default set by `cerebrium region set` is not modified. ## GPU Availability by Region GPU availability varies across regions due to infrastructure constraints and local demand. CPU workloads run in every region. | Region | Available GPUs | | - | - | | us-east-1 | BLACKWELL\_B200, HOPPER\_H200, HOPPER\_H100, AMPERE\_A100\_80GB, AMPERE\_A100\_40GB, ADA\_L40, ADA\_L4, AMPERE\_A10, TURING\_T4, INF2, TRN1 | | us-central1 | BLACKWELL\_RTX6000, BLACKWELL\_B200, HOPPER\_H200 | | eu-north1 (Finland) | HOPPER\_H200, HOPPER\_H100, ADA\_L40 | | us-west-2 (on request) | HOPPER\_H200, HOPPER\_H100, AMPERE\_A100\_80GB, AMPERE\_A100\_40GB, ADA\_L40, ADA\_L4, AMPERE\_A10, INF2, TRN1 | | eu-west-2 (on request) | HOPPER\_H100, AMPERE\_A10, ADA\_L4, TURING\_T4 | | eu-north-1 (Stockholm) | HOPPER\_H100, ADA\_L40, ADA\_L4, AMPERE\_A10, TURING\_T4, INF2, TRN1 | | ap-south-1 (on request) | ADA\_L4, AMPERE\_A10, TURING\_T4 | # Async requests Source: https://cerebrium.ai/docs/endpoints/async Run Cerebrium functions asynchronously with the async query parameter, get a run_id back instantly, and forward results via a webhook endpoint. Some apps require asynchronous "fire-and-forget" execution. In this model, Cerebrium handles running the function, while you are responsible for ensuring data leaves the function (e.g. via a webhook). Enable async execution by adding the `async=true` query parameter to the request: ```bash theme={null} curl -X POST 'https://api.cerebrium.ai/v4/p-xxxxxxxx//run?async=true' \ -H 'Content-Type: application/json' \ -H 'Authorization: Bearer ' \ --data '{"param": "hello world"}' ``` This *immediately* returns a `202 Accepted` response containing only a `run_id`: ```bash theme={null} HTTP/1.1 202 Accepted Connection: close Content-Length: 50 Content-Type: application/json Date: Tue, 12 Nov 2024 03:13:25 GMT Server: envoy Vary: Origin X-Envoy-Upstream-Service-Time: 2 X-Request-Id: 21eb3b98-4b10-9ad6-8681-a47172828024 {"run_id":"21eb3b98-4b10-9ad6-8681-a47172828024"} ``` Async functions run for a maximum of **12 hours**, bounded by the `response_grace_period` in `cerebrium.toml`. This defaults to 15 minutes. Update it to match the maximum time the task needs. Cerebrium runs the HTTP request in the background, but the function itself must still behave **synchronously**. It must complete its work and return a result. Returning a response while the application is still processing causes Cerebrium to begin terminating the container. Only return once all processing is finished. Because async calls do not return a response to the caller, the function must export any relevant data itself. Combine async execution with a `webhookEndpoint` to have Cerebrium automatically forward the function's response body: ```bash theme={null} curl -X POST 'https://api.cerebrium.ai/v4/p-xxxxxxxx//run?async=true&webhookEndpoint=https%3A%2F%2Fwebhook.site%2F' \ -H 'Content-Type: application/json' \ -H 'Authorization: Bearer ' \ --data '{"param": "hello world"}' ``` This is a proxy-level feature. No code changes are required to use webhook forwarding. In the dashboard, the function is marked **async** but still shows the status of the internal synchronous call (e.g. if the call failed, the async request state is `failure`). # REST API Source: https://cerebrium.ai/docs/endpoints/inference-api Call your Cerebrium apps over the REST API with POST requests and JWT authentication, and understand response formats and HTTP status codes. All functions on Cerebrium are accessible via POST requests, unless marked private by prefixing the function name with an underscore (e.g. `_private_function()`). Authenticate using the JWT token from the **API Keys** section of the dashboard. Endpoints require this token only when `cerebrium.toml` sets [`disable_auth = false`](/docs/toml-reference/toml-reference). Authentication is disabled by default. ## Request format The POST request follows the structure below, where `{function}` is the name of the function to invoke. This example calls `predict()` from `main.py`. ```bash theme={null} curl --location --request POST 'https://api.cerebrium.ai/v4/p-xxxxxxxx/{app-name}/{function}' \ --header 'Authorization: Bearer ' \ --header 'Content-Type: application/json' \ --data '{ "function_param": "data" }' ``` Regioned hostnames such as `https://api.aws.us-east-1.cerebrium.ai` continue to work. Requests to these hostnames are proxied to the global router, which can add latency. Use `https://api.cerebrium.ai` for the lowest-latency path. ## Response format Responses follow this format: ```json theme={null} { "run_id": "52eda406-b81b-43f5-8deb-fcf80dfsb74b", "run_time_ms": 326.34, "result": { "some": "data" } } ``` ## Status codes * `200` → the function completed successfully * `401` / `403` → authentication or authorization failed * `404` → the app or function path could not be found * `5xx` → either your app raised an exception or the platform could not complete the request To return a custom status code such as `422` or `404`, include a `status_code` field in the JSON response from `main.py`. When debugging a failed request, check the HTTP status code and the app logs in the Cerebrium dashboard. Inspecting the request payload, logs, and `run_id` together speeds up diagnosis. # OpenAI-Compatible Endpoints Source: https://cerebrium.ai/docs/endpoints/openai-compatible-endpoints Build an OpenAI compatible chat completions endpoint on Cerebrium and stream responses to the OpenAI Python client using your JWT as the API key. All Cerebrium endpoints are OpenAI-compatible, supporting both `/chat/completions` and `/embedding`. Below is a basic implementation of a streaming OpenAI-compatible endpoint. For a full example using vLLM, see the [OpenAI-compatible endpoint example](https://github.com/CerebriumAI/examples/tree/master/5-large-language-models/1-openai-compatible-endpoint). A streaming-compatible Cerebrium function must: * Specify all the parameters that OpenAI sends in the function signature. * Return `yield data`, where `yield` signals streaming and `data` is the JSON-serializable object returned to the caller. Here's a small snippet from the example listed above: ```python theme={null} def run(messages: list, model: str,...): ##existing code async for output in results_generator: prompt = output.outputs new_text = prompt[0].text[len(previous_text):] previous_text = prompt[0].text full_text += new_text response = ChatCompletionResponse( id=run_id, object="chat.completion", created=int(time.time()), model=model, choices=[{ "text": new_text, "index": 0, "logprobs": None, "finish_reason": prompt[0].finish_reason or "stop" }] ) yield json.dumps(response.model_dump()) ``` Once deployed, set the base URL to the target function and use the Cerebrium JWT (from the dashboard) as the API key. Client code: ```python theme={null} import os from openai import OpenAI client = OpenAI( # This is the default and can be omitted base_url="https://api.cerebrium.ai/v4/p-xxxxxxxx/1-openai-compatible-endpoint/run", ##This is the name of the function you are calling api_key="", ) chat_completion = client.chat.completions.create( messages=[ {"role": "user", "content": "What is a mistral?"}, {"role": "assistant", "content": "A mistral is a type of cold, dry wind that blows across the southern slopes of the Alps from the Valais region of Switzerland into the Ligurian Sea near Genoa. It is known for its strong and steady gusts, sometimes reaching up to 60 miles per hour."}, {"role": "user", "content": "How does the mistral wind form?"} ], model="meta-llama/Meta-Llama-3.1-8B-Instruct", stream=True ) print("Starting to receive chunks...") for chunk in chat_completion: print(chunk) print("Finished receiving chunks.") ``` The output then looks like this: ``` Starting to receive chunks... ChatCompletionChunk(id='412f0e25-61c4-93b8-a00f-09a5076cd9fa', choices=[Choice(delta=None, finish_reason='stop', index=0, logprobs=None, text=' The')], created=1724166657, model='gpt-3.5-turbo', object='chat.completion', service_tier=None, system_fingerprint=None, usage=None) ChatCompletionChunk(id='412f0e25-61c4-93b8-a00f-09a5076cd9fa', choices=[Choice(delta=None, finish_reason='stop', index=0, logprobs=None, text=' formation')], created=1724166657, model='gpt-3.5-turbo', object='chat.completion', service_tier=None, system_fingerprint=None, usage=None) ChatCompletionChunk(id='412f0e25-61c4-93b8-a00f-09a5076cd9fa', choices=[Choice(delta=None, finish_reason='stop', index=0, logprobs=None, text=' of')], created=1724166657, model='gpt-3.5-turbo', object='chat.completion', service_tier=None, system_fingerprint=None, usage=None) ChatCompletionChunk(id='412f0e25-61c4-93b8-a00f-09a5076cd9fa', choices=[Choice(delta=None, finish_reason='stop', index=0, logprobs=None, text=' the')], created=1724166657, model='gpt-3.5-turbo', object='chat.completion', service_tier=None, system_fingerprint=None, usage=None) ChatCompletionChunk(id='412f0e25-61c4-93b8-a00f-09a5076cd9fa', choices=[Choice(delta=None, finish_reason='stop', index=0, logprobs=None, text=' mist')], created=1724166657, model='gpt-3.5-turbo', object='chat.completion', service_tier=None, system_fingerprint=None, usage=None) ... ``` # Streaming Endpoints Source: https://cerebrium.ai/docs/endpoints/streaming Stream live model output from a Cerebrium endpoint over server-sent events by yielding results from a Python generator or iterator function. Streaming sends live output from a model over a server-sent event (SSE) stream. It works with any Python object that implements the iterator or generator protocol. The generator/iterator must `yield` data, which is sent downstream via the `text/event-stream` Content-Type. You can send data in JSON format and decode it on the client side. A minimal example: ```python theme={null} import time def run(upper_range: int): for i in range(upper_range): yield f"Number {i} " time.sleep(1) ``` Deploy this snippet and call the endpoint. SSE events appear progressively, one per second: ```bash theme={null} curl -X POST 'https://api.cerebrium.ai/v4/p-xxxxxxxx/2-streaming-endpoint/run' \ -H 'Content-Type: application/json' \ -H 'Accept: text/event-stream' \ -H 'Authorization: Bearer ' \ --data '{"upper_range": 3}' ``` This should output: ```bash theme={null} HTTP/1.1 200 OK cache-control: no-cache content-encoding: gzip content-type: text/event-stream; charset=utf-8 date: Tue, 28 May 2024 21:12:46 GMT server: envoy transfer-encoding: chunked vary: Accept-Encoding x-envoy-upstream-service-time: 198995 x-request-id: e6b55132-32af-96d7-a064-8915c4a42452 data: Number 0 ... ``` The remaining data streams in every second: ``` ... data: Number 1 data: Number 2 ``` Postman also supports SSE streams natively. Streaming For a Falcon-7B streaming example, see the [streaming endpoint example](https://github.com/CerebriumAI/examples/tree/master/5-large-language-models/2-streaming-endpoint). # Webhook Forwarding Source: https://cerebrium.ai/docs/endpoints/webhook Forward function responses to an external webhook with a query parameter, including automatic retries and HMAC signature verification on delivery. Forward function response data to an external endpoint via POST by adding the `webhookEndpoint` query parameter to any API call: ```bash theme={null} curl -X POST 'https://api.cerebrium.ai/v4/p-xxxxxxxx//run?webhookEndpoint=https%3A%2F%2Fwebhook.site%2F' \ -H 'Content-Type: application/json' \ -H 'Authorization: Bearer ' \ --data '{"param": "hello world"}' ``` The proxy forwards the response body as a POST request to the specified webhook. No code changes are required. Ensure the webhook endpoint accepts POST requests. Webhook forwarding works with both Cortex and Custom runtimes, but only for HTTP requests (not WebSockets). ## Retry Behavior Webhook requests are sent asynchronously and do not block the function's response. Cerebrium retries failed deliveries automatically: * **Maximum attempts**: 3 attempts total * **Delay between retries**: Up to 5 seconds between attempts **Retry conditions:** * Network errors (connection failures, timeouts) will trigger retries * `404 Not Found` responses will trigger retries (endpoint may be temporarily unreachable) * Other `4xx` client errors will **not** be retried (considered permanent failures) * `5xx` server errors will **not** be retried (to avoid potential duplicate side effects) * Non-2xx responses (excluding the above) will trigger retries If all attempts fail, the error is logged but will not affect your function's response to the original caller. **The webhook endpoint should be idempotent.** A webhook may be retried even if it was already processed. For example, the Cerebrium backend may not receive the success response due to network issues. Design the endpoint to handle duplicate deliveries gracefully. For example, if the function returns `{"smile": "wave"}`, the proxy sends a POST like this to the webhook: Cerebrium does not literally use `curl`. This is illustrative of the request the webhook endpoint receives. ```bash theme={null} curl -X POST 'https://webhook.site/' \ -H 'Content-Type: application/json' \ --data '{"smile": "wave"}' ``` ## Webhook Signature Verification Verify that webhooks originate from Cerebrium using signature verification: 1. Include an `X-Webhook-Secret` header with a secret value when making requests to the Cerebrium API. 2. Incoming webhooks include the following headers: * `X-Request-Id`: A unique identifier for the request * `X-Cerebrium-Webhook-Timestamp`: The Unix timestamp when the webhook was sent * `X-Cerebrium-Webhook-Signature`: The signature in the format `v1,` 3. To verify the signature: ```python theme={null} import hmac import hashlib def verify_webhook_signature(request_id, timestamp, body, signature, secret): # Remove the 'v1,' prefix from the signature signature = signature.split(',')[1] if ',' in signature else signature # Construct the signed content signed_content = f"{request_id}.{timestamp}.{body}" # Calculate expected signature expected_signature = hmac.new( secret.encode(), signed_content.encode(), hashlib.sha256 ).hexdigest() # Compare signatures return hmac.compare_digest(expected_signature, signature) ``` The body should include everything returned from Cerebrium (run\_id, function response, timestamp, etc.), not just your function response. A matching signature confirms the webhook originated from Cerebrium and has not been tampered with. # WebSocket Endpoints Source: https://cerebrium.ai/docs/endpoints/websockets Create real-time bidirectional WebSocket endpoints on Cerebrium using a custom runtime with FastAPI and connect clients over secure wss URLs. WebSocket endpoints stream responses to the client, enabling real-time, bidirectional communication. ## Required changes Setting up a WebSocket endpoint requires a custom runtime. Configure it in `cerebrium.toml`: ```toml theme={null} [cerebrium.runtime.custom] port = 5000 entrypoint = "uvicorn main:app --host 0.0.0.0 --port 5000" healthcheck_endpoint = "/health" readycheck_endpoint = "/ready" ``` Fields: * `port`: The port the app listens on inside the container. * `entrypoint`: The command to start the app. This example uses Uvicorn to run a FastAPI app in `main.py`. * `healthcheck_endpoint`: Confirms instance health. Defaults to a TCP ping on the configured port. A non-200 response marks the instance as *unhealthy*, triggering a restart if it does not recover. * `readycheck_endpoint`: Confirms the instance is ready to receive traffic. Defaults to a TCP ping on the configured port. A non-200 response removes the instance from request routing. ## Things to note * Custom Runtime Required: WebSocket endpoints require a custom runtime to control how the app runs inside the container. * WebSocket URL: Requests must use a `wss://` URL. The client must support secure WebSocket connections. ## Making a request Test the WebSocket endpoint using websocat, a command-line WebSocket client: ```bash theme={null} websocat wss://api.cerebrium.ai/v4/p-xxxxxxxx// ``` ## Implementing the WebSocket endpoint Example WebSocket endpoint using FastAPI: ```python theme={null} # In main.py: from fastapi import FastAPI, WebSocket app = FastAPI() @app.websocket("/your-websocket-function-name") async def websocket_endpoint(websocket: WebSocket): await websocket.accept() await websocket.send_text("Hello, WebSocket!") await websocket.close() ``` ## Additional info Client-side Implementation: Handle the WebSocket connection properly on the client, including error handling and reconnection logic. ```javascript theme={null} // Example using JavaScript in a browser const socket = new WebSocket( "wss://api.cerebrium.ai/v4/p-xxxxxxxx//", ); socket.onopen = function (event) { console.log("WebSocket is open now."); }; socket.onmessage = function (event) { console.log("Received data: " + event.data); }; socket.onclose = function (event) { console.log("WebSocket is closed now."); }; socket.onerror = function (error) { console.error("WebSocket error observed:", error); }; ``` # Introduction to Cerebrium for real-time AI workloads Source: https://cerebrium.ai/docs/getting-started/introduction Cerebrium is a serverless GPU platform for real-time and high-performance AI apps with low cold starts, burst scaling, and global low-latency inference. Cerebrium is the infrastructure platform for **real-time and high-performance AI workloads**. It is a strong fit when an app requires one or more of the following: * **Low latency and low cold starts** * **Bursty traffic** that should scale without wasting GPU capacity * **Multi-region deployments** for global users or data residency * **Realtime voice, video, and streaming workloads** * **GPU-heavy production inference** that has to stay reliable under load ## Why teams choose Cerebrium * Launch code in the cloud in seconds * Run [CPUs](/docs/hardware/cpu-and-memory) or [GPUs](/docs/hardware/using-gpus) with automatic scaling * Serve [REST APIs](/docs/endpoints/inference-api), [streaming endpoints](/docs/endpoints/streaming), [WebSockets](/docs/endpoints/websockets), or any [ASGI-compatible app](/docs/container-images/custom-web-servers) * Deploy once and run across [multiple regions](/docs/deployments/multi-region-deployment) for lower latency and higher availability * Tune [concurrency and batching](/docs/scaling/batching-concurrency) for real production traffic * Improve startup performance with [cold-start optimization strategies](/docs/performance/faster-cold-starts) * Store model weights and files with [persistent storage](/docs/storage/managing-files) * Pay only for the compute you use - [billed by the second](https://www.cerebrium.ai/pricing) ## Start by workload Pick the closest path below to get started: * **OpenAI-compatible LLM endpoint** → [Serve an OpenAI Compatible LLM with vLLM](/docs/v4/examples/gpt-oss) * **Voice AI / real-time speech** → [Deploy a Twilio Voice Agent with Pipecat](/docs/v4/examples/twilio-voice-agent) * **Image and video generation** → [Generate images using SDXL](/docs/v4/examples/sdxl) * **Python apps** → [Deploy Gradio Chat Interface](/docs/v4/examples/asgi-gradio-interface) For the fastest first deployment, follow the quickstart below. ## Quickstart Set up and deploy an app on Cerebrium in a few steps. ### 1. Install the CLI ```bash theme={null} pip install cerebrium ``` ```bash theme={null} brew tap cerebriumai/tap brew install cerebrium ``` ```bash theme={null} # Ubuntu/Debian wget https://github.com/CerebriumAI/cerebrium/releases/latest/download/cerebrium_linux_amd64.deb sudo dpkg -i cerebrium_linux_amd64.deb # Or binary installation curl -L https://github.com/CerebriumAI/cerebrium/releases/latest/download/cerebrium_cli_linux_amd64.tar.gz | tar xz sudo mv cerebrium /usr/local/bin/ ``` ```powershell theme={null} # PowerShell (Run as Administrator) Invoke-WebRequest -Uri "https://github.com/CerebriumAI/cerebrium/releases/latest/download/cerebrium_cli_windows_amd64.zip" -OutFile "cerebrium.zip" Expand-Archive -Path "cerebrium.zip" -DestinationPath "." # Add cerebrium.exe to PATH ``` ### 2. Log in to the CLI ```bash theme={null} cerebrium login ``` This opens your browser so you can authenticate your CLI session. If your account belongs to multiple projects, switch the CLI's active project with: ```bash theme={null} cerebrium project set ``` Find the project ID on the [Cerebrium dashboard](https://dashboard.cerebrium.ai/). ### 3. Initialize a project ```bash theme={null} cerebrium init my-first-app cd my-first-app ``` This creates a basic project with `main.py` for app code and `cerebrium.toml` for configuration. ```python theme={null} def run(prompt: str): print(f"Running on Cerebrium: {prompt}") return {"my_result": prompt} ``` ### 4. Run code remotely Run the function in the cloud and pass it a prompt: ```bash theme={null} cerebrium run main.py::run --prompt "Hello World!" ``` The prompt appears in the logs. This is useful for quick code iteration, testing snippets, or one-off scripts that need cloud CPU/GPU resources. ### 5. Deploy your app ```bash theme={null} cerebrium deploy ``` This turns the function into a persistent [REST endpoint](/docs/endpoints/inference-api) that accepts JSON input and can scale automatically. Once deployed, the app is callable at a POST endpoint: ```text theme={null} https://api.cerebrium.ai/v4/p-xxxxxxxx/{app-name}/{function-name} ``` ### 6. What to do next Useful next steps after a first deployment: * [Define container images](/docs/container-images/defining-container-images) * [Tune scaling and concurrency](/docs/scaling/scaling-apps) * [Store model weights in persistent storage](/docs/storage/managing-files) * [Deploy to multiple regions](/docs/deployments/multi-region-deployment) Join the Community [Discord](https://discord.gg/ATj6USmeE2) for support and updates. ## How Cerebrium works Cerebrium uses containerization to ensure consistent environments and reliable scaling for apps. When you deploy code, Cerebrium packages it with all necessary dependencies into a container image. This image serves as a blueprint for creating instances that handle incoming requests. The system automatically manages scaling, creating new instances when traffic increases and removing them during quiet periods. For a detailed explanation of how Cerebrium builds and manages container images, see the [defining container images guide](/docs/container-images/defining-container-images). Content-Aware Storage forms the foundation of Cerebrium's speed. This system intelligently manages container images by understanding their content structure. When launching new instances, it pulls only the specific files. This targeted approach significantly reduces cold start times and optimizes resource usage. # CPU and Memory Source: https://cerebrium.ai/docs/hardware/cpu-and-memory Set vCPU cores and memory for Cerebrium apps in cerebrium.toml, review resource limits per hardware type, and optimize usage based billing and OOM risk. ## Overview CPU and memory resources are allocated per container and billed based on actual usage. Configure each app with specific CPU and memory requirements to optimize performance and cost. ## Resource Configuration ### CPU Configuration CPU resources are specified as vCPU units (float) in the `cerebrium.toml` file: ```toml theme={null} [cerebrium.hardware] cpu = 4 # Number of vCPU cores ``` Start with 4 CPU cores for most applications. Add cores based on monitoring, performance, and resource requirements. CPU usage is throttled when exceeding the specified limit. Fractional CPUs are also supported (e.g., `0.5`). ### Memory Configuration Memory is specified in gigabytes as a floating-point number: ```toml theme={null} [cerebrium.hardware] memory = 16.0 # Memory in GB ``` Allocate system memory equal to the GPU's VRAM capacity as a baseline. This accounts for initial model loading and compilation before GPU transfer. Applications terminate with an Out of Memory (OOM) error if they exceed the specified memory limit. Memory and CPU are billed based on usage, which reduces your costs and doesn’t require the overprovisioning of an entire instance. ## Resource Limits Resource limits depend on the selected hardware configuration: | Hardware Type | Max CPU Cores | Max Memory (GB) | | - | - | - | | CPU Only | 48 | 96 | | ADA\_L40 | 16 | 128 | | AMPERE\_A100\_40GB / AMPERE\_A100\_80GB | 12 | 140 | | AMPERE\_A10 | 48 | 192 | | ADA\_L4 | 48 | 192 | | TURING\_T4 | 48 | 192 | | BLACKWELL\_RTX6000 | 24 | 218 | | HOPPER\_H100 | 24 | 256 | | HOPPER\_H200 | 24 | 256 | | BLACKWELL\_B200 | 24 | 256 | | TRN1 | 128 | 512 | ## Memory Optimization The Transformers library provides memory optimization through the `low_cpu_mem_usage` flag, which reduces memory footprint at the cost of longer initialization times. Implement lazy loading for large datasets to further reduce memory usage. Monitor memory patterns through platform metrics to identify optimization opportunities. Use memory-efficient model loading techniques for large-scale deployments. ## Resource Monitoring The platform monitors CPU utilization and throttling events to identify performance bottlenecks. The platform tracks memory usage and OOM events to prevent application failures. # Using CUDA Source: https://cerebrium.ai/docs/hardware/using-cuda Enable CUDA on Cerebrium using GPU ready Python packages or NVIDIA base images, and choose runtime over devel images to keep cold starts fast. ## Overview CUDA (Compute Unified Device Architecture) enables apps to use graphics cards (GPUs) to speed up calculations. Unlike standard processors (CPUs) that handle one task at a time, graphics cards can handle many tasks simultaneously. ## How CUDA Works CUDA connects apps to graphics cards, splitting large tasks into smaller pieces that can be processed at the same time. This makes operations like image processing and complex mathematical calculations much faster than using a standard processor alone. ## Using CUDA with App Dependencies Many Python packages include built-in CUDA support. PyTorch, for example, includes CUDA in its installation: ```toml theme={null} [cerebrium.dependencies.pip] torch = "latest" # PyTorch with graphics card support ``` ## Special Requirements Some apps need direct access to CUDA system libraries and tools. The CUDA base image provides this complete CUDA toolkit environment. This is often necessary when apps: * Compile custom CUDA code * Access low-level CUDA features * Need specific CUDA driver versions * Require CUDA development tools Set the base image in the `cerebrium.toml` file: ```toml theme={null} [cerebrium.deployment] docker_base_image_url = "nvidia/cuda:12.1.1-runtime-ubuntu22.04" ``` ## Cold-Start Optimization Image size and complexity directly impact cold-start performance: the time needed to initialize an app from an inactive state. Cerebrium uses a content-addressable file system that selectively pulls only required files, but larger images still affect startup times. ### Image Size Considerations Base image choices affect cold-start times in several ways: * Development images (like `nvidia/cuda:*-devel`) include additional tools, increasing size * Runtime images provide minimal dependencies for faster initialization * Each additional layer or installed package increases the final image size The final image size appears in the dashboard after build completion, helping track size optimization efforts. ### Balancing Tradeoffs Cold-start optimization requires balancing competing needs: ```toml theme={null} # Minimal runtime image - faster cold-starts docker_base_image_url = "debian:bookworm-slim" # Full development image - slower cold-starts, more tools docker_base_image_url = "nvidia/cuda:12.0.1-devel-ubuntu22.04" ``` Keeping instances warm avoids cold-starts entirely, at the cost of higher resource usage. # Using GPUs Source: https://cerebrium.ai/docs/hardware/using-gpus Choose from NVIDIA GPUs like H100, A100, and L40s on Cerebrium, set GPU type and count in cerebrium.toml, and check plan availability for each type. GPUs accelerate computational workloads through parallel processing. Originally designed for graphics rendering, modern GPUs are essential for AI models, large-scale data processing, and other compute-intensive applications. Cerebrium provides GPU access through configuration in the `cerebrium.toml` file, without requiring infrastructure management. ## Specifying GPUs Configure GPUs in the `[cerebrium.hardware]` section of `cerebrium.toml`, specifying the type (`compute` parameter) and quantity (`gpu_count`). Additional deployment and scaling considerations are covered in the sections below. ## Available GPUs The platform offers GPUs ranging from cost-effective development options to high-end enterprise hardware, alongside CPU-only compute and AWS accelerators. | Compute | Identifier | VRAM (GB) | Max GPUs | Plan required | | - | - | - | - | - | | NVIDIA B200 | BLACKWELL\_B200 | 180 | 8 | Standard+ | | NVIDIA H200 | HOPPER\_H200 | 141 | 8 | Standard+ | | NVIDIA H100 | HOPPER\_H100 | 80 | 8 | Standard+ | | RTX PRO 6000 | BLACKWELL\_RTX6000 | 96 | 8 | Standard+ | | NVIDIA A100 | AMPERE\_A100\_80GB | 80 | 8 | Standard+ | | NVIDIA A100 | AMPERE\_A100\_40GB | 40 | 8 | Standard+ | | NVIDIA L40s | ADA\_L40 | 48 | 8 | Hobby+ | | NVIDIA L4 | ADA\_L4 | 24 | 8 | Hobby+ | | NVIDIA A10 | AMPERE\_A10 | 24 | 8 | Hobby+ | | NVIDIA T4 | TURING\_T4 | 16 | 8 | Hobby+ | | AWS Inferentia2 | INF2 | 32 | 8 | Hobby+ | | AWS Trainium | TRN1 | 32 | 8 | Enterprise | | CPU only | CPU | - | - | Hobby+ | The identifier is used in the `cerebrium.toml` file. It consists of the GPU model generation and model name to avoid ambiguity. ### Plan Availability Compute types are gated by plan: * **Hobby** (6 types): CPU, TURING\_T4, AMPERE\_A10, ADA\_L4, ADA\_L40, INF2 * **Standard** (12 types): everything in Hobby, plus AMPERE\_A100\_40GB, AMPERE\_A100\_80GB, HOPPER\_H100, HOPPER\_H200, BLACKWELL\_B200, BLACKWELL\_RTX6000 * **Enterprise** (all 13 types): everything in Standard, plus TRN1 Deploying with a compute type outside the project's plan is rejected at deploy time. [Upgrade the plan](https://dashboard.cerebrium.ai) or contact [sales@cerebrium.ai](mailto:sales@cerebrium.ai) for per-project exceptions. GPU selection is also possible using the `--compute` and `--gpu-count` flags during application initialization. ## Multi-GPU Configuration Configure multiple GPUs in the `cerebrium.toml` file: ```toml theme={null} [cerebrium.hardware] compute = "AMPERE_A100_80GB" gpu_count = 4 # Number of GPUs needed cpu = 8 memory = 128.0 ``` ## GPU Preference Lists The `compute` parameter also accepts a list of acceptable GPU types in preference order. The platform allocates the most preferred type with available capacity and falls back to the next entry when needed: ```toml theme={null} [cerebrium.hardware] compute = ["HOPPER_H100", "HOPPER_H200", "AMPERE_A100_80GB"] gpu_count = 1 cpu = 8 memory = 128.0 ``` * Up to 5 entries, ordered from most to least preferred * All entries must belong to the same hardware family. NVIDIA GPU types cannot be mixed with CPU or AWS accelerators (INF2, TRN1) * `cerebrium run` uses only the first entry Accepting more GPU types widens the pool of capacity an app can run on and reduces the likelihood of request queuing. ## Availability GPU availability varies by region and provider. Narrowing the provider and region constraints increases the likelihood of request queuing. See [GPU availability by region](/docs/deployments/multi-region-deployment#gpu-availability-by-region). For guaranteed burst capacity, contact the [enterprise plan](mailto:sales@cerebrium.ai) team. # Exporting Metrics to Monitoring Platforms Source: https://cerebrium.ai/docs/integrations/metrics-export Export your application metrics to any OTLP-compatible observability platform including Grafana Cloud, Datadog, Prometheus, New Relic, and more Export real-time resource and execution metrics from Cerebrium applications to an existing observability platform. Monitor CPU, memory, GPU usage, request counts, and latency. Most major OTLP-compatible monitoring platforms are supported. ## What metrics are exported? ### Resource Metrics | Metric | Type | Unit | Description | | - | - | - | - | | cerebrium\_cpu\_utilization\_cores | Gauge | cores | CPU cores actively in use per app | | cerebrium\_memory\_usage\_bytes | Gauge | bytes | Memory actively in use per app | | cerebrium\_gpu\_memory\_usage\_bytes | Gauge | bytes | GPU VRAM in use per app | | cerebrium\_gpu\_compute\_utilization\_percent | Gauge | percent | GPU compute utilization (0-100) per app | | cerebrium\_containers\_running\_count | Gauge | count | Number of running containers per app | | cerebrium\_containers\_ready\_count | Gauge | count | Number of ready containers per app | ### Execution Metrics | Metric | Type | Unit | Description | | - | - | - | - | | cerebrium\_run\_execution\_time\_ms | Histogram | ms | Time spent executing user code | | cerebrium\_run\_queue\_time\_ms | Histogram | ms | Time spent waiting in queue | | cerebrium\_run\_coldstart\_time\_ms | Histogram | ms | Time for container cold start | | cerebrium\_run\_response\_time\_ms | Histogram | ms | Total end-to-end response time | | cerebrium\_run\_total | Counter | — | Total run count | | cerebrium\_run\_successes\_total | Counter | — | Successful run count | | cerebrium\_run\_errors\_total | Counter | — | Failed run count | **Prometheus metric name mapping:** When metrics are ingested by Prometheus (including Grafana Cloud), OTLP automatically appends unit suffixes to metric names. Histogram metrics will appear with `_milliseconds` appended — for example, `cerebrium_run_execution_time_ms` becomes `cerebrium_run_execution_time_ms_milliseconds_bucket`, `_count`, and `_sum`. Counter metrics with the `_total` suffix remain unchanged. The example queries throughout this guide use the Prometheus-ingested names. ### Labels Every metric includes the following labels for filtering and grouping: | Label | Description | Example | | - | - | - | | `project_id` | Your Cerebrium project ID | `p-abc12345` | | `app_id` | Full application identifier | `p-abc12345-my-model` | | `app_name` | Human-readable app name | `my-model` | | `region` | Deployment region | `us-east-1` | ## How it works Cerebrium automatically pushes metrics to the configured monitoring platform every **60 seconds** using the [OpenTelemetry Protocol (OTLP)](https://opentelemetry.io/docs/specs/otlp/). Provide an OTLP endpoint and authentication credentials through the Cerebrium dashboard — Cerebrium handles collecting resource usage and execution data, formatting it as OpenTelemetry metrics, and delivering it to the destination. * Metrics are pushed every **60 seconds** * Failed pushes are retried **3 times** with exponential backoff * If pushes fail **10 consecutive times**, export is automatically paused to avoid noise (re-enable at any time from the dashboard) * Credentials are stored encrypted and never returned in API responses ### Supported destinations * **Grafana Cloud** — Primary supported destination * **Datadog** — Via OTLP endpoint * **Prometheus** — Self-hosted with OTLP receiver enabled * **Custom** — Any OTLP-compatible endpoint (New Relic, Honeycomb, etc.) ## Setup Guide ### Step 1: Get your platform credentials Gather an OTLP endpoint and authentication credentials from the monitoring platform before configuring the Cerebrium dashboard. 1. Sign in to [Grafana Cloud](https://grafana.com) 2. Go to your stack → **Connections** → **Add new connection** 3. Search for **"OpenTelemetry"** and click **Configure** 4. Copy the **OTLP endpoint** — this will match your stack's region: * US: `https://otlp-gateway-prod-us-east-0.grafana.net/otlp` * EU: `https://otlp-gateway-prod-eu-west-0.grafana.net/otlp` * Other regions will show their specific URL on the configuration page 5. On the same page, generate an API token. Click **Generate now** and ensure the token has the **MetricsPublisher** role — this is a separate token from any Prometheus Remote Write tokens you may already have. 6. The page will show you an **Instance ID** and the generated token. Run the following in your terminal to create the Basic auth string: ```bash theme={null} echo -n "INSTANCE_ID:TOKEN" | base64 ``` Copy the output — you'll paste it in the dashboard in the next step. The API token **must** have the **MetricsPublisher** role. The default Prometheus Remote Write token will not work with the OTLP endpoint. If you're unsure, generate a new token from the OpenTelemetry configuration page — it will have the correct role by default. 1. Sign in to [Datadog](https://app.datadoghq.com) 2. Go to **Organization Settings** → **API Keys** 3. Create or copy an existing API key 4. Your OTLP endpoint depends on your [Datadog site](https://docs.datadoghq.com/getting_started/site/): | Datadog Site | OTLP Endpoint | | - | - | | US1 (datadoghq.com) | `https://api.datadoghq.com/api/v2/otlp` | | US3 (us3.datadoghq.com) | `https://api.us3.datadoghq.com/api/v2/otlp` | | US5 (us5.datadoghq.com) | `https://api.us5.datadoghq.com/api/v2/otlp` | | EU (datadoghq.eu) | `https://api.datadoghq.eu/api/v2/otlp` | | AP1 (ap1.datadoghq.com) | `https://api.ap1.datadoghq.com/api/v2/otlp` | Find the site in the Datadog URL — for example, if you log in at `app.us3.datadoghq.com`, your site is US3. Keep your API key and endpoint handy for the next step. 1. Enable the OTLP receiver in your Prometheus config: * Add `--enable-feature=otlp-write-receiver` flag * Or use an OpenTelemetry Collector as a sidecar 2. Your endpoint will be `http://YOUR_PROMETHEUS_HOST:4318` (this is the OTLP HTTP port — not `4317`, which is gRPC) — copy this for the next step Any platform that supports [OpenTelemetry OTLP over HTTP](https://opentelemetry.io/docs/specs/otlp/) will work, including New Relic, Honeycomb, Lightstep, and others. 1. Get the OTLP HTTP endpoint from your provider's documentation 2. Get the required authentication headers **Common examples:** | Platform | Auth Header Name | Auth Header Value | | - | - | - | | New Relic | `api-key` | Your New Relic license key | | Honeycomb | `x-honeycomb-team` | Your Honeycomb API key | | Lightstep | `lightstep-access-token` | Your Lightstep token | ### Step 2: Configure in the Cerebrium dashboard 1. In the [Cerebrium dashboard](https://dashboard.cerebrium.ai), go to your project → **Integrations** → **Metrics Export** 2. Paste your **OTLP endpoint** from Step 1 3. Add the **authentication headers** from Step 1: * **Header name:** `Authorization` - **Header value:** `Basic YOUR_BASE64_STRING` (the output from the terminal command in Step 1) * **Header name:** `DD-API-KEY` - **Header value:** Your Datadog API key * **Header name:** `Authorization` (if auth is enabled on your Prometheus, otherwise leave empty) - **Header value:** `Bearer your-token` (if auth is enabled) Add the authentication headers required by the platform. Add multiple headers using the **Add Header** button. 4. Click **Save & Enable** Metrics start flowing within 60 seconds. The dashboard shows a green "Connected" status with the time of the last successful export. If something looks wrong, click **Test Connection** to verify Cerebrium can reach the monitoring platform. The result includes details to help troubleshoot. ## Viewing Metrics Once connected, metrics appear in the monitoring platform within a minute or two (exact latency depends on the platform's ingestion pipeline). 1. Go to your Grafana Cloud dashboard → **Explore** 2. Select your Prometheus data source — it will be named something like **grafanacloud-yourstack-prom** (find it under **Connections** → **Data sources** if you're unsure) 3. Search for metrics starting with `cerebrium_` **Example queries:** Histogram metrics in Prometheus have `_milliseconds` appended by OTLP's unit suffix convention, so you'll see names like `cerebrium_run_execution_time_ms_milliseconds_bucket`. This is expected behavior — see the [metric name mapping note](#execution-metrics) above. ```promql theme={null} # CPU usage by app cerebrium_cpu_utilization_cores{project_id="YOUR_PROJECT_ID"} # Memory for a specific app cerebrium_memory_usage_bytes{app_name="my-model"} # Container scaling over time cerebrium_containers_running_count{project_id="YOUR_PROJECT_ID"} # Request rate (requests per second over 5 minutes) rate(cerebrium_run_total[5m]) # p99 execution latency histogram_quantile(0.99, rate(cerebrium_run_execution_time_ms_milliseconds_bucket{app_name="my-model"}[5m])) # p99 end-to-end response time histogram_quantile(0.99, rate(cerebrium_run_response_time_ms_milliseconds_bucket{app_name="my-model"}[5m])) # Error rate as a percentage rate(cerebrium_run_errors_total{app_name="my-model"}[5m]) / rate(cerebrium_run_total{app_name="my-model"}[5m]) * 100 # Average cold start time rate(cerebrium_run_coldstart_time_ms_milliseconds_sum{app_name="my-model"}[5m]) / rate(cerebrium_run_coldstart_time_ms_milliseconds_count{app_name="my-model"}[5m]) ``` 1. Go to **Metrics** → **Explorer** in your Datadog dashboard 2. Search for metrics starting with `cerebrium` 3. Filter by `project_id`, `app_name`, and other labels using the "from" field Query your Prometheus instance directly. All Cerebrium metrics are prefixed with `cerebrium_`: ```promql theme={null} # List all Cerebrium metrics {__name__=~"cerebrium_.*"} # CPU usage across all apps cerebrium_cpu_utilization_cores ``` ## Managing Metrics Export Manage metrics export configuration from the dashboard at any time under **Integrations** → **Metrics Export**. * **Disable export:** Toggle the switch off. The configuration is preserved — re-enable at any time without reconfiguring. * **Update credentials:** Enter new authentication headers and click **Save Changes**. Use this when rotating API keys. * **Change endpoint:** Update the OTLP endpoint field and click **Save Changes**. * **Check status:** The dashboard shows whether export is connected, the time of the last successful export, and any error messages. ## Troubleshooting ### Metrics not appearing 1. **Check the dashboard status.** Go to **Integrations** → **Metrics Export** and look for the connection status. If it shows "Paused," export was automatically disabled after repeated failures — click **Re-enable** after fixing the issue. 2. **Run a connection test.** Click **Test Connection** on the dashboard. Common errors: * **401 / 403 Unauthorized:** Your auth headers are wrong. For Grafana Cloud, make sure you're using a MetricsPublisher token (not a Prometheus Remote Write token). For Datadog, verify your API key is active. * **404 Not Found:** The OTLP endpoint URL is incorrect. Double-check the URL matches your platform and region. * **Connection timeout:** Your endpoint may be unreachable. For self-hosted Prometheus, confirm the host is publicly accessible and port `4318` is open. 3. **Check your platform's data source.** In Grafana Cloud, make sure you're querying the correct Prometheus data source (not a Loki or Tempo source). In Datadog, check that your site region matches the endpoint you configured. ### Metrics appear but values look wrong * **Histogram metrics have `_milliseconds` in the name.** This is normal — Prometheus appends unit suffixes from OTLP metadata. Use the full name (e.g., `cerebrium_run_execution_time_ms_milliseconds_bucket`) in your queries. * **Container counts fluctuate during deploys.** This is expected — you may see temporary spikes in `cerebrium_containers_running_count` during rolling deployments as new containers start and old ones drain. * **Gaps in metrics.** Short gaps (1-2 minutes) can occur during deployments or scaling events. If you see persistent gaps, check whether export was paused. ### Still stuck? Contact [support@cerebrium.ai](mailto:support@cerebrium.ai) with the project ID and error message from the dashboard for further investigation. # Migrating from Hugging Face Source: https://cerebrium.ai/docs/migrations/hugging-face Migrate from Hugging Face inference endpoints to Cerebrium with a step-by-step guide deploying a Llama 3.1 8B model on serverless GPUs. ## Introduction This guide covers migrating from Hugging Face inference endpoints to Cerebrium's serverless infrastructure platform, including key differences, migration benefits, and step-by-step instructions for deploying a Llama 3.1 8B model on Cerebrium. ## Comparing Hugging Face and Cerebrium Key features and performance metrics compared: | **Feature** | **Hugging Face** | **Cerebrium** | | - | - | - | | **Pricing** | \$0.000278 per second | \$0.0004676 per second | | **Minimum cooldown period** | 15m | 1s | | **First build time** | 9m25s | 49s | | **Subsequent build times** | 1m50s - 2m15s | 58s - 1m5s | | **Response time (From cold)** | 1m45s - 1m48s | 8s - 17s | | **Response time (From warm)** | 6s | 2s | | **Co-locating your models** | Requires a separate repository for each inference endpoint and mode | Co-locate multiple models from various sources in a single app | | **Response handling (From cold)** | Throws an error | Waits for infrastructure to become available and returns a response | ## Benefits of migrating to Cerebrium 1. **Faster build times**: Cerebrium significantly reduces build times by up to 95%, especially for subsequent builds (an additional 56% reduction). This can greatly improve iteration speed and the cost of running experiments with complex ML apps. 2. **Flexible cooldown period**: With a minimum cooldown period of just 1 second (compared to Hugging Face's 15 minutes), Cerebrium allows for more efficient resource utilization and cost management. 3. **Improved cold start handling**: When encountering a cold start, Cerebrium waits for the infrastructure to become available instead of throwing an error. This results in a better user experience and fewer failed requests. 4. **Model colocation flexibility**: Cerebrium doesn't require a separate repository for each inference endpoint, simplifying the management of models. Each function in your app becomes an endpoint automatically, which means that you can run multiple models from the same app to save costs. 5. **Pay-per-use model**: Cerebrium's pricing model ensures you pay only for the compute resources you actually use. This can lead to cost savings, especially for sporadic or low-volume inference needs. 6. **Competitive performance**: Cerebrium only adds up to 50ms of latency to your inference requests. This results in competitive response times from a warm start. Caching mechanisms and highly optimized orchestration pipelines help apps start from a cold state in an average of 2-5 seconds. 7. **Customizable infrastructure**: Cerebrium allows for fine-grained control over the infrastructure specifications, enabling you to optimize for your specific use case. ## Migration process The following walks through migrating a Llama 3.1 8B model from Hugging Face to Cerebrium, from configuration setup to deployment. ### 1. Cerebrium setup and configuration Set up the required files and configure the environment. #### 1.1 Install Cerebrium CLI First, install the Cerebrium CLI: ```bash theme={null} pip install cerebrium --upgrade ``` #### 1.2 Update your requirements file Scaffold your application by running `cerebrium init [PROJECT_NAME]`. During the initialization, a `cerebrium.toml` is created. This file configures the deployment, hardware, scaling, and dependencies for your Cerebrium project. Update your `cerebrium.toml` file to reflect the following: ```toml theme={null} [cerebrium.deployment] name = "llama-8b-vllm" python_version = "3.11" docker_base_image_url = "debian:bookworm-slim" include = ["./*", "main.py", "cerebrium.toml"] exclude = [".*"] [cerebrium.hardware] cpu = 2 memory = 12.0 compute = "AMPERE_A10" [cerebrium.scaling] min_replicas = 0 max_replicas = 5 cooldown = 30 [cerebrium.dependencies.pip] sentencepiece = "latest" torch = "latest" transformers = "latest" accelerate = "latest" xformers = "latest" pydantic = "latest" bitsandbytes = "latest" ``` Configuration breakdown: * `cerebrium.deployment`: Specifies the project name, Python version, base Docker image, and which files to include/exclude as project files. * `cerebrium.hardware`: Defines the CPU, memory, and GPU requirements for your deployment. * `cerebrium.scaling`: Configures auto-scaling behavior, including minimum and maximum replicas, and cooldown period. * `cerebrium.dependencies.pip`: Lists the Python packages required for your project. #### 1.3 Update your code Next, update `main.py` with the model loading and inference logic. ```python theme={null} import torch import os from huggingface_hub import login from pydantic import BaseModel from transformers import AutoModelForCausalLM, AutoTokenizer, GenerationConfig # Log into Hugging Face Hub login(token=os.environ.get("HF_AUTH_TOKEN")) model_path = "meta-llama/Meta-Llama-3.1-8B-Instruct" cache_directory = "/persistent-storage" # Set up tokenizer and model tokenizer = AutoTokenizer.from_pretrained(model_path, cache_dir=cache_directory) tokenizer.pad_token_id = 0 model = AutoModelForCausalLM.from_pretrained( model_path, load_in_8bit=True, torch_dtype=torch.float16, device_map="auto", cache_dir=cache_directory, ) class Item(BaseModel): prompt: str temperature: float top_p: float top_k: int max_tokens: int frequency_penalty: float def run( prompt, temperature=0.6, top_p=0.9, top_k=0, max_tokens=512, frequency_penalty=1 ): item = Item( prompt=prompt, temperature=temperature, top_p=top_p, top_k=top_k, max_tokens=max_tokens, frequency_penalty=frequency_penalty, ) # Place prompt in template inputs = tokenizer( item.prompt, return_tensors="pt", max_length=512, truncation=True, padding=True ) input_ids = inputs["input_ids"].to("cuda") # Set up generation config generation_config = GenerationConfig( temperature=temperature, top_p=top_p, top_k=top_k, max_tokens=max_tokens, ) with torch.no_grad(): outputs = model.generate( input_ids=input_ids, generation_config=generation_config, return_dict_in_generate=True, output_scores=True, ) result = tokenizer.decode(outputs.sequences[0], skip_special_tokens=True) return {"result": result} ``` This script: 1. Authenticates with Hugging Face using a secret token. Add this secret in the Cerebrium dashboard. 2. Initializes the Llama 3.1 8B model using vLLM for efficient inference. 3. Defines an `Item` class to structure and validate (using Pydantic) the input parameters. 4. Implements a `run` function that generates text based on the provided prompt and parameters. ### 2. Deployment To deploy your app to Cerebrium, use the following CLI command in your project directory: ```bash theme={null} cerebrium deploy ``` This command uses the configuration in `cerebrium.toml` to set up and deploy your model. ### 3. Using the deployed model Once deployed, you can use your model as follows: ```python theme={null} import requests import json url = "https://api.cerebrium.ai/v4/p-xxxxxxxx/llama-8b-vllm/run" payload = json.dumps({"prompt": "tell me about yourself"}) headers = { 'Authorization': 'Bearer [CEREBRIUM_API_KEY]', 'Content-Type': 'application/json' } response = requests.request("POST", url, headers=headers, data=payload) print(response.text) ``` Make sure to replace `[CEREBRIUM_API_KEY]` with your Inference API key, which can be found in your dashboard under API keys. This code sends a POST request to your deployed model's endpoint with a prompt, and prints the model's response. ## Additional considerations When migrating, keep the following points in mind: 1. **API structure**: The Cerebrium implementation uses a different API structure compared to Hugging Face 2. **Authentication**: Ensure you have set up the `HF_AUTH_TOKEN` secret in Cerebrium for authenticating with Hugging Face 3. **Model permissions**: The example uses the Llama 3.1 8B Instruct model. Ensure you have the necessary permissions to use this model 4. **Hardware optimization**: The `cerebrium.toml` file specifies the hardware requirements. Adjust these based on your specific model and performance needs 5. **Dependency management**: Regularly review and update the dependencies listed in `cerebrium.toml` to ensure you're using the latest compatible versions 6. **Scaling configuration**: The example sets up auto-scaling with 0 to 5 replicas and a 30-second cooldown. Monitor your usage patterns and adjust these parameters as needed 7. **Cold starts**: While Cerebrium handles cold starts more gracefully than Hugging Face, be aware that the first request after a period of inactivity may still take longer to process 8. **Monitoring and logging**: Familiarize yourself with Cerebrium's monitoring and logging capabilities to track your model's performance and usage 9. **Cost management**: Although Cerebrium's pay-per-use model can be more cost-effective, set up proper monitoring and alerts to avoid unexpected costs 10. **Testing**: Thoroughly test your migrated models to ensure they perform as expected on the new platform ## Conclusion Migrating from Hugging Face inference endpoints to Cerebrium offers faster build times, more flexible resource management, and lower costs. The migration requires some setup and code changes, but the resulting deployment provides improved performance and scalability. Continuously monitor and optimize the deployment in production. Reach out to support or join the Slack and Discord communities for questions during the migration process. Read more about Cerebrium's functionality: * [Secrets](/docs/other-topics/using-secrets) * [Model scaling](/docs/scaling/scaling-apps) * [Keeping models warm](/docs/performance/faster-cold-starts) # Migrating from Mystic Source: https://cerebrium.ai/docs/migrations/mystic Migrate your apps from Mystic to Cerebrium with this step by step guide covering config conversion, code changes, deployment, and inference. ## Introduction Mystic AI is sunsetting their services. They were an early pioneer that pushed the industry forward. This guide covers migrating apps from Mystic to Cerebrium to keep them functional. It covers converting existing Mystic code (using a stable diffusion example) and configuration to the Cerebrium platform, including deployment optimization for performance and cost efficiency. ## Key Differences Cerebrium helps teams deploy and run models efficiently. The infrastructure is designed for reliable performance: * The average model cold-starts in 2-5 seconds. * Updates to your code deploy quickly, taking only 8-14 seconds. * 99.9% uptime. Cerebrium provides precise control over computing resources. Instead of managing entire instances, select the exact CPU, memory, and GPU power needed. Billing is per-second for actual resource usage. Use the [pricing calculator](https://cerebrium.ai/pricing) for cost estimates. ## Migration Process ### 1. Project Setup and Configuration Install Cerebrium's command-line tool and create the project: ```bash theme={null} pip install cerebrium --upgrade cerebrium login # You'll be redirected to the dashboard for login cerebrium init stable-diffusion cd stable-diffusion ``` Convert the existing Mystic configuration to Cerebrium's format. A typical Mystic configuration: ```yaml theme={null} # Mystic's pipeline.yaml runtime: container_commands: - apt-get update - apt-get install -y git python: version: "3.10" requirements: - pipeline-ai - diffusers==0.24.0 - torch==2.1.1 - transformers==4.35.2 - accelerate==0.25.0 cuda_version: "11.4" accelerators: - "nvidia_a10" accelerator_memory: null pipeline_graph: sd_pipeline:pipeline_graph pipeline_name: /stable-diffusion-v1.5 extras: {} ``` Becomes this Cerebrium TOML config: ```toml theme={null} # cerebrium.toml [cerebrium.deployment] name = "stable-diffusion" python_version = "3.11" docker_base_image_url = "debian:bookworm-slim" include = ["./*", "main.py", "cerebrium.toml"] exclude = [".*"] [cerebrium.hardware] compute = "AMPERE_A10" # Choose your GPU type cpu = 4 # Number of CPU cores memory = 16.0 # Memory in GB gpu_count = 1 # Number of GPUs [cerebrium.scaling] min_replicas = 0 # Save costs when inactive and scale down your app max_replicas = 2 # Handle increased traffic and scale up where necessary cooldown = 60 # Time window at reduced concurrency before scaling down replica_concurrency = 1 # The number of requests a single container can support [cerebrium.dependencies.pip] torch = ">=2.0.0" pydantic = "latest" transformers = "latest" accelerate = "latest" diffusers = "latest" safetensors = "latest" xformers = "latest" ``` ### 2. Code Migration Convert the model implementation. A typical Mystic pipeline: ```python theme={null} import typing as t from pathlib import Path from PIL.Image import Image from pipeline.cloud.pipelines import run_pipeline from pipeline.objects.graph import InputField, InputSchema from pipeline import File, Pipeline, Variable, entity, pipe HF_MODEL_ID = "runwayml/stable-diffusion-v1-5" class ModelKwargs(InputSchema): num_images_per_prompt: int | None = InputField( title="num_images_per_prompt", description="The number of images to generate per prompt.", default=1, optional=True, ) height: int | None = InputField( title="height", description="The height in pixels of the generated image.", default=512, optional=True, multiple_of=64, ge=64, ) width: int | None = InputField( title="width", description="The width in pixels of the generated image.", default=512, optional=True, multiple_of=64, ge=64, ) num_inference_steps: int | None = InputField( title="num_inference_steps", description=( "The number of denoising steps. More denoising steps " "usually lead to a higher quality image at the expense " "of slower inference." ), default=50, optional=True, ) @entity class StableDiffusionModel: def __init__(self) -> None: self.model = None self.device = None @pipe(run_once=True, on_startup=True) def load(self) -> None: """ Load the HF model into memory""" import torch from diffusers import StableDiffusionPipeline device = torch.device("cuda") if torch.cuda.is_available() else "cpu" self.model = StableDiffusionPipeline.from_pretrained(HF_MODEL_ID) self.model.to(device) @pipe def predict(self, prompt: str, model_kwargs: ModelKwargs) -> t.List[Image]: """ Generates a list of PIL images. """ return self.model(prompt=prompt, **model_kwargs.to_dict()).images @pipe def postprocess(self, images: t.List[Image]) -> t.List[File]: """ Creates a list of Files from the `PIL` images. """ output_images = [] for i, image in enumerate(images): path = Path(f"/tmp/sd/image-{i}.jpg") path.parent.mkdir(parents=True, exist_ok=True) image.save(str(path)) output_images.append(File(path=path, allow_out_of_context_creation=True)) return output_images with Pipeline() as builder: prompt = Variable( str, title="prompt", description="The prompt to guide image generation", max_length=512, ) model_kwargs = Variable(ModelKwargs) model = StableDiffusionModel() model.load() images: t.List[Image] = model.predict(prompt, model_kwargs) output: t.List[File] = model.postprocess(images) builder.output(output) pipeline_graph = builder.get_pipeline() ``` The Cerebrium equivalent in `main.py`: ```python theme={null} import base64 import io import torch from diffusers import StableDiffusionPipeline, DPMSolverMultistepScheduler from pydantic import BaseModel # Define the structure of input parameters class Item(BaseModel): prompt: str height: int width: int num_inference_steps: int num_images_per_prompt: int # Load the model and set it up for inference model_id = "stabilityai/stable-diffusion-2-1" pipe = StableDiffusionPipeline.from_pretrained(model_id, torch_dtype=torch.float16) pipe.scheduler = DPMSolverMultistepScheduler.from_config(pipe.scheduler.config) pipe.enable_xformers_memory_efficient_attention() pipe = pipe.to("cuda") # The endpoint we'll call to make inference def predict( prompt: str, height: int = 512, width: int = 512, num_inference_steps: int = 25, num_images_per_prompt: int = 1, ): item = Item( prompt=prompt, height=height, width=width, num_inference_steps=num_inference_steps, num_images_per_prompt=num_images_per_prompt, ) images = pipe( prompt=item.prompt, height=item.height, width=item.width, num_images_per_prompt=item.num_images_per_prompt, num_inference_steps=item.num_inference_steps, ).images finished_images = [] for image in images: buffered = io.BytesIO() image.save(buffered, format="PNG") finished_images.append(base64.b64encode(buffered.getvalue()).decode("utf-8")) return finished_images ``` ### 3. Deployment Deploy your model with a single command: ```bash theme={null} cerebrium deploy ``` ### 4. Inference Once your app is deployed, you can make requests to your model using the example cURL request below: ```bash theme={null} curl --location 'https://api.cerebrium.ai/v4/p-xxxxxxxx/stable-diffusion/predict' \ --header 'Content-Type: application/json' \ --header 'Authorization: Bearer ' \ --data '{ "prompt": "a photo of an astronaut riding a horse on mars" }' ``` The Cerebrium platform provides the tools and support needed for a smooth transition. ## Join the Community Connect with other developers and the Cerebrium team for faster response and issue resolution: * Join the Discord server. * Join the Slack workspace. These communities offer migration support, quick technical answers, best practices, and feature updates. # Migrating from Replicate Source: https://cerebrium.ai/docs/migrations/replicate Migrate workloads from Replicate to Cerebrium in under 5 minutes using SDXL-Lightning-4step as a worked example for faster, cheaper inference. ### Introduction This tutorial covers migrating workloads from Replicate to Cerebrium in less than 5 minutes. This example migrates the SDXL-Lightning-4step model from ByteDance. Find it on Replicate [here](https://replicate.com/bytedance/sdxl-lightning-4step). Follow along with the code in the [GitHub repo](https://github.com/lucataco/cog-sdxl-lightning-4step). Start by creating the Cerebrium project. ```python theme={null} cerebrium init cog-migration-sdxl ``` Cerebrium and Replicate both use a setup file: **cog.yaml** and **cerebrium.toml** for Replicate and Cerebrium respectively. Based on the cog.yaml, add/change the following in `cerebrium.toml`: ```python theme={null} [cerebrium.deployment] name = "cog-migration-sdxl" python_version = "3.11" include = ["./*", "main.py", "cerebrium.toml"] exclude = ["./example_exclude"] docker_base_image_url = "nvidia/cuda:12.1.1-cudnn8-runtime-ubuntu22.04" shell_commands = [ "curl -o /usr/local/bin/pget -L 'https://github.com/replicate/pget/releases/download/v0.6.2/pget_linux_x86_64' && chmod +x /usr/local/bin/pget" ] [cerebrium.hardware] compute = "AMPERE_A10" cpu = 2 memory = 12.0 gpu_count = 1 [cerebrium.dependencies.pip] "accelerate" = "latest" "diffusers" = "latest" "torch" = "==2.0.1" "torchvision" = "==0.15.2" "transformers" = "latest" [cerebrium.dependencies.apt] "curl" = "latest" ``` The configuration above: * Uses an Nvidia base image with CUDA libraries (Cuda 12). You can see other images [here](/docs/container-images/defining-container-images). * Sets hardware based on CPU/GPU requirements. You can see the available options in the [GPU guide](/docs/hardware/using-gpus) and [CPU and memory guide](/docs/hardware/cpu-and-memory). * Copies the required pip packages * Downloads pget (used by Replicate for model weights) via curl and shell commands in cerebrium.toml The hardware and environment setup now matches. The cog.yaml indicates the endpoint file — in this case, `predict.py`. Cerebrium’s equivalent entry file is `main.py`. Start by copying all import statements and constant variables unrelated to Replicate/Cog: ```python theme={null} import os import time import torch import subprocess import numpy as np from typing import List from transformers import CLIPImageProcessor from diffusers import ( StableDiffusionXLPipeline, DDIMScheduler, DPMSolverMultistepScheduler, EulerAncestralDiscreteScheduler, EulerDiscreteScheduler, HeunDiscreteScheduler, PNDMScheduler, KDPM2AncestralDiscreteScheduler, ) from diffusers.pipelines.stable_diffusion.safety_checker import ( StableDiffusionSafetyChecker, ) UNET = "sdxl_lightning_4step_unet.pth" MODEL_BASE = "stabilityai/stable-diffusion-xl-base-1.0" UNET_CACHE = "unet-cache" BASE_CACHE = "checkpoints" SAFETY_CACHE = "safety-cache" FEATURE_EXTRACTOR = "feature-extractor" MODEL_URL = "https://weights.replicate.delivery/default/sdxl-lightning/sdxl-1.0-base-lightning.tar" SAFETY_URL = "https://weights.replicate.delivery/default/sdxl/safety-1.0.tar" UNET_URL = "https://weights.replicate.delivery/default/comfy-ui/unet/sdxl_lightning_4step_unet.pth.tar" class KarrasDPM: def from_config(config): return DPMSolverMultistepScheduler.from_config(config, use_karras_sigmas=True) SCHEDULERS = { "DDIM": DDIMScheduler, "DPMSolverMultistep": DPMSolverMultistepScheduler, "HeunDiscrete": HeunDiscreteScheduler, "KarrasDPM": KarrasDPM, "K_EULER_ANCESTRAL": EulerAncestralDiscreteScheduler, "K_EULER": EulerDiscreteScheduler, "PNDM": PNDMScheduler, "DPM++2MSDE": KDPM2AncestralDiscreteScheduler, } ``` Replicate uses classes, while Cerebrium runs standard Python code and makes each function an endpoint. Remove all `self.` references throughout the code. The repo contains a "feature-extractor" folder needed in the Cerebrium project. Since it's small, copy the folder contents directly: Folder Structure Replicate’s setup function runs on each cold start (each new app instantiation). Define it as top-level code below the import statements. ```python theme={null} def download_weights(url, dest): start = time.time() print("downloading url: ", url) print("downloading to: ", dest) subprocess.check_call(["pget", "-x", url, dest], close_fds=False) print("downloading took: ", time.time() - start) """Load the model into memory to make running multiple predictions efficient""" start = time.time() print("Loading safety checker...") if not os.path.exists(SAFETY_CACHE): download_weights(SAFETY_URL, SAFETY_CACHE) print("Loading model") if not os.path.exists(BASE_CACHE): download_weights(MODEL_URL, BASE_CACHE) print("Loading Unet") if not os.path.exists(UNET_CACHE): download_weights(UNET_URL, UNET_CACHE) self.safety_checker = StableDiffusionSafetyChecker.from_pretrained( SAFETY_CACHE, torch_dtype=torch.float16 ).to("cuda") self.feature_extractor = CLIPImageProcessor.from_pretrained(FEATURE_EXTRACTOR) print("Loading txt2img pipeline...") self.pipe = StableDiffusionXLPipeline.from_pretrained( MODEL_BASE, torch_dtype=torch.float16, variant="fp16", cache_dir=BASE_CACHE, local_files_only=True, ).to("cuda") unet_path = os.path.join(UNET_CACHE, UNET) self.pipe.unet.load_state_dict(torch.load(unet_path, map_location="cuda")) print("setup took: ", time.time() - start) ``` The code downloads model weights if they don’t exist and instantiates the models. To persist files/data on Cerebrium, store them at **/persistent-storage**. Update the paths: ```python theme={null} UNET_CACHE = "/persistent-storage/unet-cache" BASE_CACHE = "/persistent-storage/checkpoints" SAFETY_CACHE = "/persistent-storage/safety-cache" ``` Copy the remaining functions, run\_safety\_checker() and predict(). In Cerebrium, function parameters map directly to the expected JSON request data: ```python theme={null} def run_safety_checker(image): safety_checker_input = feature_extractor(image, return_tensors="pt").to( "cuda" ) np_image = [np.array(val) for val in image] image, has_nsfw_concept = safety_checker( images=np_image, clip_input=safety_checker_input.pixel_values.to(torch.float16), ) return image, has_nsfw_concept def predict( prompt: str = "A superhero smiling", negative_prompt: str = "worst quality, low quality", width: int = 1024, height: int = 1024, num_outputs: int = 1, scheduler: str = "K_EULER", num_inference_steps: int = 4, guidance_scale: float = 0, seed: int = None, disable_safety_checker: bool = False, ): """Run a single prediction on the model""" global pipe if seed is None: seed = int.from_bytes(os.urandom(4), "big") print(f"Using seed: {seed}") generator = torch.Generator("cuda").manual_seed(seed) # OOMs can leave vae in bad state if pipe.vae.dtype == torch.float32: pipe.vae.to(dtype=torch.float16) sdxl_kwargs = {} print(f"Prompt: {prompt}") sdxl_kwargs["width"] = width sdxl_kwargs["height"] = height pipe.scheduler = SCHEDULERS[scheduler].from_config( pipe.scheduler.config, timestep_spacing="trailing" ) common_args = { "prompt": [prompt] * num_outputs, "negative_prompt": [negative_prompt] * num_outputs, "guidance_scale": guidance_scale, "generator": generator, "num_inference_steps": num_inference_steps, } output = pipe(**common_args, **sdxl_kwargs) if not disable_safety_checker: _, has_nsfw_content = run_safety_checker(output.images) output_paths = [] for i, image in enumerate(output.images): if not disable_safety_checker: if has_nsfw_content[i]: print(f"NSFW content detected in image {i}") continue output_path = f"/tmp/out-{i}.png" image.save(output_path) output_paths.append(Path(output_path)) if len(output_paths) == 0: raise Exception( "NSFW content detected. Try running it again, or try a different prompt." ) return output_paths ``` The above returns a path to the generated images. To return base64-encoded images for instant rendering, use the following. Alternatively, upload images to a storage bucket. ```python theme={null} from io import BytesIO import base64 encoded_images = [] for i, image in enumerate(output.images): if not disable_safety_checker: if has_nsfw_content[i]: print(f"NSFW content detected in image {i}") continue buffered = BytesIO() image.save(buffered, format="PNG") img_b64 = base64.b64encode(buffered.getvalue()).decode("utf-8") encoded_images.append(img_b64) if len(encoded_images) == 0: raise Exception( "NSFW content detected. Try running it again, or try a different prompt." ) return encoded_images ``` Run `cerebrium deploy`. The app builds in under 90 seconds. It should output the curl statement to run your app: Curl Request Replace the end of the URL with `/predict` (the target function) and send the required JSON data. Example result: ```python theme={null} { "run_id": "c6797f2e-333a-9e89-bafa-4dd0f4fbe22a", "result": ["iVBORw0KGgoAAAANSUhEUgAABAAAAAQACAIAAADwf7zUAA...."], "run_time_ms": 43623.4176158905 } ``` Read more about Cerebrium functionality: * [Secrets](/docs/other-topics/using-secrets) * [Model scaling](/docs/scaling/scaling-apps) * [Faster cold starts](/docs/performance/faster-cold-starts) # Custom Domains Source: https://cerebrium.ai/docs/networking/custom-domains Serve Cerebrium apps from your own domain with CNAME setup, DNS validation, automatic SSL certificates, and fixes for common DNS provider issues. Custom domains serve Cerebrium apps through a custom domain instead of the default `*.cerebrium.ai` URLs. Once configured, API calls use the custom domain while keeping the same path structure: `api.example.com/v4/p-1234/my-app/run`. ## Key Features * Support for apex domains (`example.com`) and subdomains (`api.example.com`) * Automatic SSL certificate provisioning and renewal * Project-level domains: one domain serves all apps in a project * Multiple domains can point to the same project * Professional branding with custom domains instead of `*.cerebrium.ai` URLs ## How Custom Domains Work * **Domains are region-specific** meaning each will always resolve to the selected region * **Domains are app-agnostic** enabling connection to any number of deployed apps within a project via custom domain * **Multiple domains** can be configured on the same project (useful for apps within the project deployed in different regions) ## Getting Started ### Step 1: Create a Custom Domain 1. Navigate to project settings in the Cerebrium dashboard 2. Click the "Custom Domains" tab 3. Click "Add Custom Domain" 4. Enter the domain name and select the target region (the region closest to users provides the lowest latency) 5. Click "Create Domain" ### Step 2: Configure DNS Records After you create the domain, the dashboard displays DNS configuration instructions. Create a CNAME record at the DNS provider using the DNS record Cerebrium generates. DNS record details can also be found later by clicking "DNS Record" in the Custom Domains list. **Benefits:** * Easier DNS configuration * Allows other subdomains for different purposes ``` Type: CNAME Name: api Value: proxy.aws.{region}.cerebrium.ai ``` **Benefits:** * Routes all traffic from the domain root * Best for dedicated API domains **Note:** ALIAS records are preferred over CNAME since only one CNAME is allowed per domain ``` Type: CNAME (or ALIAS if supported) Name: @ (or leave blank) Value: proxy.aws.{region}.cerebrium.ai ``` A `{region}` must be selected when creating a custom domain. ### Step 3: Domain Validation 1. After configuring DNS, return to the Cerebrium dashboard 2. Cerebrium will automatically attempt to validate DNS records every 30 minutes for up to 2 days 3. To trigger an immediate validation attempt, click "Validation Status" -> "Validate Domain" (this works even after the 2-day window has elapsed) 4. If validation fails, the dialog will show the last known error 5. Once validated, Cerebrium automatically provisions SSL certificates ### Step 4: Start Using the Custom Domain Once validation is complete, the custom domain can be used immediately with all apps in the project: ```bash theme={null} curl -X POST https://api.example.com/v4/p-1234/my-app/run \ -H "Authorization: Bearer {YOUR_API_KEY}" \ -H "Content-Type: application/json" \ -d "{'inputs': {'prompt': 'Hello world'}}" ``` ```bash theme={null} curl https://api.example.com/v4/p-1234/my-app/run \ -H "Authorization: Bearer {YOUR_API_KEY}" ``` ## Managing Domains ### Domain Validation Statuses Cerebrium attempts to validate domains once every 30 minutes for up to 2 days. If validation has not succeeded, click "Validation Status" to see the last attempt and any errors. * **Pending**: Domain is waiting for DNS validation * **Validated**: Domain is successfully validated and ready to use * **Failed**: DNS validation failed. Click "Validation Status" to see the specific error ### Validation Errors If a domain shows "Failed" status, the DNS record has been misconfigured. Common issues include: * **Wrong target**: CNAME points to incorrect proxy URL * **Multiple conflicting records**: Remove any existing ALIAS records for the same hostname ### Common DNS Provider Issues * **Cloudflare**: Disable proxy (set to "DNS only" - gray cloud icon) * **Route53**: Use simple routing policy, not weighted or latency-based * **Namecheap**: Use "@" for apex domains, not "www" or blank * **GoDaddy**: CNAME records cannot be used with apex domains. Consider using a subdomain or an ALIAS record ### Common DNS Mistakes * **Wrong proxy URL**: Use the exact URL shown in the dashboard * **TTL too high**: Consider lowering TTL before making DNS changes * **Multiple records**: Remove conflicting A or CNAME records for the same name ## Security Considerations * Custom domains automatically receive SSL certificates * All traffic is encrypted end-to-end * Certificates auto-renew before expiration * Domain validation ensures control of the domain # Inter-cluster routing Source: https://cerebrium.ai/docs/networking/inter-cluster-routing Use inter-cluster routing to connect Cerebrium apps privately within a region for sub-millisecond latency without traversing the public internet. Inter-cluster routing enables direct, low-latency communication between Cerebrium apps within the same region. Traffic stays off the public internet, reducing latency and improving performance. Each application scales independently based on its configured scaling parameters. Inter-cluster routing provides: * **Low latency**: Direct container-to-container communication within the same region (\~0.3–1 ms typical) * **High bandwidth**: Up to 50 Gbps between containers * **No public internet**: Apps communicate directly without external routing * **Observability**: All requests appear in the Cerebrium dashboard with full logs, payloads, and latency metrics ## How it works Inter-cluster Routing When one app communicates with another within the same region, the request is routed through Cerebrium’s local proxy layer. This proxy keeps every request inside the regional cluster while enforcing authentication, observability, and scaling. All communication follows the same security standards as the public API server — every request is authenticated unless authentication is explicitly disabled. If authentication is disabled without custom security in place, other apps within the same cluster can access the endpoint. The proxy enforces configured [scaling parameters](/docs/scaling/scaling-apps), including concurrency and RPS-based autoscaling, so services scale predictably as traffic grows. Multiple communication protocols are supported — apps can interact over HTTP, WebSocket, or batch job execution for streaming data, chaining models, or triggering asynchronous workloads. Apps communicate using a consistent internal endpoint format: `http://api.aws/v4///` This endpoint pattern remains the same across all regions, so URLs do not change when deploying to multiple locations. Inter-cluster routing only works between applications deployed within the same region. Requests never traverse the public internet — they stay fully contained within the cluster network, achieving typical latencies of 0.3–1 ms and bandwidth up to 50 Gbps between containers. gRPC is not currently supported but is on the roadmap. # Audit Log Source: https://cerebrium.ai/docs/other-topics/audit-log Review who changed what in a Cerebrium project — app, secret, credential, member, and volume actions, with retention and API access by plan. The audit log records actions taken against a project: who took the action, what it acted on, when, and whether it succeeded. Use it to answer who deleted an app, who read a secret, and who invited a member. Inference requests are not audited. They appear in [app logs](/docs/other-topics/request-response-logging) instead. ## Availability The audit log is included on the Standard and Enterprise plans. Retention sets how far back the log is readable: | Plan | Retention | | - | - | | Hobby | Not included | | Standard | 7 days | | Enterprise | 30 days | ## Viewing the Log Open a project in the [dashboard](https://dashboard.cerebrium.ai/login) and select **Audit Log**. Select a row to expand the full detail of the entry. Filter by date range, action, or outcome. Only project owners can read the audit log. ## Recorded Actions | Action | Records | | - | - | | `app.create` | An app was created | | `app.update` | An app's configuration was changed | | `app.delete` | An app was deleted | | `build.download` | An app's source was downloaded from a build | | `build.cancel` | A build was cancelled | | `container.stop` | A running container was stopped | | `run.cancel` | An async run was cancelled | | `secrets.read` | Secret values were retrieved, at project or app level | | `secrets.update` | Secrets were changed, at project or app level | | `apikey.create` | An API key was created | | `apikey.read` | API keys were listed | | `apikey.delete` | An API key was deleted | | `serviceaccount.create` | A service account was created | | `serviceaccount.update` | A service account was changed | | `serviceaccount.delete` | A service account was deleted | | `serviceaccount.keys_list` | A service account's tokens were listed | | `project.member_invite` | A user was invited to the project | | `project.member_remove` | A user was removed from the project | | `volume.file_download` | A file was downloaded from a volume | | `volume.file_delete` | A file was deleted from a volume | | `volume.resize` | A volume was resized | | `project.delete` | The project was deleted | ## Entry Contents Each entry records: * **Time** — when the action was taken * **Action** — the action name from the table above * **Actor** — the user or service account that took the action, with the email address, IP address, and user agent of the request * **Target** — what was acted on, as a type and ID, such as `app:my-app` * **Outcome** — `success` or `failure`, with the HTTP status code * **Details** — action-specific fields, such as the invited email address and role on `project.member_invite`, or the file path and region on `volume.file_delete` Failed attempts are recorded alongside successful ones. ## Reading the Log Through the API The log is also available over the REST API, authenticated as a project owner: ```bash theme={null} curl -X GET "https://rest.cerebrium.ai/v2/projects/{project_id}/audit-log?limit=50" \ -H "Authorization: Bearer " ``` Supported query parameters: | Parameter | Description | | - | - | | `from` | Start of the window, RFC3339. Clamped to the plan's retention window | | `to` | End of the window, RFC3339. Defaults to now | | `action` | Return only this action, for example `secrets.read` | | `actor` | Return only actions taken by this actor ID | | `outcome` | Return only `success` or `failure` | | `limit` | Page size, 1–200. Defaults to 50 | | `nextToken` | Page token from the previous response | Entries are returned newest first. When `hasMore` is true, pass the response's `nextPageToken` as `nextToken` to fetch the next page. # Request and Response Logging Source: https://cerebrium.ai/docs/other-topics/request-response-logging Disable request and response logging in the Cortex runtime using secrets to protect sensitive data or reduce overhead in Cerebrium apps. By default, the Cortex runtime logs all requests and responses. These logs appear in the app dashboard and are useful for debugging and monitoring. Disable them when privacy, security, or performance requires it. ## Controlling Logs Two settings control logging behavior in the Cortex runtime: * `DISABLE_REQUEST_LOGS`: When enabled, prevents logging of incoming request data * `DISABLE_RESPONSE_LOGS`: When enabled, prevents logging of response data from your app These settings only affect the default Cortex runtime. If you are using a [custom runtime](/docs/container-images/custom-web-servers), you will need to handle logging behavior in your own server implementation. ## Configuration Configure these settings as [secrets](/docs/other-topics/using-secrets) at either the app or project level: 1. Navigate to your project dashboard 2. Go to the "Secrets" section 3. Add one or both of the settings as a key with a value of `"true"` to disable the corresponding logs Setting the value to anything other than `true` (or leaving it undefined) will keep logging enabled, which is the default behavior. ## Use Cases Common reasons to disable logs: * **Processing sensitive data**: If your app processes personal or sensitive information that shouldn't be stored in logs * **Observability optimization**: For high-throughput apps where logging adds unnecessary overhead making it difficult to monitor your app ## Example To disable both request and response logging, add the following secrets to your app: | Key | Value | | - | - | | DISABLE\_REQUEST\_LOGS | true | | DISABLE\_RESPONSE\_LOGS | true | Secrets are loaded on container startup. After updating a secret, restart the app container for the change to take effect. # Using Secrets Source: https://cerebrium.ai/docs/other-topics/using-secrets Store API keys and credentials as encrypted secrets in Cerebrium, expose them as environment variables, and manage them at project or app level. Secrets store API keys, passwords, and other sensitive information outside of code. Secrets are encrypted at rest (256-bit AES) and decrypted only at runtime. You can manage secrets at both project and app levels. Project-level secrets are shared across all apps in your project, while app-level secrets are specific to an individual app. App secrets take precedence over project-wide secrets. Each secret is exposed to the app as an environment variable. Secrets are available in every region an app runs in. Apps deployed to [multiple regions](/docs/deployments/multi-region-deployment) require no per-region secret setup. Secrets are loaded on container startup. If you update a secret, you must restart your app container for the changes to take effect. ```python theme={null} def predict(run_id): print(f"Run ID: {run_id}") hf_token = os.environ.get("HF_TOKEN") logger.info(f"HF_TOKEN: {hf_token}") return {"result": f"Your HF_TOKEN is {hf_token}"} ``` Secrets are stored as strings. If your secret is a JSON payload or similar, remember to convert it to the correct format using `json.loads(os.environ.get("MY_JSON_SECRET"))`. ### Managing Secrets You create, update, and delete secrets in your dashboard. Secrets Secrets are loaded on model start. You will need to wait for your app container to restart or deploy your app before the new secret is available. ### Automatic Environment Variables Cerebrium automatically sets the following environment variables for your app: * APP\_NAME: The name of your application * HF\_HOME: Set to '/persistent-storage/.cache/huggingface' for caching HuggingFace models * PROJECT\_ID: The ID of your Cerebrium project * BUILD\_ID: The unique identifier for the current build The app\_id is a composite of PROJECT\_ID + '\_' + APP\_NAME. ### Local Development For local development, use an `.env` file. Add the same secrets to the dashboard before deploying. ```python theme={null} import os from dotenv import load_dotenv load_dotenv() hf_token = os.environ.get("HF_TOKEN") ``` # Self-hosted Deepgram speech-to-text on Cerebrium Source: https://cerebrium.ai/docs/partner-services/deepgram Run self hosted Deepgram speech to text on Cerebrium with model file uploads, engine and api TOML setup, GPU scaling, and low latency voice agents. Cerebrium's partnership with [Deepgram](https://www.deepgram.com/) enables simple deployment of speech-to-text (STT) services with simplified configuration and independent scaling. Using Deepgram services requires an Enterprise Deepgram account and API key for self-hosted models. Contact Deepgram support to access this feature. Obtain links to the Deepgram model files referenced below (file extension `.dg`) from your Deepgram Account Representative. Consult your Deepgram representative on how to achieve parity with the Deepgram API. Deepgram Partner Service is available from CLI version 1.39.0 and greater Deployments currently run Deepgram self-hosted release `260728`. Model files and `engine.toml`/`api.toml` settings should match that release. Check with your Deepgram Account Representative if you are unsure whether your model files are current. The Deepgram Partner Service is in beta. It's available to all users and ready for production workloads, but expect occasional rough edges while the integration matures. Reach out to [support](mailto:support@cerebrium.ai) if you hit any issues. ## Setup 1. Create a Cerebrium app with the CLI: ```bash theme={null} cerebrium init deepgram ``` 2. Create a self-hosted API key from the **Deepgram** dashboard. Navigate to the **Secrets** tab in the **Cerebrium** dashboard and add the API key with the name `DEEPGRAM_API_KEY`. This secret automatically becomes available as an environment variable in the deployment. 3. Download model files from Deepgram's self-hosted section in the **Deepgram** dashboard using the guide available [here](https://developers.deepgram.com/docs/deploy-deepgram-services#pull-deepgram-container-images). Select the 'license proxy' deployment type. Upload downloaded model files using the links provided by your Account Representative (with `.dg` extension) to persistent storage in the `/deepgram-models` folder. This folder automatically attaches to the engine container. Use this command to upload the files: ```bash theme={null} cerebrium cp .dg deepgram-models/.dg #example cerebrium cp nova-3-general.en.streaming.123456.dg deepgram-models/nova-3-general.en.streaming.123456.dg ``` The `deepgram-models` directory remains at the root level of persistent storage and is shared across all Deepgram apps in the project. The configuration files (`api.toml` and `engine.toml`), however, belong under the app name directory (see steps 4 and 5 below). 4. Create a file named `engine.toml` with the following content and upload to your persistent storage under the app name directory (e.g., `{appName}/engine.toml`). These are the default settings. Adjust as needed. For example, if your app is named `deepgram`: ```bash theme={null} cerebrium cp engine.toml deepgram/engine.toml ``` Keep the Engine's `[server]` port set to **8055**, and the `driver_pool` URL in `api.toml` pointing at it. Cerebrium uses this port to tell when the Engine is ready, so requests that arrive during startup are queued instead of failing with `503 Please try again later`. ```bash theme={null} ### Keep in mind that all paths are in-container paths and do not need to exist ### on the host machine. ### Limit the number of active requests handled by a single Engine container. ### Engine will reject additional requests from API beyond this limit, and the ### API container will continue with the retry logic configured in `api.toml`. ### ### The default is no limit. # max_active_requests = ### Configure license validation by passing in a DEEPGRAM_API_KEY environment variable ### See https://developers.deepgram.com/docs/deploy-deepgram-services#credentials [license] server_url = ["https://license.deepgram.com"] ### Configure the server to listen for requests from the API. [server] ### The IP address to listen on. Since this is likely running in a Docker ### container, you will probably want to listen on all interfaces. host = "0.0.0.0" ### The port to listen on. On Cerebrium, keep 8055 so requests are queued while ### the Engine starts up. port = 8055 ### To support metrics we need to expose an Engine endpoint. ### See https://developers.deepgram.com/docs/metrics-guide#deepgram-engine [metrics_server] host = "0.0.0.0" port = 9991 [model_manager] ### The number of models to have concurrently loaded in system memory. ### If managing a deployment with dozens of models this setting will ### help prevent instances where models consume too much memory and ### offload the models to disk as needed on a least-recently-used basis. ### ### The default is no limit. # max_concurrently_loaded_models = 20 ### Inference models. You can place these in one or multiple directories. search_paths = ["/models"] ### Enable ancillary features [features] ### Allow multichannel requests by setting this to true, set to false to disable multichannel = true # or false ### Enables language detection *if* a valid language detection model is available language_detection = true # or false ### Enables streaming entity formatting *if* a valid NER model is available streaming_ner = false # or true ### Size of audio chunks to process in seconds. [chunking.batch] # min_duration = # max_duration = [chunking.streaming] # min_duration = # max_duration = ### How often to return interim results, in seconds. Default is 1.0s. ### ### This value may be lowered to increase the frequency of interim results. ### However, this may cause a significant decrease in number of concurrent ### streams supported by a single GPU. Please contact your Deepgram Account ### representative for more details. step = 0.2 ``` The Engine sets half precision automatically based on GPU support, so `engine.toml` needs no `[half_precision]` section. Deepgram dropped that setting in self-hosted release `260305`; if your `engine.toml` still carries it, remove it when you next update the file. 5. Create a file named `api.toml` with the following content and upload to your persistent storage under the app name directory (e.g., `{appName}/api.toml`). These are the default settings. Adjust as needed. For example, if your app is named `deepgram`: ```bash theme={null} cerebrium cp api.toml deepgram/api.toml ``` ```bash theme={null} ### Keep in mind that all paths are in-container paths, and do not need to exist ### on the host machine. ### Configure license validation by passing in a DEEPGRAM_API_KEY environment variable ### See https://developers.deepgram.com/docs/deploy-deepgram-services#credentials [license] server_url = ["https://license.deepgram.com"] ### Configure how the API will listen for your requests [server] ### The base URL (prefix) for requests to the API. base_url = "/v1" ### The IP address to listen on. Since this is likely running in a Docker ### container, you will probably want to listen on all interfaces. host = "0.0.0.0" ### The port to listen on port = 8082 ### How long to wait for a connection to a callback URL. callback_conn_timeout = "1s" ### How long to wait for a response to a callback URL. callback_timeout = "10s" ### How long to wait for a connection to a fetch URL. fetch_conn_timeout = "1s" ### How long to wait for a response to a fetch URL. fetch_timeout = "60s" ### By default, the API listens over HTTP. By passing both a certificate ### and a key file, the API will instead listen over HTTPS. ### ### This performs TLS termination only, and does not provide any ### additional authentication. [server.https] # cert_file = "/path/to/cert.pem" # key_file = "/path/to/key.pem" ### Specify custom DNS resolution options. [resolver] ### Specify custom domain name server(s). ### Format is "{IP} {PORT} {PROTOCOL (tcp or udp)}" # nameservers = ["127.0.0.1 53 udp"] ### If specifying a custom DNS nameserver, set the DNS TTL value. # max_ttl = 10 ### Limit the number of active requests handled by a single API container. ### If additional requests beyond the limit are sent, API will return ### a 429 HTTP status code. Default is no limit. [concurrency_limit] # active_requests = ### Enable ancillary features [features] ### Enables topic detection *if* a valid topic detection model is available topic_detection = true # or false ### Enables summarization *if* a valid summarization model is available summarization = true # or false ### Enables pre-recorded entity detection *if* a valid entity detection model is available entity_detection = false # or true ### Enables pre-recorded entity-based redaction *if* a valid entity detection model is available entity_redaction = false # or true ### Enables pre-recorded entity formatting *if* a valid NER model is available format_entity_tags = false # or true ### If API is receiving requests faster than Engine can process them, a request ### queue will form. By default, this queue is stored in memory. Under high load, ### the queue may grow too large and cause Out-Of-Memory errors. To avoid this, ### set a disk_buffer_path to buffer the overflow on the request queue to disk. ### ### WARN: This is only to temporarily buffer requests during high load. ### If there is not enough Engine capacity to process the queued requests over time, ### the queue (and response time) will grow indefinitely. # disk_buffer_path = "/path/to/disk/buffer/directory" ### Enables streaming TTS *if* a valid Aura TTS model is available speak_streaming = true # or false ### Configure the backend pool of speech engines (generically referred to as ### "drivers" here). The API will load-balance among drivers in the standard ### pool; if one standard driver fails, the next one will be tried. ### ### Each driver URL will have its hostname resolved to an IP address. If a domain ### name resolves to multiple IP addresses, the API will load-balance across each ### IP address. ### ### This behavior is provided for convenience, and in a production environment ### other tools can be used, such as HAProxy. ### ### Below is a new Speech Engine ("driver") in the "standard" pool. [[driver_pool.standard]] ### Host to connect to. If you are using a different method of orchestrating, ### then adjust the IP address accordingly. ### ### WARN: This must be HTTPS. ### ### Docker Compose and Podman Compose create a dedicated network that allows inter-container communication by app name. ### See [Networking in Compose](https://docs.docker.com/compose/networking/) for details. ### ### On Cerebrium, keep port 8055 here so it matches `[server]` in engine.toml. url = "https://0.0.0.0:8055/v2" ### Factor to increase the timeout by for each additional retry (for ### exponential backoff). timeout_backoff = 1.2 ### Before attempting a retry, sleep for this long (in seconds) retry_sleep = "2s" ### Factor to increase the retry sleep by for each additional retry (for ### exponential backoff). retry_backoff = 1.6 ### Maximum response to deserialize from Driver (in bytes) max_response_size = 1073741824 # 1GB ``` 6. Update `cerebrium.toml` with the following configuration for hardware, scaling, region, and other settings: ```toml theme={null} [cerebrium.deployment] name = "deepgram" # Enable below in production environments disable_auth = true [cerebrium.runtime.deepgram] [cerebrium.hardware] cpu = 4 memory = 32 compute = "AMPERE_A10" gpu_count = 1 [cerebrium.scaling] min_replicas = 0 max_replicas = 2 cooldown = 120 replica_concurrency = 150 ``` 7. Run `cerebrium deploy`. After deployment, an endpoint for the Deepgram services appears in the terminal output. The URL is also available on the app's overview page in the dashboard. 8. Download an example audio file for use with the Deepgram service: ```bash theme={null} wget https://dpgr.am/bueller.wav ``` 9. Call the Deepgram endpoint with appropriate parameters: ```curl theme={null} curl -X POST --data-binary @bueller.wav "https://api.cerebrium.ai/v4/p-xxxxxxxx/deepgram/v1/listen?model=nova-3&smart_format=true" ``` You can find the parameters accepted by the Deepgram service in the [speech-to-text API reference](https://developers.deepgram.com/reference/speech-to-text-api/listen-streaming). If `disable_auth` in `cerebrium.toml` is set to false, include the inference token in the Authorization header to authenticate with the Cerebrium service. Cerebrium pulls the Deepgram API key automatically from secrets. ## API Key Configuration To use Deepgram services: 1. Sign up at [deepgram.com](https://www.deepgram.com/) 2. Create an API key in the Deepgram dashboard 3. Add the API key to Cerebrium: * Navigate to the **Secrets** tab in the Cerebrium dashboard * Add the Deepgram API key as an app-specific or project-wide secret named `DEEPGRAM_API_KEY` * This secret automatically becomes available as an environment variable in the deployment ## Scaling and Concurrency Deepgram services support independent scaling configurations: * **min\_replicas**: Minimum number of instances to maintain (0 for scale-to-zero) * **max\_replicas**: Maximum number of instances that can be created during high load * **replica\_concurrency**: Number of concurrent requests each instance can handle * **cooldown**: Time window (in seconds) that must pass at reduced concurrency before scaling down Adjust these parameters based on traffic patterns and latency requirements. ## Usage Examples Cerebrium runs both Deepgram STT models and applications on the same network alongside LiveKit workers, reducing latency by approximately 400ms. This is a significant advantage for voice agent applications. For a complete implementation reference, see the [LiveKit Outbound Agent example](/docs/v4/examples/livekit-outbound-agent). # Partner Services on Cerebrium with Deepgram and Rime Source: https://cerebrium.ai/docs/partner-services/index Learn how Cerebrium partner services like Deepgram and Rime offer quick deployment, independent scaling, and lower latency for AI workloads. Partner Services are available from CLI version 1.39.0 and greater Partner Services are in beta. They're available to all users and ready for production workloads, but expect occasional rough edges while the integrations mature. Reach out to [support](mailto:support@cerebrium.ai) if you hit any issues. Cerebrium offers specialized services in partnership with leading AI companies, featuring simplified configuration, independent scaling, and quick deployment. Available Partner Services: * [Deepgram](/docs/partner-services/deepgram): Speech-to-text (STT) services * [Rime](/docs/partner-services/rime): Text-to-speech (TTS) services ## Benefits of Partner Services Partner Services provide: * Quick and easy deployment * Independent scaling of each service * Reduced costs by running models on Cerebrium's optimized runtime * Reduced latency by running models on the same network as the app * Deploy to specific regions for data compliance and latency requirements ## Getting Started Configure service-specific requirements through the Cerebrium platform. Refer to individual service pages linked above for detailed requirements, which may include: 1. API keys and authentication details 2. Service-specific configuration parameters 3. Resource requirements and limitations ## Scaling and Concurrency Partner Services support independent scaling configurations: * Use the `min_replicas` and `max_replicas` parameters to control the number of instances * The `replica_concurrency` parameter determines how many concurrent requests each instance can handle * Adjust the `cooldown` parameter to control the time window that must pass at reduced concurrency before scaling down * Adjust the `hardware` section to control the instance type which affects performance and/or cost For more information on specific Partner Services, see: * [Deepgram](/docs/partner-services/deepgram) * [Rime](/docs/partner-services/rime) # Rime Source: https://cerebrium.ai/docs/partner-services/rime Deploy Rime text to speech on Cerebrium with a simple TOML runtime config, HTTP and WebSocket endpoints, and scaling tuned for low latency TTS. Rime Partner Service is available from CLI version 1.39.0 and greater The Rime Partner Service is in beta. It's available to all users and ready for production workloads, but expect occasional rough edges while the integration matures. Reach out to [support](mailto:support@cerebrium.ai) if you hit any issues. Cerebrium's partnership with [Rime](https://www.rime.ai/) enables text-to-speech (TTS) deployment with low latency and region selection for data privacy compliance. ## Setup 1. Create a [Rime](https://www.rime.ai/) account and get an API key. Add the key as a secret in Cerebrium with the name "RIME\_API\_KEY". 2. Create a Cerebrium app with the CLI: ```bash theme={null} cerebrium init rime ``` 3. Rime services use a simplified TOML configuration with the `[cerebrium.runtime.rime]` section. Create a `cerebrium.toml` file with the following: ```toml theme={null} [cerebrium.deployment] name = "rime" disable_auth = true [cerebrium.runtime.rime] port = 8001 # model_name = "arcana" # Optional: specify a Rime model (e.g. "arcana", "mist", "mistv2") # language = "en" # Optional: specify language code (e.g. "en", "es") [cerebrium.hardware] cpu = 4 memory = 30 compute = "AMPERE_A10" gpu_count = 1 [cerebrium.scaling] min_replicas = 1 max_replicas = 2 cooldown = 120 replica_concurrency = 50 ``` Disable auth because the Rime API key in the header handles authentication. The Rime Server validates the API key directly. 4. Run `cerebrium deploy` to deploy the Rime service. The output should appear as follows: ``` App Dashboard: https://dashboard.cerebrium.ai/projects/p-xxxxxxxx/apps/p-xxxxxxxx-rime ``` 5. Send requests to the HTTP Rime service using the deployment URL from the output: ``` curl --location 'https://api.cerebrium.ai/v4/p-xxxxxxxx/rime' \ --header 'Authorization: Bearer ' \ --header 'Content-Type: application/json' \ --header 'Accept: audio/pcm' \ --data '{ "text": "I would love to have a conversation with you.", "speaker": "joy", "modelId": "mist" }' ``` For Websockets, send the following ``` wss://api.cerebrium.ai/v4/p-xxxxxxxx/rime/ws2?audioFormat=mp3&speaker=cove&modelId=mistv2&phonemizeBetweenBrackets=true Authorization Bearer #With a message like: {"text": "This "}, {"text": "is "}, {"text": "a "}, {"text": "test against the "}, {"text": "websockets endpoint of the "}, {"text": "api image. "}, {"operation": "flush"}, {"text": "This "}, {"text": "is "}, {"text": "an "}, {"text": "incomplete "}, {"text": "phrase "}, {"operation": "eos"} ``` ## Runtime Configuration The `[cerebrium.runtime.rime]` section supports the following parameters: | Option | Type | Default | Description | | - | - | - | - | | `port` | integer | required | Port the Rime server listens on. Typically `8001`. | | `model_name` | string | — | Rime model to load (e.g. `"arcana"`, `"mist"`, `"mistv2"`). Defaults to Rime's server default if not set. | | `language` | string | — | Language code for the model (e.g. `"en"`, `"es"`). Defaults to Rime's server default if not set. | Example with optional parameters: ```toml theme={null} [cerebrium.runtime.rime] port = 8001 model_name = "arcana" language = "en" ``` ## Scaling and Concurrency Rime services support independent scaling configurations: * **min\_replicas**: Minimum instances to maintain (0 for scale-to-zero). Recommended: 1. * **max\_replicas**: Maximum instances during high load. * **replica\_concurrency**: Concurrent requests per instance. Recommended: 3. * **cooldown**: Time window (in seconds) that must pass at reduced concurrency before scaling down. Recommended: 50. * **compute**: Instance type. Recommended: `AMPERE_A10`. Adjust these parameters based on traffic patterns and latency requirements. Consult the Rime team for concurrency and scalability guidance. For further documentation on Rime, see the [Rime documentation](https://docs.rime.ai/). # Memory and GPU Checkpointing (Beta) Source: https://cerebrium.ai/docs/performance/checkpointing Snapshot CPU and GPU memory to skip imports, model loading, and CUDA kernel compilation, drastically cutting Cerebrium container cold start times. Checkpointing is in early beta. Some workloads may fail to checkpoint or restore, and behavior may change as the feature matures. Report issues via the [Discord Community](https://discord.gg/ATj6USmeE2) or [email](mailto:support@cerebrium.ai). ## Introduction Memory checkpointing takes a snapshot of a container's CPU memory and GPU memory, and uses it to speed up the startup of future containers. Applications that perform a large amount of work at container start time benefit the most from this process. This is useful for both CPU-only and GPU workloads. For CPU applications, checkpointing can preserve expensive initialization work such as imports, dependency loading, configuration setup, and in-memory state. For GPU applications, it can also preserve model weights, CUDA state, and compiled kernels. For example, ML and LLM frameworks often load large model weights and compile CUDA kernels at container start time, which can take many seconds or minutes. Loading from a checkpoint that already contains this initialized state can skip most of that delay. Since this feature is still in beta, please report all issues to the team via our [Discord Community](https://discord.gg/ATj6USmeE2) or via [Email](mailto:support@cerebrium.ai). ## How to use Checkpointing is available in early beta to our customer base. ### 1. Enable checkpointing in `cerebrium.toml` ```toml theme={null} [cerebrium.experimental] checkpointing = true ``` ### 2. Send the checkpoint trigger You have full control over when to create a checkpoint - preferably it's after the majority of your application's initialization work is complete and your desired state is reached. The state of your application at this exact moment is what will be restored for future container launches. Once you reach this optimal point in your startup logic, send a POST request from inside the container to instruct Cerebrium to capture the checkpoint: ```python theme={null} import json import urllib.request # Trigger the checkpoint only after finishing all major initialization tasks. req = urllib.request.Request( "http://169.254.169.253:8234/checkpoint", method="POST", ) with urllib.request.urlopen(req, timeout=300) as response: print(json.loads(response.read())) ``` The endpoint uses `169.254.169.253`, a link-local address that routes to the Cerebrium runtime sidecar inside the container. The address is reachable only from inside the container, not from external networks. Set the HTTP client timeout to at least 300 seconds. Checkpoint duration scales with the amount of memory captured - from a few seconds for small CPU workloads to several minutes for large GPU snapshots. ### 3. When the runtime creates a checkpoint When the runtime receives the POST, it checks whether a new checkpoint is required. To save resources, the system skips checkpoint creation if: 1. A checkpoint already exists for the current build version. 2. Another container instance is already undergoing the checkpointing process. If a checkpoint should occur, the container is frozen for the duration of the process. GPU memory is copied to CPU memory, and then all container memory is written to storage. The saved checkpoint is then distributed throughout the region. Checkpoint size roughly equals CPU memory in use at trigger time plus GPU memory copied during the freeze. For GPU workloads, expect the snapshot to be on the order of model weights plus runtime overhead unless caches are dropped first. Allocate enough container memory to hold the GPU dump in addition to normal usage — see **Memory overhead** under Limitations. ### 4. Verify restoration If checkpoint creation succeeds, subsequent containers restore from that snapshot. A restored container logs `CEREBRIUM_RESTORED: container restored from checkpoint` as its first log line. ### 5. Disable checkpointing A checkpoint is tightly coupled to a single deployment. To stop restoring from checkpoints, remove the POST request and redeploy the application. You can find several implementations in our [Examples repository on Github](https://github.com/CerebriumAI/examples). ### vLLM example ```python theme={null} from vllm import AsyncLLMEngine from vllm.engine.arg_utils import AsyncEngineArgs import http import urllib # Init vLLM engine engine_args = AsyncEngineArgs( model="Qwen/Qwen2.5-0.5B-Instruct", async_scheduling=False, # Async scheduling is incompatible with checkpoint restore sleep_mode=True, # Required for engine.sleep() / wake_up() to drop KV cache before checkpoint ) engine = AsyncLLMEngine.from_engine_args(engine_args) # Drop KV cache for reduced GPU memory footprint. engine.sleep(level=1) # Trigger checkpoint try: import json req = urllib.request.Request("http://169.254.169.253:8234/checkpoint", method="POST") with urllib.request.urlopen(req, timeout=300) as response: result = response.read() print(json.loads(result)) except http.client.RemoteDisconnected: # TCP connections disconnect on restore and throw remote pass # Restore KV cache engine.wake_up() ``` ## Limitations **Memory overhead:** The container memory allocation must be large enough to contain the GPU memory dump in addition to your regular memory use. **Execution lifecycle:** When a container is restored from a checkpoint, execution continues from the point where the HTTP request was sent. Any environment variables read before this point remain the same as they were at the time of the checkpoint. **Network connections:** Any TCP connections made before the checkpoint will have disconnected. For example, if you connected to a database before the checkpoint, you must reestablish that connection after restore. **Ephemeral filesystem:** Any files written to disk before the checkpoint are not copied to the restored container. Only memory is checkpointed. ## Platform-specific recommendations ### vLLM vLLM checkpointing support is not complete but is still possible. See [vllm-project/vllm#34303](https://github.com/vllm-project/vllm/issues/34303) and related issues. The larger the size of the memory checkpoint, the slower the restore is. Reduce the size of the snapshot substantially and improve startup times by dropping the KV Cache before checkpoint and recreating it after restore. vLLM has functionality that does this built in as part of [vLLM Sleep Mode](https://docs.vllm.ai/en/latest/features/sleep_mode/). # Faster Cold Start Performance Source: https://cerebrium.ai/docs/performance/faster-cold-starts Diagnose and reduce cold start latency on Cerebrium by cutting queueing delay, speeding model weight loading, and keeping warm containers available. Cold starts happen when Cerebrium must boot a new container to serve a request. That adds latency in two phases: 1. **Queueing** - No warm container is available, so the request waits until a new one starts and becomes ready. 2. **Initialization** - The new container runs startup work before it accepts traffic: importing dependencies, loading model weights into GPU memory, compiling CUDA kernels. Use metrics and request logs from your Cerebrium dashboard to see which phase dominates and how you can improve it. However, most production workloads benefit from both the reduction of initialization and then configuring scaling to keep warm capacity available. ## Reducing initialization time Most cold-start time in ML workloads comes from loading model weights into GPU memory. For large models, a standard Hugging Face load can take 40+ seconds even when reading from storage at \~2 GB/s. Work through the techniques below in order. Each step targets a different part of startup and can be combined with the others. ### Store model weights on persistent storage Store model weights on [persistent storage](/docs/storage/managing-files) at `/persistent-storage` rather than baking them into the container image. Cerebrium caches reads from persistent storage within each region. After weights are loaded once, future cold starts can reuse the cached copy and load faster. For apps running in [multiple regions](/docs/deployments/multi-region-deployment), each region fills its own cache independently. Store weights on the global volume at `/global-persistent-storage` to make them available in every region without a per-region copy. This is usually the best default for large models. Baking weights into the container increases image size, which means Cerebrium has to pull and restore a larger image before your application can start initialization. Only include weights in the container image when they are small enough that the image remains lightweight. Increasing CPU core count can parallelise reads from storage and improve pull-through times for large files. Multiple cores process different parts simultaneously, reducing overall transfer time. ### Run initialization at module scope Move as much initialization work as possible out of the request path and into module scope so it runs once at container start, before the container accepts traffic. ```python theme={null} # Runs once at container start, not on every request model = load_model("/persistent-storage/models/my-model/") tokenizer = load_tokenizer("/persistent-storage/models/my-model/") def predict(prompt: str): return model.generate(prompt) ``` For multiple independent models or weight files, load them concurrently rather than sequentially. Use `ThreadPoolExecutor` or similar patterns to read files in parallel and take full advantage of storage bandwidth. ### Load weights directly to GPU Standard PyTorch and Hugging Face loading paths copy weights through CPU memory. Libraries that stream weights directly from disk to GPU reduce this overhead. Use one of these when model loading remains the bottleneck after moving work to module scope. #### Tensorizer [Tensorizer](https://github.com/coreweave/tensorizer) serialises model weights into a format optimised for fast transfer and loads them directly into GPU memory in a single step. It works with Cerebrium persistent storage at nearly 2 GB/s read speed. For large models (20B+ parameters), loading time typically decreases by 30–50%, with greater improvements on larger models. Tensorizer works with Transformers, Diffusers, scikit-learn, or custom PyTorch modules. The only requirement is the ability to initialise an empty model before the deserializer restores weights into it. #### FlashPack [FlashPack](https://github.com/fal-ai/flashpack) loads PyTorch tensors from disk to GPU at high throughput without requiring GPUDirect Storage. Convert a model once, store the `.flashpack` file on persistent storage, then load directly into GPU memory on startup. FlashPack also provides integration mixins for Transformers and Diffusers models. See the [FlashPack repository](https://github.com/fal-ai/flashpack) for conversion and loading patterns. ### Restore from a checkpoint When initialization includes work that does not change between deployments - compiled CUDA kernels, large weight loads, framework setup - [memory checkpointing](/docs/performance/checkpointing) captures CPU and GPU memory state after initialization and restores it on future cold starts. Checkpointing skips repeated initialization work entirely. A container restored from a checkpoint resumes from the point where the checkpoint trigger was sent, with model weights and compiled kernels already in memory. Use checkpointing when Tensorizer or FlashPack still leave multi-minute startup times, or when compiled kernels dominate initialization. Enable checkpointing in `cerebrium.toml` and trigger it after initialization completes. See the [Memory Checkpointing guide](/docs/performance/checkpointing) for configuration, trigger endpoints, and framework-specific recommendations. ## Reduce queueing with scaling When initialization is already optimised, keep warm containers available so requests do not wait for new ones to boot. See [Scaling Apps](/docs/scaling/scaling-apps) for full parameter reference. Use these scaling options based on traffic pattern: | Goal | Parameter | When to use | | - | - | - | | Eliminate cold starts from scaling to zero | `min_replicas` | Latency-sensitive production workloads that cannot tolerate startup delay | | Handle bursty traffic without waiting for scale-up | `scaling_buffer` | Traffic arrives in bursts where one request is followed by several more | | Keep containers warm through brief dips | `cooldown` | Steady workloads with occasional gaps before traffic returns | | Maintain headroom before autoscaler adds replicas | `scaling_target` | Workloads using `concurrency_utilization` that need spare capacity per replica | ### Keep containers warm with `min_replicas` Set `min_replicas` to maintain a floor of running instances at all times. This eliminates cold starts from scaling to zero but increases cost while idle. ```toml theme={null} [cerebrium.scaling] min_replicas = 1 ``` Use `min_replicas = 1` or higher for latency-sensitive production workloads that cannot tolerate cold-start delay. ### Buffer capacity with `scaling_buffer` `scaling_buffer` provisions extra idle replicas above what the scaling metric recommends. This helps with bursty traffic - when one request arrives, additional warm containers are already available for the requests that follow. ```toml theme={null} [cerebrium.scaling] min_replicas = 0 max_replicas = 10 replica_concurrency = 1 scaling_metric = "concurrency_utilization" scaling_target = 100 scaling_buffer = 3 ``` `scaling_buffer` is available with `concurrency_utilization` and `requests_per_second` metrics. ### Tune cooldown for traffic patterns The `cooldown` parameter sets how long reduced concurrency must persist before a container scales down. A longer cooldown keeps containers warm through brief traffic dips and reduces cold starts when traffic returns quickly. ```toml theme={null} [cerebrium.scaling] cooldown = 600 # Keep containers warm for 10 minutes after traffic drops ``` Match cooldown to traffic patterns. Steady workloads with occasional gaps benefit from longer cooldowns. Highly intermittent workloads may accept shorter cooldowns to reduce idle cost. ### Leave headroom with `scaling_target` With `concurrency_utilization`, set `scaling_target` below 100 to maintain excess capacity before the autoscaler adds replicas. For example, `scaling_target = 70` with `replica_concurrency = 1` keeps containers at 70% utilisation, leaving room for new requests without waiting for a scale-up event. ```toml theme={null} [cerebrium.scaling] replica_concurrency = 1 scaling_metric = "concurrency_utilization" scaling_target = 70 ``` All scaling strategies trade cost for latency. Monitor cold start frequency and request latency in the dashboard to find the right balance. # Batching and Concurrency Source: https://cerebrium.ai/docs/scaling/batching-concurrency Tune replica_concurrency and use framework native or custom batching with vLLM or LitServe to boost GPU throughput and cut costs on Cerebrium. ## Understanding Concurrency Each instance can process multiple requests simultaneously. The `replica_concurrency` setting in `cerebrium.toml` determines how many requests each instance handles in parallel: ```toml theme={null} [cerebrium.scaling] replica_concurrency = 4 # Process up to 4 requests simultaneously. ``` Requests arriving at an instance below its concurrency limit begin processing immediately. Once an instance reaches its maximum, additional requests queue until capacity becomes available. GPUs excel at parallel processing, so concurrent request handling utilizes GPU resources more efficiently than sequential processing. ## Understanding Batching Batching determines how concurrent requests are grouped and executed within an instance. Concurrency controls the number of simultaneous requests; batching controls how those requests are processed together. The default concurrency is 1 request per container for GPU apps and 100 for CPU-only apps. Cerebrium supports two approaches to request batching. ### Framework-native Batching Many frameworks handle batched processing natively. vLLM, for example, automatically batches model inference requests: ```toml theme={null} [cerebrium.scaling] min_replicas = 0 max_replicas = 2 cooldown = 10 replica_concurrency = 4 # Each container can now handle multiple requests. [cerebrium.dependencies.pip] sentencepiece = "latest" torch = "latest" vllm = "latest" transformers = "latest" accelerate = "latest" xformers = "latest" ``` When multiple requests arrive, vLLM combines them into optimal batch sizes and processes them together, maximizing GPU utilization. Check out the complete [vLLM batching example](https://github.com/CerebriumAI/examples/tree/master/10-batching/3-vllm-batching-gpu) for more information. ### Custom Batching Implement custom batching through Cerebrium's [custom runtime feature](/docs/container-images/defining-container-images#custom-runtimes) for precise control over request processing and custom batching strategies. LitServe implementation requires additional configuration in `cerebrium.toml`: ```toml theme={null} [cerebrium.runtime.custom] port = 8000 entrypoint = ["python", "app/main.py"] healthcheck_endpoint = "/health" readycheck_endpoint = "/ready" [cerebrium.dependencies.pip] litserve = "latest" fastapi = "latest" ``` Check out the complete [Litserve example](https://github.com/CerebriumAI/examples/tree/master/10-batching/2-litserve-batching-gpu) for more information. Custom batching provides full control over request grouping and processing. This is particularly useful for frameworks without native batching support. The [Container Images Guide](/docs/container-images/defining-container-images#custom-runtimes) provides detailed implementation instructions. Concurrency enables parallel request handling; batching optimizes how those requests are processed. Together, they improve resource utilization and throughput. # Preemption and Graceful Termination Source: https://cerebrium.ai/docs/scaling/graceful-termination Handle SIGTERM signals in custom runtimes on Cerebrium to finish in flight requests, avoid 502 errors, and shut down instances gracefully. ## Graceful Termination Cerebrium runs in a shared, multi-tenant environment. The platform continuously adjusts capacity: spinning down nodes and launching new ones to scale, optimize compute usage, and roll out updates. The platform migrates workloads to new nodes during this process. Applications also have metric-based autoscaling criteria that dictate when instances scale, remain active, or shift during deployments. Implement graceful termination to prevent requests from ending prematurely when instances are marked for termination. ## Understanding Instance Termination For both application autoscaling and internal node scaling, the platform sends a SIGTERM signal to warn the application of an impending shutdown. Cortex applications (Cerebrium's default runtime) handle this automatically. Custom runtimes must catch and handle this signal to shut down gracefully. Once `response_grace_period` elapses, the platform sends a SIGKILL signal, terminating the instance immediately. When Cerebrium terminates a container, the following sequence occurs: 1. Stop routing new requests to the container. 2. Send a SIGTERM signal to the container. 3. Wait for `response_grace_period` seconds to elapse. 4. Send SIGKILL if the container hasn't stopped. The following diagram illustrates this flow: ```mermaid theme={null} flowchart TD A[SIGTERM sent] --> B[Cortex] A --> C[Custom Runtime] B --> D[automatically captured] C --> E[User needs to capture] D --> F[request finishes] D --> G[response_grace_period reached] E --> H[User logic] F --> I[Graceful termination] G --> J[SIGKILL] H --> O[Graceful termination] H --> G[response_grace_period reached] J --> K[Gateway Timeout Error] ``` Without SIGTERM handling in a custom runtime, Cerebrium terminates containers immediately after sending `SIGTERM`, which interrupts in-flight requests and causes **502 errors**. ## Implementation For custom runtimes using FastAPI, implement the [`lifespan` pattern](https://fastapi.tiangolo.com/advanced/events/) to respond to SIGTERM. ### main.py Create a new file named `main.py` with the following: ```python theme={null} from fastapi import FastAPI, HTTPException, Request from contextlib import asynccontextmanager import asyncio from dataclasses import dataclass @dataclass class AppState: active_requests: int = 0 shutting_down: bool = False shutdown_event: asyncio.Event = None lock: asyncio.Lock = None def __post_init__(self): if self.shutdown_event is None: self.shutdown_event = asyncio.Event() if self.lock is None: self.lock = asyncio.Lock() @asynccontextmanager async def lifespan(app: FastAPI): app.state.app_state = AppState() yield state = app.state.app_state state.shutting_down = True if state.active_requests > 0: try: await asyncio.wait_for(state.shutdown_event.wait(), timeout=30.0) except asyncio.TimeoutError: pass app = FastAPI(lifespan=lifespan) @app.middleware("http") async def track_requests(request: Request, call_next): state: AppState = request.app.state.app_state async with state.lock: if state.shutting_down: raise HTTPException(status_code=503, detail="Service shutting down") state.active_requests += 1 try: response = await call_next(request) return response finally: async with state.lock: state.active_requests -= 1 if state.shutting_down and state.active_requests == 0: state.shutdown_event.set() @app.get("/ready") async def ready(request: Request): state: AppState = request.app.state.app_state async with state.lock: if state.shutting_down: raise HTTPException(status_code=503, detail="Not ready - shutting down") return {"status": "ready", "active_requests": state.active_requests} @app.post("/hello") def hello(): return {"message": "Hello Cerebrium!"} ``` ### cerebrium.toml If you already have a `cerebrium.toml` file, add or update these sections. If you don't have one, create a new file with the following: ```toml theme={null} [cerebrium.deployment] name = "your-app-name" include = ["./*"] exclude = [".*"] [cerebrium.hardware] cpu = 1 memory = 1.0 compute = "CPU" gpu_count = 0 [cerebrium.scaling] min_replicas = 0 max_replicas = 2 cooldown = 30 replica_concurrency = 1 # Must match max_concurrency in your AppState [cerebrium.runtime.custom] port = 8000 dockerfile_path = "./Dockerfile" ``` **Key configuration notes:** * `replica_concurrency` should match `max_concurrency` in your `AppState` class (if you add that field) * `port` must match the port in your Dockerfile CMD * Adjust hardware settings based on your application needs ### Dockerfile Create a new file named `Dockerfile` with the following: ```dockerfile theme={null} FROM python:3.12-bookworm COPY . . RUN pip install -r requirements.txt EXPOSE 8000 CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"] ``` ### requirements.txt Add to your existing `requirements.txt` or create a new file with: ``` fastapi[standard] ``` ## Key Points * The `/ready` endpoint is essential for proper load balancing during scaling events * Without a proper `/ready` endpoint, Cerebrium uses TCP ping which only checks if the port is open, potentially routing traffic to replicas that are shutting down * All request tracking uses asyncio locks to ensure thread safety Test SIGTERM handling locally before deploying: start your app, send SIGTERM with `Ctrl+C`, and verify you see graceful shutdown logs. # Autoscaling apps on Cerebrium Source: https://cerebrium.ai/docs/scaling/scaling-apps Configure replicas, scaling metrics, load balancing, and compute tiers to scale Cerebrium apps for cost, latency, and availability. Cerebrium's scaling system automatically manages computing resources to match app demand, from single requests to multiple simultaneous requests, optimizing for both performance and cost. ## How Autoscaling Works The scaling system monitors a configurable metric - such as concurrency utilization, requests per second, CPU usage, or memory usage - and compares it against a target threshold. When the metric exceeds the target, new instances start within seconds. See [Scaling Metrics](#using-scaling-metrics) for details on each option. ## Scaling Configuration The `cerebrium.toml` file controls scaling behavior through several key parameters: ```toml theme={null} [cerebrium.scaling] min_replicas = 0 # Minimum running instances max_replicas = 3 # Maximum concurrent instances cooldown = 60 # Cooldown period in seconds replica_concurrency = 1 # The maximum number of requests each replica of an app can accept ``` ### Minimum Instances The `min_replicas` parameter defines how many instances remain active at all times. Setting this to 1 or higher maintains warm instances for immediate response, eliminating cold starts but increasing costs. Use this for apps requiring consistent response times or specific SLA guarantees. ### Maximum Instances The `max_replicas` parameter sets an upper limit on concurrent instances, controlling costs and protecting backend systems. When traffic increases, new instances start automatically up to this configured maximum. For apps running in multiple regions, `min_replicas` and `max_replicas` apply to the app as a whole rather than per region. The platform distributes instances across eligible regions based on traffic and available capacity. See [Multi-Region Deployment](/docs/deployments/multi-region-deployment). ### Cooldown Period The `cooldown` parameter specifies the time window (in seconds) that must pass at reduced concurrency before an instance scales down. This prevents premature scale-down during brief traffic dips that might be followed by more requests. A longer cooldown period helps handle bursty traffic patterns but increases instance running time and cost. ### Replica Concurrency The number of requests an app instance can handle concurrently is dictated by the `replica_concurrency` parameter. This is a hard limit, and an individual replica will not accept more than this limit at a time. By default, once this concurrency limit is reached on an instance and there are still requests to be processed in-flight, the system will scale out by the number of new instances required to fulfil the in-flight requests. For example, if `replica_concurrency=1` and there are *3* requests in flight with no replicas currently available, Cerebrium will scale out 3 instances of the application to meet that demand. Typically most GPU applications will require that `replica_concurrency` is set to **1**. If the workload requires GPU but higher throughput is desired, `replica_concurrency` may be increased so long as access to GPU resources is controlled within the application through batching. ## Processing Multiple Requests Apps can process multiple requests simultaneously through batching and concurrency. Cerebrium supports frameworks with built-in batching and enables custom implementations through the [custom runtime](/docs/container-images/defining-container-images#custom-runtimes) feature. See the [Batching & Concurrency Guide](/docs/scaling/batching-concurrency) for details. ## Instance Management Cerebrium automatically restarts failed instances, starts new instances to maintain capacity, and monitors instance health continuously. Apps requiring maximum reliability combine several scaling features: ```toml theme={null} [cerebrium.scaling] min_replicas = 2 # Maintain redundant instances cooldown = 600 # Extended warm period max_replicas = 10 # Room for traffic spikes response_grace_period = 1200 # Maximum request lifespan ensuring graceful exit ``` The `response_grace_period` parameter stipulates how long in seconds a request would need at most to finish, and provides time for instances to complete active requests during normal operation and shutdown. During normal replica operation, this acts as a request timeout value. During replica shutdown, the Cerebrium system sends a SIGTERM signal to the replica, waits for the specified grace period, issues a SIGKILL command if the instance has not stopped, and kills any active requests with a GatewayTimeout error. When using the Cortex runtime (default), SIGTERM signals are automatically handled to allow graceful termination of requests. For custom runtimes, you'll need to implement SIGTERM handling yourself to ensure requests complete gracefully before termination. See our [Graceful Termination guide](/docs/scaling/graceful-termination) for detailed implementation examples, including FastAPI patterns for tracking and completing in-flight requests during shutdown. Performance metrics available through the dashboard help monitor scaling behavior: * Request processing times * Active instance count * Cold start frequency * Resource usage patterns ## Using Scaling Metrics Cerebrium supports multiple scaling metrics beyond the default `replica_concurrency`. Four scaling metrics are available: * `concurrency_utilization` * `requests_per_second` * `cpu_utilization` * `memory_utilization` Specify a metric and target in the `cerebrium.scaling` section: ```toml theme={null} [cerebrium.scaling] min_replicas = 0 cooldown = 600 max_replicas = 10 response_grace_period = 120 replica_concurrency = 1 scaling_metric = "concurrency_utilization" scaling_target = 100 ``` ### Concurrency Utilization `concurrency_utilization` is the default scaling metric, with a default target of *100%*. This metric maintains a maximum percentage of `replica_concurrency` averaged across all instances. For example, with `replica_concurrency=1` and `scaling_target=70`, Cerebrium maintains *0.7* requests per instance, ensuring 30% excess capacity. With `replica_concurrency=200` and `scaling_target=80`, Cerebrium maintains *160* requests per instance and scales out once that target is exceeded. ### Requests per Second `requests_per_second` maintains a maximum application throughput measured in requests per second, averaged over all instances. This metric is more effective than `concurrency_utilization` when application throughput has been benchmarked. It does not enforce concurrency limits, so it is not recommended for most GPU applications. For example, `scaling_target=5` maintains a 5 requests/s average across all instances. ### CPU Utilization `cpu_utilization` scales based on maximum CPU percentage utilization averaged over all instances, relative to the `cerebrium.hardware.cpu` value. For example, with `cpu=2` and `scaling_target=80`, Cerebrium maintains *80%* CPU utilization (1.6 CPUs) per instance. This metric requires `min_replicas=1` since scaling relative to 0 CPU units is undefined. ### Memory Utilization `memory_utilization` scales based on maximum RAM percentage utilization averaged over all instances, relative to `cerebrium.hardware.memory`. This refers to RAM, **not** GPU VRAM. For example, with `memory=10` and `scaling_target=80`, Cerebrium maintains *80%* memory utilization (8GB) per instance. This metric requires `min_replicas=1` since scaling relative to 0GB of memory is undefined. ## Keeping a Scaling Buffer For apps with long startup times or predictable traffic, a replica buffer maintains consistent excess capacity above what the scaling metric suggests. The `scaling_buffer` option adds a fixed number of extra replicas to the autoscaler's recommendation. This is available with the following scaling metrics: * `concurrency_utilization` * `requests_per_second` Add `scaling_buffer` to the `cerebrium.scaling` section: ```toml theme={null} [cerebrium.scaling] min_replicas = 1 cooldown = 600 max_replicas = 10 response_grace_period = 120 replica_concurrency = 1 scaling_metric = "concurrency_utilization" scaling_target = 100 scaling_buffer = 3 ``` With the above config: when no traffic is received, the app runs **1 replica** as a baseline. The buffer scales based on requests. With `concurrency_utilization` at `100` and `replica_concurrency=1`, receiving 1 request causes the autoscaler to suggest 1 replica. With `scaling_buffer=3`, the app scales to **(1+3)=4** replicas. The buffer adds a static number of replicas on top of the autoscaler's recommendation. After the request completes, the `cooldown` period applies and the replica count scales back to the **1 replica** baseline. ## Evaluation Interval Requires CLI version 2.1.5 or higher. The `evaluation_interval` parameter controls the time window (in seconds) over which the autoscaler evaluates metrics before making scaling decisions. The default is 30 seconds, with a valid range of 6-300 seconds. ```toml theme={null} [cerebrium.scaling] evaluation_interval_seconds = 30 # Evaluate metrics over 30-second windows ``` A shorter interval makes the autoscaler more responsive to traffic spikes but may cause more frequent scaling events. A longer interval smooths out transient spikes but may delay scaling responses. For bursty workloads, a shorter `evaluation_interval` (e.g., 10-15 seconds) helps the system respond quickly to demand. For steady workloads, a longer interval (e.g., 60 seconds) reduces unnecessary scaling churn. ## Load Balancing Requires CLI version 2.1.5 or higher. The `load_balancing_algorithm` parameter controls how incoming requests are distributed across your replicas. When not specified, the system automatically selects the best algorithm based on your `replica_concurrency` setting. ```toml theme={null} [cerebrium.scaling] load_balancing_algorithm = "min-connections" # Explicitly set load balancing algorithm ``` **Default behavior**: When `load_balancing_algorithm` is not set, the system uses `first-available` for `replica_concurrency <= 3` (typical for GPU workloads) and `round-robin` for higher concurrency. ### Available Algorithms #### round-robin Cycles through replicas starting from the last successful target. Each replica's concurrency limit is respected - if a replica is at capacity, the algorithm proceeds to the next one in rotation. | Characteristic | Value | | - | - | | Selection complexity | O(1) typical, O(N) worst case when scanning for available capacity | | Latency profile | Consistent p50, good p90 under uniform load | | Strategy | Stateful index rotation with mutex synchronization; skips full replicas | **Best for**: Workloads with predictable request times where you want even distribution across replicas over time. #### first-available Scans replicas from the start of the list and selects the first one with available capacity. | Characteristic | Value | | - | - | | Selection complexity | O(1) typical, O(N) worst case | | Latency profile | Optimal p50 when load is light, may degrade p90 under high load | | Strategy | Linear scan from list start; returns first replica that accepts via Reserve() | **Best for**: GPU workloads with low concurrency (`replica_concurrency <= 3`). Maximizes utilization of warm replicas before spreading load, reducing cold starts and keeping models in VRAM. **Tradeoff**: Earlier replicas in the list handle more traffic. This is desirable for GPU workloads but may cause uneven distribution for CPU workloads. #### min-connections Linear scan to find the replica with the fewest in-flight requests, then attempts to reserve it. If that replica cannot accept (at capacity), falls back to trying other replicas in iteration order. | Characteristic | Value | | - | - | | Selection complexity | Θ(N) - always scans all replicas to find minimum | | Latency profile | Best p90/p99 tail latency | | Strategy | Single pass to find minimum in-flight; fallback in iteration order | **Best for**: Workloads with variable request times (e.g., LLM inference where output length varies). Routes new requests to the least busy replica, preventing fast requests from queuing behind slow ones. #### random-choice-2 Implements the "Power of Two Choices" algorithm: randomly samples two replicas and routes to the one with lower weight (based on active request tracking). Ties are broken randomly. | Characteristic | Value | | - | - | | Selection complexity | Θ(1) - constant time regardless of replica count | | Latency profile | Good balance of p50 and p90 | | Strategy | Sample 2 random replicas, compare weights, pick lighter one | **Best for**: High-throughput scenarios with many replicas where selection overhead matters. Research shows this achieves exponentially better load distribution than pure random selection. **Note**: Uses weight-based tracking rather than reservation-based concurrency limiting, making it suitable for unlimited concurrency scenarios. ### Choosing an Algorithm | Scenario | Recommended | Reason | | - | - | - | | GPU inference, `replica_concurrency=1` | `first-available` (default) | Maximizes GPU utilization, keeps models warm | | LLMs with variable output lengths | `min-connections` | Prevents head-of-line blocking, best tail latency | | High-throughput, many replicas | `random-choice-2` | Θ(1) selection with near-optimal distribution | | Uniform request times, even distribution | `round-robin` | Predictable rotation, no hot spots over time | | Latency-sensitive with variable load | `min-connections` | Minimizes p90/p99 by routing to least busy replica | ## Compute Tier Requires CLI version 2.1.6 or higher. The `compute_tier` parameter sets the interruption guarantee for an app's instances: whether they run only on capacity that cannot be reclaimed (protected) or on any available capacity, including preemptible instances (interruptible). This directly affects cost and availability. ```toml theme={null} [cerebrium.scaling] compute_tier = "protected" # No interruptions, billed at 2x the interruptible rate ``` ### Available Tiers | Tier | Description | Price | | - | - | - | | `interruptible` | **(Default)** Runs on any available capacity, preferring preemptible instances. Lower cost, but may be interrupted by the cloud provider or relocated during consolidation. | Base rate | | `protected` | Runs only on capacity that cannot be reclaimed. No interruptions, excluded from capacity consolidation. | 2x the interruptible rate | ### Pricing Protected compute is billed at **2x the interruptible rate**. The multiplier applies to all compute on the instance: GPU, CPU, and memory. Persistent storage is priced independently of the compute tier. The rates on the [pricing page](https://www.cerebrium.ai/pricing) are interruptible rates, and the cost estimate shown on the dashboard reflects the configured tier. The tier is fixed when an instance starts. Changing `compute_tier` applies to instances started after the next deployment; instances already running keep the tier, and the rate, they started with. Interruptible instances run on preemptible capacity when it is available but can be scheduled on on-demand capacity during shortages. They are billed at the interruptible rate regardless of the underlying capacity. ### Interruptions and Consolidation Interruptible instances can be interrupted in two ways: the cloud provider can reclaim preemptible capacity during periods of high demand, and the platform periodically consolidates workloads onto fewer nodes to reduce cost, which can relocate interruptible instances regardless of the capacity they run on. In both cases instances are drained gracefully within `response_grace_period`; see [Graceful Termination](/docs/scaling/graceful-termination). Protected instances are excluded from capacity consolidation and never run on preemptible capacity, so they are not interrupted in either case. ### Node Maintenance and Updates The `compute_tier` setting governs preemption and consolidation only. It does not opt an instance out of node lifecycle events such as OS patches, kernel upgrades, or platform rollouts. Both interruptible and protected instances can be migrated to a new node when the underlying node is cycled for maintenance. In every case the instance receives SIGTERM and is drained within `response_grace_period` before SIGKILL. See [Graceful Termination](/docs/scaling/graceful-termination) for the shutdown sequence. Custom runtimes must handle SIGTERM themselves so in-flight requests complete instead of returning 502s. To minimize disruption during a migration, run with `min_replicas >= 2` so at least one replica remains available while another is drained. Set `response_grace_period` to cover your worst-case request duration. **Choosing a tier:** * Use `interruptible` (default) for batch workloads, development, or cost-sensitive applications that can tolerate occasional interruptions. * Use `protected` for production services with strict availability requirements or long-running requests where interruption would be costly. # Security & Data Privacy Source: https://cerebrium.ai/docs/security Cerebrium's security practices, data privacy commitments, and support for SOC 2, HIPAA, GDPR (DPA), and ISO compliance requirements. Cerebrium is SOC 2 Type I, HIPAA-compliant, GDPR and ISO compliant, enforcing strict security standards and protocols. Compliance is continually monitored through Vanta and a dedicated team. Visit the [trust center](https://trust.cerebrium.ai) for compliance reports, or contact [security@cerebrium.ai](mailto:security@cerebrium.ai) for additional information. ## Infrastructure Security * Cerebrium frequently performs vulnerability scans, with remediation following the incident response plan timelines. * Cerebrium conducts annual business continuity and security incident exercises as required for SOC 2 compliance. * Cerebrium has daily database backups enabled. * Employee computers are frequently monitored via the Vanta agent. * Multi-Factor Authentication (MFA) is enforced across all platforms relating to Cerebrium. * Cerebrium uses logging and metrics observability providers, including Datadog and BugSnag. ## Organizational Security * Cerebrium employees are subject to a general security awareness training during their onboarding period. * Cerebrium regularly audits employee access to internal systems. * Employee computers are frequently monitored via the Vanta agent. * Multi-Factor Authentication (MFA) is enforced across all platforms relating to Cerebrium. ## Product Security * Cerebrium enforces HTTPS for all services using TLS (SSL), including the Cerebrium Dashboard and Python package. * Cerebrium maintains access logs across all its infrastructure services. * Software dependencies are audited by GitHub's Dependabot. * User data is encrypted at rest. ## Internal Security Procedures * Cerebrium performs regular vulnerability scans, with remediation following incident response plan timelines. * Cerebrium regularly audits employee access to internal systems. * Cerebrium conducts annual business continuity and security incident exercises as part of SOC 2 compliance requirements. ## Data and Privacy * Cerebrium does not use customer data to train machine learning models. * For customers on the Hobby and Standard plans, request/log data is automatically deleted after 7 and 30 days, respectively. * Cerebrium deletes customer data upon request. A purge request endpoint is available for immediate deletion. * All user data is encrypted at rest. ## GDPR Compliance Cerebrium supports customer obligations under the General Data Protection Regulation (GDPR). ### Data Processing Agreements (DPA) * Cerebrium offers a standardized DPA to customers who process personal data of individuals in the EU, UK, or other jurisdictions with equivalent requirements. * The DPA covers the roles of controller and processor, the categories of data processed, security measures, sub-processors, and international data transfer safeguards. * Customers can request a DPA by contacting [compliance@cerebrium.ai](mailto:compliance@cerebrium.ai). * Executed DPAs and other compliance documents are also available through the [trust center](https://trust.cerebrium.ai). ## HIPAA Compliance As a business associate to covered entities in the healthcare sector, Cerebrium implements the following measures to support HIPAA compliance: ### Business Associate Agreements (BAA) * Cerebrium offers a standardized BAA to all customers who require HIPAA compliance. * The BAA outlines the responsibilities and obligations of both parties in protecting Protected Health Information (PHI). * Customers can initiate the BAA process by contacting [compliance@cerebrium.ai](mailto:compliance@cerebrium.ai). ### PHI Handling and Storage * Cerebrium's infrastructure is designed to handle PHI securely, with encryption at rest and in transit. * Cerebrium does not access, use, or disclose PHI unless explicitly required for service delivery. * Customers are responsible for de-identifying PHI before transmission to Cerebrium's systems, if de-identification is required for their use case. ### Access Controls * Strict access controls are in place to ensure that only authorized personnel can access systems that may contain PHI. * Role-based access controls are used to limit access to PHI based on job responsibilities and the principle of least privilege. ### Audit Logging * Comprehensive audit logs are maintained for all activities that could potentially involve PHI. * These logs are available to support customers' accounting of disclosures requirements. ### Breach Notification * Cerebrium maintains an incident response plan that includes HIPAA-compliant breach notification procedures. * Any potential breaches involving PHI are promptly investigated and reported to affected customers within required timeframes. ### Employee Training * All Cerebrium employees undergo HIPAA awareness training as part of their onboarding process. * Regular refresher training is conducted to ensure ongoing HIPAA compliance. ### Risk Assessments * Cerebrium conducts regular risk assessments to identify and address potential vulnerabilities in PHI handling. * These assessments help maintain a secure environment for customer data. ### Subcontractors * Any subcontractors who may have access to PHI are required to sign a BAA and comply with the same HIPAA requirements as Cerebrium. ### Data Retention and Destruction * Cerebrium adheres to HIPAA-compliant data retention policies. * Secure data destruction processes are in place for when PHI needs to be deleted or when a customer relationship ends. ### Compliance Monitoring * HIPAA compliance measures are continuously monitored and updated to align with changes in regulations and best practices. For more information on HIPAA compliance or specific compliance needs, contact the compliance team at [compliance@cerebrium.ai](mailto:compliance@cerebrium.ai). # Managing Files Source: https://cerebrium.ai/docs/storage/managing-files Manage files on Cerebrium persistent storage with CLI upload, download, and list commands, plus global volumes and resizing options up to 1TB. Cerebrium offers file management through a persistent volume that's available to all apps in a project. This storage mounts at `/persistent-storage` and helps store model weights and files efficiently across deployments. The `/persistent-storage` directory is not directly accessible during build time. Use the Cerebrium CLI commands or the mounted volume at runtime to manage these files. ## Including Files in Deployments The `cerebrium.toml` configuration file controls which files become part of the app: ```toml theme={null} [cerebrium.deployment] include = [ "src/*.py", # Python files in src. "config/*.json", # JSON files in config. "requirements.txt" # Specific files. ] exclude = [ "tests/*", # Skip test files. "*.log" # Skip log files. ] ``` Files included in deployments must be under 2GB each, with deployments working best for files under 1GB. Larger files should use persistent storage instead. ## Managing Persistent Storage The CLI provides four commands for working with persistent storage. At runtime, the volume is mounted at `/persistent-storage`. When using these commands, the Cerebrium CLI does not display the `/persistent-storage/` portion of the path. Persistent storage is managed per region. Files written to `/persistent-storage` in one region are not guaranteed to be available in other regions; store data that must be available everywhere on [global storage](#global-storage). File commands operate on the default region set by `cerebrium region set` (`us-east-1` for new accounts). To target a different region for a single command, pass `--region` (or `-r`): ```bash theme={null} cerebrium ls --region eu-north-1 cerebrium cp model.bin -r eu-north-1 ``` 1. Upload files with `cerebrium cp`: ```bash theme={null} # Upload to root directory cerebrium cp src_file_name.txt # Upload to specific location cerebrium cp src_file_name.txt dest_file_name.txt # Upload to directory cerebrium cp dir_name sub_folder/ ``` 2. List files with `cerebrium ls`: ```bash theme={null} # List root contents cerebrium ls # List specific folder cerebrium ls sub_folder/ ``` 3. Remove files with `cerebrium rm`: ```bash theme={null} # Remove a file cerebrium rm file_name.txt # Remove a directory (directory paths must end with a forward slash) cerebrium rm folder_name/ ``` 4. Download files with `cerebrium download`: ```bash theme={null} # Download to current directory with same filename cerebrium download file_name.txt # Download with a different local filename cerebrium download file_name.txt local_file_name.txt # Download from a subdirectory cerebrium download sub_folder/file_name.txt ``` ## Using Stored Files Access files in persistent storage at runtime: ```python theme={null} import os import torch # Load a model from persistent storage. file_path = "/persistent-storage/segment-anything/sam_vit_h_4b8939.pth" model = torch.jit.load(file_path) ``` Applications access files using the full `/persistent-storage/` path at runtime: ```python theme={null} # Read a file from persistent storage with open("/persistent-storage/data/config.json", "r") as f: config = json.load(f) # Write to persistent storage with open("/persistent-storage/data/results.json", "w") as f: json.dump(results, f) # Check if file exists in persistent storage if os.path.exists("/persistent-storage/models/mymodel.pt"): model = torch.load("/persistent-storage/models/mymodel.pt") ``` Remember that while the CLI commands don't display the `/persistent-storage/` prefix in their output, your code must use the full path to access these files at runtime. ## Global Storage Apps deployed with `region = "global"` mount an additional volume at `/global-persistent-storage`. It is a single volume shared across every region the app runs in: files written there are visible from all regions, and reads are cached per region. See [Multi-Region Deployment](/docs/deployments/multi-region-deployment). Manage files on the global volume by passing `--region global` to the file commands: ```bash theme={null} cerebrium cp model.bin --region global cerebrium ls --region global ``` Use the global volume for data that must be available everywhere, such as model weights for a global app. Use `/persistent-storage` for region-local data. The global volume is 1TB. It does not appear in the dashboard Volumes list and cannot be resized. ## Increasing Storage Capacity Persistent volumes can be resized up to 1TB through self-service. Storage costs are listed on the [pricing page](https://www.cerebrium.ai/pricing). 1. **Via Dashboard**: Navigate to the project and select **Volumes**. Click the three-dot menu on the volume and select **Increase Size**. Volume resize option in the Cerebrium dashboard 2. **Via API**: Use the resize volume endpoint. The volume ID is `default-`. Sizes range from 50GB to 1TB, and shrinking a volume is not supported: ```bash theme={null} curl -X PATCH "https://rest.cerebrium.ai/v2/projects/{project_id}/volumes/default-us-east-1/resize" \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{"newSizeGb": 200}' ``` For storage needs beyond 1TB, contact support on [Discord](https://discord.gg/ATj6USmeE2) or via [email](mailto:support@cerebrium.ai). ### Storage Quota Errors If you hit your storage quota, you'll see errors like `No space left on device` when trying to download or write files. To resolve this: 1. Remove unnecessary files from `/persistent-storage` using `cerebrium rm` 2. Resize the volume to a larger capacity (up to 1TB self-service) 3. Contact support for storage beyond 1TB # TOML Reference Source: https://cerebrium.ai/docs/toml-reference/toml-reference Reference for every cerebrium.toml parameter covering deployment, hardware, scaling, dependencies, and custom runtime settings on Cerebrium. The configuration is organized into the following main sections: * **\[cerebrium.deployment]** Core settings like app name, Python version, and file inclusion rules * **\[cerebrium.runtime.custom]** Custom web server settings and app startup behavior * **\[cerebrium.hardware]** Compute resources including CPU, memory, and GPU specifications * **\[cerebrium.scaling]** Auto-scaling behavior and replica management * **\[cerebrium.dependencies]** Package management for Python (pip), system (apt), and Conda dependencies ## Deployment Configuration The `[cerebrium.deployment]` section defines core deployment settings. | Option | Type | Default | Description | | - | - | - | - | | name | string | required | Desired app name | | python\_version | string | "3.12" | Python version to use (3.10, 3.11, 3.12) | | disable\_auth | boolean | true | Disable token-based authentication on app endpoints | | include | string\[] | \["\*"] | Files/patterns to include in deployment | | exclude | string\[] | \[".\*"] | Files/patterns to exclude from deployment | | shell\_commands | string\[] | \[] | Commands to run at the end of the build | | pre\_build\_commands | string\[] | \[] | Commands to run before dependencies install | | docker\_base\_image\_url | string | "debian:bookworm-slim" | Base Docker image | | use\_uv | boolean | false | Use UV for faster Python package installation | | deployment\_initialization\_timeout | integer | 600 (10 minutes) | Max time to wait for app initialisation during build before timing out. Value must be between 60 and 7200 | `disable_auth` defaults to `true`, so omitting it leaves app endpoints callable without a token. Set `disable_auth = false` to require the JWT token described in the [REST API reference](/docs/endpoints/inference-api). Changes to python\_version or docker\_base\_image\_url trigger full rebuilds since they affect the base environment. ### UV Package Manager UV is a fast Python package installer written in Rust that significantly speeds up deployment times. When enabled, UV replaces pip for installing Python dependencies. UV typically installs packages 10-100x faster than pip, especially beneficial for: * Large dependency trees * Multiple packages * Clean builds without cache **Example with UV enabled:** ```toml theme={null} [cerebrium.deployment] use_uv = true ``` ### Monitoring UV Usage Check your build logs for these indicators: * **UV\_PIP\_INSTALL\_STARTED**: UV is successfully being used * **PIP\_INSTALL\_STARTED**: Standard pip installation (when `use_uv` is `false`) While UV is compatible with most packages, some edge cases may cause build failures, such as legacy packages with non-standard metadata. ### Deploying with UV Lock Files Read only if you're using `pyproject.toml` and `uv.lock`. Generate your lock file locally. This creates a uv.lock file with exact dependency versions. ```bash theme={null} # In your project directory with pyproject.toml uv sync ``` Export your locked dependencies to requirements.txt ```bash theme={null} uv pip compile pyproject.toml -o requirements.txt # Or if you want to use the lock file: uv pip compile uv.lock -o requirements.txt ``` Include in your deployment: * Ensure requirements.txt is in your project directory * Deploy with UV enabled ## Runtime Configuration The `[cerebrium.runtime.custom]` section configures custom web servers and runtime behavior. | Option | Type | Default | Description | | - | - | - | - | | port | integer | required | Port the application listens on | | entrypoint | string\[] | required | Command to start the application | | healthcheck\_endpoint | string | "" | HTTP path for health checks (empty uses TCP). Failure causes the instance to restart | | readycheck\_endpoint | string | "" | HTTP path for readiness checks (empty uses TCP). Failure ensures the load balancer does not route to the instance | The port specified in entrypoint must match the port parameter. All endpoints will be available at `https://api.cerebrium.ai/v4/p-xxxxxxxx/your-app-name/your/endpoint` ## Hardware Configuration The `[cerebrium.hardware]` section defines compute resources. | Option | Type | Default | Description | | - | - | - | - | | cpu | float | required | Number of CPU cores | | memory | float | required | Memory allocation in GB | | compute | string or string\[] | "CPU" | Compute type, or a list of acceptable compute types in preference order | | gpu\_count | integer | 0 | Number of GPUs | | provider | string | platform-selected | Cloud provider. Omit to let the platform choose | | region | string | platform-selected | "global" to run in any region with available capacity, or a specific region to pin | Memory refers to RAM, not GPU VRAM. Ensure sufficient memory for your workload. `compute` accepts a single type or a preference-ordered list, e.g. `compute = ["HOPPER_H100", "AMPERE_A100_80GB"]`. See [GPU preference lists](/docs/hardware/using-gpus#gpu-preference-lists). Placement is covered in [Multi-Region Deployment](/docs/deployments/multi-region-deployment). ## Scaling Configuration The `[cerebrium.scaling]` section controls auto-scaling behavior. | Option | Type | Default | CLI Requirement | Description | | - | - | - | - | - | | min\_replicas | integer | 0 | 2.1.2+ | Minimum running instances | | max\_replicas | integer | 1 | 2.1.2+ | Maximum running instances | | replica\_concurrency | integer | 1 (GPU), 100 (CPU) | 2.1.2+ | Concurrent requests per replica. Defaults to 1 for GPU compute types and 100 when compute is CPU | | response\_grace\_period | integer | 900 | 2.1.2+ | Grace period in seconds | | cooldown | integer | 10 | 2.1.2+ | Time window (seconds) that must pass at reduced concurrency before scaling down. Helps avoid cold starts from brief traffic dips. | | scaling\_metric | string | "concurrency\_utilization" | 2.1.2+ | Metric for scaling decisions (concurrency\_utilization, requests\_per\_second, cpu\_utilization, memory\_utilization) | | scaling\_target | integer | 100 | 2.1.2+ | Target value for scaling metric (percentage for utilization metrics, absolute value for requests\_per\_second) | | scaling\_buffer | integer | optional | 2.1.2+ | Additional replica capacity above what scaling metric suggests | | evaluation\_interval\_seconds | integer | 30 | 2.1.5+ | Time window in seconds over which metrics are evaluated before scaling decisions (6-300s) | | load\_balancing\_algorithm | string | "" | 2.1.5+ | Algorithm for distributing traffic across replicas. Default: round-robin if replica\_concurrency > 3, first-available otherwise. Options: round-robin, first-available, min-connections, random-choice-2 | | compute\_tier | string | "interruptible" | 2.1.6+ | Sets the interruption guarantee for app instances. Options: interruptible (may be interrupted or relocated, base rate), protected (no interruptions, billed at 2x the interruptible rate) | | roll\_out\_duration\_seconds | integer | 0 | 2.1.2+ | Gradually send traffic to new revision after successful build. Max 600s. Keep at 0 during development. | Setting min\_replicas > 0 maintains warm instances for immediate response but increases costs. The `scaling_metric` options are: * **concurrency\_utilization**: Maintains a percentage of your replica\_concurrency across instances. For example, with `replica_concurrency=200` and `scaling_target=80`, maintains 160 requests per instance. * **requests\_per\_second**: Maintains a specific request rate across all instances. For example, `scaling_target=5` maintains 5 requests/s average across instances. * **cpu\_utilization**: Maintains CPU usage as a percentage of cerebrium.hardware.cpu. For example, with `cpu=2` and `scaling_target=80`, maintains 80% CPU utilization (1.6 CPUs) per instance. * **memory\_utilization**: Maintains RAM usage as a percentage of cerebrium.hardware.memory. For example, with `memory=10` and `scaling_target=80`, maintains 80% memory utilization (8GB) per instance. The scaling\_buffer option is only available with concurrency\_utilization and requests\_per\_second metrics. It ensures extra capacity is maintained above what the scaling metric suggests. For example, with `min_replicas=0` and `scaling_buffer=3`, the system will maintain 3 replicas as baseline capacity. ## Dependencies ### Pip Dependencies The `[cerebrium.dependencies.pip]` section lists Python package requirements. ```toml theme={null} [cerebrium.dependencies.pip] torch = "==2.0.0" # Exact version numpy = "latest" # Latest version pandas = ">=1.5.0" # Minimum version ``` ### APT Dependencies The `[cerebrium.dependencies.apt]` section specifies system packages. ```toml theme={null} [cerebrium.dependencies.apt] ffmpeg = "latest" libopenblas-base = "latest" ``` ### Conda Dependencies The `[cerebrium.dependencies.conda]` section manages Conda packages. ```toml theme={null} [cerebrium.dependencies.conda] cuda = ">=11.7" cudatoolkit = "11.7" ``` ### Dependency Files The `[cerebrium.dependencies.paths]` section allows using requirement files. ```toml theme={null} [cerebrium.dependencies.paths] pip = "requirements.txt" apt = "pkglist.txt" conda = "conda_pkglist.txt" ``` ## Complete Example ```toml theme={null} [cerebrium.deployment] name = "llm-inference" python_version = "3.12" disable_auth = false include = ["*"] exclude = [".*"] shell_commands = [] pre_build_commands = [] docker_base_image_url = "debian:bookworm-slim" use_uv = true # Enable fast package installation with UV (omit or set to false if you want to use pip) [cerebrium.runtime.custom] port = 8000 entrypoint = ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"] healthcheck_endpoint = "/health" readycheck_endpoint = "/ready" [cerebrium.hardware] cpu = 4 memory = 16.0 compute = "AMPERE_A10" gpu_count = 1 region = "global" [cerebrium.scaling] min_replicas = 0 max_replicas = 2 replica_concurrency = 10 response_grace_period = 900 cooldown = 10 scaling_metric = "concurrency_utilization" scaling_target = 100 evaluation_interval_seconds = 30 # load_balancing_algorithm = "" # Auto-selects based on replica_concurrency # compute_tier = "interruptible" # Use "protected" for no interruptions at 2x the interruptible rate roll_out_duration_seconds = 0 [cerebrium.dependencies.pip] torch = "latest" transformers = "latest" uvicorn = "latest" [cerebrium.dependencies.apt] ffmpeg = "latest" [cerebrium.dependencies.conda] # Optional conda dependencies [cerebrium.dependencies.paths] # Optional paths to dependency files # pip = "requirements.txt" # apt = "pkglist.txt" # conda = "conda_pkglist.txt" ``` # Gradio Chat Interface Source: https://cerebrium.ai/docs/v4/examples/asgi-gradio-interface Deploy a Gradio chat UI for a Llama LLM on Cerebrium with FastAPI and a custom ASGI runtime, running the frontend on CPU while the model scales on GPU. This tutorial covers creating and deploying a Gradio chat interface connected to a Llama 8B language model using Cerebrium's custom ASGI runtime. The architecture runs the frontend on CPU instances while the model runs separately on GPU instances for optimal resource utilization. You can find the full codebase for deploying your Gradio frontend [here](https://github.com/CerebriumAI/examples/tree/master/11-python-apps/2-asgi-gradio-interface). ## Architecture Overview The application consists of two main components: 1. A frontend interface running on CPU instances using FastAPI and Gradio. 2. A separate Llama model endpoint running on GPU instances. (For a comprehensive example for deploying Llama 8B with TensorRT [here](https://docs.cerebrium.ai/v4/examples/tensorRT).) This separation enables: * Keep the frontend always available while minimizing costs (CPU-only). * Scale the GPU-intensive model independently based on demand. * Optimize resource allocation for different components. ## Prerequisites Before starting, you'll need: * A Cerebrium account (sign up [here](https://dashboard.cerebrium.ai/register)). * The Cerebrium CLI installed: `pip install --upgrade cerebrium`. * A Llama model endpoint (or other LLM API endpoint). ## Basic Setup First, create a new directory for your project and initialize it: ``` cerebrium init 2-gradio-interface ``` Add the following configuration to `cerebrium.toml`: ```toml theme={null} [cerebrium.deployment] name = "2-gradio-interface" python_version = "3.12" disable_auth = true include = ['./*', 'main.py', 'cerebrium.toml'] exclude = ['.*'] [cerebrium.runtime.custom] entrypoint = ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8080"] port = 8080 healthcheck_endpoint = "/health" [cerebrium.hardware] cpu = 2 memory = 4.0 compute = "CPU" [cerebrium.scaling] min_replicas = 0 max_replicas = 2 cooldown = 30 replica_concurrency = 10 [cerebrium.dependencies.pip] gradio = "latest" fastapi = "latest" requests = "latest" httpx = "latest" uvicorn = "latest" starlette = "latest" ``` This configuration: * Disables default JWT authentication, making the Gradio interface publicly accessible. * Sets the ASGI server entrypoint to Uvicorn. * Sets the default port to 8080. * Sets the health endpoint to `/health` for availability checks. * Configures hardware settings for the CPU instance. * Defines scaling with min/max replicas, cooldown, and concurrency (10 requests per replica). * Specifies required dependencies: Gradio, FastAPI, Requests, HTTPX, Uvicorn, and Starlette. Set up the main entrypoint file (`main.py`). Start by creating the FastAPI application: ```python theme={null} # at the top of your main.py file import requests from typing import Optional, List import httpx from fastapi import FastAPI, Request from starlette.responses import Response as StarletteResponse app = FastAPI() # Get the Gradio app URL (when running on Cerebrium) GRADIO_HOST = os.getenv("GRADIO_HOST", "127.0.0.1") GRADIO_PORT = int(os.getenv("GRADIO_PORT", "7860")) GRADIO_URL = os.getenv("GRADIO_SERVER_URL", f"http://{GRADIO_HOST}:{GRADIO_PORT}") # Health check endpoint @app.get("/health") async def health_check(): return {"status": "healthy"} @app.route("/{path:path}", include_in_schema=False, methods=["GET", "POST"]) async def gradio(request: Request): print(f"Forwarding request path: {request.url.path}") headers = dict(request.headers) # Construct the full URL to Gradio, preserving the original path target_url = f"{GRADIO_URL}{request.url.path}" async with httpx.AsyncClient() as client: response = await client.request( request.method, target_url, headers=headers, data=await request.body(), params=request.query_params, ) content = await response.aread() response_headers = dict(response.headers) return StarletteResponse( content=content, status_code=response.status_code, headers=response_headers, ) ``` The above code: * Initializes a FastAPI application that forwards requests to the Gradio app running as a subprocess on a different port. * Sets up a health check endpoint at `/health`. * Creates a catchall proxy that routes all requests to Gradio, including headers. Next, set up the Gradio application. Add the following code to `main.py`: ```python theme={null} import multiprocessing import os import sys import time import gradio as gr # Global variable for the Gradio server gradio_server = None # Configure the Llama endpoint URL LLAMA_ENDPOINT = os.getenv("LLAMA_ENDPOINT", "") # Update with your endpoint class GradioServer: def __init__(self): self.host = GRADIO_HOST self.port = GRADIO_PORT self.process: Optional[multiprocessing.Process] = None self.url = GRADIO_URL async def chat_with_llama(self, message: str, history: List[List[str]]) -> str: """Make a request to the Llama endpoint""" # Convert history and new message into OpenAI chat format messages = [] for h in history: messages.extend([ {"role": "user", "content": h[0]}, {"role": "assistant", "content": h[1]} ]) messages.append({"role": "user", "content": message}) async with httpx.AsyncClient() as client: try: response = await client.post( f"{LLAMA_ENDPOINT}/v1/chat/completions", json={ "messages": messages, "model": "meta-llama/Meta-Llama-3.1-8B-Instruct", "stream": False, "temperature": 0.7, "top_p": 0.95 }, timeout=30.0 ) if response.status_code == 200: response_data = response.json() return response_data['choices'][0]['text'] else: return f"Error: Received status code {response.status_code} from Llama endpoint" except Exception as e: return f"Error communicating with Llama endpoint: {str(e)}" def run_server(self): interface = gr.ChatInterface( fn=self.chat_with_llama, type="messages", title="Chat with Llama", description="This is a chat interface powered by Llama 3.1 8B Instruct", examples=[ ["What is the capital of France?"], ["Explain quantum computing in simple terms"], ["Write a short poem about technology"] ], ) interface.launch( server_name=self.host, server_port=self.port, root_path=f"https://api.cerebrium.ai/v4/{os.getenv('PROJECT_ID')}/{os.getenv('APP_NAME')}/", quiet=True ) def start(self): print(f"Starting Gradio server at {self.url} port {self.port}") # Start Gradio in a separate process self.process = multiprocessing.Process(target=self.run_server) self.process.start() # Wait for Gradio to become ready max_retries = 30 retry_delay = 1.0 for _ in range(max_retries): try: response = requests.get(f"{self.url}/") if response.status_code == 200: print(f"Gradio server is ready at {self.url}") return True except requests.exceptions.ConnectionError: time.sleep(retry_delay) print("Failed to start Gradio server") self.stop() return False def stop(self): if self.process: self.process.terminate() self.process.join() self.process = None @app.on_event("startup") async def startup_event(): global gradio_server if not os.getenv("GRADIO_SERVER_URL"): # Only start local server if no external URL provided gradio_server = GradioServer() if not gradio_server.start(): sys.exit(1) @app.on_event("shutdown") async def shutdown_event(): global gradio_server if gradio_server: gradio_server.stop() ``` The code above defines: * `GradioServer`: handles communication with the Llama model endpoint * `chat_with_llama`: sends a message to the Llama model and returns the response * `run_server`: creates a Gradio chat interface * `start`: starts the Gradio server in a separate process * `stop`: stops the Gradio server * `on_event` startup/shutdown: starts and stops the Gradio server respectively The final `main.py` file: ```python theme={null} import multiprocessing import os import sys import time from typing import Optional, List import gradio as gr import httpx import requests from fastapi import FastAPI, Request from starlette.responses import Response as StarletteResponse # Initialize FastAPI app = FastAPI() # Server configuration gradio_server = None GRADIO_HOST = os.getenv("GRADIO_HOST", "127.0.0.1") GRADIO_PORT = int(os.getenv("GRADIO_PORT", "7860")) GRADIO_URL = os.getenv("GRADIO_SERVER_URL", f"http://{GRADIO_HOST}:{GRADIO_PORT}") LLAMA_ENDPOINT = os.getenv("LLAMA_ENDPOINT", "") class GradioServer: def __init__(self): self.host = GRADIO_HOST self.port = GRADIO_PORT self.process: Optional[multiprocessing.Process] = None self.url = GRADIO_URL async def chat_with_llama(self, message: str, history: List[List[str]]) -> str: """Make a request to the Llama endpoint""" # Convert history and new message into OpenAI chat format messages = [] for h in history: messages.extend([ {"role": "user", "content": h[0]}, {"role": "assistant", "content": h[1]} ]) messages.append({"role": "user", "content": message}) async with httpx.AsyncClient() as client: try: response = await client.post( f"{LLAMA_ENDPOINT}/v1/chat/completions", json={ "messages": messages, "model": "meta-llama/Meta-Llama-3.1-8B-Instruct", "stream": False, "temperature": 0.7, "top_p": 0.95 }, timeout=30.0 ) if response.status_code == 200: response_data = response.json() return response_data['choices'][0]['text'] else: return f"Error: Received status code {response.status_code} from Llama endpoint" except Exception as e: return f"Error communicating with Llama endpoint: {str(e)}" def run_server(self): interface = gr.ChatInterface( fn=self.chat_with_llama, type="messages", title="Chat with Llama", description="This is a chat interface powered by Llama 3.1 8B Instruct", examples=[ ["What is the capital of France?"], ["Explain quantum computing in simple terms"], ["Write a short poem about technology"] ], ) interface.launch( server_name=self.host, server_port=self.port, root_path=f"https://api.cerebrium.ai/v4/{os.getenv('PROJECT_ID')}/{os.getenv('APP_NAME')}/", quiet=True ) def start(self): print(f"Starting Gradio server at {self.url} port {self.port}") # Start Gradio in a separate process self.process = multiprocessing.Process(target=self.run_server) self.process.start() # Wait for Gradio to become ready max_retries = 30 retry_delay = 1.0 for _ in range(max_retries): try: response = requests.get(f"{self.url}/") if response.status_code == 200: print(f"Gradio server is ready at {self.url}") return True except requests.exceptions.ConnectionError: time.sleep(retry_delay) print("Failed to start Gradio server") self.stop() return False def stop(self): if self.process: self.process.terminate() self.process.join() self.process = None @app.get("/health") async def health_check(): return {"status": "healthy"} # Catchall proxy endpoint for Gradio @app.route("/{path:path}", include_in_schema=False, methods=["GET", "POST"]) async def gradio(request: Request): print(f"Forwarding request path: {request.url.path}") headers = dict(request.headers) # Construct the full URL to Gradio, preserving the original path target_url = f"{GRADIO_URL}{request.url.path}" async with httpx.AsyncClient() as client: response = await client.request( request.method, target_url, headers=headers, data=await request.body(), params=request.query_params, ) content = await response.aread() response_headers = dict(response.headers) return StarletteResponse( content=content, status_code=response.status_code, headers=response_headers, ) @app.on_event("startup") async def startup_event(): global gradio_server if not os.getenv("GRADIO_SERVER_URL"): # Only start local server if no external URL provided gradio_server = GradioServer() if not gradio_server.start(): sys.exit(1) @app.on_event("shutdown") async def shutdown_event(): global gradio_server if gradio_server: gradio_server.stop() ``` ## Deploy Deploy the app: ```bash theme={null} cerebrium deploy -y ``` Once deployed, navigate to the following URL in your browser: ``` https://api.cerebrium.ai/v4/p-xxxxxxxx/2-gradio-interface/ ``` The Gradio chat interface appears at this URL. ## Conclusion This architecture provides a scalable chat app using the ASGI custom runtime. The frontend/backend separation improves performance and cost management while maintaining scaling flexibility. Share feedback, challenges, or Gradio apps in the [Discord](https://discord.com/invite/ATj6USmeE2) community. # ComfyUI Application at Scale Source: https://cerebrium.ai/docs/v4/examples/comfyUI Turn ComfyUI stable diffusion workflows into autoscaling API endpoints on Cerebrium, from exporting the workflow JSON to serving generated images. ### Introduction ComfyUI is a popular no-code interface for building complex stable diffusion workflows. Its modular setup and intuitive flowchart interface have produced an extensive collection of community workflows. Several websites offer shared workflows: * [https://comfyworkflows.com/](https://comfyworkflows.com/) * [https://openart.ai/workflows/home‍](https://openart.ai/workflows/home‍) Production-scale deployment guidance for ComfyUI is limited. This tutorial covers deploying ComfyUI pipelines on Cerebrium as autoscaling API endpoints with pay-as-you-go compute. Find the full example code [here](https://github.com/CerebriumAI/examples/tree/master/7-image-and-video/1-comfyui). ### Creating Your ComfyUI Workflow Locally Create the workflow locally or using a rented GPU from [Lambda Labs](https://lambdalabs.com/). Ensure [ComfyUI is installed](https://github.com/comfyanonymous/ComfyUI#installing) in the local environment. This tutorial uses Stable Diffusion XL and ControlNet to create custom QR codes. Skip to "Export ComfyUI Workflow" if a workflow already exists. 1. Create the Cerebrium project: `cerebrium init 1-comfyui` 2. Clone the ComfyUI GitHub project inside the project directory: `git clone https://github.com/comfyanonymous/ComfyUI` 3. Download the following models and install them in the appropriate folders within the ComfyUI folder: * SDXL base in models/checkpoints. * ControlNet in models/ControlNet. 4. Run ComfyUI locally from inside the cloned ComfyUI folder: python main.py --force-fp16 on MacOS. 5. A server should be loaded locally at [http://127.0.0.1:8188/‍](http://127.0.0.1:8188/‍) ‍ The default ComfyUI workflow interface appears in this view. Use this locally running instance to build the image generation pipeline. ### Export ComfyUI Workflow The example GitHub repository contains a workflow\.json file. Click the “Load” button on the right to load the workflow. It should appear populated. ComfyUI Workflow Don’t worry about the pre-filled values and prompts. They get overridden at inference time. To export the workflow in API format, click the gear icon (settings) in the top-right hovering panel. Ensure that “Enable dev mode” is selected, then close the popup. Save ComfyUI API format A button appears in the right hover panel labeled “Save (API format)”. Use it to save the workflow as “workflow\_api.json”. ### ComfyUI Application The main code in `main.py` uses the exported ComfyUI API workflow to create an endpoint. The code: 1. Initializes the ComfyUI server 2. Loads the workflow API template 3. Processes inputs (prompts, images) through the `run` function 4. Generates base64-encoded image outputs Note: Cerebrium runs code outside the `run` function only during initialization, while the `run` function handles all subsequent requests. Alter the workflow\_api.json file to include placeholders for user values at inference time: * Replace line 4, the seed input, with: "\{\{seed}}" * Replace line 45, the input text of node 6 with: "\{\{positive\_prompt}}" * Replace line 58, the input text of node 7 with: "\{\{negative\_prompt}}" * Replace line 108, the image of node 11 with: "\{\{controlnet\_image}}" ### FastAPI App In addition to the `main.py` file below (which runs the FastAPI server and initializes ComfyUI on application start), a separate `helpers.py` file contains utility functions for working with ComfyUI. Find the helper code [here](https://github.com/CerebriumAI/examples/blob/master/7-image-and-video/1-comfyui/helpers.py). Create a file named `helpers.py` and copy the code into it. ```python theme={null} import copy import json import logging import os import signal import time import uuid from contextlib import contextmanager from multiprocessing import Process import websocket from fastapi import FastAPI, BackgroundTasks, HTTPException from fastapi.middleware.cors import CORSMiddleware from helpers import ( convert_outputs_to_base64, convert_request_file_url_to_path, fill_template, get_images, setup_comfyui, ) # Configure logging logging.basicConfig(level=logging.INFO) logger = logging.getLogger("comfyui-api") # Initialize FastAPI app app = FastAPI(title="ComfyUI API") app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"], allow_headers=["*"]) # Define configuration server_address = "127.0.0.1:8188" original_working_directory = os.getcwd() json_workflow = None side_process = None WEBSOCKET_TIMEOUT = 60 @contextmanager def websocket_connection(): """Establish WebSocket connection to ComfyUI server with proper cleanup.""" client_id = str(uuid.uuid4()) ws = None try: ws = websocket.WebSocket() ws.settimeout(WEBSOCKET_TIMEOUT) ws.connect(f"ws://{server_address}/ws?clientId={client_id}") logger.info("WebSocket connected successfully") yield ws, client_id except Exception as e: logger.error(f"WebSocket connection error: {str(e)}") raise HTTPException(status_code=503, detail=f"ComfyUI server connection error: {str(e)}") finally: if ws: try: ws.close() except Exception: pass def load_workflow_file(file_path: str) -> dict: """Load workflow JSON file.""" try: with open(file_path, "r") as json_file: return json.load(json_file) except FileNotFoundError: logger.error(f"Workflow file not found: {file_path}") raise HTTPException(status_code=500, detail=f"Workflow file not found: {file_path}") except json.JSONDecodeError: logger.error(f"Invalid JSON in workflow file: {file_path}") raise HTTPException(status_code=500, detail=f"Invalid JSON in workflow file: {file_path}") def cleanup_tempfiles(files): """Clean up temporary files.""" for file in files: try: if hasattr(file, 'name') and os.path.exists(file.name): os.unlink(file.name) except Exception as e: logger.warning(f"Error cleaning up temp file: {str(e)}") def terminate_process(): """Terminate the ComfyUI process.""" global side_process if side_process and side_process.is_alive(): logger.info("Terminating ComfyUI process...") side_process.terminate() side_process.join(timeout=5) if side_process.is_alive(): side_process.kill() @app.on_event("startup") async def startup_event(): """Start ComfyUI server on application startup.""" global json_workflow, side_process # Load workflow JSON json_workflow = load_workflow_file("workflow_api.json") logger.info("Loaded workflow from workflow_api.json") # Start ComfyUI process if side_process is None: side_process = Process( target=setup_comfyui, kwargs=dict(original_working_directory=original_working_directory, data_dir=""), daemon=True, ) side_process.start() logger.info(f"Started ComfyUI process (PID: {side_process.pid})") for sig in [signal.SIGINT, signal.SIGTERM]: signal.signal(sig, lambda s, f: terminate_process()) # Wait for ComfyUI to start max_attempts = 30 for attempt in range(max_attempts): try: with websocket_connection() as (ws, _): logger.info("Successfully connected to ComfyUI!") break except Exception: logger.info(f"Waiting for ComfyUI to start... ({attempt + 1}/{max_attempts})") time.sleep(2) else: logger.warning("Could not confirm ComfyUI is running after multiple attempts") @app.post("/run") async def run(workflow_values: dict, background_tasks: BackgroundTasks): """Run a workflow with the provided template values.""" # Process input values template_values, tempfiles = convert_request_file_url_to_path(workflow_values) background_tasks.add_task(cleanup_tempfiles, tempfiles) try: with websocket_connection() as (ws, client_id): # Apply template values to workflow json_workflow_copy = copy.deepcopy(json_workflow) json_workflow_copy = fill_template(json_workflow_copy, template_values) # Run workflow and get outputs outputs = get_images(ws, json_workflow_copy, client_id, server_address) # Process outputs to base64 result = [] for node_id in outputs: for unit in outputs[node_id]: try: file_name = unit.get("filename") file_data = unit.get("data") output = convert_outputs_to_base64( node_id=node_id, file_name=file_name, file_data=file_data ) result.append(output) except Exception as e: result.append({ "node_id": node_id, "error": f"Failed to process output: {str(e)}", "format": "error" }) return {"result": result, "status": "success"} except Exception as e: logger.error(f"Error running workflow: {str(e)}") raise HTTPException(status_code=500, detail=str(e)) @app.get("/health") async def health_check(): """Health check endpoint.""" global side_process process_status = "running" if side_process and side_process.is_alive() else "not running" return { "status": "ok", "comfyui_process": process_status, "timestamp": time.time() } @app.on_event("shutdown") def shutdown_event(): """Clean up on application shutdown.""" logger.info("Application shutting down, cleaning up resources...") terminate_process() ``` ### Deploy ComfyUI Application Before deploying, define the deployment configuration in `cerebrium.toml`. This specifies dependencies, hardware, scaling, and any pre-run scripts: ``` [cerebrium.deployment] name = "1-comfyui" python_version = "3.11" include = ["./*", "main.py", "cerebrium.toml", "workflow.json", "workflow_api.json", "helpers.py", "model.json"] exclude = ["./example_exclude", "./ComfyUI", "./ComfyUI/models/checkpoints/sd_xl_base_1.0.safetensors", "./ComfyUI/models/controlnet/diffusion_pytorch_model.fp16.safetensors"] shell_commands = ["git clone https://github.com/comfyanonymous/ComfyUI", "pip install -r ComfyUI/requirements.txt"] [cerebrium.runtime.custom] port = 8765 entrypoint = ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8765"] healthcheck_endpoint = "/health" [cerebrium.hardware] compute = "AMPERE_A10" cpu = 4 memory = 16.0 [cerebrium.scaling] min_replicas = 0 max_replicas = 2 cooldown = 30 replica_concurrency = 1 response_grace_period = 900 scaling_metric = "concurrency_utilization" scaling_target = 100 scaling_buffer = 0 [cerebrium.dependencies.pip] uvicorn = "latest" fastapi = "latest" requests = "latest" channels = "latest" websockets = "latest" websocket-client = "==1.6.4" accelerate = "==0.23.0" opencv-python = "latest" pydantic = "latest" pillow = "latest" safetensors = "latest" torch = "latest" torchvision = "latest" transformers = "latest" torchsde = "latest" einops = "latest" aiohttp = "latest" pyyaml = "latest" Pillow = "latest" scipy = "latest" tqdm = "latest" psutil = "latest" kornia = ">=0.7.1" [cerebrium.dependencies.apt] git = "latest" ``` Running `cerebrium deploy` uploads both the ComfyUI directory and \~10GB of model weights. For slower connections, use the helper file to download model weights directly to the appropriate folders. Create a file called `model.json` with the following contents: ```json theme={null} [ { "url": "https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0/resolve/main/sd_xl_base_1.0.safetensors", "path": "models/checkpoints/sd_xl_base_1.0.safetensors" }, { "url": "https://huggingface.co/diffusers/controlnet-canny-sdxl-1.0/resolve/main/diffusion_pytorch_model.fp16.safetensors", "path": "models/controlnet/diffusers_xl_canny_full.safetensors" } ] ``` This configuration tells the helper function where to download models from and where to save them. Models download only on first deployment, with subsequent deploys checking for existing files. This is what your file folder structure should look like: ComfyUI folder structure Deploy the application: `cerebrium deploy` After successful deployment, make a request to the endpoint with the following JSON payload: ```curl theme={null} curl --location 'https://api.cerebrium.ai/v4/p-xxxxxxxx/1-comfyui/run' \ --header 'Content-Type: application/json' \ --header 'Authorization: Bearer ' \ --data '{"workflow_values": { "positive_prompt": "A top down view of a mountain with large trees and green plants", "negative_prompt": "blurry, text, low quality", "controlnet_image": "https://cerebrium-assets.s3.eu-west-1.amazonaws.com/qr-code.png", "seed": 1000 } }' ``` The output contains two responses: * A base64-encoded image of the ControlNet input outline, used as input in the flow. * A base64-encoded image of the final result. ComfyUI generated QR code ### Conclusion Cerebrium enables production-ready ComfyUI workflows that automatically scale with demand, with costs aligned to actual compute usage. Share creations by tagging @cerebriumai. # Deploy a Vision Language Model with SGLang Source: https://cerebrium.ai/docs/v4/examples/deploy-a-vision-language-model-with-sglang Deploy a vision language model with SGLang on Cerebrium and build an ad analysis system that scores advertisements across multiple criteria. This tutorial deploys a Vision Language Model (VLM) using SGLang on Cerebrium. A VLM combines a large language model (LLM) with a vision encoder, enabling it to understand and process both images and text. The example builds an intelligent ad analysis system that evaluates advertisements across multiple dimensions. It scores how the advertisement relates to the business in question and how it performs on the given criteria. SGLang (Structured Generation Language) differs from other inference frameworks such as vLLM and TensorRT by focusing on structured generation and complex multi-step LLM workflows. Teams at xAI and Deepseek use SGLang in production to power their core language model capabilities, making it a trusted choice. ### SGLang Architecture SGLang isn't just a domain-specific language (DSL). It's a complete, integrated execution system with a clear separation of functionality: | Layer | What it does | Why it matters | | - | - | - | | Frontend | Where you define your LLM logic (with gen, fork, join, etc.) | This keeps your code clean, readable, and your workflows easily reusable. | | Backend | Where SGLang intelligently figures out how to run your logic most efficiently. | This is where the speed, scalability, and optimized inference truly come to life. | Here are some frontend primitives for creating multi-step workflows: | Primitive | What it does | Example | | - | - | - | | `gen()` | Generates a text span | `gen("title", stop="\n")` | | `fork()` | Splits execution into multiple branches | For parallel sub-tasks | | `join()` | Merges branches back together | For combining outputs | | `select()` | Chooses one option from many | For controlled logic, like multiple choice | SGLang Architecture Here is a summary of key advantages over traditional inference engines: | Feature | Traditional Engines (vLLM, TGI) | SGLang | | - | - | - | | **Programming Model** | Sequential API calls with manual prompt chaining | Native structured logic with `gen()`, `fork()`, `join()`, `select()` | | **Memory Management** | Basic KV caching, often discarded between calls | **RadixAttention**: Intelligent prefix-aware cache reuse (up to 6x faster) | | **Output Control** | Hope and pray for correct formatting | **Compressed FSMs**: Guaranteed structured output (JSON, XML, etc.) | | **Parallel Processing** | Manual batching and coordination | Built-in `fork()` and `join()` for parallel execution | | **Performance** | Standard inference optimization | PyTorch-native with `torch.compile()`, quantization, sparse inference | For more details, see this [article](https://huggingface.co/blog/paresh2806/sglang-efficient-llm-workflows). You can see the final code sample [here](https://github.com/CerebriumAI/examples/tree/master/5-large-language-models/7-vision-language-sglang). ## Tutorial ### Step 1: Project Setup Create the project structure: ```bash theme={null} cerebrium init 7-vision-language-sglang cd 7-vision-language-sglang ``` ### Step 2: Configure Dependencies The VLM is [Qwen3-VL-30B-A3B-Instruct-FP8](https://huggingface.co/Qwen/Qwen3-VL-30B-A3B-Instruct-FP8), which requires significant GPU memory. The `cerebrium.toml` defines the environment, hardware, and scaling settings. This configuration uses an ADA\_L40 GPU and includes: * Hardware settings for GPU, CPU, and memory allocation * Scaling parameters to control instance counts * Required pip packages: SGLang, flashinfer (the chosen backend), and PyTorch * APT system dependencies * FastAPI server configuration for hosting the API For a complete reference of all available TOML settings, see the [TOML Reference](/docs/toml-reference/toml-reference). This example uses flashinfer as the backend, but other options like flash attention are also available. Update `cerebrium.toml` with: ```toml theme={null} [cerebrium.deployment] name = "7-vision-language-sglang" python_version = "3.11" docker_base_image_url = "nvidia/cuda:12.8.0-devel-ubuntu22.04" deployment_initialization_timeout = 860 [cerebrium.hardware] cpu = 6.0 memory = 60.0 compute = "ADA_L40" [cerebrium.scaling] min_replicas = 0 max_replicas = 2 [cerebrium.build] use_uv = true [cerebrium.dependencies.pip] transformers = "latest" huggingface_hub = "latest" pydantic = "latest" pillow = "latest" requests = "latest" torch = "latest" "sglang[all]" = "latest" "sgl-kernel" = "latest" "flashinfer-python" = "latest" [cerebrium.dependencies.apt] libnuma-dev = "latest" [cerebrium.runtime.custom] port = 8000 entrypoint = ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"] ``` ### Step 3: Implement the Ad Analysis Logic Cerebrium does not enforce any special class design or application architecture. Write Python code as if running locally. The code below sets up the SGLang Runtime Engine (Backend) with FastAPI and loads the model on container startup. The first request incurs a model load, but subsequent requests execute instantaneously. In your `main.py` file: ```python theme={null} import sglang as sgl from sglang import function from fastapi import FastAPI, HTTPException from transformers import AutoProcessor app = FastAPI(title="Vision Language SGLang API") model_path = "Qwen/Qwen3-VL-30B-A3B-Instruct-FP8" processor = AutoProcessor.from_pretrained(model_path) @app.on_event("startup") def _startup_warmup(): # Initialize engine on main thread during app startup runtime = sgl.Runtime( model_path=model_path, enable_multimodal=True, mem_fraction_static=0.8, tp_size=1, attention_backend="flashinfer", ) runtime.endpoint.chat_template = sgl.lang.chat_template.get_chat_template( "qwen2-vl" ) sgl.set_default_backend(runtime) @app.get("/health") def health(): return { "status": "healthy", } ``` To score the advertisement, the code uses one of SGLang's core differentiators: `fork`, which runs many prompts in parallel and brings the results together. This enables many simultaneous evaluations with no increase in total latency. The results are then structured in a specific format for the response. ```python theme={null} @function def analyze_ad(s, image, ad_description, dimensions): s += sgl.system("Evaluate an advertisement about an company's description.") s += sgl.user(sgl.image(image) + "Company Description: " + ad_description) s += sgl.assistant("Sure!") s += sgl.user("Is the company description related to the image?") s += sgl.assistant(sgl.select("related", choices=["yes", "no"])) if s["related"] == "no": return forks = s.fork(len(dimensions)) for i, (f, dim) in enumerate(zip(forks, dimensions)): f += sgl.user("Evaluate based on the following dimension: " + dim + ". End your judgment with the word 'END'") # Use unique slot names per dimension to avoid collisions f += sgl.assistant("Judgment: " + sgl.gen(f"judgment_{i}", stop="END")) s += sgl.user("Provide a one-sentence synthesis of the overall evaluation, then we will output JSON.") s += sgl.assistant(sgl.gen("summary_one_liner", stop=".")) schema = r'^\{"summary": ".{1,400}", "grade": "[ABCD][+\-]?"\}$' s += sgl.user("Return only a 3 line parapgrah JSON object with keys summary and grade (A, B, C, D, +, -), where summary briefly synthesizes the above judgments.") s += sgl.assistant(sgl.gen("output", regex=schema)) ``` Bring it all together in an endpoint: ```python theme={null} from pydantic import BaseModel import base64 import io import json from PIL import Image class AnalyzeRequest(BaseModel): image_base64: str ad_description: str dimensions: list def process_image(image_base64: str) -> Image.Image: image_data = base64.b64decode(image_base64) return Image.open(io.BytesIO(image_data)) @app.post("/analyze") def analyze_advertisement(req: AnalyzeRequest): try: image = process_image(req.image_base64) state = analyze_ad.run(image, req.ad_description, req.dimensions) try: print(state) output = state["output"] except KeyError: output = None if isinstance(output, str): start = output.find("{") end = output.rfind("}") + 1 if start != -1 and end > start: return { "success": True, "analysis": json.loads(output[start:end]), "dimensions_evaluated": req.dimensions } return { "success": True, "analysis": output, "dimensions_evaluated": req.dimensions } except Exception as e: raise HTTPException(status_code=500, detail=str(e)) ``` Deploy the application to create a scalable inference endpoint. ### Step 4: Deploy Your Application Run: ```bash theme={null} cerebrium deploy ``` Once deployed, test with a sample request: ```bash theme={null} curl -X POST "https://api.cerebrium.ai/v4/p-xxxxxxxx/7-vision-language-sglang/analyze" \ -H "Content-Type: application/json" \ -d '{ "company_description": "Nike is a global leader in athletic footwear, apparel, and sports equipment known for its innovative designs and the iconic “swoosh” logo. The brand embodies performance, style, and inspiration, empowering athletes worldwide to Just Do It.", "image_base64": "", "dimensions": ["Effectiveness","Clarity", "Appeal","Credibility"] }' ``` Nike AD ### Example Response ```json theme={null} { "success": true, "analysis": { "summary": "The company description is relevant to the image because it accurately reflects Nike's branding, which is showcased through the advertised sneaker and logo. The ad promotes Nike's core products—athletic footwear—and its values of performance, style, and inspiration, aligning with the brand's identity. The collaboration with a superhero theme further emphasizes innovation and empowerment, core ", "grade": "A" }, "dimensions_evaluated": ["Effectiveness", "Clarity", "Appeal", "Credibility"] } ``` This example demonstrates how to leverage SGLang's structured generation capabilities to build an ad analysis system, using features like `fork()` for parallel processing and SGLang's built-in output control. You can find the complete code for this tutorial in our [examples repository](https://github.com/CerebriumAI/examples/tree/master/5-large-language-models/7-vision-language-sglang). # Deploy Triton Inference server and TensorRT-LLM Source: https://cerebrium.ai/docs/v4/examples/deploy-an-llm-with-tensorrtllm-tritonserver Serve Llama 3.2 with NVIDIA Triton Inference Server and TensorRT-LLM on Cerebrium for up to 15x higher throughput and much lower inference latency. This tutorial deploys Llama 3.2 3B using TensorRT-LLM's PyTorch backend served through Nvidia Triton Inference Server. The TensorRT + Triton setup delivers **15x higher throughput** with **100% reliability** compared to the baseline (vanilla deployment), while reducing latency by **7-9x** across all percentiles. See the [Performance Analysis](#performance-analysis) section for detailed test methodology and results. You can view the final implementation [here](https://github.com/CerebriumAI/examples/tree/master/5-large-language-models/8-faster-inference-with-triton-tensorrt). ## Why TensorRT + Triton? ### Why TensorRT? NVIDIA TensorRT is a software development kit for high-performance deep learning inference. It compiles model weights into optimized engines that run more efficiently on specific GPU hardware through CUDA-level optimizations, custom kernels, and optional quantization. TensorRT requires you to specify optimization parameters upfront: GPU architecture, batch size, precision (FP8, INT8, etc.), and input/output shapes. This specialization allows TensorRT to generate highly optimized inference engines that maximize GPU utilization, reduce latency, and lower inference costs compared to serving raw model weights. ### Why Triton? NVIDIA Triton Inference Server streamlines production AI deployment by handling operational concerns that are critical for serving models at scale. It provides automatic request batching, health checks, metrics collection, and standardized HTTP/gRPC APIs out of the box. Triton supports multiple frameworks (TensorRT, PyTorch, TensorFlow, ONNX, etc.), offers built-in Prometheus metrics for observability, and integrates seamlessly with Kubernetes for auto-scaling. It also supports model versioning, A/B testing, and can chain multiple models into pipelines. [Here](https://substackcdn.com/image/fetch/\$s_!FEPb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d4460ad-0e7e-4545-aee6-274b93dd5959_2300x2304.gif) is a diagram of how Triton works. Below is the process of how the two work together in terms of handling requests: 1. Client sends text via HTTP/gRPC to Triton 2. Triton queues the request in the scheduler 3. Triton batches incoming requests (waits for more or timeout) 4. When batch is ready, Triton calls your Python backend 5. TensorRT-LLM generates tokens for the entire batch in parallel on GPU 6. Triton returns responses to clients This setup allows multiple concurrent requests to be processed together on the GPU for maximum throughput. The following sections combine Triton and TensorRT-LLM into a working deployment. ## Basic Setup Install the Cerebrium CLI: ```bash theme={null} pip install cerebrium cerebrium login ``` Create your project: ```bash theme={null} cerebrium init tensorrt-triton-demo cd tensorrt-triton-demo ``` To download the model, [request access](https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct) on Hugging Face. Then add the [HuggingFace token](https://huggingface.co/settings/tokens) to Cerebrium project secrets as `HF_AUTH_TOKEN` through the dashboard for authentication during download. ## Implementation Place all files in the same project directory. ### Triton Model Configuration Create `config.pbtxt` to define Triton's model interface. See the [full configuration reference](https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/model_configuration.html#dynamic-batcher) for all available options. ```protobuf theme={null} name: "llama3_2" backend: "python" max_batch_size: 128 dynamic_batching { max_queue_delay_microseconds: 100 } instance_group [ { count: 1 kind: KIND_GPU } ] input [ { name: "text_input" data_type: TYPE_STRING dims: [ 1 ] }, { name: "max_tokens" data_type: TYPE_INT32 dims: [ 1 ] optional: true }, { name: "temperature" data_type: TYPE_FP32 dims: [ 1 ] optional: true }, { name: "top_p" data_type: TYPE_FP32 dims: [ 1 ] optional: true } ] output [ { name: "text_output" data_type: TYPE_STRING dims: [ 1 ] } ] ``` This configuration tells Triton: * Use Python backend (runs our model.py) * Automatically batch up to 128 requests together for efficient GPU utilization * Use dynamic batching with a 100 microsecond queue delay to maximize batch sizes * Accept text input with optional sampling parameters * Run on a single GPU instance * Return generated text as output ### Python Backend Implementation Triton's Python backend requires implementing a `TritonPythonModel` class with three key methods: * **`initialize(args)`**: Called once when Triton loads the model. This is where you load the tokenizer and initialize TensorRT-LLM with your build configuration. * **`execute(requests)`**: Called every time Triton has a batch ready. Triton automatically batches incoming requests (up to your configured `max_batch_size`) and passes them here. This method extracts prompts from each request, runs batch inference with TensorRT-LLM, and returns responses. * **`finalize()`**: Called when the model is being unloaded. Use this to clean up GPU memory and shut down the TensorRT-LLM engine. Create `model.py` implementing Triton's Python backend interface: ```python theme={null} """ Triton Python Backend for TensorRT-LLM. """ import numpy as np import triton_python_backend_utils as pb_utils import torch from tensorrt_llm import LLM, SamplingParams, BuildConfig from tensorrt_llm.plugin.plugin import PluginConfig from transformers import AutoTokenizer MODEL_ID = "meta-llama/Llama-3.2-3B-Instruct" MODEL_DIR = f"/persistent-storage/models/{MODEL_ID}" class TritonPythonModel: def initialize(self, args): """Initialize TensorRT-LLM with PyTorch backend.""" print("Loading tokenizer...") self.tokenizer = AutoTokenizer.from_pretrained(MODEL_DIR) print("Initializing TensorRT-LLM...") plugin_config = PluginConfig.from_dict({ "paged_kv_cache": True, }) build_config = BuildConfig( plugin_config=plugin_config, max_input_len=4096, max_batch_size=128, # Matches Triton max_batch_size in config.pbtxt ) self.llm = LLM( model=MODEL_DIR, build_config=build_config, tensor_parallel_size=torch.cuda.device_count(), ) print("✓ Model ready") def execute(self, requests): """ Execute inference on batched requests. Triton automatically batches requests (up to max_batch_size: 128). This function processes the batch that Triton provides. """ try: prompts = [] sampling_params_list = [] original_prompts = [] # Extract data from each request in the batch. We need to look through requests: https://github.com/triton-inference-server/python_backend?tab=readme-ov-file#execute for request in requests: try: # Get input text - handle batched tensor structures input_tensor = pb_utils.get_input_tensor_by_name(request, "text_input") text_array = input_tensor.as_numpy() # Extract text handling different array structures if text_array.ndim == 0: text = text_array.item() elif text_array.dtype == object: text = text_array.flat[0] if text_array.size > 0 else text_array.item() else: text = text_array.flat[0] if text_array.size > 0 else text_array.item() # Decode if bytes if isinstance(text, bytes): text = text.decode('utf-8') elif isinstance(text, np.str_): text = str(text) # Get optional parameters with defaults max_tokens = 1024 if pb_utils.get_input_tensor_by_name(request, "max_tokens") is not None: max_tokens_array = pb_utils.get_input_tensor_by_name(request, "max_tokens").as_numpy() max_tokens = int(max_tokens_array.item() if max_tokens_array.ndim == 0 else max_tokens_array.flat[0]) temperature = 0.8 if pb_utils.get_input_tensor_by_name(request, "temperature") is not None: temp_array = pb_utils.get_input_tensor_by_name(request, "temperature").as_numpy() temperature = float(temp_array.item() if temp_array.ndim == 0 else temp_array.flat[0]) top_p = 0.95 if pb_utils.get_input_tensor_by_name(request, "top_p") is not None: top_p_array = pb_utils.get_input_tensor_by_name(request, "top_p").as_numpy() top_p = float(top_p_array.item() if top_p_array.ndim == 0 else top_p_array.flat[0]) # Format prompt using chat template prompt = self.tokenizer.apply_chat_template( [{"role": "user", "content": text}], tokenize=False, add_generation_prompt=True ) prompts.append(prompt) original_prompts.append(prompt) sampling_params_list.append(SamplingParams( temperature=temperature, top_p=top_p, max_tokens=max_tokens, )) except Exception as e: print(f"Error processing request: {e}", flush=True) prompts.append("") original_prompts.append("") sampling_params_list.append(SamplingParams(max_tokens=1024)) # Batch inference if not prompts: return [] outputs = self.llm.generate(prompts, sampling_params_list) # Create responses responses = [] for i, output in enumerate(outputs): try: generated_text = output.outputs[0].text # Strip prompt from output if included if original_prompts[i] and original_prompts[i] in generated_text: generated_text = generated_text.replace(original_prompts[i], "").strip() responses.append(pb_utils.InferenceResponse( output_tensors=[pb_utils.Tensor( "text_output", np.array([generated_text.encode('utf-8')], dtype=object) )] )) except Exception as e: print(f"Error creating response {i}: {e}", flush=True) responses.append(pb_utils.InferenceResponse( output_tensors=[pb_utils.Tensor( "text_output", np.array([f"Error: {str(e)}".encode('utf-8')], dtype=object) )] )) return responses except Exception as e: print(f"Error in execute: {e}", flush=True) return [ pb_utils.InferenceResponse( output_tensors=[pb_utils.Tensor( "text_output", np.array([f"Batch error: {str(e)}".encode('utf-8')], dtype=object) )] ) for _ in requests ] def finalize(self): """Cleanup on shutdown.""" if hasattr(self, 'llm'): self.llm.shutdown() torch.cuda.empty_cache() ``` ### Model Download Script Create `download_model.py` to download the model: ```python theme={null} #!/usr/bin/env python3 """Download HuggingFace model to persistent storage.""" import os from pathlib import Path from huggingface_hub import snapshot_download, login MODEL_ID = "meta-llama/Llama-3.2-3B-Instruct" MODEL_DIR = Path("/persistent-storage/models") / MODEL_ID def download_model(): """Download model if not already present.""" hf_token = os.environ.get("HF_AUTH_TOKEN") if not hf_token: print("WARNING: HF_AUTH_TOKEN not set") return if MODEL_DIR.exists() and any(MODEL_DIR.iterdir()): print("✓ Model already exists") return print("Downloading model...") login(token=hf_token) snapshot_download( MODEL_ID, local_dir=str(MODEL_DIR), token=hf_token ) print("✓ Model downloaded") if __name__ == "__main__": download_model() ``` This script checks if the model exists in persistent storage before downloading to avoid redundant downloads on subsequent deployments. ### Container Setup Create `Dockerfile` extending Nvidia's Triton container: ```dockerfile theme={null} FROM nvcr.io/nvidia/tritonserver:25.10-trtllm-python-py3 ENV PYTHONPATH=/usr/local/lib/python3.12/dist-packages:$PYTHONPATH ENV PYTHONDONTWRITEBYTECODE=1 ENV DEBIAN_FRONTEND=noninteractive ENV HF_HOME=/persistent-storage/models ENV TORCH_CUDA_ARCH_LIST=8.6 # Install dependencies RUN apt-get update && apt-get install -y \ git \ git-lfs \ && rm -rf /var/lib/apt/lists/* WORKDIR /app RUN pip install --break-system-packages \ huggingface_hub \ transformers \ || true # Create directories RUN mkdir -p \ /app/model_repository/llama3_2/1 \ /persistent-storage/models \ /persistent-storage/engines # Copy files COPY model.py /app/model_repository/llama3_2/1/ COPY config.pbtxt /app/model_repository/llama3_2/ EXPOSE 8000 8001 8002 CMD ["tritonserver", "--model-repository=/app/model_repository", "--http-port=8000", "--grpc-port=8001", "--metrics-port=8002"] ``` The Dockerfile uses Nvidia's official Triton container with TensorRT-LLM pre-installed, creates the model repository structure that Triton expects, and copies our application files to the correct locations. ### Deployment Configuration Configure the container and autoscaling environment in `cerebrium.toml`: ```toml theme={null} [cerebrium.deployment] name = "tensorrt-triton-demo" python_version = "3.12" disable_auth = true include = ['./*', 'cerebrium.toml'] exclude = ['.*'] deployment_initialization_timeout = 830 [cerebrium.hardware] cpu = 4.0 memory = 40.0 compute = "AMPERE_A10" gpu_count = 1 [cerebrium.scaling] min_replicas = 0 max_replicas = 5 cooldown = 300 replica_concurrency = 128 scaling_metric = "concurrency_utilization" [cerebrium.runtime.custom] port = 8000 healthcheck_endpoint = "/v2/health/live" readycheck_endpoint = "/v2/health/ready" dockerfile_path = "./Dockerfile" ``` Key configuration details: * `replica_concurrency = 128`: Each replica can handle up to 128 concurrent requests, matching our Triton batch size * `max_replicas = 5`: Scale up to 5 replicas for peak load ## Deploy ### Download Model to Persistent Storage Before deploying, download the model to Cerebrium's persistent storage. This ensures the model is available across all deployments and avoids redundant downloads during container startup. The `cerebrium run` command executes a Python script in a temporary container with the same environment and hardware configuration as the deployment. It has access to persistent storage at `/persistent-storage`, so any files written there are available to deployed containers. Run the download script: ```bash theme={null} cerebrium run download_model.py ``` The logs confirm whether the model already exists or has been downloaded successfully. ### Deploy the Model Deploy the model: ```bash theme={null} cerebrium deploy ``` After successful deployment, the base endpoint URL appears in the output. Use this URL in the next section. ## Test Send a request to your deployed endpoint: ```bash theme={null} curl -X POST https://api.cerebrium.ai/v4/p-xxxxxxxx//v2/models/llama3_2/infer \ -H "Content-Type: application/json" \ -d '{ "inputs": [ { "name": "text_input", "shape": [1, 1], "datatype": "BYTES", "data": ["What is machine learning?"] } ], "outputs": [{"name": "text_output"}] }' ``` The endpoint returns results in this format: ```json theme={null} { "outputs": [ { "name": "text_output", "datatype": "BYTES", "shape": [1], "data": [ "Machine learning is a subset of artificial intelligence (AI) that involves training algorithms..." ] } ] } ``` The response follows Triton's standard inference protocol format with the generated text in the `data` field of the output tensor. ## Performance Analysis ### Test Setup To validate performance improvements, TensorRT + Triton was compared against a vanilla HuggingFace baseline serving the same Llama 3.2 3B Instruct model. Both deployments used identical hardware (NVIDIA A10 GPU) and were tested under the same load conditions. **Vanilla Baseline Setup:** * Model served directly using HuggingFace Transformers with PyTorch * Single request processing (no batching) * Standard FastAPI endpoint * Same hardware configuration (A10 GPU, 4 CPU cores, 40GB memory) **TensorRT + Triton Setup:** * TensorRT-LLM with PyTorch backend * Triton Inference Server with dynamic batching (max batch size: 128) * Automatic request queuing and batching * Same hardware configuration (A10 GPU, 4 CPU cores, 40GB memory) Both deployments were tested with the same load testing parameters to ensure fair comparison. ### Results | Metric | Vanilla Baseline | TensorRT + Triton | Improvement | | - | - | - | - | | **Requests Per Second (RPS)** | 0.83 | 12.46 | **15x faster** | | **Success Rate** | 61.6% | 100.0% | **38.4% increase** | | **P50 Latency** | 297.7s | 41.7s | **7.1x faster** | | **P99 Latency** | 593.2s | 79.3s | **7.5x faster** | | **Average Latency** | 376.2s | 42.4s | **8.9x faster** | The TensorRT + Triton setup delivers **15x higher throughput** with **100% reliability** compared to the baseline, while reducing latency by **7-9x** across all percentiles. The baseline's 61.6% success rate and high latency come from processing requests sequentially without batching, leading to GPU underutilization and request timeouts. TensorRT + Triton eliminates these issues by keeping the GPU fully utilized with batched, optimized inference, resulting in 100% success rate and consistent, predictable latency. These results demonstrate that TensorRT + Triton is not just faster, but also more reliable and cost-effective for production LLM serving at scale. ## Get Started The complete implementation, including all configuration files and deployment scripts, is available in our [GitHub repository](https://github.com/CerebriumAI/examples/tree/master/5-large-language-models/8-faster-inference-with-triton-tensorrt). Clone the repository and follow this tutorial to deploy Llama 3.2 3B (or adapt it for other models) with TensorRT-LLM and Triton Inference Server. # Featured Examples Source: https://cerebrium.ai/docs/v4/examples/featured Browse featured Cerebrium examples and tutorials covering LLM endpoints, voice agents, image generation, integrations and other AI applications. Create chat interfaces with Gradio and ASGI Build conversational voice agents using Twilio Deploy real-time voice AI agents for interactive conversations Set up OpenAI API-compatible endpoints using vLLM Set up OpenAI API-compatible endpoints using vLLM Integrate with LangChain and LangSmith Transcribe audio using Whisper Build conversational voice agents using Twilio Deploy real-time voice AI agents for interactive conversations Deploy and use ComfyUI Generate images with SDXL Create Gradio interfaces with ASGI # Serving GPT-OSS with vLLM Source: https://cerebrium.ai/docs/v4/examples/gpt-oss Serve OpenAI GPT-OSS open weight models with vLLM on Cerebrium, covering MoE architecture, MXFP4 quantization and H100 GPU deployment setup. GPT recently released GPT-OSS ([gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b) and [gpt-oss-120b](https://huggingface.co/openai/gpt-oss-120b)), two state-of-the-art open-weight language models that deliver strong real-world performance at low cost. Available under the flexible Apache 2.0 license, these models outperform similarly sized open models on reasoning tasks and demonstrate strong tool use capabilities. They are also optimized for efficient deployment on consumer hardware. ## What Makes GPT-OSS Special? GPT-OSS introduces capabilities that set it apart from other open-source LLMs: * **Mixture of Experts (MoE) Architecture**: The model comes in 20B and 120B parameter variants, but uses MoE to keep active parameters low while maintaining strong capabilities * **MXFP4 Quantization**: A novel 4-bit floating point format specifically designed for MoE layers, enabling efficient serving * **Attention Sinks**: Special attention mechanism that allows for longer context lengths without degrading output quality * **Harmony Response Format**: Built-in support for structured outputs like chain-of-thought reasoning and tool use. See examples from OpenAI [here](https://cookbook.openai.com/articles/openai-harmony) Please note that in vLLM, you can only run it on NVIDIA H100, H200, B200 as well as MI300x, MI325x, MI355x and Radeon AI PRO R9700 as of 6th August 2025 This tutorial covers the simplest variation of deploying this model using `vllm serve`. For more control, see the [OpenAI compatible endpoint with vLLM guide](/docs/v4/examples/openai-compatible-endpoint-vllm). ### Project Setup 1. Run the command, `cerebrium init gpt-oss` 2. Edit your toml file with the following settings ``` [cerebrium.deployment] name = "7-openai-gpt-oss" python_version = "3.12" docker_base_image_url = "nvidia/cuda:12.8.1-devel-ubuntu22.04" disable_auth = true include = ['./*', 'main.py', 'cerebrium.toml'] exclude = ['.*'] pre_build_commands = [ "apt-get update", "apt-get install -y curl", "curl -LsSf https://astral.sh/uv/install.sh | sh", "export PATH=\"$HOME/.local/bin:$PATH\" && uv pip install --pre vllm==0.10.1+gptoss --extra-index-url https://wheels.vllm.ai/gpt-oss/ --extra-index-url https://download.pytorch.org/whl/nightly/cu128 --index-strategy unsafe-best-match", "uv pip install huggingface_hub[hf_transfer]==0.34" ] [cerebrium.hardware] cpu = 8.0 memory = 18.0 compute = "HOPPER_H100" [cerebrium.scaling] min_replicas = 0 max_replicas = 5 cooldown = 30 replica_concurrency = 32 scaling_metric = "concurrency_utilization" [cerebrium.runtime.custom] port = 8000 entrypoint = ["sh", "-c", "export HF_HUB_ENABLE_HF_TRANSFER=1 && export VLLM_USER_V1=1 && vllm serve openai/gpt-oss-20b --enforce-eager"] ``` Key configuration details: * The `docker_base_image_url` is set to the cuda:12.8.1-devel image, which is large but necessary to include all required packages/libraries. * Pre-build commands install uv (a faster Python package installer) and the required vLLM packages. These commands execute at the start of the build process, before dependency installation begins, making them essential for setting up the build environment. Read more [here](https://docs.cerebrium.ai/container-images/defining-container-images#pre-build-commands). * Hardware is set to an H100 in the us-east-1 region. * Replica concurrency is 32, meaning a single H100 container handles 32 concurrent requests. * `vllm serve` turns the container into an OpenAI compatible server running on port 8000. ### Deploy & Test Deploy by running `cerebrium deploy`. Cerebrium creates the environment and downloads the model. Test the endpoint with the following request: ``` curl --location 'https://api.cerebrium.ai/v4/p-xxxxxxxx/7-openai-gpt-oss/v1/chat/completions' \ --header 'Content-Type: application/json' \ --header 'Accept: text/event-stream' \ --data '{"messages": [{"role": "user", "content": "hello how are you"}], "model": "Qwen/Qwen2.5-1.5B-Instruct", "stream": true}' ``` On the first request, a container spins up, loads the model, and streams the output. As of 6th August, this generates roughly 30 tokens per second. # Deploy a High Throughput Server for Embeddings and Reranking Source: https://cerebrium.ai/docs/v4/examples/high-throughput-embeddings Serve text embeddings, reranking, CLIP and ColPali models through a high throughput REST API on Cerebrium using the open source Infinity framework. This tutorial covers deploying a high-throughput, low-latency REST API for serving text-embeddings, reranking models, clip, clap, and colpali using the open-source framework [infinity](https://github.com/michaelfeil/infinity/tree/main). Infinity supports multiple GPUs/CPUs and frameworks. The inference server is built on PyTorch, optimum (ONNX/TensorRT), and CTranslate2, using FlashAttention for NVIDIA CUDA, AMD ROCM, CPU, AWS INF2, and APPLE MPS accelerators. It uses dynamic batching and dedicated tokenization worker threads. Find the final working version [here](https://github.com/CerebriumAI/examples/tree/master/14-embeddings/1-high-throughput) on GitHub. ### Project Setup Complete the [quickstart](/docs/getting-started/introduction) to install the CLI and create an account. 1. Run the command: `cerebrium init infinity-throughput` This creates two files: * main.py: The entrypoint code * cerebrium.toml: Container image and auto-scaling parameters Start by defining the container environment. Infinity has a public Docker image on [Dockerhub](https://hub.docker.com/r/michaelf34/infinity). Cerebrium requires Dockerhub authentication to pull images (even public ones). Sign in with the following command: ``` docker login -u your-dockerhub-username # Enter your password or access token when prompted ``` Add the following to cerebrium.toml ``` [cerebrium.deployment] name = "1-high-throughput" python_version = "3.11" docker_base_image_url = "michaelf34/infinity:0.0.77" disable_auth = true include = ['./*', 'main.py', 'cerebrium.toml'] exclude = ['.*'] ``` Autoscaling criteria vary by hardware type and model selection. Define them in the following `cerebrium.toml` sections: ``` [cerebrium.hardware] cpu = 6.0 memory = 12.0 compute = "AMPERE_A10" [cerebrium.scaling] min_replicas = 0 max_replicas = 2 cooldown = 30 replica_concurrency = 500 scaling_metric = "concurrency_utilization" [cerebrium.dependencies.pip] numpy = "latest" "infinity-emb[all]" = "0.0.77" optimum = ">=1.24.0,<2.0.0" transformers = "<4.49" click = "==8.1.8" fastapi = "latest" uvicorn = "latest" pandas = "latest" ``` The model runs on an Ampere A10, which handles up to 500 concurrent inputs. In main.py, create a class that handles embedding model functionality using the Infinity framework. This example uses multiple models to demonstrate the range of supported functionality. ```python theme={null} from infinity_emb import AsyncEngineArray, EngineArgs class InfinityModel: def __init__(self): self.model_ids = [ "jinaai/jina-clip-v1", "michaelfeil/bge-small-en-v1.5", "mixedbread-ai/mxbai-rerank-xsmall-v1", "philschmid/tiny-bert-sst2-distilled" ] self.engine_array = None def _get_array(self): return AsyncEngineArray.from_args([ EngineArgs(model_name_or_path=model, model_warmup=False) for model in self.model_ids ]) async def setup(self): print(f"Setting up models: {self.model_ids}") self.engine_array = self._get_array() await self.engine_array.astart() print("All models loaded successfully!") model = InfinityModel() ``` Model loading can take time, so FastAPI provides greater control over readiness. Cerebrium supports custom ASGI servers. Add the following to main.py ```python theme={null} from fastapi import FastAPI, Body app = FastAPI(title="High-Throughput Embedding Service") @app.on_event("startup") async def startup_event(): """Initialize models on container startup""" await model.setup() @app.get("/health") async def health(): return {"status": "healthy"} @app.get("/ready") async def ready(): """Readiness endpoint to report model initialization state.""" is_ready = model.engine_array is not None return {"ready": is_ready} ``` Infinity supports text embeddings, image embeddings, reranking, and classification. Create separate endpoints for each: ```python theme={null} def embeddings_to_list(embeddings: list) -> list: """Convert list of numpy arrays to list of lists.""" return [e.tolist() for e in embeddings] @app.post("/embed") async def embed(sentences: list[str] = Body(...), model_index: int = Body(1)): """Generate embeddings using the specified model.""" engine = model.engine_array[model_index] embeddings, usage = await engine.embed(sentences=sentences) return { "embeddings": to_json(embeddings), "usage": to_json(usage), "model": model.model_ids[model_index] } @app.post("/image_embed") async def image_embed(image_urls: list[str] = Body(...), model_index: int = Body(0)): """Generate embeddings for images using CLIP model.""" engine = model.engine_array[model_index] embeddings, usage = await engine.image_embed(images=image_urls) return { "embeddings": to_json(embeddings), "usage": to_json(usage), "model": model.model_ids[model_index] } @app.post("/rerank") async def rerank(query: str = Body(...), docs: list[str] = Body(...), model_index: int = Body(2)): """Rerank documents based on query relevance.""" engine = model.engine_array[model_index] rankings, usage = await engine.rerank(query=query, docs=docs) return { "rankings": to_json(rankings), "usage": to_json(usage), "model": model.model_ids[model_index] } @app.post("/classify") async def classify(sentences: list[str] = Body(...), model_index: int = Body(3)): """Classify text sentiment.""" engine = model.engine_array[model_index] classes, usage = await engine.classify(sentences=sentences) return { "classifications": to_json(classes), "usage": to_json(usage), "model": model.model_ids[model_index] } ``` This creates a multi-purpose embedding server. Update cerebrium.toml to point to the FastAPI server by adding the following section: ``` [cerebrium.runtime.custom] port = 5000 entrypoint = ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "5000"] healthcheck_endpoint = "/health" readycheck_endpoint = "/ready" ``` Deploy with `cerebrium deploy`. After deployment, run inference with a command like: ``` curl --location 'https://api.cerebrium.ai/v4/p-xxxxxxxx/infinity-throughput/image_embed' \ --header 'Content-Type: application/json' \ --data '{"image_urls": ["https://www.borrowmydoggy.com/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F4ij0poqn%2Fproduction%2Fe24bfbd855cda99e303975f2bd2a1bf43079b320-800x600.jpg&w=1080&q=80"]}' ``` The response looks like: ``` { "embeddings": [ [ -0.05284368246793747, 0.0011637501884251833, -0.029046623036265373, .... ] ] } ``` The result is a scalable, multi-purpose embedding/reranking server. # Build a LangChain agent with LangSmith monitoring Source: https://cerebrium.ai/docs/v4/examples/langchain-langsmith Build an executive assistant agent with LangChain tool calling, monitor it in LangSmith and deploy it on Cerebrium to manage Cal.com bookings. This tutorial builds Cal-vin, an executive assistant that manages calendar appointments (via Cal.com) with employees, customers, partners, and friends. It uses the LangChain SDK for agent creation and the LangSmith platform for monitoring scheduling activities and identifying failure points. The app deploys on Cerebrium for seamless scaling. You can find the final version of the code [here](https://github.com/CerebriumAI/examples/tree/master/4-integrations/2-tool-calling-langsmith). ### Concepts This app requires calendar interaction based on user instructions: an ideal use case for an agent with function (tool) calling capabilities. LangChain provides extensive agent support, and its companion tool LangSmith makes monitoring integration straightforward. A tool refers to any framework, utility, or system with defined functionality for specific use cases, such as searching Google or retrieving credit card transactions. Key LangChain concepts: `ChatModel.bind_tools()`: Attaches tool definitions to model calls. While providers have different tool definition formats, LangChain provides a standard interface for versatility. Accepts tool definitions as dictionaries, Pydantic classes, LangChain tools, or functions, telling the LLM how to use each tool. ```python theme={null} @tool def exponentiate(x: float, y: float) -> float: """Raise 'x' to the 'y'.""" return x**y ``` `AIMessage.tool_calls`: An attribute on AIMessage that provides easy access to model-initiated tool calls, specifying invocations in the bind\_tools format: ```python theme={null} # -> AIMessage( # content=..., # additional_kwargs={...}, # tool_calls=[{'name': 'exponentiate', 'args': {'y': 2.743, 'x': 5.0}, 'id': '54c166b2-f81a-481a-9289-eea68fc84e4f'}] # response_metadata={...}, # id='...' # ) ``` `create_tool_calling_agent()`: Unifies the above concepts to work across different provider formats, enabling easy model switching. ```python theme={null} agent = create_tool_calling_agent(llm, tools, prompt) agent_executor = AgentExecutor(agent=agent, tools=tools, verbose=True) agent_executor.invoke({"input": "what's 3 plus 5 raised to the 2.743. also what's 17.24 - 918.1241", }) ``` ### Cal.com Setup [Cal.com](https://cal.com) provides the calendar management foundation. Create an account [here](https://app.cal.com/signup) if needed. Cal serves as the source of truth. Updates to time zones or working hours automatically reflect in the assistant's responses. After creating your account: 1. Navigate to "API keys" in the sidebar 2. Create an API key without expiration 3. Test the setup with a CURL request (replace these variables): * Username * API key * dateFrom and dateTo Cal.com API Keys ```curl theme={null} curl --location 'https://api.cal.com/v1/availability?apiKey=cal_live_xxxxxxxxxxxxxx&dateFrom=2024-04-15T00%3A00%3A00.000Z&dateTo=2024-04-22T00%3A00%3A00.000Z&username=michael-louis-xxxx' ``` You should get a response similar to the following: ``` { "busy": [ { "start": "2024-04-15T13:00:00.000Z", "end": "2024-04-15T13:30:00.000Z" }, { "start": "2024-04-22T13:00:00.000Z", "end": "2024-04-22T13:30:00.000Z" }, { "start": "2024-04-29T13:00:00.000Z", "end": "2024-04-29T13:30:00.000Z" }, .... ], "timeZone": "America/New_York", "dateRanges": [ { "start": "2024-04-15T13:45:00.000Z", "end": "2024-04-15T16:00:00.000Z" }, { "start": "2024-04-15T16:45:00.000Z", "end": "2024-04-15T19:45:00.000Z" }, .... { "start": "2024-04-19T18:45:00.000Z", "end": "2024-04-19T21:00:00.000Z" } ], "oooExcludedDateRanges": [ ], "workingHours": [ { "days": [ 1, 2, 3, 4, 5 ], "startTime": 780, "endTime": 1260, "userId": xxxx } ], "dateOverrides": [], "currentSeats": null, "datesOutOfOffice": {} } ``` The API key is now confirmed working and pulling calendar information. The API calls used later in this tutorial are: * **/availability**: Get your availability * **/bookings**: Book a slot ### Cerebrium Setup Set up Cerebrium: 1. Sign up [here](https://dashboard.cerebrium.ai/register) 2. Follow installation docs [here](https://docs.cerebrium.ai/getting-started/installation) 3. Create a starter project: ```bash theme={null} cerebrium init agent-tool-calling ``` This creates: * `main.py`: Entrypoint file * `cerebrium.toml`: Build and environment configuration Add these pip packages to your `cerebrium.toml`: ``` [cerebrium.dependencies.pip] pydantic = "latest" langchain = "latest" pytz = "latest" ##this is used for timezones openai = "latest" langchain_openai = "latest" ``` Set up API keys: 1. OpenAI GPT-3.5: * Sign up at [OpenAI](https://openai.com/) * Create API key [here](https://platform.openai.com/api-keys) (format: sk\_xxxxx) 2. Add secrets in Cerebrium dashboard: * Navigate to "Secrets" * Add keys: * `CAL_API_KEY`: Your Cal.com API key * `OPENAI_API_KEY`: Your OpenAI API key Cerebrium Secrets Dashboard ### Agent Setup Create two tool functions in `main.py` for calendar management: 1. Get availability tool 2. Book slot tool The Cal.com API provides: * Busy time slots * Working hours per day Below is the code to achieve this: ```python theme={null} from langchain_core.tools import tool import os import requests from cal import find_available_slots @tool def get_availability(fromDate: str, toDate: str) -> float: """Get my calendar availability using the 'fromDate' and 'toDate' variables in the date format '%Y-%m-%dT%H:%M:%S.%fZ'""" url = "https://api.cal.com/v1/availability" params = { "apiKey": os.environ.get("CAL_API_KEY"), "username": "xxxxx", "dateFrom": fromDate, "dateTo": toDate } response = requests.get(url, params=params) if response.status_code == 200: availability_data = response.json() available_slots = find_available_slots(availability_data, fromDate, toDate) return available_slots else: return {} ``` The code above: 1. Uses `@tool` decorator to identify functions as LangChain tools 2. Includes docstrings explaining functionality and required inputs 3. Uses `find_available_slots` helper function to format Cal.com API responses into readable time slots ‍ The book\_slot tool follows a similar pattern. It books a slot based on the selected time/day. Get the eventTypeId from the dashboard by selecting an event and grabbing the ID from the URL. ```python theme={null} @tool def book_slot(datetime: str, name: str, email: str, title: str, description: str) -> float: """Book a meeting on my calendar at the requested date and time using the 'datetime' variable. Get a description about what the meeting is about and make a title for it""" url = "https://api.cal.com/v1/bookings" params = { "apiKey": os.environ.get("CAL_API_KEY"), "username": "xxxx", "eventTypeId": "xxx", "start": datetime, "responses": { "name": name, "email": email, "guests": [], "metadata": {}, "location": { "value": "inPerson", "optionValue": "" } }, "timeZone": "America/New York", "language": "en", "status": "PENDING", "title": title, "description": description, } response = requests.post(url, params=params) if response.status_code == 200: booking_data = response.json() return booking_data else: print('error') print(response) return {} ``` With both tools created, set up the agent in `main.py`: ```python theme={null} from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder from langchain_core.tools import tool from langchain.agents import create_tool_calling_agent, AgentExecutor from langchain_openai import ChatOpenAI prompt = ChatPromptTemplate.from_messages([ ("system", "you're a helpful assistant managing the calendar of Michael Louis. You need to book appointments for a user based on available capacity and their preference. You need to find out if the user is: From Michaels team, a customer of Cerebrium or a friend or entrepreneur. If the person is from his team, book a morning slot. If its a potential customer for Cerebrium, book an afternoon slot. If its a friend or entrepreneur needing help or advice, book a night time slot. If none of these are available, book the earliest slot. Do not book a slot without asking the user what their preferred time is. Find out from the user, their name and email address."), MessagesPlaceholder(variable_name="chat_history"), ("human", "{input}"), MessagesPlaceholder(variable_name="agent_scratchpad"), ]) tools = [get_availability, book_slot] llm = ChatOpenAI(model="gpt-3.5-turbo-0125", temperature=0, api_key=os.environ.get("OPENAI_API_KEY")) agent = create_tool_calling_agent(llm, tools, prompt) agent_executor = AgentExecutor(agent=agent, tools=tools, verbose=True) ``` The agent executor consists of: * The prompt template: * Defines the agent’s role, goals, and situational behavior. More precise instructions yield better results. * Chat History stores previous messages for conversation context. * Input receives new input from the end user. * The GPT-3.5 model serves as the LLM. Swap to Anthropic or any other provider by replacing this one line. LangChain makes this seamless. * Finally, these components combine with the tools to create an agent executor. ### Chatbot Setup The above code only handles a single question. Finding a mutually suitable time requires a multi-turn conversation. LangChain’s RunnableWithMessageHistory() adds tool calling capabilities and message memory. It stores previous replies in the chat\_history variable (from the prompt template) and ties them to a session identifier, so the API remembers information per user/session: ```python theme={null} from langchain.memory import ChatMessageHistory from langchain_core.runnables.history import RunnableWithMessageHistory demo_ephemeral_chat_history_for_chain = ChatMessageHistory() conversational_agent_executor = RunnableWithMessageHistory( agent_executor, lambda session_id: demo_ephemeral_chat_history_for_chain, input_messages_key="input", output_messages_key="output", history_messages_key="chat_history", ) ``` Run a local test to verify everything works: ```python theme={null} class Item(BaseModel): prompt: str session_id: str def predict(item, run_id, logger): item = Item(**item) output = conversational_agent_executor.invoke( { "input": user_input, }, {"configurable": {"session_id": item.session_id}}, ) return {"result": output} # return your results if __name__ == "__main__": while True: user_input = input("Enter the input (or type 'exit' to stop): ") if user_input.lower() == 'exit': break result = predict({"prompt": user_input, "session_id": "12345"}, "test", logger=None) print(result) ``` This code: * Defines a Pydantic object specifying the expected API parameters: user prompt and session ID. * The predict function (Cerebrium’s API entry point) passes the prompt and session ID to the agent and returns results. ‍ Install pip dependencies locally: `pip install pydantic langchain pytz openai langchain_openai langchain-community`, then run `python main.py`. Replace secrets with actual values when running locally. Output looks similar to: Langchain Agent Continuing the conversation eventually results in a booked slot. ### Integrate LangSmith Production monitoring is crucial for agent applications with nondeterministic workflows. LangSmith, a LangChain tool for logging, debugging, and monitoring, tracks performance and surfaces edge cases. Learn more [here](https://docs.smith.langchain.com/monitoring). Set up LangSmith monitoring: 1. Add LangSmith to `cerebrium.toml` dependencies 2. Create a free LangSmith account [here](https://smith.langchain.com/) 3. Generate API key (click gear icon in bottom left) Set the following environment variables at the top of main.py. Add the API key to Cerebrium secrets. ```python theme={null} import os os.environ['LANGCHAIN_TRACING_V2']="true" os.environ['LANGCHAIN_API_KEY']=os.environ.get("LANGCHAIN_API_KEY") ``` Enable tracing by adding the `@traceable` decorator to functions. LangSmith automatically tracks tool invocations and OpenAI responses through function traversal. Add the decorator to the `predict` function and any independently instantiated functions: ``` from langsmith import traceable @traceable def predict(item, run_id, logger): ``` LangSmith is now set up. Run `python main.py` and test booking an appointment. After a successful test run, data populates in LangSmith: LangSmith Runs Dashboard The Runs tab shows all runs (invocations/API requests). In 1 above, the function name appears, with input set to the Cerebrium RunID (set to “test”). The input and total latency of the run are also visible. LangSmith supports various data automations: * Data annotation for positive/negative case labeling * Dataset creation for model training * Online LLM-based evaluation (rudeness, topic analysis) * Webhook endpoint triggers * Additional features Set automations by clicking the “Add rule” button above (2) and specifying conditions. Rule options include a filter, sampling rate, and action. Section 3 shows overall project metrics: run count, error rate, latency, etc. LangSmith Threads provide clean conversation tracking between agents and users. Track conversation evolution and investigate anomalies through trace analysis. Each thread links to its session ID. LangSmith Threads The Monitor tab shows agent performance metrics: trace count, LLM call success rate, time to first token, and more. LangSmith Performance Monitoring LangSmith offers straightforward integration with extensive functionality. Beyond the basics covered here, it supports the full application feedback loop: data collection/annotation → monitoring → iteration. ### Deploy to Cerebrium Deploy to Cerebrium by running `cerebrium deploy`. Delete the `name == "main"` block first (used only for local testing). After successful deployment: Cerebrium Deployment The API endpoint is now live. The agent remembers conversations as long as the session ID remains the same. Cerebrium automatically scales the application based on demand, with pay-per-use compute. ``` { "run_id": "UHCJ_GkTKh451R_nKUd3bDxp8UJrcNoPWfEZ3AYiqdY85UQkZ6S1vg==", "status_code": 200, "result": { "result": { "input": "Hi! I would like to book a time with Michael the 18th of April 2024.", "chat_history": [], "output": "Michael is available on the 18th of April 2024 at the following times:\n1. 13:00 - 13:30\n2. 14:45 - 17:00\n3. 17:45 - 19:00\n\nPlease let me know your preferred time slot. Are you from Michael's team, a potential customer of Cerebrium, or a friend/entrepreneur seeking advice?" } }, "run_time_ms": 6728.828907012939, "process_time_ms": 6730.178117752075 } ``` You can find the final version of the code [here](https://github.com/CerebriumAI/examples/tree/master/4-integrations/2-tool-calling-langsmith). ### Future Enhancements Consider implementing: 1. Response streaming for seamless user experience 2. Email integration for context-aware scheduling when Cal-vin is tagged 3. Voice capabilities for phone-based scheduling ### Conclusion LangChain, LangSmith, and Cerebrium together enable scalable agent deployment. LangChain handles LLM orchestration, tooling, and memory management. LangSmith provides production monitoring and edge case identification. Cerebrium offers pay-as-you-go scaling across hundreds or thousands of CPU/GPUs. Tag **@cerebriumai** in extensions to the code repository to share with the community. # Outbound Agent with LiveKit Source: https://cerebrium.ai/docs/v4/examples/livekit-outbound-agent Build an outbound AI voice agent with LiveKit and Twilio SIP trunking on Cerebrium that makes calls and warm transfers callers to human agents Voice agents are transforming business operations by introducing efficiencies and personalization for each customer interaction. While most use cases focus on agents receiving calls, this tutorial covers outbound voice AI agents and the use cases they unlock. Outbound agents handle tasks like appointment scheduling, lead qualification, and customer follow-ups while maintaining a human-like, conversational tone. They replace manual efforts with intelligent automation, improving customer engagement in sales, healthcare, hospitality, and collections. This tutorial sets up an outbound calling agent that does a warm transfer: handing off the call to a real person using LiveKit. A fitting example: a support center where an agent collects data before connecting to a real person. The final code implementation and complete example can be found in our examples GitHub repository [here](https://github.com/CerebriumAI/examples/tree/master/6-voice/7-outbound-livekit-agent). ### Cerebrium Setup Install the Cerebrium CLI and create a project: ``` pip install cerebrium --upgrade cerebrium login cerebrium init livekit-voice-agent ``` This creates a folder with two files: * **main.py** - The entrypoint file where application code lives. * **cerebrium.toml** - A configuration file that contains all build and environment settings. These files are populated later in the tutorial. ### LiveKit Setup [LiveKit](https://livekit.io/kitt) is a platform for building real-time audio and video applications. Start by creating a LiveKit account. You can do this by signing up to an account [here](https://cloud.livekit.io/login?r=%2F) (they have a generous free tier). Once the account is created, run the following commands to install the CLI package (Homebrew can be installed from [here](https://brew.sh/)). ``` brew update && brew install livekit-cli ``` Authenticate the CLI with the newly created account: ``` lk cloud auth ``` Create a **.env** file with LiveKit credentials. Get these from the [LiveKit dashboard](https://cloud.livekit.io/): click the Settings tab for the project URL, then the Keys tab to create a new key and copy the API key and secret: ``` LIVEKIT_URL= LIVEKIT_API_KEY= LIVEKIT_API_SECRET= ``` LiveKit Dashboard The LiveKit instance is now set up as the transport layer for the agent. ### Twilio Setup Setting up an outbound calling agent requires a SIP trunk in Twilio. A SIP trunk is a virtual phone line that connects the app to the traditional phone network, enabling outbound calls to any phone number. It routes calls properly and reliably, acting as the bridge between the app and the global telephony system. 1. **Log in to your Twilio Console**: Go to [Twilio Console](https://console.twilio.com/) and log in with your credentials. If you don’t already have an account, create one. 2. **Add Phone Numbers**: Under the “Develop” tab is the "Phone numbers" section. Navigate to “Active numbers” and purchase a number if you don’t already have one. Toll free numbers won't work for this use case, so ensure that a "Local" number exists. 3. **Install the Twilio CLI and authenticate your CLI:** ``` brew tap twilio/brew && brew install twilio twilio login ``` 4. **Create a SIP trunk:** The domain name for your SIP trunk must end in [pstn.twilio.com](http://pstn.twilio.com/). For example to create a trunk named My test trunk with the domain name [my-test-trunk.pstn.twilio.com](http://my-test-trunk.pstn.twilio.com/), run the following command: ``` twilio api trunking v1 trunks create \ --friendly-name "My test trunk" \ --domain-name "my-test-trunk.pstn.twilio.com" ``` 5. **Create your credentials:** To secure the SIP trunk, create a credential list in the Twilio console dashboard. Navigate to “Voice”, then “Manage”, then “Credential lists”. Click the plus icon and add a friendly name, username, and password. Save these credentials. They are required in a later step. Twilio Dashboard 6. Associate your SIP with your newly created credentials. Copy the values for your Account SID and Auth Token from the Twilio console. ``` export TWILIO_ACCOUNT_SID="" export TWILIO_AUTH_TOKEN="" ``` With the trunk configured, register Twilio’s outbound trunk with LiveKit. SIP trunking providers typically require authentication when accepting outbound SIP requests to ensure only authorized users can make calls. Create a file named `outbound-trunk.json` using the Twilio phone number you purchased in the previous step, trunk domain name, and `username` and `password`. ``` { "trunk": { "name": "My outbound trunk", "address": ".pstn.twilio.com", "numbers": [""], "auth_username": "cerebrium", "auth_password": "" } } ``` Run the following in your CLI: `lk sip outbound create outbound-trunk.json` The output returns the trunk ID, which is required in a later step. ### Services Setup The following services power the outbound AI agent: * [Deepgram](https://www.deepgram.com/) for the Speech-to-Text service * [OpenAI](https://openai.com/) for the LLM * [Cartesia](https://www.cartesia.com/) for the Text-to-Speech service Each service offers a generous free tier: * **Deepgram** You can sign up for a Deepgram account [here](https://www.notion.so/Livekit-Outbound-agent-13ebfd664fb180c5b02bf9c71f72b23d?pvs=21). Straight from the dashboard, you can create an API key. Store this value for later. Deepgram Dashboard * **OpenAI** You can sign up for an OpenAI account [here](https://platform.openai.com/signup). You can then click “Dashboard” top right and then “API keys” in the left sidebar. Create an API key and store this value to use later. OpenAI Dashboard * **Cartesia** You can sign up for a Cartesia account [here](https://www.notion.so/Livekit-Outbound-agent-13ebfd664fb180c5b02bf9c71f72b23d?pvs=21). Once you log in, you can click API keys in the left sidebar and create a new API key. Cartesia Dashboard Finally, create a **.env** file that contains the environment variables from all the steps above. ``` LIVEKIT_API_KEY="" LIVEKIT_API_SECRET="" LIVEKIT_URL="" DEEPGRAM_API_KEY="" OPENAI_API_KEY="" CARTESIA_API_KEY="" SIP_TRUNK_ID="" ``` Create a **requirements.txt** with the following Python packages: ``` livekit-agents>=0.11.1 livekit-plugins-openai>=0.10.5 livekit-plugins-deepgram livekit-plugins-cartesia livekit-plugins-silero>=0.7.3 livekit-plugins-rag>=0.2.2 python-dotenv~=1.0 aiofile~=3.8.8 fastapi uvicorn ``` Run `pip install -r requirements.txt`. FastAPI is explained later. Create a file called `main.py` for the agent code. The directory structure should look like this: * main.py * .env * requirements.txt * outbound-trunk.json Add the following code to your main.py ``` from fastapi import FastAPI from dotenv import load_dotenv from livekit.agents import ( AutoSubscribe, JobContext, JobProcess, WorkerOptions, WorkerType, cli, llm, metrics ) from livekit import api from livekit.agents.pipeline import VoicePipelineAgent from livekit.plugins import openai, deepgram, silero, cartesia import os import asyncio import sys import logging load_dotenv() logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s') app = FastAPI() async def entrypoint(ctx: JobContext): initial_ctx = llm.ChatContext().append( role="system", text="You are a voice assistant created by Cerebrium. Your interface with users will be voice. You should use short and concise responses, and avoiding usage of unpronouncable punctuation.", ) await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY) agent = VoicePipelineAgent( vad=silero.VAD.load(), # flexibility to use any models stt=deepgram.STT(model="nova-2-general"), llm=openai.LLM( model="gpt-4o-mini", temperature=0.5, ), tts=cartesia.TTS(), # initial ChatContext with system prompt chat_ctx=initial_ctx, # whether the agent can be interrupted allow_interruptions=True, # sensitivity of when to interrupt interrupt_speech_duration=0.5, interrupt_min_words=0, # minimal silence duration to consider end of turn min_endpointing_delay=0.3, fnc_ctx=fnc_ctx ) usage_collector = metrics.UsageCollector() @agent.on("metrics_collected") def _on_metrics_collected(mtrcs: metrics.AgentMetrics): metrics.log_metrics(mtrcs) usage_collector.collect(mtrcs) async def log_usage(): summary = usage_collector.get_summary() print(f"Usage: ${summary}") ctx.add_shutdown_callback(log_usage) agent.start(ctx.room) await asyncio.sleep(1.2) await agent.say("Hey, how can I help you today?", allow_interruptions=True) if __name__ == '__main__': if len(sys.argv) == 1: sys.argv.append('start') cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint, worker_type=WorkerType.ROOM, port=8600)) ``` This LiveKit setup does the following: * Specifies an initial system prompt to guide the AI agent * Lists the services and configures response timing settings * Collects and displays call metrics throughout, with a summary logged when the call ends * Starts the agent with an initial greeting To connect the customer to a real person when a specific stage is reached or the user requests it, use function calling. Add the following code to the entrypoint: ``` #Add this below the initial_ctx line fnc_ctx = llm.FunctionContext() @fnc_ctx.ai_callable() async def transfer_call( # by using the Annotated type, arg description and type are available to the LLM ): """Called when the receiver would like to be transferred to a real person. This function will add another participant to the call.""" await create_sip_participant("", "Test SIP Room") await agent.say("Connecting you to my colleague - please hold on", allow_interruptions=True) await ctx.room.disconnect() ``` The **@ai\_callable decorator** tells the LLM this function is available, and the docstring describes when to use it. When triggered, the function makes an outbound call to a phone number and disconnects the agent from the room, leaving just the two users in the call. The `create_sip_participant()` function is defined in the next step. To create an outbound call and add a user to the call, add the following code to `main.py` above `entrypoint()`: ``` async def create_sip_participant(phone_number, room_name): LIVEKIT_URL = os.getenv('LIVEKIT_URL') LIVEKIT_API_KEY = os.getenv('LIVEKIT_API_KEY') LIVEKIT_API_SECRET = os.getenv('LIVEKIT_API_SECRET') SIP_TRUNK_ID = os.getenv('SIP_TRUNK_ID') livekit_api = api.LiveKitAPI( LIVEKIT_URL, LIVEKIT_API_KEY, LIVEKIT_API_SECRET ) sip_trunk_id = SIP_TRUNK_ID try: await livekit_api.sip.create_sip_participant( api.CreateSIPParticipantRequest( sip_trunk_id=sip_trunk_id, sip_call_to=phone_number, room_name=room_name, participant_identity=f"sip_{phone_number}", participant_name="SIP Caller" ) ) await livekit_api.aclose() return f"Call initiated to {phone_number}" except Exception as e: await livekit_api.aclose() return f"Error: {str(e)}" ``` To test locally, create a `test.py` file: ``` import asyncio from main import create_sip_participant async def main(number: str): room_name = "Test SIP Room" result = await create_sip_participant(number, room_name) print(result) if __name__ == '__main__': asyncio.run(main("")) ``` To test locally, run `python main.py` in one terminal. This keeps the agents running as an open process: LiveKit processes Once the job processes initialize, open a separate terminal and run `python test.py`. This initiates a call to the provided phone number. Note: calls are limited to the region the number is purchased from. With local testing working, deploy on Cerebrium for production-ready autoscaling. ### Deploy on Cerebrium Specify the Cerebrium environment using a Dockerfile: ``` FROM debian:bookworm-slim # Install Python and pip RUN apt-get update && apt-get install -y \ python3.11 \ python3-pip \ && rm -rf /var/lib/apt/lists/* # Create and set working directory WORKDIR /cortex # Copy requirements and install dependencies COPY requirements.txt . RUN pip install --break-system-packages -r requirements.txt # Copy application code COPY . . # Download model files RUN python3 main.py download-files # Set environment variables ENV HF_HOME=/cortex/.cache/ # Expose port EXPOSE 8600 # Set entrypoint CMD ["python3", "main.py", "start"] ``` Edit `cerebrium.toml` with the following: ``` [cerebrium.deployment] name = "outbound-livekit-agent" python_version = "3.11" docker_base_image_url = "" disable_auth = false include = ['./*', 'main.py', 'cerebrium.toml'] exclude = ['.*'] [cerebrium.hardware] cpu = 2 memory = 8.0 compute = "CPU" [cerebrium.scaling] min_replicas = 1 max_replicas = 5 cooldown = 30 replica_concurrency = 1 [cerebrium.dependencies.paths] pip = "requirements.txt" [cerebrium.runtime.custom] port = 8600 dockerfile_path = "./Dockerfile" ``` This defines the hardware, scaling parameters, dependencies, and the port and Dockerfile for the application. In the Cerebrium dashboard, go to the Secrets tab in the left sidebar and upload the `.env` file. Cerebrium imports secrets as standard environment variables, making local and cloud testing seamless. Cerebrium Secrets Run the following command: ``` cerebrium deploy ``` This installs all necessary packages and deploys the application. LiveKit workers run in the cloud. Test by running the `test.py` script locally and checking logs on the Cerebrium dashboard. ### Further Improvements To reduce latency, Cerebrium partners with Deepgram and Rime to run STT and TTS models locally alongside the LiveKit worker, reducing latency by \~400ms. This tutorial set up an outbound calling agent that performs warm transfers to live agents using LiveKit. The solution supports diverse use cases such as pre-call data collection, sales outreach, and customer follow-ups by integrating LiveKit, Twilio, and AI-driven STT, TTS, and LLM services. Find the final example on [GitHub](https://github.com/CerebriumAI/examples/tree/master/6-voice/7-outbound-livekit-agent). Ask questions or give feedback in the [Discord](https://discord.gg/ATj6USmeE2) community. # OpenAI compatible vLLM endpoint Source: https://cerebrium.ai/docs/v4/examples/openai-compatible-endpoint-vllm Deploy an OpenAI compatible endpoint for open source LLMs like Llama 3.1 using vLLM on Cerebrium serverless GPUs with streaming responses This tutorial creates an OpenAI-compatible endpoint that works with any open-source model. Use existing OpenAI code with Cerebrium serverless functions by changing just two lines of code. To see the final code implementation, you can view it [here](https://github.com/CerebriumAI/examples/tree/master/5-large-language-models/1-openai-compatible-endpoint) ### Cerebrium Setup Create a Cerebrium account by signing up [here](https://dashboard.cerebrium.ai/register) and follow the [installation docs](https://docs.cerebrium.ai/getting-started/installation). Run the following command to create the Cerebrium starter project: `cerebrium init 1-openai-compatible-endpoint`. This creates two files: * `main.py`: The entrypoint file where application code lives * `cerebrium.toml`: A configuration file that contains all build and environment settings Add the following pip packages and hardware requirements to your `cerebrium.toml` to create your deployment environment: ```toml theme={null} [cerebrium.hardware] cpu = 2 memory = 12.0 compute = "AMPERE_A10" [cerebrium.dependencies.pip] vllm = "latest" pydantic = "latest" ``` Define the imports and initialize the model. This example uses Meta's Llama 3.1 model, which requires Hugging Face authorization. Add the HF token to secrets in the Cerebrium dashboard, then add this code to `main.py`: ```python theme={null} from vllm import SamplingParams, AsyncLLMEngine from vllm.engine.arg_utils import AsyncEngineArgs from pydantic import BaseModel from typing import Any import time import json import os from huggingface_hub import login # Your huggingface token (HF_AUTH_TOKEN) should be stored in your project secrets on your dashboard login(token=os.environ.get("HF_AUTH_TOKEN")) engine_args = AsyncEngineArgs( model="meta-llama/Meta-Llama-3.1-8B-Instruct", gpu_memory_utilization=0.9, # Increase GPU memory utilization max_model_len=8192 # Decrease max model length ) engine = AsyncLLMEngine.from_engine_args(engine_args) ``` Next, define the required output format for OpenAI endpoints using Pydantic: ```python theme={null} class Message(BaseModel): role: str content: str class ChatCompletionResponse(BaseModel): id: str object: str created: int model: str choices: List[Any] async def run(messages: list, model: str, run_id: str, stream: bool = True, temperature: float = 0.8, top_p: float = 0.95): prompt = " ".join([f"{Message(**msg).role}: {Message(**msg).content}" for msg in messages]) sampling_params = SamplingParams(temperature=temperature, top_p=top_p) results_generator = engine.generate(prompt, sampling_params, run_id) previous_text = "" full_text = "" # Collect all generated text here async for output in results_generator: prompt = output.outputs new_text = prompt[0].text[len(previous_text):] previous_text = prompt[0].text full_text += new_text # Append new text to full_text response = ChatCompletionResponse( id=run_id, object="chat.completion", created=int(time.time()), model=model, choices=[{ "text": new_text, "index": 0, "logprobs": None, "finish_reason": prompt[0].finish_reason or "stop" }] ) print(response.model_dump()) yield f"data: {json.dumps(response.model_dump())}\n\n" # Send the final [DONE] message yield "data: [DONE]\n\n" ``` The function: * Takes parameters through its signature, with optional and default values available * Automatically receives a unique `run_id` for each request * Processes the entire prompt through the model * Streams results when `stream=True` using async functionality * Returns the complete result at the end if streaming is disabled ## Deploy & Inference To deploy the model use the following command: ```bash theme={null} cerebrium deploy ``` After deployment, a curl command like this appears: ```curl theme={null} curl --location 'https://api.cerebrium.ai/v4/p-xxxxxxxx/5-openai-compatible-endpoint/{function}' \ --header 'Content-Type: application/json' \ --header 'Authorization: Bearer ' \ --data '{"..."}' ``` In Cerebrium, each function name becomes an endpoint (ending with `/run`). While OpenAI-compatible endpoints typically end with `/chat/completions`, all Cerebrium endpoints are OpenAI-compatible. Call the endpoint as follows: ```python theme={null} import os from openai import OpenAI client = OpenAI( base_url="https://api.cerebrium.ai/v4/p-xxxxxxxx/5-openai-compatible-endpoint/run", api_key="", ) chat_completion = client.chat.completions.create( messages=[ {"role": "user", "content": "What is a mistral?"}, {"role": "assistant", "content": "A mistral is a type of cold, dry wind that blows across the southern slopes of the Alps from the Valais region of Switzerland into the Ligurian Sea near Genoa. It is known for its strong and steady gusts, sometimes reaching up to 60 miles per hour."}, {"role": "user", "content": "How does the mistral wind form?"} ], model="meta-llama/Meta-Llama-3.1-8B-Instruct", stream=True ) for chunk in chat_completion: print(chunk) print("Finished receiving chunks.") ``` Set the base URL to the one from the deploy command (ending in `/run`). Use the JWT token from either the curl command or the Cerebrium dashboard's API Keys section. # Real-time Voice Agent Source: https://cerebrium.ai/docs/v4/examples/realtime-voice-agents Build a low latency voice AI agent with PipeCat, Deepgram and a self hosted vLLM Llama endpoint on Cerebrium that responds in about 500ms This tutorial creates a real-time voice agent that responds to queries via speech in \~500ms. The implementation supports swapping in any Large Language Model (LLM) or Text-to-Speech (TTS) model, making it ideal for voice-based use cases like customer support bots and receptionists. The app uses [PipeCat](https://www.pipecat.ai/), a framework that handles component integration, user interruptions, and audio data processing. The example joins a meeting room with a voice agent using [Daily](https://daily.co) (PipeCat's creators) and deploys on Cerebrium for scaling. The application has 3–4 parts: * A Pipecat agent that acts as the orchestrator * A Deepgram TTS/STT service (requires a Deepgram Enterprise account) * A self-hosted LLM using the vLLM framework Low latency is achieved because each service is hosted within Cerebrium. Communication across containers is less than 10ms with no network latency overhead. Realtime Voice Agents You can find the final version of the code [here](https://github.com/CerebriumAI/examples/tree/master/6-voice/2-realtime-voice-agent) Create a Cerebrium account by signing up [here](https://dashboard.cerebrium.ai/register) and follow the [installation docs](https://docs.cerebrium.ai/getting-started/installation). ### Deepgram Deployment See the [Partner Services page](/docs/partner-services/deepgram) to deploy a Deepgram service on Cerebrium. You need a Deepgram Enterprise License to deploy Deepgram on Cerebrium. Otherwise, use their API endpoint below. ### LLM Deployment The LLM is an OpenAI-compatible Llama-3 endpoint using the vLLM framework. For low TTFT, a quantized version is used (RedHatAI/Meta-Llama-3.1-8B-Instruct-quantized.w8a8). Run `cerebrium init llama-llm` and add the following to `cerebrium.toml`: ``` [cerebrium.deployment] name = "llama-llm" python_version = "3.11" docker_base_image_url = "debian:bookworm-slim" include = ["./*", "main.py", "cerebrium.toml"] exclude = [".*"] [cerebrium.hardware] cpu = 4 memory = 12.0 compute = "ADA_L40" [cerebrium.scaling] min_replicas = 1 max_replicas = 5 cooldown = 60 [cerebrium.dependencies.pip] vllm = "latest" pydantic = "latest" ``` Add the following code to `main.py`. This uses the vLLM framework and makes it OpenAI compatible: ``` import os import time import json from huggingface_hub import login from pydantic import BaseModel from vllm import SamplingParams, AsyncLLMEngine from vllm.engine.arg_utils import AsyncEngineArgs login(token=os.environ.get("HF_TOKEN")) engine_args = AsyncEngineArgs( model="RedHatAI/Meta-Llama-3.1-8B-Instruct-quantized.w8a8", gpu_memory_utilization=0.9, max_model_len=8192, ) engine = AsyncLLMEngine.from_engine_args(engine_args) class Message(BaseModel): role: str content: str def format_chat_prompt(messages: list) -> str: formatted_messages = [] for msg in messages: msg_obj = Message(**msg) formatted_messages.append( f"<|start_header_id|>{msg_obj.role}<|end_header_id|>\n{msg_obj.content}<|eot_id|>" ) return "<|begin_of_text|>" + "".join(formatted_messages) + "<|start_header_id|>assistant<|end_header_id|>" async def run( messages: list, model: str, run_id: str, stream: bool = True, temperature: float = 0.8, top_p: float = 0.95, max_tokens: int = 256, stream_options: dict = None, ): # Format your prompt for llama-friendly usage: prompt = format_chat_prompt(messages) sampling_params = SamplingParams(temperature=temperature, top_p=top_p, max_tokens=max_tokens) results_generator = engine.generate(prompt, sampling_params, run_id) previous_text = "" first_chunk = True async for output in results_generator: prompt_output = output.outputs new_text = prompt_output[0].text[len(previous_text) :] previous_text = prompt_output[0].text # Construct OpenAI-compatible chunk chunk = { "id": run_id, "object": "chat.completion.chunk", "created": int(time.time()), "model": model, "choices": [ { "index": 0, "delta": {}, "finish_reason": None, } ], } # Include the role in the first chunk if first_chunk: chunk["choices"][0]["delta"]["role"] = "assistant" first_chunk = False # Add new text to the delta if any if new_text: chunk["choices"][0]["delta"]["content"] = new_text # Capture a finish reason if it's provided finish_reason = prompt_output[0].finish_reason or None if finish_reason and finish_reason != "none": chunk["choices"][0]["finish_reason"] = finish_reason yield f"data: {json.dumps(chunk)}\n\n" # Send the final [DONE] message yield "data: [DONE]\n\n" ``` Add the HuggingFace token to Secrets on Cerebrium as `HF_TOKEN`. Run `cerebrium deploy` to make it live. The deployment URL appears in the dashboard and is used in the next step. Adjust the GPU hardware and `replica_concurrency` in `cerebrium.toml` to control how many concurrent calls the LLM handles. ### Pipecat Setup Run the following command to create the pipecat-agent: `cerebrium init pipecat-agent`. The [Pipecat framework](https://docs.pipecat.ai/getting-started/overview) orchestrates the services to create a voice agent. Add the following pip packages to `cerebrium.toml`: ``` [cerebrium.deployment] # This file was automatically generated by Cerebrium as a starting point for your project. # You can edit it as you wish. # If you would like to learn more about your Cerebrium config, please visit https://docs.cerebrium.ai/environments/config-files#config-file-example [cerebrium.deployment] name = "pipecat-agent" python_version = "3.11" include = ["./*", "main.py", "cerebrium.toml"] exclude = ["./example_exclude"] [cerebrium.hardware] compute = "CPU" cpu = 6 memory = 12.0 [cerebrium.scaling] min_replicas = 1 # Note: This incurs a constant cost since at least one instance is always running. max_replicas = 2 cooldown = 180 [cerebrium.dependencies.pip] torch = ">=2.0.0" "pipecat-ai[silero, daily, openai, deepgram, cartesia]" = "==0.0.67" aiohttp = ">=3.9.4" torchaudio = ">=2.3.0" channels = ">=4.0.0" requests = "==2.32.2" dotenv = "latest" ``` Add the following code to `main.py`: ```python theme={null} import asyncio import os import subprocess import sys import time from multiprocessing import Process import aiohttp import requests from loguru import logger from pipecat.frames.frames import LLMMessagesFrame, EndFrame from pipecat.pipeline.pipeline import Pipeline from pipecat.pipeline.runner import PipelineRunner from pipecat.pipeline.task import PipelineParams, PipelineTask from pipecat.processors.aggregators.llm_response import ( LLMAssistantResponseAggregator, LLMUserResponseAggregator, ) from pipecat.services.deepgram.stt import DeepgramSTTService from pipecat.services.cartesia.tts import CartesiaTTSService from deepgram import LiveOptions from pipecat.services.rime.tts import RimeHttpTTSService, RimeTTSService, Language from pipecat.services.openai import OpenAILLMService from pipecat.transports.services.daily import DailyParams, DailyTransport from pipecat.audio.vad.silero import SileroVADAnalyzer from pipecat.audio.vad.vad_analyzer import VADParams from helpers import ( CustomDeepgramSTTService, ) from dotenv import load_dotenv load_dotenv() logger.remove(0) logger.add(sys.stderr, level="DEBUG") deepgram_voice: str = "aura-asteria-en" async def main(room_url: str, token: str): async with aiohttp.ClientSession() as session: transport = DailyTransport( room_url, token, "Respond bot", DailyParams( audio_out_enabled=True, audio_in_enabled=True, transcription_enabled=False, vad_enabled=True, vad_analyzer=SileroVADAnalyzer(params=VADParams(stop_secs=0.15)), vad_audio_passthrough=True, ), ) stt = CustomDeepgramSTTService( api_key=os.environ.get("DEEPGRAM_API_KEY"), websocket_url="ws://api.aws/v4/p-xxxxxxxx/deepgram/v1/listen", live_options=LiveOptions( model="nova-2-general", language="en-US", smart_format=True, vad_events=True ) ) tts = CartesiaTTSService( api_key=os.environ.get("CARTESIA_API_KEY"), voice_id='97f4b8fb-f2fe-444b-bb9a-c109783a857a', ) llm = OpenAILLMService( name="LLM", model="RedHatAI/Meta-Llama-3.1-8B-Instruct-quantized.w8a8", base_url="http://api.aws/v4/p-xxxxxxxx/llama-llm/run" ) messages = [ { "role": "system", "content": "You are a fast, low-latency chatbot. Your goal is to demonstrate voice-driven AI capabilities at human-like speeds. The technology powering you is Daily for transport, Cerebrium for serverless infrastructure, Llama 3 (8-B version) LLM, and Deepgram for speech-to-text and text-to-speech. You are hosted on the east coast of the United States. Respond to what the user said in a creative and helpful way, but keep responses short and legible. Ensure responses contain only words. Check again that you have not included special characters other than '?' or '!'.", }, ] tma_in = LLMUserResponseAggregator(messages) tma_out = LLMAssistantResponseAggregator(messages) pipeline = Pipeline( [ transport.input(), # Transport user input stt, # Speech-to-text tma_in, # User responses llm, # LLM tts, # TTS transport.output(), # Transport bot output tma_out, # Assistant spoken responses ] ) task = PipelineTask( pipeline, params=PipelineParams( allow_interruptions=True, enable_metrics=True ), ) # When the first participant joins, the bot should introduce itself. @transport.event_handler("on_first_participant_joined") async def on_first_participant_joined(transport, participant): # Kick off the conversation. time.sleep(1.5) messages.append( { "role": "system", "content": "Introduce yourself by saying 'hello, I'm FastBot, how can I help you today?'", } ) await task.queue_frame(LLMMessagesFrame(messages)) # When the participant leaves, we exit the bot. @transport.event_handler("on_participant_left") async def on_participant_left(transport, participant, reason): await task.queue_frame(EndFrame()) # If the call is ended make sure we quit as well. @transport.event_handler("on_call_state_updated") async def on_call_state_updated(transport, state): if state == "left": await task.queue_frame(EndFrame()) runner = PipelineRunner() await runner.run(task) await session.close() async def start_bot(room_url: str, token: str = None): try: await main(room_url, token) except Exception as e: logger.error(f"Exception in main: {e}") sys.exit(1) # Exit with a non-zero status code return {"message": "session finished"} def create_room(): url = "https://api.daily.co/v1/rooms/" headers = { "Content-Type": "application/json", "Authorization": f"Bearer {os.environ.get('DAILY_TOKEN')}", } data = { "properties": { "exp": int(time.time()) + 60 * 5, ##5 mins "eject_at_room_exp": True, } } response = requests.post(url, headers=headers, json=data) if response.status_code == 200: room_info = response.json() token = create_token(room_info["name"]) if token and "token" in token: room_info["token"] = token["token"] else: logger.error("Failed to create token") return { "message": "There was an error creating your room", "status_code": 500, } return room_info else: data = response.json() if data.get("error") == "invalid-request-error" and "rooms reached" in data.get( "info", "" ): logger.error( "We are currently at capacity for this demo. Please try again later." ) return { "message": "We are currently at capacity for this demo. Please try again later.", "status_code": 429, } logger.error(f"Failed to create room: {response.status_code}") return {"message": "There was an error creating your room", "status_code": 500} def create_token(room_name: str): url = "https://api.daily.co/v1/meeting-tokens" headers = { "Content-Type": "application/json", "Authorization": f"Bearer {os.environ.get('DAILY_TOKEN')}", } data = {"properties": {"room_name": room_name, "is_owner": True}} response = requests.post(url, headers=headers, json=data) if response.status_code == 200: token_info = response.json() return token_info else: logger.error(f"Failed to create token: {response.status_code}") return None # if __name__ == "__main__": # room = create_room() # if room and "token" in room: # asyncio.run(main(room["url"], room["token"])) # else: # logger.error("Failed to create room") ``` Summary of the code above: * WebRTC functionality from Daily creates the room (swappable for Twilio/Telenyx). Two functions handle room creation and authentication: `create_room()` and `create_token()`. * The Deepgram and LLM services use a local URL to connect within the Cerebrium cluster. Edit the project key in the URL as needed. * TTS uses the Cartesia service to demonstrate Pipecat's versatility, but the Deepgram TTS service works as well. The Daily Python SDK provides event webhooks to trigger functionality based on events like users joining or leaving calls. Add this event handling code to the `main()` function: This code handles these events: * First participant joins: Bot introduces itself via a conversation message * Additional participants join: Bot listens and responds to all participants * Participant leaves or call ends: Bot terminates itself Adjust the CPU hardware and `replica_concurrency` in `cerebrium.toml` to control how many concurrent calls the Pipecat agent handles. Create a `.env` file in the pipecat-agent folder with the following: ``` DAILY_TOKEN= HF_TOKEN= DEEPGRAM_API_KEY= CARTESIA_API_KEY= ``` Get the Daily developer token from the profile page. Sign up [here](https://dashboard.daily.co/u/signup) if needed (generous free tier available). Navigate to the "developers" tab for the API key and add it to Cerebrium Secrets. Daily API Key To test the voice bot locally, uncomment the main code at the bottom and run `python main.py`. The result is a fully functioning AI bot that interacts with users through speech in \~500ms. The next section creates a user interface for it. ### Deploy to Cerebrium Deploy to Cerebrium by running `cerebrium deploy`. The endpoints are used in the frontend interface below. ## Connect Frontend A public fork of the PipeCat frontend demonstrates this application. Clone the repo [here](https://github.com/CerebriumAI/web-client-ui). Follow the instructions in the README.md and then populate the following variables in your .env.development.local. ``` VITE_SERVER_URL=https://api.cerebrium.ai/v4/p-xxxxxxxx/ #This is the base url of your pipecat-agent. Do not include the function names VITE_SERVER_AUTH= #This is the JWT token you can get from the API Keys section of your Cerebrium Dashboard. ``` You can now run yarn dev and go to the URL: [http://localhost:5173/](http://localhost:5173/) to test your application! ### Conclusion This tutorial provides a foundation for implementing voice in applications and extending into image and vision capabilities. PipeCat is an extensible, open-source framework for building voice-enabled apps. Cerebrium handles deployment and autoscaling with pay-as-you-go compute. Tag **@cerebriumai** to showcase work and join the [Discord](https://discord.gg/ATj6USmeE2) community for questions and feedback. # Generate Images using SDXL Source: https://cerebrium.ai/docs/v4/examples/sdxl Deploy the Stability AI SDXL refiner model on Cerebrium serverless GPUs to generate high quality images from prompts with a Diffusers pipeline This example is only compatible with CLI v1.20 and later. Should you be making use of an older version of the CLI, please run `pip install --upgrade cerebrium` to upgrade it to the latest version. This tutorial shows you how to generate high-quality images using the SDXL refiner model from Stability AI, available on [Hugging Face](https://huggingface.co/stabilityai/stable-diffusion-xl-refiner-1.0). To see the final implementation, you can view it [here](https://github.com/CerebriumAI/examples/tree/master/7-image-and-video/3-sdxl-refiner) ## Basic Setup Developing on Cerebrium is similar to a virtual machine or Google Colab. Install the Cerebrium package and log in before proceeding. See the [installation docs](https://docs.cerebrium.ai/getting-started/installation) for details. First, create your project: ``` cerebrium init 2-sdxl-refiner ``` Configure your compute and environment settings in `cerebrium.toml`: ```toml theme={null} [cerebrium.deployment] name = "3-sdxl-refiner" python_version = "3.10" include = ["./*", "main.py", "cerebrium.toml"] exclude = ["./.*", "./__*"] [cerebrium.hardware] compute = "AMPERE_A10" cpu = 2 memory = 16.0 gpu_count = 1 [cerebrium.scaling] min_replicas = 0 max_replicas = 5 cooldown = 60 [cerebrium.dependencies.pip] accelerate = "latest" transformers = ">=4.35.0" safetensors = "latest" opencv-python = "latest" diffusers = "latest" [cerebrium.dependencies.conda] [cerebrium.dependencies.apt] ffmpeg = "latest" ``` Create a `main.py` file. This implementation fits in a single file. Start by defining the request object: ```python theme={null} from typing import Optional from pydantic import BaseModel import torch from diffusers import StableDiffusionXLImg2ImgPipeline from diffusers.utils import load_image import io import base64 class Item(BaseModel): prompt: str url: str negative_prompt: Optional[str] conditioning_scale: float height: int width: int num_inference_steps: int guidance_scale: float num_images_per_prompt: int ``` The code uses Pydantic for data validation. The `prompt` and `url` parameters are required; all others are optional. Missing required parameters trigger an automatic error message. ## Instantiate Model The SDXL model loads outside the `predict` function since it only needs to load once at startup. The model downloads during initial deployment and is automatically cached in persistent storage for subsequent use. ```python theme={null} pipe = StableDiffusionXLImg2ImgPipeline.from_pretrained( "stabilityai/stable-diffusion-xl-refiner-1.0", torch_dtype=torch.float16, variant="fp16", use_safetensors=True ) pipe = pipe.to("cuda") ``` ## Predict Function The `predict` function takes parameters from the request, passes them to the SDXL model, and returns base64-encoded images for direct JSON-serializable responses. ```python theme={null} def predict(prompt, url, negative_prompt=None, conditioning_scale=0.5, height=512, width=512, num_inference_steps=20, guidance_scale=7.5, num_images_per_prompt=1): item = Item( prompt=prompt, url=url, negative_prompt=negative_prompt, conditioning_scale=conditioning_scale, height=height, width=width, num_inference_steps=num_inference_steps, guidance_scale=guidance_scale, num_images_per_prompt=num_images_per_prompt ) init_image = load_image(item.url).convert("RGB") images = pipe( item.prompt, negative_prompt=item.negative_prompt, controlnet_conditioning_scale=item.conditioning_scale, height=item.height, width=item.width, num_inference_steps=item.num_inference_steps, guidance_scale=item.guidance_scale, num_images_per_prompt=item.num_images_per_prompt, image=init_image ).images finished_images = [] for image in images: buffered = io.BytesIO() image.save(buffered, format="PNG") finished_images.append(base64.b64encode(buffered.getvalue()).decode("utf-8")) return {"images": finished_images} ``` ## Deploy Deploy the model using this command: ```bash theme={null} cerebrium deploy ``` After deployment, make this request: ```curl theme={null} curl --location 'https://api.cerebrium.ai/v4/p-xxxxxxxx/3-sdxl-refiner/predict' \ --header 'Content-Type: application/json' \ --header 'Authorization: Bearer ' \ --data '{ "url": "https://huggingface.co/datasets/patrickvonplaten/images/resolve/main/aa_xl/000000009.png", "prompt": "a photo of an astronaut riding a horse on mars" }' ``` The endpoint returns results in this format: ```json theme={null} { "run_id": "Gd2fLvweh1sHpdEQd4XnxYRvtGmghFxSg2rpbchK7wWAFeso9-sOVg==", "message": "Finished inference request with run_id: `Gd2fLvweh1sHpdEQd4XnxYRvtGmghFxSg2rpbchK7wWAFeso9-sOVg==`", "result": { "images": [ ] }, "status_code": 200, "run_time_ms": 4388.460874557495 } ``` Example output: SDXL # Transcribe 1 hour podcast Source: https://cerebrium.ai/docs/v4/examples/transcribe-whisper Transcribe hour long podcasts and audio files with Distil Whisper on Cerebrium using base64 uploads or file URLs and webhooks for long jobs This tutorial transcribes an hour-long audio file using Distill Whisper: an optimized version of Whisper-large-v2 that's 60% faster while maintaining accuracy within 1% of the original. The endpoint accepts either a base64-encoded string of the audio file or a URL to download the audio file. To see the final implementation, you can view it [here](https://github.com/CerebriumAI/examples/tree/master/6-voice/1-whisper-transcription) ## Basic Setup Developing models with Cerebrium is similar to developing on a virtual machine or Google Colab. Install the Cerebrium package and log in. See the [installation docs](https://docs.cerebrium.ai/getting-started/installation) for details. Create the project: ``` cerebrium init 1-whisper-transcription ``` Add the following packages to the `[cerebrium.dependencies.pip]` section of your `cerebrium.toml` file: ```toml theme={null} [cerebrium.dependencies.pip] accelerate = "latest" transformers = ">=4.35.0" openai-whisper = "latest" pydantic = "latest" ``` Create a `util.py` file for utility functions that download a file from a URL or convert a base64 string to a file: ```python theme={null} import base64 import uuid DOWNLOAD_ROOT = "/tmp/" def download_file_from_url(url: str, filename: str): response = requests.get(url) if response.status_code == 200: with open(filename, "wb") as f: f.write(response.content) return filename else: raise Exception("Download failed") def save_base64_string_to_file(audio: str): decoded_data = base64.b64decode(audio) filename = f"{DOWNLOAD_ROOT}/{uuid.uuid4()}" with open(filename, "wb") as file: file.write(decoded_data) return filename ``` With the utility functions complete, update `main.py` with the main application code. The endpoint accepts either a base64-encoded string or a public URL of the audio file, passes it to the model, and returns the output. Define the request object: ```python theme={null} from typing import Optional from pydantic import BaseModel, HttpUrl class Item(BaseModel): audio: Optional[str] file_url: Optional[HttpUrl] webhook_endpoint: Optional[HttpUrl] ``` Pydantic handles data validation. While `audio` and `file_url` are optional parameters, you must provide at least one. The `webhook_endpoint` parameter, automatically included by Cerebrium in every request, is useful for long-running requests. Note: Cerebrium has a 3-minute timeout for each inference request. For long audio files (2+ hours) that take several minutes to process, use a `webhook_endpoint`: a URL where Cerebrium sends a POST request with the function's results. ## Setup Model and Inference Import the required packages and load the Whisper model. The model downloads during initial deployment and is automatically cached in persistent storage for subsequent use. Loading the model outside the `predict` function ensures this code only runs on cold start (startup). For warm containers, only the `predict` function executes for inference. ```python theme={null} from huggingface_hub import hf_hub_download from whisper import load_model, transcribe from util import download_file_from_url, save_base64_string_to_file distil_large_v2 = hf_hub_download(repo_id="distil-whisper/distil-large-v3", filename="original-model.bin") model = load_model(distil_large_v2) def predict(run_id, audio=None, file_url=None, webhook_endpoint=None): item = Item(audio=audio, file_url=file_url, webhook_endpoint=webhook_endpoint) input_filename = f"{run_id}.mp3" if audio is None and file_url is None: raise 'Either audio or file_url must be provided' else: if item.audio is not None: file = save_base64_string_to_file(item.audio) elif item.file_url is not None: file = download_file_from_url(item.file_url, input_filename) print("Transcribing file...") result = transcribe(model, audio=file) return result ``` The `predict` function, which runs only on inference requests, creates an audio file from either the download URL or base64 string, transcribes it, and returns the output. ## Deploy Configure your compute and environment settings in `cerebrium.toml`: ```toml theme={null} [cerebrium.deployment] name = "1-whisper-transcription" python_version = "3.11" include = ["./*", "main.py", "cerebrium.toml"] exclude = ["./example_exclude"] docker_base_image_url = "nvidia/cuda:12.1.1-runtime-ubuntu22.04" [cerebrium.hardware] compute = "AMPERE_A10" cpu = 3 memory = 12.0 gpu_count = 1 [cerebrium.scaling] min_replicas = 0 max_replicas = 5 cooldown = 60 [cerebrium.dependencies.pip] accelerate = "latest" transformers = ">=4.35.0" openai-whisper = "latest" pydantic = "latest" [cerebrium.dependencies.conda] [cerebrium.dependencies.apt] "ffmpeg" = "latest" ``` Deploy the app using this command: ```bash theme={null} cerebrium deploy ``` After deployment, make this request: ```curl theme={null} curl --location 'https://api.cerebrium.ai/v4/p-xxxxxxxx/1-whisper-transcription/predict' \ --header 'Content-Type: application/json' \ --header 'Authorization: Bearer ' \ --data '{"file_url": "https://your-public-url.com/test.mp3"}' ``` The response returns immediately with a 202 status code and a `run_id`: a unique identifier to correlate the result with the initial workload. The endpoint returns results in this format: ```json theme={null} { "run_id": "2R5PnHprwNqiS5tcFMor-4c6rSrxuzrVtBU1JfjT5iWFG6s4pHo1Ug==", "message": "Finished inference request with run_id: `2R5PnHprwNqiS5tcFMor-4c6rSrxuzrVtBU1JfjT5iWFG6s4pHo1Ug==`", "result": { "text": " Testing, one, two, three, testing.", "segments": [ { "id": 0, "seek": 0, "start": 0, "end": 4, "text": " Testing, one, two, three, testing.", "tokens": [ 50364, 45517, 11, 472, 11, 220, 20534, 11, 220, 27583, 11, 220, 83, 8714, 13, 50564 ], "temperature": 0, "avg_logprob": -0.3824356023003073, "compression_ratio": 1, "no_speech_prob": 0.019467202946543694 } ], "language": "en" }, "status_code": 200, "run_time_ms": 2053.8525581359863 } ``` # Twilio Voice Agent with PipeCat Source: https://cerebrium.ai/docs/v4/examples/twilio-voice-agent Connect a real time AI voice agent to phone calls using Twilio, PipeCat and FastAPI WebSockets on Cerebrium for support bots and receptionists This tutorial creates a real-time voice agent that responds to phone calls via Twilio. The implementation supports any LLM or Text-to-Speech (TTS) model, making it ideal for voice applications like customer support bots and receptionists. The example uses [PipeCat](https://www.pipecat.ai/) to handle component integration, user interruptions, and audio data processing. You can find the final version of the code [here](https://github.com/CerebriumAI/examples/tree/master/6-voice/4-twilio-voice-agent) ### Cerebrium Setup Set up Cerebrium: 1. Sign up [here](https://dashboard.cerebrium.ai/register) 2. Follow installation docs [here](https://docs.cerebrium.ai/getting-started/installation) 3. Create a starter project: ```bash theme={null} cerebrium init 4-twilio-voice-agent ``` This creates: * `main.py`: Entrypoint file * `cerebrium.toml`: Build and environment configuration Add these pip packages to your `cerebrium.toml`: ``` [cerebrium.dependencies.pip] torch = ">=2.0.0" "pipecat-ai[silero, daily, openai, deepgram, cartesia, twilio]" = "0.0.47" aiohttp = ">=3.9.4" torchaudio = ">=2.3.0" channels = ">=4.0.0" requests = "==2.32.2" twilio = "latest" fastapi = "latest" uvicorn = "latest" python-dotenv = "latest" loguru = "latest" ``` Set up a FastAPI server to handle Twilio calls and upgrade to WebSocket connections for real-time communication. Add this code to `main.py`: ```python theme={null} import json from fastapi import FastAPI, WebSocket from fastapi.middleware.cors import CORSMiddleware from starlette.responses import HTMLResponse from bot import main app = FastAPI() app.add_middleware( CORSMiddleware, allow_origins=["*"], # Allow all origins for testing allow_credentials=True, allow_methods=["*"], allow_headers=["*"], ) @app.post("/") async def start_call(): print("POST TwiML") return HTMLResponse(content=open("templates/streams.xml").read(), media_type="application/xml") @app.websocket("/ws") async def websocket_endpoint(websocket: WebSocket): await websocket.accept() start_data = websocket.iter_text() await start_data.__anext__() call_data = json.loads(await start_data.__anext__()) print(call_data, flush=True) stream_sid = call_data["start"]["streamSid"] print("WebSocket connection accepted") await main(websocket, stream_sid) ``` Create a `templates` folder with `stream.xml` inside. This XML response tells Twilio to upgrade to a WebSocket connection: ```xml theme={null} ``` Replace the stream URL with the deployment's base endpoint, using the correct project ID. Configure Cerebrium to run the FastAPI server by adding this to `cerebrium.toml`: ``` [cerebrium.runtime.custom] port = 8765 entrypoint = ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8765"] healthcheck_endpoint = "/health" ``` You can read more about running custom web servers [here](/docs/container-images/custom-web-servers). ### Twilio Setup Twilio provides cloud communications APIs for messaging, voice, video, and authentication. Other providers work as well. Sign up for a free account [here](https://www.twilio.com/try-twilio). Purchase a local number (not toll-free) from the [phone numbers page](https://console.twilio.com/us1/develop/phone-numbers/manage/search?isoCountry=US\&types\[]=Local\&capabilities\[]=Sms\&capabilities\[]=Mms\&capabilities\[]=Voice\&capabilities\[]=Fax\&searchTerm=\&searchFilter=left\&searchType=number). Then set up a webhook to connect calls to your agent. Buy a number Setup Webhook Save the changes and proceed to set up the AI agent. ### AI Agent Setup Create `bot.py` to set up the AI agent using PipeCat for component integration, interruption handling, and audio processing: ```python theme={null} import os import sys from loguru import logger from pipecat.frames.frames import LLMMessagesFrame, EndFrame from pipecat.pipeline.pipeline import Pipeline from pipecat.pipeline.runner import PipelineRunner from pipecat.pipeline.task import PipelineParams, PipelineTask from pipecat.services.openai import OpenAILLMService from pipecat.processors.aggregators.openai_llm_context import ( OpenAILLMContext, ) from pipecat.services.deepgram import DeepgramSTTService from pipecat.vad.silero import SileroVADAnalyzer from twilio.rest import Client from twilio.twiml.voice_response import VoiceResponse from pipecat.transports.network.fastapi_websocket import ( FastAPIWebsocketTransport, FastAPIWebsocketParams, ) from pipecat.serializers.twilio import TwilioFrameSerializer from pipecat.services.cartesia import CartesiaTTSService logger.remove(0) logger.add(sys.stderr, level="DEBUG") twilio = Client( os.environ.get("TWILIO_ACCOUNT_SID"), os.environ.get("TWILIO_AUTH_TOKEN") ) async def main(websocket_client, stream_sid): transport = FastAPIWebsocketTransport( websocket=websocket_client, params=FastAPIWebsocketParams( audio_out_enabled=True, add_wav_header=False, vad_enabled=True, vad_analyzer=SileroVADAnalyzer(), vad_audio_passthrough=True, serializer=TwilioFrameSerializer(stream_sid), ), ) stt = DeepgramSTTService(api_key=os.getenv("DEEPGRAM_API_KEY")) llm = OpenAILLMService( name="LLM", api_key=os.environ.get("OPENAI_API_KEY"), model="gpt-4", ) tts = CartesiaTTSService( api_key=os.getenv("CARTESIA_API_KEY"), voice_id="79a125e8-cd45-4c13-8a67-188112f4dd22", # British Lady ) messages = [ { "role": "system", "content": "You are a helpful LLM in an audio call. Your goal is to demonstrate your capabilities in a succinct way. Your output will be converted to audio so don't include special characters in your answers. Respond to what the user said in a creative and helpful way.", }, ] print('here', flush=True) context = OpenAILLMContext(messages=messages) context_aggregator = llm.create_context_aggregator(context) pipeline = Pipeline( [ transport.input(), # Websocket input from client stt, # Speech-To-Text context_aggregator.user(), llm, # LLM tts, # Text-To-Speech transport.output(), # Websocket output to client context_aggregator.assistant(), ] ) task = PipelineTask(pipeline, params=PipelineParams(allow_interruptions=True)) @transport.event_handler("on_client_connected") async def on_client_connected(transport, client): # Kick off the conversation. messages.append({"role": "system", "content": "Please introduce yourself to the user."}) await task.queue_frames([LLMMessagesFrame(messages)]) @transport.event_handler("on_client_disconnected") async def on_client_disconnected(transport, client): await task.queue_frames([EndFrame()]) runner = PipelineRunner(handle_sigint=False) await runner.run(task) ``` The code: 1. Connects to WebSocket transport for audio I/O 2. Sets up services: * [Deepgram](https://deepgram.com/) for STT * [OpenAI](https://openai.com/) for LLM * [Cartesia](https://www.cartesia.ai/) for TTS * Other providers supported via PipeCat 3. Uses [Secrets](https://docs.cerebrium.ai/environments/using-secrets) for authentication 4. Creates a customizable PipelineTask supporting: * Image and Vision use cases ([docs](https://docs.pipecat.ai/docs/category/services)) * Built-in interruption handling * Easy model swapping 5. Handles call events (join/leave) via webhooks For lower latency (\~500ms end-to-end), run parts or all of the pipeline locally. See the [voice agents guide](/docs/v4/examples/realtime-voice-agents) and [RAG voice agent blog post](https://www.cerebrium.ai/blog/creating-a-realtime-rag-voice-agent). ### Deploy to Cerebrium Deploy to Cerebrium by running `cerebrium deploy` in the terminal. A successful deployment looks like this: Cerebrium Deployment Test by calling the Twilio number. The agent responds automatically. ### Scaling Pipecat For scaling PipeCat on Cerebrium: 1. Use large CPU instances (10 CPUs, 8GB memory) for Twilio's less than 1s response requirement 2. Run concurrent PipeCat processes: * Each process uses \~0.5 CPUs * 10 CPU instance handles 20 concurrent calls * Adjust based on traffic needs For scaling criteria, use Cerebrium's `replica_concurrency` setting to spawn new containers based on utilization, preventing cold starts for subsequent calls. Update `cerebrium.toml` with the following: ``` [cerebrium.hardware] compute = "CPU" cpu = 10 memory = 8.0 [cerebrium.scaling] min_replicas = 1 max_replicas = 3 cooldown = 30 replica_concurrency = 20 scaling_metric = "concurrency_utilization" scaling_target = 80 ``` ### Conclusion This tutorial provides a foundation for implementing voice features and expanding into image and vision capabilities. PipeCat is an extensible, open-source framework for building voice-enabled apps. Cerebrium handles deployment and autoscaling with pay-as-you-go compute. Tag **@cerebriumai** to showcase work and join the [Discord](https://discord.gg/ATj6USmeE2) community for questions and feedback. # Hyperparameter Sweep training Llama 3.2 with WandB Source: https://cerebrium.ai/docs/v4/examples/wandb-sweep Fine tune Llama 3.2 with Weights and Biases hyperparameter sweeps, running training experiments in parallel across Cerebrium serverless GPUs Hyperparameter sweeps systematically test parameter combinations to find the best-performing model for the least compute or training time. This tutorial covers training Llama 3.2, using [Wandb](https://wandb.ai/site) (Weights and Biases) to run hyperparameter sweeps and Cerebrium to scale experiments across serverless GPUs. View the final version [on GitHub](https://github.com/CerebriumAI/examples/tree/master/12-training/1-wandb-sweep). Read this section if you're unfamiliar with sweeps. ### Analogy: Pizza Topping Sweep Forget about ML for a second. Imagine making pizzas to discover the best combination of toppings. Three variables are available: • **Type of Cheese** (mozzarella, cheddar, parmesan) • **Type of Sauce** (tomato, pesto) • **Extra Topping** (pepperoni, mushrooms, olives) There are 12 possible combinations. One of them will taste the best. To find the tastiest pizza, try all combinations and rate them. This process is a **hyperparameter sweep**. The three hyperparameters are cheese, sauce, and extra topping. One pizza at a time takes hours. With 12 ovens, all pizzas bake at once and the best one emerges in minutes. If a kitchen is a GPU, 12 GPUs run all experiments in parallel. Cerebrium enables sweeps across 12 GPUs (or 1,000) to find the best model version fast. ### Setup Cerebrium If you don’t have a Cerebrium account, run the following: ```bash theme={null} pip install cerebrium --upgrade cerebrium login cerebrium init wandb-sweep ``` This creates a folder with two files: * **main.py** - The entrypoint file where application code lives. * **cerebrium.toml** - A configuration file for build and environment settings. The rest of this tutorial continues in this folder. ### Setup Wandb **Weights & Biases (Wandb)** tracks, visualizes, and manages machine learning experiments in real-time. It logs hyperparameters, metrics, and results for comparing models and optimizing performance. 1. [Sign up](https://wandb.ai/site) for a free account and then log in to your wandb account by running the following in your CLI. ``` pip install wandb wandb login ``` A link prints in the terminal. Click it and copy the API key back into the terminal. Add the W\&B API key to Cerebrium secrets. In the [Cerebrium Dashboard](https://dashboard.cerebrium.ai/), navigate to the “secrets” tab in the left sidebar. Add the following: * **Key**: WANDB\_API\_KEY * **Value**: The value you copied from the Wandb website. Click the “Save All Changes” button. Cerebrium Secrets Wandb authentication is now complete. ### Training Script To train with Llama 3.2, you'll need: 1. Model access permission: * Visit the [Llama 3.2 model page](https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct) on Hugging Face * Accept all permissions 2. Hugging Face token: * Click your profile image (top right) * Select "Access token" * Create a new token if needed * Add to Cerebrium Secrets: * Key: `HF_TOKEN` * Value: Your Hugging Face token * Click "Save All Changes" Hugging Face Token The training script adapts [this Kaggle notebook](https://www.kaggle.com/code/scratchpad/notebookad2fe9fab1/edit). Create a `requirements.txt` file with these dependencies: ``` transformers datasets accelerate peft trl bitsandbytes wandb ``` These packages are required both locally and on Cerebrium. Update your `cerebrium.toml` to include: * The requirements.txt path * Hardware requirements for training * A 1-hour max timeout using `response_grace_period` Add this configuration: ``` #existing configuration [cerebrium.hardware] cpu = 6 memory = 30.0 compute = "ADA_L40" [cerebrium.scaling] #existing configuration response_grace_period = 3600 [cerebrium.dependencies.paths] pip = "requirements.txt" ``` Install the dependencies locally: ```bash theme={null} pip install -r requirements.txt ``` Add this code to `main.py`: ```python theme={null} import os from typing import Dict import wandb from transformers import ( AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig, TrainingArguments, ) from peft import LoraConfig, get_peft_model from datasets import load_dataset from trl import SFTTrainer, setup_chat_format import torch import bitsandbytes as bnb from huggingface_hub import login login(token=os.getenv("HF_TOKEN")) def find_all_linear_names(model): cls = bnb.nn.Linear4bit lora_module_names = set() for name, module in model.named_modules(): if isinstance(module, cls): names = name.split('.') lora_module_names.add(names[0] if len(names) == 1 else names[-1]) if 'lm_head' in lora_module_names: lora_module_names.remove('lm_head') return list(lora_module_names) def train_model(params: Dict): """ Training function that receives parameters from the Cerebrium endpoint """ # Initialize wandb wandb.login(key=os.getenv("WANDB_API_KEY")) wandb.init( project=params.get("wandb_project", "Llama-3.2-Customer-Support"), name=params.get("run_name", None), config=params, ) # Model configuration base_model = params.get("base_model", "meta-llama/Llama-3.2-3B-Instruct") new_model = params.get("output_model_name", f"/persistent-storage/llama-3.2-3b-it-Customer-Support-{wandb.run.id}") # Set torch dtype and attention implementation torch_dtype = torch.bfloat16 if torch.cuda.get_device_capability()[0] >= 8 else torch.float16 attn_implementation = "eager" # QLoRA config bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch_dtype, bnb_4bit_use_double_quant=True, ) # Load model and tokenizer model = AutoModelForCausalLM.from_pretrained( base_model, quantization_config=bnb_config, device_map="auto", attn_implementation=attn_implementation ) tokenizer = AutoTokenizer.from_pretrained(base_model, trust_remote_code=True) # Find linear modules modules = find_all_linear_names(model) # LoRA config peft_config = LoraConfig( r=params.get("lora_r", 16), lora_alpha=params.get("lora_alpha", 32), lora_dropout=params.get("lora_dropout", 0.05), bias="none", task_type="CAUSAL_LM", target_modules=modules ) # Setup model tokenizer.chat_template = None # Clear existing chat template model, tokenizer = setup_chat_format(model, tokenizer) model = get_peft_model(model, peft_config) # Load and prepare dataset dataset = load_dataset( params.get("dataset_name", "bitext/Bitext-customer-support-llm-chatbot-training-dataset"), split="train" ) dataset = dataset.shuffle(seed=params.get("seed", 65)) if params.get("max_samples"): dataset = dataset.select(range(params.get("max_samples"))) instruction = params.get("instruction", """You are a top-rated customer service agent named John. Be polite to customers and answer all their questions.""" ) def format_chat_template(row): row_json = [ {"role": "system", "content": instruction}, {"role": "user", "content": row["instruction"]}, {"role": "assistant", "content": row["response"]} ] row["text"] = tokenizer.apply_chat_template(row_json, tokenize=False) return row dataset = dataset.map(format_chat_template, num_proc=4) dataset = dataset.train_test_split(test_size=params.get("test_size", 0.1)) # Training arguments training_arguments = TrainingArguments( output_dir=new_model, per_device_train_batch_size=params.get("batch_size", 1), per_device_eval_batch_size=params.get("batch_size", 1), gradient_accumulation_steps=params.get("gradient_accumulation_steps", 2), optim="paged_adamw_32bit", num_train_epochs=params.get("epochs", 1), eval_strategy="steps", eval_steps=params.get("eval_steps", 0.2), logging_steps=params.get("logging_steps", 1), warmup_steps=params.get("warmup_steps", 10), learning_rate=params.get("learning_rate", 2e-4), fp16=params.get("fp16", False), bf16=params.get("bf16", False), group_by_length=params.get("group_by_length", True), report_to="wandb" ) # Initialize trainer trainer = SFTTrainer( model=model, train_dataset=dataset["train"], eval_dataset=dataset["test"], peft_config=peft_config, args=training_arguments, ) # Train and save model.config.use_cache = False trainer.train() model.config.use_cache = True # Save model trainer.model.save_pretrained(new_model) wandb.finish() return {"status": "success", "model_path": new_model} ``` You can read a deeper explanation of the training script [here](https://www.datacamp.com/tutorial/fine-tuning-llama-3-2) but here's a high-level explanation of the code in bullet points: * This code sets up a fine-tuning pipeline for a Large Language Model (specifically Llama 3.2) using several modern training techniques: * Takes a dictionary of parameters for flexible training configurations: the hyperparameter sweep. * Loads a customer support dataset from Hugging Face and formats it into chat template format * Implements QLoRA (Quantized Low-Rank Adaptation) for efficient fine-tuning. * Uses Weights & Biases (Wandb) for experiment tracking, logging results to the Wandb dashboard. * Saves the final model to a Cerebrium volume and returns a “success” message. Deploy the training endpoint: ```bash theme={null} cerebrium deploy ``` This command: 1. Sets up the environment with required packages 2. Deploys the training script as an endpoint 3. Returns a POST URL (save this for later) Cerebrium requires no special decorators or syntax. Wrap the training code in a function. The endpoint automatically scales based on request volume, making it ideal for hyperparameter sweeps. ### Hyperparameter Sweep Create a **run.py** file for running locally. Add the following code: ```python theme={null} import wandb import requests import os from typing import Dict, Any from dotenv import load_dotenv load_dotenv() CEREBRIUM_API_KEY = os.getenv("CEREBRIUM_API_KEY") ENDPOINT_URL = "https://api.cerebrium.ai/v4/p-xxxxxxxx/wandb-sweep/train_model?async=true" def train_with_params(params: Dict[str, Any]): """ Send training parameters to Cerebrium endpoint """ headers = { "Authorization": f"Bearer {CEREBRIUM_API_KEY}", "Content-Type": "application/json" } response = requests.post(ENDPOINT_URL, json={"params": params}, headers=headers) if response.status_code != 202: raise Exception(f"Training failed: {response.text}") return response.json() # Define the sweep configuration sweep_config = { "method": "bayes", # Bayesian optimization "metric": { "name": "eval/loss", "goal": "minimize" }, "parameters": { "learning_rate": { "distribution": "log_uniform", "min": -10, # exp(-10) ≈ 4.54e-5 "max": -7, # exp(-7) ≈ 9.12e-4 }, "batch_size": { "values": [1, 2, 4] }, "gradient_accumulation_steps": { "values": [2, 4, 8] }, "lora_r": { "values": [8, 16, 32] }, "lora_alpha": { "values": [16, 32, 64] }, "lora_dropout": { "distribution": "uniform", "min": 0.01, "max": 0.1 }, "max_seq_length": { "values": [512, 1024] } } } def main(): # Initialize the sweep sweep_id = wandb.sweep(sweep_config, project="Llama-3.2-Customer-Support") def sweep_iteration(): # Initialize a new W&B run with wandb.init() as run: # Get the parameters for this run params = wandb.config.as_dict() # Add any fixed parameters params.update({ "wandb_project": "Llama-3.2-Customer-Support", "base_model": "meta-llama/Llama-3.2-3B-Instruct", "dataset_name": "bitext/Bitext-customer-support-llm-chatbot-training-dataset", "run_name": f"sweep-{run.id}", "epochs": 1, "test_size": 0.1, }) # Call the Cerebrium endpoint with these parameters try: result = train_with_params(params) print(f"Training completed: {result}") except Exception as e: print(f"Training failed: {str(e)}") run.finish(exit_code=1) # Run the sweep wandb.agent(sweep_id, function=sweep_iteration, count=10) # Run 10 experiments if __name__ == "__main__": main() ``` This code implements a hyperparameter sweep using Wandb sweeps to train a Llama 3.2 model for customer support. Here's what it does: * Create a .env file and add the Inference API key from the Cerebrium Dashboard. ``` CEREBRIUM_API_KEY=eyJ.... ``` * Update the Cerebrium endpoint with the correct project ID and function name. The URL is appended with “?async=true”. This makes it a fire-and-forget request that can run up to 12 hours. Read more [here](https://docs.cerebrium.ai/endpoints/async). * The Bayesian optimization sweep configuration searches through these hyperparameters: * Learning rate (log uniform distribution between \~4.54e-5 and \~9.12e-4) * Batch size (1, 2, or 4) * Gradient accumulation steps (2, 4, or 8) * LoRA parameters (r, alpha, and dropout) * Maximum sequence length (512 or 1024) * The sweep is created in the "Llama-3.2-Customer-Support" W\&B project * For each sweep iteration: * Initializes a new W\&B run * Combines the sweep's hyperparameters with fixed parameters (like model name and dataset) * Sends the parameters to a Cerebrium endpoint for training that happens asynchronously. * Logs the results back to W\&B * Runs 10 experiments (10 concurrent GPUs is the limit on Cerebrium’s Hobby plan) Run the script: ```bash theme={null} python run.py ``` Cerebrium Runs Monitor training progress in the W\&B dashboard: Hugging Face Token ### Next Steps 1. Export model: * Copy to AWS S3 using Boto3 * Download locally using Cerebrium Python package 2. Quality assurance: * Run CI/CD tests on model outputs * Use Cerebrium's [webhook functionality](https://docs.cerebrium.ai/endpoints/webhook) 3. Deployment: * Create inference endpoint * Load model directly from Cerebrium volume ### Conclusion Combining W\&B for tracking and Cerebrium for serverless compute enables efficient hyperparameter sweeps for Llama 3.2, optimizing model performance with minimal effort. View the complete code in the [GitHub repository](https://github.com/CerebriumAI/examples/tree/master/12-training/1-wandb-sweep) # List Audit Logs Source: https://cerebrium.ai/docs/api-reference/audit-log/list-audit-logs https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/audit-log List the project's audit log, newest first. Readable by the project owner only, over the window the project's plan allows. # Apply Coupon Source: https://cerebrium.ai/docs/api-reference/billing/apply-coupon https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json post /v2/projects/{project_id}/billing/coupons/apply Apply a promotional coupon code to the project's billing account. # Get Billing Graph Source: https://cerebrium.ai/docs/api-reference/billing/get-billing-graph https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/billing/graph Retrieve daily billing cost data per app as time series for charting. # Run App Source: https://cerebrium.ai/docs/api-reference/cerebrium-run/run-app https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json post /v3/projects/{project_id}/apps/{app_id}/run Execute code on a deployed app. # Assign Domain to App Source: https://cerebrium.ai/docs/api-reference/custom-domains/assign-domain-to-app https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json post /v2/projects/{project_id}/domains/{domain_id}/assign Assign a validated custom domain to a specific app. # Create Custom Domain Source: https://cerebrium.ai/docs/api-reference/custom-domains/create-custom-domain https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json post /v2/projects/{project_id}/domains Create a new custom domain for a project with DNS validation records. # Delete Custom Domain Source: https://cerebrium.ai/docs/api-reference/custom-domains/delete-custom-domain https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json delete /v2/projects/{project_id}/domains/{domain_id} Remove a custom domain from a project. Domain must be unassigned from all apps first. # Get Custom Domain Source: https://cerebrium.ai/docs/api-reference/custom-domains/get-custom-domain https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/domains/{domain_id} Retrieve details for a specific custom domain. # List Custom Domains Source: https://cerebrium.ai/docs/api-reference/custom-domains/list-custom-domains https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/domains Retrieve a list of custom domains for a specific project. # Unassign Domain from App Source: https://cerebrium.ai/docs/api-reference/custom-domains/unassign-domain-from-app https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json delete /v2/projects/{project_id}/domains/{domain_id}/assign Remove the assignment between a domain and an app. # Validate Custom Domain Source: https://cerebrium.ai/docs/api-reference/custom-domains/validate-custom-domain https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json post /v2/projects/{project_id}/domains/{domain_id}/validate Manually trigger DNS validation for a custom domain. Useful for retrying failed validations after DNS configuration changes. # Complete File Upload Source: https://cerebrium.ai/docs/api-reference/files/complete-file-upload https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json post /v2/projects/{project_id}/volumes/{volume_id}/cp/complete Finalize the file upload process to a specific volume. Intended for internal use - rather use the `cerebrium cp` command. # Delete File Source: https://cerebrium.ai/docs/api-reference/files/delete-file https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json delete /v2/projects/{project_id}/volumes/{volume_id}/rm Remove a file from a specific volume. Intended for internal use - rather use the `cerebrium rm` command. # Download File Source: https://cerebrium.ai/docs/api-reference/files/download-file https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/volumes/{volume_id}/download Download a file from a specific volume. # Initialize File Upload Source: https://cerebrium.ai/docs/api-reference/files/initialize-file-upload https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json post /v2/projects/{project_id}/volumes/{volume_id}/cp/initialize Begin the file upload process to a specific volume. Intended for internal use - rather use the `cerebrium cp` command. # List Files Source: https://cerebrium.ai/docs/api-reference/files/list-files https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/volumes/{volume_id}/ls Retrieve a list of files in a specified volume. Intended for internal use - rather use the `cerebrium ls` command. # Get GitHub Install URL Source: https://cerebrium.ai/docs/api-reference/integrations/get-github-install-url https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v4/projects/{project_id}/integrations/github/install-url Generate a GitHub App installation URL for a project. # Get Repo File Tree Source: https://cerebrium.ai/docs/api-reference/integrations/get-repo-file-tree https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v4/projects/{project_id}/integrations/github/repos/{owner}/{repo}/tree Retrieve the file tree for a GitHub repository. # Get Repo Metadata Source: https://cerebrium.ai/docs/api-reference/integrations/get-repo-metadata https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v4/projects/{project_id}/integrations/github/repos/{owner}/{repo} Retrieve metadata for a GitHub repository. # Get TOML Config Source: https://cerebrium.ai/docs/api-reference/integrations/get-toml-config https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v4/projects/{project_id}/integrations/github/repos/{owner}/{repo}/toml Parse and retrieve the cerebrium.toml configuration from a GitHub repository. # List GitHub Branches Source: https://cerebrium.ai/docs/api-reference/integrations/list-github-branches https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v4/projects/{project_id}/integrations/github/repos/{owner}/{repo}/branches List branches for a GitHub repository. # List GitHub Repos Source: https://cerebrium.ai/docs/api-reference/integrations/list-github-repos https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v4/projects/{project_id}/integrations/github/repos List accessible GitHub repositories for the project's GitHub integration. # List Integrations Source: https://cerebrium.ai/docs/api-reference/integrations/list-integrations https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v4/projects/{project_id}/integrations List all integrations for a project. # Remove GitHub Integration Source: https://cerebrium.ai/docs/api-reference/integrations/remove-github-integration https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json delete /v4/projects/{project_id}/integrations/github Remove the GitHub integration from a project. # Cancel Run Source: https://cerebrium.ai/docs/api-reference/runs/cancel-run https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json delete /v2/projects/{project_id}/apps/{app_id}/runs/{run_id} Cancel an ongoing run for an app. # List App Secrets Source: https://cerebrium.ai/docs/api-reference/secrets/list-app-secrets https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/apps/{app_id}/secrets Retrieve a list of secrets for a specific app. # List Secrets Source: https://cerebrium.ai/docs/api-reference/secrets/list-secrets https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/secrets Retrieve a list of secrets for a specific project. # Update App Secrets Source: https://cerebrium.ai/docs/api-reference/secrets/update-app-secrets https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json patch /v2/projects/{project_id}/apps/{app_id}/secrets Modify secrets for a specific app. The request body is a JSON object mapping secret names to string values. # Update Secrets Source: https://cerebrium.ai/docs/api-reference/secrets/update-secrets https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json patch /v2/projects/{project_id}/secrets Modify secrets for a specific project. The request body is a JSON object mapping secret names to string values. # Create Service Account Source: https://cerebrium.ai/docs/api-reference/service-accounts/create-service-account https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json post /v2/projects/{project_id}/service-accounts Create a new service account for a project. # Delete Service Account Source: https://cerebrium.ai/docs/api-reference/service-accounts/delete-service-account https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json delete /v2/projects/{project_id}/service-accounts/{service_account_id} Delete a service account. # List Service Account Keys Source: https://cerebrium.ai/docs/api-reference/service-accounts/list-service-account-keys https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/service-accounts/{service_account_id}/keys List all tokens for a service account. # List Service Accounts Source: https://cerebrium.ai/docs/api-reference/service-accounts/list-service-accounts https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/service-accounts List all service accounts for a project. # Update Service Account Source: https://cerebrium.ai/docs/api-reference/service-accounts/update-service-account https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json patch /v2/projects/{project_id}/service-accounts/{service_account_id} Update service account grants. # Change Plan Source: https://cerebrium.ai/docs/api-reference/subscriptions/change-plan https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json patch /v2/projects/{project_id}/subscription Upgrade or downgrade the subscription plan for a project. # Get Payment URL Source: https://cerebrium.ai/docs/api-reference/subscriptions/get-payment-url https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/payment-url Generate a Stripe checkout URL for adding a payment method. # Get Subscription Source: https://cerebrium.ai/docs/api-reference/subscriptions/get-subscription https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/subscription Retrieve the current subscription plan and status for a project. # List Invoices Source: https://cerebrium.ai/docs/api-reference/subscriptions/list-invoices https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/invoices Retrieve billing invoices for a project. # List Payment Methods Source: https://cerebrium.ai/docs/api-reference/subscriptions/list-payment-methods https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/payment-methods Retrieve saved payment methods for a project. # Remove Payment Method Source: https://cerebrium.ai/docs/api-reference/subscriptions/remove-payment-method https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json delete /v2/projects/{project_id}/payment-methods/{payment_method_id} Delete a saved payment method from a project. # Invite User Source: https://cerebrium.ai/docs/api-reference/users/invite-user https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json post /v2/projects/{project_id}/users Invite a user to join a project. Any member of the project may invite users. # List Users Source: https://cerebrium.ai/docs/api-reference/users/list-users https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/users Retrieve all users with access to a project. # Remove User Source: https://cerebrium.ai/docs/api-reference/users/remove-user https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json delete /v2/projects/{project_id}/users Remove a user from a project. Only project owners may remove users, and the last owner cannot be removed. # Get Volume Alerts Source: https://cerebrium.ai/docs/api-reference/volumes/get-volume-alerts https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/volumes/alerts Retrieve the storage alert thresholds configured for the project. # List Volumes Source: https://cerebrium.ai/docs/api-reference/volumes/list-volumes https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json get /v2/projects/{project_id}/volumes Retrieve a list of volumes for a specific project. # Resize Volume Source: https://cerebrium.ai/docs/api-reference/volumes/resize-volume https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json patch /v2/projects/{project_id}/volumes/{volume_id}/resize Modify the size of a specific volume. # Update Volume Alerts Source: https://cerebrium.ai/docs/api-reference/volumes/update-volume-alerts https://s3.eu-west-1.amazonaws.com/www.cerebrium.ai/openapi_spec.json put /v2/projects/{project_id}/volumes/alerts Replace the storage alert thresholds for the project. The project owner is emailed when a volume's used space crosses a configured threshold.