How Linkup Scales AI-Native Web Search with Cerebrium

How Linkup Scales AI-Native Web Search with Cerebrium

Introduction

AI agents are only as useful as the information they can access.

A sales agent researching a prospect needs to know what that company announced this week. A legal copilot may need the latest regulatory filing. An investment agent needs to understand what happened in the market minutes ago. The underlying models weren’t trained on any of that information - they need a reliable way to search and reason over the live web.

That is what Linkup is building.

Linkup provides a web search API purpose-built for AI applications and agents. Rather than sitting on top of a traditional search engine designed for humans clicking links, Linkup operates its own crawling, indexing and retrieval infrastructure, returning fresh, structured information that models can consume directly.

Today, companies including Cohere, Legora, SNCF, Artisan and others use Linkup to bring real-time web intelligence into production AI systems - from autonomous sales agents and legal copilots to enterprise search and research workflows.

The company has grown quickly. Within months of launching its API, Linkup was already serving thousands of users having scaled globally by 20x, and its search infrastructure has since expanded into multiple search modes ranging from sub-second retrieval to multi-step agentic research.

That growth creates a demanding infrastructure problem.

A single Linkup query can trigger embedding, retrieval and reranking workloads that need to complete in milliseconds. Traffic can change rapidly as agents fan out into many parallel searches, while the economics of running search at high volume make idle GPU capacity expensive. Linkup therefore needed infrastructure that could combine the elasticity of serverless compute with the latency and reliability expected from a production search engine.

That’s where Cerebrium came in.

The Challenge

Search is a real-time, globally distributed workload: every query can trigger multiple stages of retrieval, embedding and reranking, and the response needs to come back quickly regardless of where the user is located.

As Linkup grew, four infrastructure requirements became increasingly important:

  1. GPU availability and fast scaling. Linkup’s workloads are well suited to smaller accelerators scaled horizontally. The team needed the right GPUs to be available when traffic increased, without permanently provisioning enough capacity for peak demand.

  2. Low network latency. Model inference is only one component of search latency. Sending requests across regions can add hundreds of milliseconds to every query, which compounds quickly as AI agents perform multiple searches in parallel. Linkup needed workloads running close to its users and customers.

  3. Global availability and data residency. Linkup serves customers across the US and Europe, including enterprises with requirements around where their data is processed. The team needed the ability to run the same search infrastructure across multiple regions while keeping traffic within the appropriate geography.

  4. Simple deployment. Linkup’s ML team continuously experiments with new embedding models, rerankers, configurations and custom Docker images. They needed to move quickly without turning every model change into an infrastructure project.

Linkup had previously used platforms including Baseten and Google Cloud Run, but wanted more flexibility around capacity, regional deployment, scaling and cost.

Why Cerebrium

Scaling GPU capacity without provisioning for peak traffic

Cerebrium gave Linkup access to a range of GPU types, allowing the team to match different embedding, reranking, and inference workloads to the most appropriate hardware while scaling each of them independently as traffic changed.

Checkpointing made that serverless architecture significantly more practical. Rather than repeating the full model initialization process every time new GPU capacity came online, Cerebrium could restore workers from an already initialized state.

For Linkup, this reduced startup times by more than 60%, bringing new GPU workers online in just a few seconds.

That meant Linkup could respond quickly to sudden increases in search traffic without permanently provisioning infrastructure for peak demand. The team could keep enough baseline capacity for normal usage, then rapidly scale additional workers as needed while maintaining much higher infrastructure utilization.

Running search closer to users

For a search engine, compute latency is only part of the equation. Network latency matters too.

Cerebrium’s multi-region infrastructure allows Linkup to run workloads closer to the customers generating the search request. Cerebrium has previously measured requests between Europe and a US deployment adding roughly 100–250ms of network latency, compared with 30–70ms when the workload runs locally in the EU, CA or US.

At the routing layer, Cerebrium’s distributed networking infrastructure is designed to direct requests to healthy capacity across regions and clouds. This means Linkup can serve US traffic from US infrastructure and European traffic from European infrastructure, minimizing unnecessary network hops while maintaining a single deployment workflow.

That regional architecture also helps Linkup support customers with data residency requirements. European workloads can remain within European infrastructure while US workloads run in the US, without Linkup having to maintain separate infrastructure stacks for each geography. Cerebrium supports multi-region deployments specifically for latency, data residency and fault-tolerance requirements.

It also adds another layer of resilience. If capacity in one region becomes unavailable, traffic can be redirected toward healthy capacity elsewhere rather than the failure taking down the search workload entirely.

Giving the ML team control without the infrastructure overhead

Cerebrium’s custom container support also gave Linkup control over how its search stack was packaged and deployed.

Instead of treating every embedding or reranking model as an isolated hosted endpoint, Linkup could combine models and application logic inside its own custom application. The ML team could test different models, dependencies and configurations while Cerebrium handled the underlying provisioning, scaling and regional infrastructure.

For Linkup, that combination - elastic GPU capacity, fast startup, low-latency regional routing and flexible deployments - meant the team could focus on improving search quality rather than operating the infrastructure underneath it.

And the infrastructure wasn’t the only differentiator.

“The various performative features of the Cerebrium platform were important for us, but the support has also been amazing. The engineers are often there answering us within just minutes, and when we’ve needed engineering support, we’ve been able to work directly with the team.”

Denis Charrier - CTO & Co Founder

The Results

With Cerebrium, Linkup has been able to scale its search infrastructure without building a large GPU operations layer internally. New GPU capacity can come online in seconds, workloads can be deployed across regions closer to customers, and the team can experiment with different models and hardware without reworking its infrastructure each time.

That gives Linkup more room to optimize for the things that matter most to its product: search quality, latency, reliability, and cost. Most importantly, its ML and engineering teams can spend less time thinking about where and how workloads run, and more time improving the search engine itself. As Linkup continues to grow and AI agents generate more search traffic across geographies, Cerebrium gives the team an infrastructure layer that can grow with it.


Related case studies

See all
  • Case Study
Read Case Study
Scaling Resemble AI’s Real-Time Deepfake Detection Models with Cerebrium
  • Case Study
  • Video
  • Generative AI
Read Case Study
How DistilLabs is Delivering 50% Lower Inference Costs with Production-Grade Autoscaling on Cerebrium
  • Case Study
  • LLMs
  • Generative AI
Read Case Study
Lelapa AI uses Cerebrium to Break Language Barriers