logo

How Cloudflare runs more AI models on fewer GPUs: A technical deep-dive

ID: 2bc69c0c-d328-5a65-bf28-0821316f3202

STIX ID: report--2bc69c0c-d328-5a65-bf28-0821316f3202

Feed Name: Cloudflare Blog

Date Published: 2025-08-27

Date Updated: 2026-04-27

Author: Sven Sauleau

...
...

Cloudflare presents Omni, a platform for running many AI models on edge GPUs from a single control plane, featuring lightweight process and filesystem isolation, per-model Python environments, and controlled GPU memory overcommit via a CUDA stub and unified memory. Omni routes and buffers inference requests, enforces per-model CPU/GPU limits (including a FUSE-backed /proc/meminfo and overridden cudaMemGetInfo/cuMemGetInfo), and supports multiple backends (vLLM, Python, Infire) through a unified Python API for Workers AI features like batching and function calling. By packing small or low-traffic models per GPU and swapping them between CPU and GPU as needed, Omni improves utilization, availability, and latency while reducing idle GPU cost.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.