Workers AI gets a speed boost, batch workload support, more LoRAs, new models, and a refreshed dashboard
ID: df0c23b5-a207-5702-9691-96b47fb6fb27
STIX ID: report--df0c23b5-a207-5702-9691-96b47fb6fb27
Feed Name: Cloudflare Blog
Cloudflare announces major Workers AI updates: 2–4x faster inference via speculative decoding, prefix caching, and an updated backend; an asynchronous batch API for large, non-latency-sensitive workloads; expanded LoRA support (more models, higher ranks, larger adapters); quality-of-life improvements including updated pricing and a new usage dashboard; and over 10 new or updated models such as @cf/deepseek-ai/deepseek-r1-distill-qwen-32b, @cf/baai/bge-m3, @cf/openai/whisper-large-v3-turbo, @cf/mistralai/mistral-small-3.1-24b-instruct, @cf/google/gemma-3-12b-it, @cf/qwen/qwq-32b, and @cf/qwen/qwen2.5-coder-32b-instruct. These changes target better performance, reliability, customization, and usability for distributed inference at scale.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
