Most teams still rent a GPU box and let it idle around the clock for a workload that spikes for five minutes a day.
Cloudflare Workers AI removes that idle-GPU tax by moving inference itself onto the edge — serverless GPUs deployed to 180+ cities, called the same way you'd call any function in a Worker. 60+ open-source and partner models behind one API, nothing to provision.
Inference runs on GPUs deployed to more than 180 cities on Cloudflare's network. A request is typically served from a nearby node instead of a round trip to one distant region — the difference between an agent that feels instant and one that feels laggy.
No driver management or capacity planning. Workers AI is invoked the same way you'd call any function inside a Worker — provisioning and scaling are handled for you.
Pairs with AI Gateway for observability, caching, and rate limiting on top of any AI request, and with Vectorize as the vector database behind RAG and semantic search.
Inference is billed per request in a normalized unit called a Neuron — not per hour of rented hardware. Industry GPU utilization commonly sits well under 50%; Workers AI removes the cost of that idle time entirely.
Data sent to Workers AI for inference is not used by Cloudflare to train models, which helps teams meet data-handling expectations without extra tooling.
Compare a traditional self-hosted GPU deployment against the Workers AI architecture. Select the ⓘ badge on any node for a plain-English and a technical explanation.
The catalog spans text, vision, audio, and embeddings, sourced from Meta, Mistral, Google, OpenAI, DeepSeek, Moonshot AI, Black Forest Labs, Deepgram, and others, alongside Cloudflare-hosted fine-tuning bases. Every model is called the same way, by name, with an @cf/ or @hf/ prefix. Rate limits below are Workers AI defaults per task type; individual models can carry their own higher or lower limit.
Chat, reasoning, coding, and agentic language models, from 1B-parameter edge models up to trillion-parameter mixture-of-experts systems.
Prompt-to-image generation and editing, billed per megapixel tile and per diffusion step so cost tracks output resolution.
Dense vector representations for semantic search, RAG retrieval, and clustering — designed to pair with Vectorize.
Transcription and translation of spoken audio, including a low-latency WebSocket model built for live conversational agents.
Speech synthesis for narration and voice-agent responses, priced per thousand input characters.
Many-to-many machine translation, including a model trained specifically for Indic-language pairs.
Sentiment analysis, reranking for RAG pipelines, and safety classification for moderating prompts or model output.
Label prediction over large class sets using convolutional networks trained on ImageNet-scale data.
Vision-language models that describe, caption, or answer questions about image input.
Bounding-box localization and labeling of multiple objects within a single frame.
Workers AI ships with the primitives production and agentic workloads actually need, not just a prompt-in, text-out call.
Two modes: a traditional request/response format for defining tool schemas, and an embedded mode where the model itself drives multi-step tool execution inside a single call.
Constrain text-generation output to a schema using response_format, so downstream code can parse a model's answer without guarding against malformed JSON.
Attach your own trained LoRA adapter to a compatible base model at inference time, or start from Cloudflare's catalog of public adapters.
Submit large volumes of inference requests as a single job and retrieve results later — built for bulk and offline workloads rather than interactive latency.
Reuse computation from previously seen prompt prefixes to reduce latency and cost on repeated or templated requests.
Convert HTML, images, and other document formats to clean Markdown via the toMarkdown method — a common pre-processing step for RAG ingestion.
Workers AI bills in a unified unit called a Neuron, which normalizes cost across every model type: tokens, image megapixels, and audio minutes are all converted into the same measure.
10,000 Neurons per day, available on both the Workers Free and Workers Paid plans, resetting daily at 00:00 UTC.
On Workers Paid, usage beyond the free allocation is billed at $0.011 per 1,000 Neurons. Workers Free accounts hit a hard rate limit instead of billing.
Each model has its own Neuron cost by task, size, and unit — tokens, resolution, or audio minutes — published on that model's catalog page.
| Model | Task | Price |
|---|---|---|
| @cf/meta/llama-4-scout-17b-16e-instruct | Text Generation | $0.27 / M input · $0.85 / M output tokens |
| @cf/openai/gpt-oss-120b | Text Generation | $0.35 / M input · $0.75 / M output tokens |
| @cf/meta/llama-3.1-8b-instruct-fp8-fast | Text Generation | $0.045 / M input · $0.384 / M output tokens |
| @cf/baai/bge-m3 | Text Embeddings | $0.012 / M input tokens |
| @cf/black-forest-labs/flux-2-klein-9b | Text-to-Image | $0.015 per first megapixel (1024×1024) |
| @cf/openai/whisper-large-v3-turbo | Speech Recognition | $0.0005 per audio minute |
| @cf/deepgram/aura-2-en | Text-to-Speech | $0.030 per 1,000 input characters |
| @cf/huggingface/distilbert-sst-2-int8 | Text Classification | $0.026 / M input tokens |
| Task Type | Default Limit |
|---|---|
| Automatic Speech Recognition | 720 / min |
| Image Classification | 3,000 / min |
| Image-to-Text | 720 / min |
| Object Detection | 3,000 / min |
| Summarization | 1,500 / min |
| Text Classification | 2,000 / min |
| Text Embeddings | 3,000 / min (1,500 on bge-large-en-v1.5) |
| Text Generation | 300 / min (up to 1,500 on select smaller models) |
| Text-to-Image | 720 / min |
| Translation | 720 / min |
Pick whichever fits the stack — none of them require a server to run or manage.
Choose a simulated user location and run the request. This is an illustrative simulation, not live measured data — it visualizes why routing inference to the nearest of Workers AI's 180+ GPU cities tends to beat a round trip to one distant cloud region.