01 — System Positioning

Most teams still rent a GPU box and let it idle around the clock for a workload that spikes for five minutes a day.

Inference,
measured in nanoseconds,
not regions.

Cloudflare Workers AI removes that idle-GPU tax by moving inference itself onto the edge — serverless GPUs deployed to 180+ cities, called the same way you'd call any function in a Worker. 60+ open-source and partner models behind one API, nothing to provision.

180+GPU cities
10kfree Neurons / day
60+models, one API

Strategic Value Pillars

The core trade

Global, Low-Latency Inference

Inference runs on GPUs deployed to more than 180 cities on Cloudflare's network. A request is typically served from a nearby node instead of a round trip to one distant region — the difference between an agent that feels instant and one that feels laggy.

Zero Infrastructure to Manage

No driver management or capacity planning. Workers AI is invoked the same way you'd call any function inside a Worker — provisioning and scaling are handled for you.

A Full-Stack AI Toolkit

Pairs with AI Gateway for observability, caching, and rate limiting on top of any AI request, and with Vectorize as the vector database behind RAG and semantic search.

Why the Billing Model Is Different

Pay-As-You-Go, in Neurons

Inference is billed per request in a normalized unit called a Neuron — not per hour of rented hardware. Industry GPU utilization commonly sits well under 50%; Workers AI removes the cost of that idle time entirely.

Data Not Used for Training

Data sent to Workers AI for inference is not used by Cloudflare to train models, which helps teams meet data-handling expectations without extra tooling.

02 — Architecture & Topology

Topological Comparison

Compare a traditional self-hosted GPU deployment against the Workers AI architecture. Select the ⓘ badge on any node for a plain-English and a technical explanation.

03 — Model Catalog

60+ Models Across 10 Task Types

The catalog spans text, vision, audio, and embeddings, sourced from Meta, Mistral, Google, OpenAI, DeepSeek, Moonshot AI, Black Forest Labs, Deepgram, and others, alongside Cloudflare-hosted fine-tuning bases. Every model is called the same way, by name, with an @cf/ or @hf/ prefix. Rate limits below are Workers AI defaults per task type; individual models can carry their own higher or lower limit.

💬 Text Generation

Chat, reasoning, coding, and agentic language models, from 1B-parameter edge models up to trillion-parameter mixture-of-experts systems.

Llama 4 ScoutGPT-OSS 120BKimi K2.7DeepSeek-R1 Distill Qwen 32B
300 req/min (up to 1,500 on select smaller models)

🎨 Text-to-Image

Prompt-to-image generation and editing, billed per megapixel tile and per diffusion step so cost tracks output resolution.

FLUX.2 [dev]FLUX.2 KleinFLUX.1 [schnell]Leonardo Phoenix
720 req/min

🧬 Text Embeddings

Dense vector representations for semantic search, RAG retrieval, and clustering — designed to pair with Vectorize.

BGE-M3Qwen3 EmbeddingBGE base/largePLaMo-Embedding-1B
3,000 req/min (1,500 on BGE-large)

🎙️ Automatic Speech Recognition

Transcription and translation of spoken audio, including a low-latency WebSocket model built for live conversational agents.

Whisper Large v3 TurboDeepgram Nova-3Deepgram Flux
720 req/min

🔊 Text-to-Speech

Speech synthesis for narration and voice-agent responses, priced per thousand input characters.

Deepgram Aura-2MeloTTS
720 req/min

🌐 Translation

Many-to-many machine translation, including a model trained specifically for Indic-language pairs.

M2M100IndicTrans2
720 req/min

🏷️ Text Classification

Sentiment analysis, reranking for RAG pipelines, and safety classification for moderating prompts or model output.

DistilBERT SST-2BGE RerankerLlama Guard 3
2,000 req/min

🖼️ Image Classification

Label prediction over large class sets using convolutional networks trained on ImageNet-scale data.

ResNet-50
3,000 req/min

📝 Image-to-Text

Vision-language models that describe, caption, or answer questions about image input.

Moondream
720 req/min

📦 Object Detection

Bounding-box localization and labeling of multiple objects within a single frame.

Task type — see docs for current models
3,000 req/min
04 — Platform Capabilities

Beyond Single-Shot Inference

Workers AI ships with the primitives production and agentic workloads actually need, not just a prompt-in, text-out call.

Function CallingBeta

Two modes: a traditional request/response format for defining tool schemas, and an embedded mode where the model itself drives multi-step tool execution inside a single call.

JSON Mode

Constrain text-generation output to a schema using response_format, so downstream code can parse a model's answer without guarding against malformed JSON.

LoRA Fine-TunesBeta

Attach your own trained LoRA adapter to a compatible base model at inference time, or start from Cloudflare's catalog of public adapters.

Asynchronous Batch APIBeta

Submit large volumes of inference requests as a single job and retrieve results later — built for bulk and offline workloads rather than interactive latency.

Prompt Caching

Reuse computation from previously seen prompt prefixes to reduce latency and cost on repeated or templated requests.

Markdown ConversionBeta

Convert HTML, images, and other document formats to clean Markdown via the toMarkdown method — a common pre-processing step for RAG ingestion.

05 — Pricing & Limits

Neuron-Based Pricing

Workers AI bills in a unified unit called a Neuron, which normalizes cost across every model type: tokens, image megapixels, and audio minutes are all converted into the same measure.

Free Allocation

10,000 Neurons per day, available on both the Workers Free and Workers Paid plans, resetting daily at 00:00 UTC.

Paid Overage

On Workers Paid, usage beyond the free allocation is billed at $0.011 per 1,000 Neurons. Workers Free accounts hit a hard rate limit instead of billing.

Per-Model Pricing

Each model has its own Neuron cost by task, size, and unit — tokens, resolution, or audio minutes — published on that model's catalog page.

Example Model Pricing

ModelTaskPrice
@cf/meta/llama-4-scout-17b-16e-instructText Generation$0.27 / M input · $0.85 / M output tokens
@cf/openai/gpt-oss-120bText Generation$0.35 / M input · $0.75 / M output tokens
@cf/meta/llama-3.1-8b-instruct-fp8-fastText Generation$0.045 / M input · $0.384 / M output tokens
@cf/baai/bge-m3Text Embeddings$0.012 / M input tokens
@cf/black-forest-labs/flux-2-klein-9bText-to-Image$0.015 per first megapixel (1024×1024)
@cf/openai/whisper-large-v3-turboSpeech Recognition$0.0005 per audio minute
@cf/deepgram/aura-2-enText-to-Speech$0.030 per 1,000 input characters
@cf/huggingface/distilbert-sst-2-int8Text Classification$0.026 / M input tokens

Rate Limits by Task Type

Task TypeDefault Limit
Automatic Speech Recognition720 / min
Image Classification3,000 / min
Image-to-Text720 / min
Object Detection3,000 / min
Summarization1,500 / min
Text Classification2,000 / min
Text Embeddings3,000 / min (1,500 on bge-large-en-v1.5)
Text Generation300 / min (up to 1,500 on select smaller models)
Text-to-Image720 / min
Translation720 / min
06 — Implementation Pattern

Three Ways to Call Workers AI

Pick whichever fits the stack — none of them require a server to run or manage.

1. Workers Binding (native, recommended)

// wrangler.jsonc { "ai": { "binding": "AI" } } // src/index.ts export default { async fetch(request, env) { const response = await env.AI.run("@cf/meta/llama-3.1-8b-instruct", { prompt: "Explain what a Cloudflare Worker is", }); return new Response(JSON.stringify(response)); }, };

2. REST API

curl https://api.cloudflare.com/client/v4/accounts/{ACCOUNT_ID}/ai/run/@cf/meta/llama-3.1-8b-instruct \ -X POST \ -H "Authorization: Bearer {API_TOKEN}" \ -d '{ "messages": [{ "role": "user", "content": "Hello!" }] }'

3. OpenAI-Compatible Endpoint

from openai import OpenAI client = OpenAI( api_key="CLOUDFLARE_API_TOKEN", base_url="https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/v1", ) response = client.chat.completions.create( model="@cf/meta/llama-3.1-8b-instruct", messages=[{"role": "user", "content": "Hello!"}], )
07 — Interactive Simulation

Edge vs. Single-Region Latency Lab

Choose a simulated user location and run the request. This is an illustrative simulation, not live measured data — it visualizes why routing inference to the nearest of Workers AI's 180+ GPU cities tends to beat a round trip to one distant cloud region.

Disclaimer: All figures shown in this simulator are approximate, illustrative estimates generated for demonstration purposes only. They are not live measurements, are not guaranteed, and do not represent a commitment, warranty, or SLA regarding actual latency, performance, or availability of Cloudflare Workers AI or any other service. Actual results will vary by network conditions, model, payload size, and region. Nanosek accepts no liability for decisions made based on these figures.
SINGLE-REGION CLOUD (us-east-1) — ms 📍
🏢
CLOUDFLARE WORKERS AI (nearest of 180+ cities) — ms 📍
⚡
Simulated Round-Trip Telemetry
[SYSTEM] Ready. Pick a location and simulate a request.