
Cloudflare AI Search Goes GA—and Billing Starts November 1

Cloudflare made AI Search generally available on October 1, 2026, and usage-based billing begins November 1, 2026. The release adds native image embeddings and OCR for scanned PDFs; newly created instances use hybrid search by default. For developers moving from preview, the immediate questions are what an existing index will do and which parts of its workload will appear on the bill.
AI Search’s general availability does not require every existing instance to switch retrieval methods. Teams need to check the method configured on each index before assuming it has keyword search, while budgeting for ingestion, storage and queries beyond the monthly allowance. Workers AI embeddings and reranking performed within AI Search are included in its pricing; generation, query rewriting and external model providers follow their applicable service billing.
New hybrid defaults do not rebuild existing indexes
Hybrid retrieval combines semantic vector search with full-text matching. That matters when a query contains an exact product code or phrase as well as a broader concept: the keyword and vector methods can contribute different matches. The GA default applies to new instances, so it does not by itself establish the configuration of an index created during preview.
Cloudflare’s keyword indexing documentation specifies that `index_method.keyword` can be enabled when an instance is created or updated, and that changing `index_method` triggers a full reindex. An existing vector-only instance can therefore stay with its current retrieval method. Moving it to hybrid is a deliberate indexing change, with time and content-processing work to plan around.
Capacity may decide whether that change is practical. On Workers Paid, an instance with keyword search enabled supports up to 500,000 files, compared with 1 million for a vector-only instance. Teams close to the higher limit cannot treat hybrid as a simple query setting; the keyword index must exist, and its file limit applies to the corpus being indexed. For smaller collections, the decision rests on whether exact-term matching adds value to the semantic results already in use.
Images and scanned PDFs need separate decisions
Native image retrieval changes what the index can represent. AI Search can embed image pixels directly while retaining captions for text-based understanding. A caption may describe the subject of an image but omit a distinctive pattern, layout or other visual detail that a later query needs. Native embeddings keep a route to those details when the instance uses a multimodal embedding model.
The model choice matters at query time too. With a supported multimodal model, an image query can be embedded directly into the same vector space as indexed images and text. With a text-only embedding model, AI Search converts the query image to a caption before searching. Both paths accept an image input, but they offer different kinds of matching; an existing text-only index should not be assumed to gain native pixel-based retrieval merely because the service reached GA.
OCR solves a different problem. A scanned PDF can consist of page images with little or no extractable text. Turning on OCR extracts text before the usual chunking and embedding steps, making those pages searchable as text. It is available to every account, but the image-processing work enters the ingestion calculation. An OCR decision can therefore change both the content users can find and the volume used to estimate a monthly bill.
The allowance and rates behind the November bill
Cloudflare’s GA release and pricing account sets a monthly allowance on all Workers plans of 5 million ingestion tokens, 10 GB of stored data, 1,000 semantic queries and 1,000 full-text queries. It also sets the larger-file boundary: plain-text and code files and PDFs with OCR enabled can reach 10 MiB, while PDFs without OCR and other supported formats remain limited to 4 MiB. The higher ceiling is therefore tied to file type and, for PDFs, the OCR setting.
Beyond the allowance, base ingestion is $0.75 per million tokens, with an additional $0.50 per million tokens for image processing. Storage is $2 per GB-month. Semantic, vector and hybrid queries cost $0.75 per 1,000 queries; full-text queries cost $0.10 per 1,000. The semantic and full-text allowances are separate pools, replacing the shared query allowance described during preview.
Image processing does not bring a second free token pool: its tokens share the monthly ingestion allowance. For images and scanned documents processed with OCR, extracted text can count toward image processing as well as base ingestion. A cost estimate based only on uploaded file size or the number of documents could miss that distinction, especially when OCR makes previously inaccessible page content available for indexing.
Parsing, chunking, keyword indexing, and Workers AI embedding and reranking performed by AI Search are included in its usage pricing. Answer generation, query rewriting and third-party model providers remain on the applicable account or gateway bill. Thus, an application that retrieves passages and then asks a model to write an answer has a generation cost to account for beyond its AI Search query usage.
What to settle before billing starts
For an existing deployment, the compatibility decision starts with its stored index method. A switch from vector-only to hybrid requires a full reindex and must fit the keyword-enabled file limit. For collections of images, the embedding model determines whether image queries use native visual representations or a text caption. For scanned PDFs, OCR determines whether page text becomes searchable and whether image-processing ingestion enters the usage estimate.
The cost estimate then needs distinct lines for ingestion tokens, stored index size, semantic or hybrid queries, full-text queries, and any image processing. Generation and external-provider calls belong on their own service bills. Billing activation is the next concrete change: once it begins, a running index with usage above its monthly allowances can incur charges even if its retrieval configuration stays exactly as it was during preview.
Also read:
Related articles


Keenable Raises $26M—but Its 100B-Page Index Has No Named Customers

7 Best Brand Protection Platforms in 2026 (Ranked & Reviewed)

Cloudflare R2 vs Amazon S3: Egress Can Overwhelm Storage Price

Beehiiv Adds Enforced AI Crawler Controls—but Only on Higher Tiers

Track Instagram and TikTok in Search Console—No Website Required
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.