Cloudflare Workers vs Vercel: $5 Entry Price Meets a 128MB Ceiling

|Author: QUASA Editorial Team|6 min read| 1
Cloudflare Workers vs Vercel: $5 Entry Price Meets a 128MB Ceiling

For short APIs serving users across regions, Cloudflare Workers Paid is a strong cost fit. Cloudflare’s Workers pricing sets a $5 monthly minimum per account, includes 10 million requests and 30 million CPU milliseconds, and charges $0.30 and $0.02 per additional million respectively, with no separate data-transfer fee. The constraint is a fixed 128 MB memory limit per isolate, shared by work in that isolate rather than granted afresh to every request.

Vercel Functions are a better fit when a route needs more memory, a full Node.js runtime, or Next.js features with less deployment configuration. Vercel’s framework comparison puts its configurable function memory at up to 4 GB and contrasts regional function placement with Workers’ default execution near the visitor. Cloudflare can run Next.js through the OpenNext adapter, so the choice turns on the route’s memory use, dependencies and distance from its data, as well as its bill.

What the compute bill actually measures

Workers Paid meters incoming dynamic requests and active CPU time. Waiting on a database query or an upstream API does not consume Worker CPU milliseconds. The monthly minimum is an account charge, so adding a second small Worker does not automatically mean another base fee; use beyond the included allowances does increase the bill. Other services such as storage have their own meters, which the function-only examples below exclude.

Vercel’s Fluid compute pricing separates Pro invocations at $0.60 per million, active CPU, and provisioned memory; its monthly Pro usage credit can offset usage charges. For a function in Washington, D.C., the listed rates are $0.128 per CPU-hour and $0.0106 per GB-hour of provisioned memory. CPU billing pauses during I/O, but memory remains allocated and billable until the last request on an instance finishes. An instance can handle concurrent requests, so multiplying elapsed time by request count overstates memory use if requests share instances; these estimates also exclude the Pro plan fee and other services.

Short global API: location and allowances favor Workers

Take a hypothetical API with 10 million dynamic requests per month and 2 milliseconds of CPU work per request. Its 20 million CPU milliseconds and request count both fit the Workers Paid allowances, yielding the $5 monthly account minimum for this compute workload. A small authentication check, redirect, or response transformation could fit this pattern if its real memory use stays within the isolate limit.

At the stated Vercel Pro rates, those requests imply $6 in invocation charges and about $0.71 for active CPU, before provisioned memory or the plan fee and before applying the usage credit. The amount of memory and the time instances stay active are needed for a fuller estimate. A Worker near the visitor can shorten the trip to the application code, but a remote database call may still dominate response time; this case favors Workers most clearly when the handler has little back-end I/O.

Database-heavy Next.js rendering: the cheap estimate can be impossible

Now suppose a hypothetical route renders 1 million pages a month, uses 50 milliseconds of CPU and 200 milliseconds of elapsed time per render, and has a 1 GB resident working set. Pure Workers arithmetic gives $5.40: the request count is covered, while 50 million CPU milliseconds exceed the included amount by 20 million. Yet that version of the route cannot run within an isolate’s fixed memory ceiling. Streaming data, moving state to storage, or changing the rendering design might make a smaller Worker viable, but those are application changes with their own engineering cost.

Assume the same route uses a 2 GB Vercel allocation in Washington, D.C. Its invocation and active CPU charges would be about $2.38 in total. If every render occupied a separate instance for exactly 200 milliseconds, provisioned memory would amount to about 111 GB-hours and $1.18. Real instance lifetimes and concurrent requests can change that figure, and the example omits caching, bandwidth, storage and the Pro plan fee. The numerical result is a meter calculation, not a measured performance result or a complete hosting bill.

Placement can matter more than the compute difference. Several sequential queries to a database in one region repeatedly pay the route-to-database round trip when code runs near a faraway visitor. A Vercel function in the database’s region may improve that path; Workers can also use Smart Placement to move execution toward back-end services. Next.js compatibility adds a separate cost: on Cloudflare, OpenNext translates the build for the Workers runtime, and features such as incremental caching and revalidation need additional configuration. Vercel supplies those framework features in its native deployment path.

Long AI calls: waiting is cheap for CPU, but not always for memory

Consider a hypothetical 100,000 HTTP requests that each wait 20 seconds for an AI service while using 5 milliseconds of local CPU. The Workers total is 500,000 CPU milliseconds, comfortably inside Paid allowances, provided the code and any buffered data fit in the isolate. Cloudflare puts no hard wall-clock limit on an HTTP-triggered Worker while the client stays connected, although CPU limits still apply and work associated with a disconnected request may be canceled. This makes streaming a response while awaiting an upstream model technically plausible.

Vercel’s corresponding Pro compute charges would be about $0.06 for invocations and $0.02 for CPU at the chosen regional rates. With a hypothetical 0.5 GB allocation and one instance occupied for each full 20-second wait, provisioned memory would be about 278 GB-hours, or $2.94; concurrency can reduce the instance time needed to serve the same traffic. For much longer calls, Vercel’s duration update allows configured Node.js and Python Functions on Pro and Enterprise to run for up to 30 minutes, while durations above 800 seconds remain in beta and require Fluid compute. Duration and memory rules therefore matter even when local CPU time is tiny.

Choose for the route with the binding constraint

Workers makes the clearest case for lightweight global endpoints that fit the isolate and its Node.js compatibility layer. Vercel earns its larger resource envelope when rendering needs more resident memory, a dependency expects full Node.js, or a Next.js application relies on native caching and revalidation. A database-heavy route may also favor regional execution close to its data, whereas a thin API can benefit more from code close to users. For a mixed application, the decisive comparison is the request path with the most memory, back-end round trips, or time spent waiting, rather than the platform’s cheapest published request price.

Also read:

Share:

Subscribe to our newsletter

Get the latest Web3, AI, and crypto news delivered straight to your inbox.

0