
Most teams shipping AI in 2026 pay for the wrong Ollama tier. They set one sticker price beside another, skip the amortization entirely, and conclude that a managed plan costs too much next to hardware that quietly bills more per request once electricity and depreciation land. The figure that settles the argument is never the monthly fee. It is the daily request volume where the two cost lines cross, and almost nobody calculates it before signing up.
The Ollama download counter passed fifty-two million per month in Q1 2026, and the questions hitting search engines shifted with that scale. People stopped asking whether local AI works. They now ask what Ollama Cloud costs, what hardware a specific model demands, and at what volume running your own box starts to win.
Every number on this page traces back to a source. The plan and model prices come from ollama.com/pricing, checked 2026-09-25. By then Ollama had replaced the session and weekly limits of its old plans with monthly usage credits. The throughput figures come from measured tokens per second on an Apple M4 Max and an RTX 4090 rather than vendor claims. The cost model amortizes each build across thirty-six months and adds electricity at typical metropolitan rates, which is the step most comparisons skip.
Three tables and a chart carry that data. The plan table shows what each tier includes, and the model table shows what a million tokens costs on each model. The hardware chart maps model size against the memory it demands and the tokens per second it returns. The crossover table shows the monthly token bill where a self-hosted box starts to win.
Subscribe to the newsletter for more local AI cost analyses and infrastructure deep dives.
The Short Answer
Ollama Cloud sells four self-serve plans, checked on ollama.com/pricing on 2026-09-25. The Free plan costs nothing and includes starter credits for a small set of starter models. Cloud Pro costs twenty dollars per month and includes sixty dollars of usage credits. Cloud Max costs one hundred dollars per month and includes three hundred dollars of credits, 10 concurrent requests and early access to the newest models. Team, still in early access, costs five hundred dollars for a thousand dollars of credits shared across unlimited users.
| Plan | Price | Included usage | Concurrent requests |
|---|---|---|---|
| Free | $0 | Starter credits, starter models | 1 |
| Pro | $20 / month or $200 / year | $60 of credits a month | 3 |
| Max | $100 / month | $300 of credits a month | 10 |
| Team (early access) | $500 / month | $1,000 of credits a month, shared | 10 |
| Enterprise | Custom | Volume usage pricing | Not published |
Credits pay for tokens at each model's own rate. The table below lists a selection of the published prices per million tokens.
| Model | Input | Cached input | Output |
|---|---|---|---|
| nemotron-3-nano | $0.06 | none listed | $0.24 |
| gpt-oss:20b | $0.07 | $0.035 | $0.30 |
| gemma4 | $0.14 | $0.05 | $0.40 |
| gpt-oss:120b | $0.15 | $0.014 | $0.60 |
| deepseek-v4-pro | $1.32 | $0.044 | $3.96 |
| deepseek-v4-pro (off-peak) | $0.66 | $0.022 | $1.98 |
| glm-5.3 | $1.40 | $0.26 | $4.40 |
| kimi-k3 | $3.00 | $0.30 | $15.00 |
Off-peak rates apply outside 12:00 to 18:00 UTC on weekdays and all day on weekends. Only the DeepSeek models list them.
Always confirm the live prices on the official site. Included credits reset monthly and never roll over, and usage past them draws on credits you buy at the same token rates. Concurrency matters more than the headline price for most production workloads, because requests over your limit wait in a queue.
Ollama hardware requirements move in three steps rather than a smooth curve. A 7B Q4 model fits in eight gigabytes of unified RAM or VRAM. A 32B Q4 model needs thirty-two gigabytes. A 70B Q4 model needs sixty-four gigabytes, which rules out every consumer NVIDIA card and leaves Apple Silicon as the only consumer path to that tier.
Self-hosting overtakes Ollama Cloud at a measurable monthly token bill. An RTX 4090 build wins once usage passes about $110 a month at list rates. A Mac Studio M4 Max wins past about $355. Below those lines, the managed plan wins on operational simplicity alone. The rest of this page shows the work behind each of those figures.
What Ollama Cloud Actually Is
Ollama Cloud is the managed-inference companion to the local Ollama runtime. It serves a list of cloud-enabled open-weight models behind a hosted endpoint, with the same OpenAI-compatible HTTP surface that local Ollama exposes. You point your client at a different base URL and the rest of your code does not change. That portability is the entire pitch. Prompts, agents, and RAG pipelines that run on a laptop work identically on Ollama Cloud and on a self-hosted GPU box.
Ollama pitches Pro at shorter, well-defined day-to-day tasks. Max targets power users who run several agents at once. Team adds shared credits and central billing for a whole team. The Free plan exists to let you validate a prompt before you commit a card.
Hardware Requirements by Model Size
Ollama hardware requirements are not a mystery. A model needs to fit in memory before it can serve a token. Quantization (Q4 by default for most models in the registry) reduces the disk and memory footprint to roughly twenty-five percent of the original full-precision weight. The disk file scales linearly with parameter count. RAM and VRAM jump in tiers because models must fit entirely in memory for usable throughput.
Three practical takeaways from this curve.
A 7B model is the universal floor. Eight gigabytes of unified RAM or VRAM is enough, which makes any modern laptop with Apple Silicon or an NVIDIA card with 8 GB of VRAM a viable target. Forty tokens per second on an M4 is faster than human reading speed, which means streaming UX feels instant.
A 32B model is the production sweet spot. Thirty-two gigabytes of unified memory delivers Qwen 2.5 32B at fifteen tokens per second on an M4 Max, with MMLU scores within striking distance of GPT-4. This is the tier where local inference stops being a hobbyist's compromise and starts being a serious cloud-API replacement.
A 70B+ model is unified-memory territory. The 70B Q4 tier needs sixty-four gigabytes of memory, which rules out every consumer NVIDIA card. Apple Silicon's unified memory architecture (M2 Ultra at 192 GB, M4 Max at 128 GB) is the only consumer path to running this class of model locally. Beyond 120B parameters, Ollama Cloud is the right answer unless you have an actual GPU server.
Where Self-Hosting Beats Cloud
The pricing-versus-volume question is where most teams get the math wrong. Cloud Max looks expensive at one hundred dollars per month until you price the alternative. A GPU box carries electricity, depreciation, and the operational tax of running your own runtime. The crossover depends on your monthly token bill, and daily request volume drives that bill.
A single RTX 4090 build amortizes to roughly seventy dollars per month over thirty-six months, plus power. Pro's twenty dollars buys sixty dollars of usage, and every dollar past that bills at list rates. So Pro plus top-ups reaches seventy dollars at $110 of monthly usage, and Max already costs one hundred. Above $110, the 4090 beats every Cloud plan.
A Mac Studio M4 Max amortizes to about one hundred and fifty-five dollars per month. It also runs 70B models that the 4090 cannot load. Max covers three hundred dollars of usage for its one hundred dollar fee. Max plus top-ups reaches one hundred and fifty-five dollars at $355 of monthly usage, so the Mac Studio pulls ahead past that line.
| Setup | Amortized monthly cost | Beats Ollama Cloud above | Worked example |
|---|---|---|---|
| RTX 4090 build | about $70 | about $110 of monthly usage | about 16,700 daily requests on gpt-oss:20b |
| Mac Studio M4 Max | about $155 | about $355 of monthly usage | about 26,300 daily requests on gpt-oss:120b |
The worked example assumes 1,000 input and 500 output tokens per request, a 30-day month and no cached-input discount. Swap in your own token counts and model rate, since the dollar thresholds do not depend on the model. Check that one box can sustain that throughput before you buy it.
Below $110 of monthly usage, Cloud Pro is the right answer for most teams. The operational simplicity, zero hardware capex, and built-in geographic redundancy make the unit-cost argument for self-hosting irrelevant.
Far above those lines, self-hosting wins by a wide margin. Every dollar of usage past Max's three hundred dollars of credits bills at list rates, while the rig's cost stays flat. Once you top up Max every month, you sit close to the Mac Studio's break-even and well past the 4090's.
Ollama 2026 Update Timeline
Ollama is now a real platform, not a wrapper script. Two and a half years of compounding releases have taken the project from a hundred thousand downloads to fifty-two million per month and from twelve thousand GitHub stars to one hundred and fifty-eight thousand.
The updates that matter most for production work in 2026:
Native vision support across Qwen-VL, Llama 3.2 Vision, and the Phi-4 multimodal lines. Vision models now run with the same ollama run command as text-only models, with no extra adapter installation.
OpenAI-compatible structured outputs with JSON Schema validation. The runtime enforces the schema during decoding, which eliminates entire classes of retry loops in agentic workflows. This was the single biggest quality-of-life improvement in 2026.
Tool calling parity with the OpenAI Chat Completions API. Models that support tool calling (Qwen 2.5, Llama 3.1+, Mistral Large, DeepSeek-V2.5) now expose the exact same tools and tool_choice shape, so frameworks like Mastra, LangGraph, and CrewAI work without provider-specific adapters.
Ollama Cloud GA. The Cloud product moved out of beta and now exposes the same HTTP surface as the local runtime, which makes it a drop-in deployment target.
For a deeper look at how these changes affect agent frameworks, see Local AI Agent Frameworks 2026 and GitHub Copilot + Ollama for Agentic Local LLMs.
A Practical Decision Tree
The cost and hardware data above collapses into a short decision tree.
Building a side project or solo agent. Start with local Ollama on whatever hardware you already own. A 7B model on an M-series MacBook or an 8 GB consumer GPU covers ninety percent of personal use cases at zero recurring cost.
Building a startup MVP without provisioning hardware. Ollama Cloud Pro at twenty dollars per month is the right entry point. You get sixty dollars of monthly credits, the larger pro models, the same API surface as local, and zero ops. Migrate later when volume justifies it.
Running production within Max's three hundred dollars of monthly credits. Cloud Max. The operational simplicity beats self-hosting on TCO once you account for monitoring, on-call, and replacement hardware budgets.
Running production past Max's included credits every month, or any regulated workload. Self-host. A single RTX 4090 box covers up to 32B models with room to spare. Add a Mac Studio for 70B+ workloads and you have a two-machine cluster that handles most enterprise scenarios. Pair the rig with a Cloud Max account as a failover lane.
Need 120B+ MoE models. Ollama Cloud is the only sane option unless you have a GPU server. Any plan reaches them once you buy credits, and Max adds early access to the newest models. The hardware required to self-host these models exceeds the lifetime cost of Max for most teams.
When Cloud APIs Still Win
Ollama and Ollama Cloud do not replace every workload. Frontier reasoning tasks (long chain-of-thought on novel problems, complex multi-step coding agents) still favor GPT-5.3-Codex and Claude Opus 4.6 by a noticeable margin. The gap is narrowing every quarter, but it is real today. For a side-by-side comparison, see Claude Opus 4.6 vs GPT-5.3 Codex.
The right architecture in 2026 is hybrid. Use Ollama (local or Cloud) as the default for high-volume cheap inference: classification, summarization, RAG synthesis, agent tool selection. Reserve frontier cloud APIs for the few requests that genuinely need frontier capability. This pattern cuts most teams' inference bill by sixty to eighty percent without quality loss.
Closing Numbers
Ollama Cloud Pro costs twenty dollars per month with sixty dollars of credits. Max costs one hundred with three hundred of credits. A self-hosted RTX 4090 amortizes to seventy and beats every Cloud plan past about $110 of monthly usage. A Mac Studio M4 Max amortizes to one hundred and fifty-five and beats Max past about $355. Hardware requirements are linear in disk space and tiered in RAM. The 7B floor is eight gigabytes, the 32B production tier is thirty-two, the 70B unified-memory tier is sixty-four.
Those are the numbers, checked against ollama.com/pricing on 2026-09-25. Pick the row in the decision tree that matches your monthly token bill, then run the math against what you pay today. Most teams shipping AI in 2026 are paying for the wrong tier.
Subscribe for the next deep dive on running production agents on a hybrid local plus cloud stack.