Maintained buyer guide

Free AI Inference in 2026: Credits, Limits and Catches

Free inference is not one market. Some providers give repeatable daily capacity, some expose a prototype sandbox, and others rotate zero-price models until demand changes. This guide compares what you can actually build before paying.

7 current providersClaims checked 2026-08-11
Hetzner

not stated; experimental availability may change or stop during high demand

Verified
Useful monthly capacity
≈15B input + 150M output tokens per 30 days, if continuously available
Models / fit
Qwen/Qwen3.6-35B-A3B-FP8
Card
unknown
After free use
Requests are rate-limited. Hetzner does not promise continued availability.
Official source · Aug 11, 2026
Nous Portal

not stated; availability can rotate

Verified
Useful monthly capacity
Not publicly quantified; catalog and account limits can rotate
Models / fit
poolside/laguna-s-2.1:free, poolside/laguna-xs-2.1:free, stepfun/step-3.7-flash:free, tencent/hy3:free, upstage/solar-pro4:free
Card
unknown
After free use
Choose a paid model or wait for limits to reset; free-model availability can rotate.
Official source · Aug 11, 2026
OpenRouter

Ongoing; model availability rotates

Verified
Useful monthly capacity
Request-based; current account cap applies across free models
Models / fit
openrouter/free, Models ending in :free
Card
no
After free use
Wait for the daily reset, purchase credits to change account limits, or use paid models.
Official source · Aug 11, 2026
Cloudflare Workers AI

Ongoing daily allocation

Verified
Useful monthly capacity
300,000 Neurons per 30 days
Models / fit
Cloudflare Workers AI model catalog
Card
no for Workers Free
After free use
Free-plan operations fail until reset. Workers Paid charges $0.011 per 1,000 Neurons above the free allocation.
Official source · Aug 11, 2026
NVIDIA

not stated

Verified
Useful monthly capacity
Not stated as a monthly allowance; rate-limited prototype access
Models / fit
NVIDIA NIM hosted catalog
Card
no
After free use
Use a partner endpoint or deploy a NIM locally when the hosted prototype limit is insufficient.
Official source · Aug 11, 2026
Google Gemini API

Ongoing for selected models

Verified
Useful monthly capacity
Model- and project-specific quota shown in AI Studio
Models / fit
Selected Gemini models marked Free Tier on the official pricing page
Card
no for Free tier
After free use
Wait for quota reset or link billing and move to a paid usage tier.
Official source · Aug 11, 2026
GroqCloud

Ongoing free account tier

Verified
Useful monthly capacity
Model-specific daily request and token limits
Models / fit
GroqCloud production and preview catalog
Card
no for Free tier
After free use
Wait for limits to reset or upgrade to the Developer plan.
Official source · Aug 11, 2026

Start with workload shape, not the biggest number

A huge daily token ceiling is useful only if the endpoint supports your model, input type, latency target, and reliability needs. Hetzner's experiment is unusually generous but explicitly unstable. Cloudflare's allowance is durable but measured in Neurons, so model choice changes how far it goes. NVIDIA Build is a prototype environment rather than a promised monthly bucket.

For a small evaluation harness, a rotating free-model router may be enough. For a customer-facing feature, treat every free path as disposable capacity and keep a paid fallback.

Permanent free tier vs promotional access

Repeatable free tier

A documented allowance resets on a known cadence. Cloudflare's daily Neuron allocation is the clearest example in this snapshot.

Experiment or rotating catalog

The provider can remove a model, throttle access, or end the experiment without a published date. Hetzner and rotating :free catalogs belong here.

The privacy catch is part of the price

A zero-dollar request can still carry a data trade-off. Google marks free-tier usage differently from paid-tier usage on its pricing table. Router services may select providers with different retention rules. Never send secrets, production customer data, or regulated records until you have checked the current provider and model terms.

  • • Prefer synthetic evaluation data for first tests.
  • • Pin a provider when router privacy policies differ.
  • • Re-check terms before moving from prototype to production.

A practical selection order

  1. 1. Define the minimum model and modality.

    Text-only extraction, vision, coding agents, and embeddings need different catalogs.

  2. 2. Convert limits into your own requests.

    Estimate input, output, retries, and concurrency—not just headline tokens.

  3. 3. Check card, region, and data terms.

    These can disqualify an otherwise attractive tier.

  4. 4. Price the fallback before launch.

    Know what happens after the allowance and how quickly you can switch.

See only the offers that are current now

The guide explains the trade-offs. The live Deals page handles expiry, verification windows, and claim links.

Open verified inference deals