LLM API buying looks like a price-per-token comparison until the first invoice and the first data-protection review.
This guide ranks twelve providers on gateway fees, whether free allowances survive production, where prompts are processed and retained, and how much of the OpenAI interface actually works, with smaller independent vendors placed ahead of larger platforms.
AuthorHannah ReiterSenior Analyst, Business Applications
Vendors can pay for visibility on this page. It never changes what an entry
says about a product, including the criticism, and we earn nothing when you click through to a
vendor. How that works.
In short
What LLM API providers software does
An LLM API provider hosts large language models behind an HTTP interface, so applications send prompts and receive completions, billed by tokens, time or subscription, without the buyer running GPUs.
In this order: setup effort, what it really costs, how your data comes back out, whether
you can leave, and who each LLM API providers tool is built for. Why those five, and why there is no
score out of ten, is on the how we work page.
Provider list price, no markup; free tier with a subset of models
Yes
Teams already deploying on Vercel who want one billing line for model calls
Teams that want a gateway without a Vercel account
Country is where the vendor is headquartered or contracts from, which is a
different question from where your data is hosted. Where the two tell different stories, the
entry says so.
Swedish inference API with prepaid monthly token plans
Ranked #1 of 12 in Best LLM API Providers in 2026.
Published pricingEurope
Berget separates two products on one pricing page: serverless inference billed in tokens against a monthly prepaid plan, and seat-based coding subscriptions that never charge per token. For an API buyer, only the first matters.
The stated terms come from a single Swedish limited company, and the site says inference runs under EU data protection rules with no retention of prompts. Dedicated endpoints exist for custom models, but the page gave no price for them.
What stands out
Zero retention claim
Prepaid plans
Trial credit
Where it costs you
The model catalogue is limited to what Berget hosts; there is no routing to other vendors
Plans bundle prepaid tokens, so a quiet month still bills the plan
Right for
Teams that want predictable monthly spend and a stated zero-retention policy
Wrong for
Anyone who needs every frontier model through one key
SwedenPrepaid monthly token plans, published; trial credit on signup
One API key routing to hundreds of models, at provider list prices
Ranked #2 of 12 in Best LLM API Providers in 2026.
Free tierPublished pricingNorth America
OpenRouter sits between your code and the vendors that serve the models, so you gain breadth and fallback options and add a hop. It is an independent New York company whose only product is this router.
Zero data retention routing is offered on every plan, and the site says it does not train on your data. Bringing your own provider keys is allowed with a monthly fee-free allowance, after which a small percentage applies.
What stands out
Model router
Platform fee
Free plan
Where it costs you
The platform fee on credits applies to every purchase, and rises on the Business plan
Unused credits may expire 365 days after purchase
Right for
Prototyping across many models, or teams that want one invoice for several vendors
Wrong for
Teams that must avoid an extra intermediary between app and model
United StatesPay-as-you-go credits with a percentage platform fee; free plan with rate limits, published
Per-token inference for open models with flexible priority tiers
Ranked #3 of 12 in Best LLM API Providers in 2026.
Published pricingNorth America
DeepInfra publishes three scheduling levels on one price list: Standard at the base rate, Priority at one and a half times that for faster first tokens at peak, and Flex at a discount for asynchronous work.
Dedicated GPU rentals run from A100 to B300 cards at stated hourly rates. The privacy policy says API inputs and outputs are not stored, sold or used for training without consent, and that the company is Deep Infra Inc. in Palo Alto.
What stands out
Per-token billing
Dedicated GPUs
Priority and flex tiers
Where it costs you
Data may be stored and processed in the United States, per the privacy policy
Non-language models bill by execution time, which is harder to forecast than tokens
Right for
Cost-sensitive teams running open models who can accept US processing
Wrong for
Buyers who need processing kept inside the EU
United StatesPer token for language models, per execution time for others, published
Dutch OpenAI-compatible API hosted in Paris on renewable power
Ranked #4 of 12 in Best LLM API Providers in 2026.
Published pricingEurope
GreenPT bundles a chat product, a no-code chatbot builder and an API under one company. The plans page lists subscriptions with a 14-day trial and says data is stored in the EU and is not used for model training.
The named models include a Mistral Small variant, GPT-OSS 120B and a Gemma model. Because API pricing is documented separately from the subscriptions, check which of the two your usage will actually be billed on.
What stands out
EU hosting
OpenAI-compatible
Subscription
Where it costs you
API access needs a paid subscription; token prices are in the documentation, not on the plans page
The headline models are a small open-weight set
Right for
Teams that want EU hosting and a sustainability reporting angle
Wrong for
Teams needing proprietary frontier models from the big labs
NetherlandsSubscription per month, published; 14-day trial
Model APIs plus per-minute dedicated GPU deployments for your own models
Ranked #5 of 12 in Best LLM API Providers in 2026.
Self-hostablePublished pricingNorth America
The pricing page splits Baseten into Basic, Pro and Enterprise. Model APIs carry per-million-token rates for models such as GPT OSS 120B and DeepSeek V4.1 Flash, while dedicated deployments charge per minute from a T4 up to a B200, with no charge for idle time stated as a design point.
Pro adds priority access to high-demand GPUs. The governing law in the terms is California, and the company is in San Francisco.
What stands out
Dedicated GPUs
Model APIs
Per-minute billing
Where it costs you
Data residency control and self-hosted deployments are Enterprise plan features
Dedicated deployments need sizing decisions that a token API does not
Right for
Teams serving fine-tuned or custom models who want per-minute GPU billing
Wrong for
Teams that only need a drop-in chat completion endpoint
United StatesPer token for model APIs, per minute for dedicated GPUs, published; starter credits
Serverless token billing plus per-second dedicated GPU deployments
Ranked #6 of 12 in Best LLM API Providers in 2026.
Published pricingElsewhere
Fireworks separates three meters: serverless tokens for input, output and cached tokens, training priced by model size and method, and per-second GPU time for dedicated capacity. Metadata such as token counts is retained, and logging is opt-in for features like FireOptimizer.
The OpenAI-compatible base URL lets existing client code move over, with differences such as usage statistics arriving in the final streaming chunk and max_tokens truncating when prompts exceed context.
What stands out
Prepaid credits
OpenAI-compatible
Zero retention default
Where it costs you
The pricing page mentions no free credits, so serverless use starts with a purchase
The Responses API stores conversation data for 30 days unless store is set to false
Right for
Teams that want serverless tokens now and dedicated GPUs once volume grows
Wrong for
Buyers who need a stated company location before contracting
Not stated by the vendorPrepaid credits per token for serverless; per GPU second for on-demand, published
Inference on the company's own chips, with a small starting credit
Ranked #7 of 12 in Best LLM API Providers in 2026.
Pricing on requestNorth America
Cerebras also lists partner routes to its inference through AWS Marketplace, OpenRouter, Hugging Face and Vercel, which gives you a way to buy it with an existing contract.
Going direct means a Developer tier with a payment method on file and a short promotional window. The company states California governing law in its terms. For a first test, the clock on the promotional credit is the constraint to plan around.
What stands out
Own hardware
Promo credit
Enterprise tier
Where it costs you
The promotional credit expires 30 days after activation and access then pauses
Model prices were not visible on the pricing page we read
Right for
Teams that want to test fast inference on a vendor-owned chip platform
Wrong for
Anyone needing a published rate card before signing up
United StatesPay-as-you-go credits; one-time promotional credit; enterprise by agreement
Three-plan inference cloud with a free plan and enterprise subscription
Ranked #8 of 12 in Best LLM API Providers in 2026.
Free tierPublished pricingNorth America
The plans page gives the three tiers without model prices in what we could read. The Developer plan includes all production and preview models, while Enterprise adds production-grade limits and optional custom features.
The documentation lists an OpenAI compatibility route and custom checkpoints. Community credits are offered for developers building on the platform. The company is headquartered in San Jose, and its terms are governed by California law.
What stands out
Free plan
Pay as you go
Enterprise subscription
Where it costs you
Production rate limits sit behind the Enterprise subscription
Two product names, SambaCloud and SambaStack, appear in the documentation
Right for
Teams comparing a free entry plan against a paid developer tier
Wrong for
Teams that need production limits without an enterprise contract
United StatesFree plan; pay-as-you-go developer plan; enterprise subscription
Fast inference with a free plan and rate-limited developer tier
Ranked #9 of 12 in Best LLM API Providers in 2026.
Free tierPublished pricingNorth America
Limits are applied per organisation and measured in requests, tokens and audio seconds, and you hit whichever cap comes first. The documentation points to the Developer plan, Batch and Flex processing, and enterprise arrangements for higher capacity.
The terms name Groq LLC with a Mountain View address and California law. The pricing page returned no readable price list, so the per-model rates must be confirmed in the console.
What stands out
Free plan
OpenAI-compatible
Batch discount
Where it costs you
The OpenAI compatibility layer does not support logprobs or logit_bias
The privacy policy excludes customer API data and points to a separate agreement
Right for
Latency-sensitive applications that can live within the documented rate limits
Wrong for
Teams that need logprobs or a retention statement upfront
United StatesFree plan with rate limits; pay-as-you-go developer plan
One token to reach hundreds of models across partner providers
Ranked #10 of 12 in Best LLM API Providers in 2026.
Published pricingNorth America
The pricing page lists 200+ models from partner providers, billed on your Hugging Face account when routed through the router endpoint, which works with the OpenAI client.
Team and Enterprise organisations get a credit pool per seat and can set spending limits. The terms name Hugging Face, Inc., a Delaware company under New York law, and the privacy policy names a French SAS as its main EU establishment.
What stands out
No markup claimed
Monthly credits
Provider routing
Where it costs you
The privacy policy gives no retention period for inference requests
Monthly credits for free users are small and stated as subject to change
Right for
Developers already on Hugging Face who want to try many hosted models
Wrong for
Buyers who need a stated retention period for prompts
United StatesProvider pass-through rates, published; small monthly credits on free accounts
Chinese lab's own API, in OpenAI and Anthropic formats
Ranked #11 of 12 in Best LLM API Providers in 2026.
Published pricingAsia-Pacific
The pricing page lists two models and bills input, cached input and output tokens separately, with off-peak hours priced at half. The documentation gives separate base URLs for OpenAI and Anthropic clients, so existing SDK code can usually be repointed.
The privacy question is the real filter: the terms reference de-identified processing but no location, and the governing law is that of the People's Republic of China.
What stands out
Own models
Dual API formats
Off-peak discount
Where it costs you
The terms do not state a data storage location and are governed by PRC law
You buy only DeepSeek's own two models, with no routing to others
Right for
Cost-driven workloads that can tolerate a vendor in mainland China
Wrong for
Regulated buyers who need a stated data location
ChinaPer token with cache and peak pricing, published
A gateway billed at provider list prices, tied to a Vercel team
Ranked #12 of 12 in Best LLM API Providers in 2026.
Free tierPublished pricingNorth America
The pricing documentation separates free credits, paid credits and add-on surcharges. Custom reporting, a team-wide provider allowlist and zero data retention all carry their own charges, while the per-request retention option has none on Pro and Enterprise.
You also bear payment processing fees unless you invoice as an Enterprise customer. Vercel's terms name Vercel Inc. as the contracting entity, with California law and courts in San Francisco County, which is worth knowing before you sign.
What stands out
No markup
Free tier
Bring your own key
Where it costs you
Access needs a Vercel team account
Team-wide zero data retention costs extra per request and requires Pro or Enterprise
Right for
Teams already deploying on Vercel who want one billing line for model calls
Wrong for
Teams that want a gateway without a Vercel account
United StatesProvider list price, no markup; free tier with a subset of models
An LLM API provider hosts large language models behind an HTTP interface, so applications send prompts and receive completions, billed by tokens, time or subscription, without the buyer running GPUs. The differences that matter are rarely in the feature list, so this is
the order we would work through them.
01
Decide whether you need a published price
11 of the 12 tools here publish what they cost; the other 1 quote per organisation. The ones you can compare without a sales call: Berget AI, OpenRouter, DeepInfra, GreenPT, Baseten, Fireworks AI, SambaNova, Groq, Hugging Face Inference Providers, DeepSeek, Vercel AI Gateway.
02
Decide how much the jurisdiction matters
These 12 vendors are established in 5 countries across 4 regions (North America 8, Europe 2, Asia-Pacific 1, Elsewhere 1). That decides whose disclosure law applies to what the vendor holds, wherever the servers are.
Gateway fees: OpenRouter, Vercel AI Gateway and Hugging Face Inference Providers
Three products in this list sit in front of other people's models, and they make money in different places. OpenRouter bills at provider list price and takes a percentage platform fee when you buy credits, 5.5 per cent on Standard and 8 per cent on Business.
Vercel AI Gateway states no markup and no platform fee on tokens, but charges per request for add-ons such as team-wide zero data retention and leaves payment processing fees with you. Hugging Face says it passes provider rates through, and gives small monthly credits instead. A gateway that looks free on tokens can still cost money once you switch on retention controls, so price the controls you need, not the base rate.
Ask which add-ons your compliance team will switch on, and price each per thousand requests.
Check whether credits expire; OpenRouter may expire unused credits after 365 days.
Test bring-your-own-key billing, since the fee-free allowance differs between gateways.
Retention claims: Heabsy, Berget AI and DeepInfra
Every vendor says it respects your data, and the useful question is what the text commits to. Heabsy separates models that run on its own EEA hardware with zero retention from models routed through global providers, so one API key can carry two different promises. Berget AI states that prompts and outputs are not retained, under Swedish company terms.
DeepInfra's privacy policy says it does not store, sell or train on API data without consent, and that data may be processed in the United States. Groq's privacy policy excludes customer API data altogether and defers to a services agreement. Read the contract that governs customer data, not the marketing page, and ask which of the promises survive when a request is routed to another provider.
Find the document that governs API data, which may differ from the privacy policy.
Ask whether a retention promise applies to every model or only to some.
Get the processing location in writing for the exact model you will call.
Free allowances: Groq, Cerebras and SambaNova
A free plan in this market is a rate limit, and a promotional credit is a clock. Groq documents a Free plan and a Developer plan, and gives gpt-oss-120b a free base of 30 requests and 8K tokens per minute, which is enough for a demo and not for a queue of users. Cerebras gives a one-time promotional credit after you add a payment method, expiring after 30 days, and pauses access when it runs out.
SambaNova lists a Free plan, a pay-as-you-go Developer plan and an Enterprise subscription for production limits. Treat free usage as a way to test latency and output quality, then price the paid tier separately, because its limits and its cost are what you will actually live with in production.
Multiply the free rate limit by your expected concurrent users before relying on it.
Note the expiry date of any promotional credit when you create the account.
Ask what limit applies on the first paid tier, not only on the free one.
Compatibility claims: DeepSeek, Groq and GreenPT
Pointing the OpenAI SDK at a new base URL is the usual migration, and it works until you rely on a feature the vendor left out. Groq says its API is mostly compatible and lists logprobs and logit_bias as unsupported. DeepSeek documents separate base URLs for OpenAI and Anthropic formats, which helps if your code uses either.
GreenPT offers an OpenAI-compatible API for open-weight models, with chat, embeddings and reranking endpoints in its documentation. Compatibility describes the request shape, not the model, so the same prompt can behave differently on different models. Run your own test set against each candidate before moving production traffic, covering streaming, tool calls and structured output as well as plain chat, since those are where gaps usually appear.
List every request parameter your code sets, and test each against the vendor.
Check whether tool calling and structured output behave the same on the target model.
Keep the old provider configured as a fallback for the first month.
What goes wrong most often when buying LLM API providers software
Choosing a provider by headline token price while ignoring gateway fees, add-on charges and expiring credits that change the invoice.
Treating a free plan as capacity planning, when its rate limit and expiry decide whether it survives real traffic.
Accepting a data-residency claim without checking which models it covers; routed models can be served elsewhere.
Moving production traffic after a compatibility check on one endpoint, without testing tool calls and parameters the application uses.
07
Frequently asked questions
7 answers
What is the best LLM API providers in 2026?
Berget AI leads our ranking of 12. A small Swedish company, Berget AI AB, that sells serverless inference through an OpenAI-compatible API and says prompts and outputs are not retained.
Plans are fixed monthly sums that include prepaid tokens, which makes spend predictable and the trial starts with starting credit. The catch is that the model list is whatever Berget chooses to host, and the prepaid structure means unused allowance is a cost you carry.
Which LLM API providers tools publish their pricing?
11 of the 12, with the pricing model each one publishes:
Groq: Free plan with rate limits; pay-as-you-go developer plan.
Hugging Face Inference Providers: Provider pass-through rates, published; small monthly credits on free accounts.
DeepSeek: Per token with cache and peak pricing, published.
Vercel AI Gateway: Provider list price, no markup; free tier with a subset of models.
The other 1 quote per organisation.
Is there a free LLM API providers tool?
OpenRouter, SambaNova, Groq, Vercel AI Gateway offer a free tier or a free self-hosted edition.
Where are these LLM API providers vendors established?
In 5 countries across 4 regions: North America 8, Europe 2, Asia-Pacific 1, Elsewhere 1.
Berget AI: Sweden.
OpenRouter: United States.
DeepInfra: United States.
GreenPT: Netherlands.
Baseten: United States.
Fireworks AI: Not stated by the vendor.
Cerebras: United States.
SambaNova: United States.
Groq: United States.
Hugging Face Inference Providers: United States.
DeepSeek: China.
Vercel AI Gateway: United States.
Which LLM API providers tools can you host yourself?
Baseten. The other 11 are hosted by the vendor only.
What should you use instead of Berget AI?
OpenRouter and DeepInfra are the next two on this page. OpenRouter is for Prototyping across many models, or teams that want one invoice for several vendors; DeepInfra is for Cost-sensitive teams running open models who can accept US processing.
Who should not buy Berget AI?
Anyone who needs every frontier model through one key. The model catalogue is limited to what Berget hosts; there is no routing to other vendors.
If your LLM API providers product belongs among these 12, tell us what it does and who it is for. Inclusion is an editorial call; what a listing is and is not is set out under software advice.