Where to Fine-Tune an LLM Now: Every Option, Priced (2026)
Abhay Khant
Jan 1, 1970 • 13 min read
You can still fine-tune an LLM in 2026, but not where you used to. OpenAI's notice reads: "OpenAI is winding down the fine-tuning platform. The platform is no longer accessible to new users, but existing users of the fine-tuning platform will be able to create training jobs for the coming months" (OpenAI docs). The hard stop for new jobs lands January 6, 2027. Most other closed-model vendors already left: Mistral and Cohere deprecated their features outright, and Google's Gemini API has had nothing tunable since May 2025 (Google AI docs).
This post lists every door still open, with prices checked this week: Together AI, Fireworks, Tinker, Nebius, Alibaba Model Studio, AWS Bedrock, Vertex's supervised tuning on gemini-2.5-pro, and the run-it-yourself routes. The dead ends were verified dead too, so you can skip the 404 hunt.
Key Takeaways
- OpenAI blocks all new fine-tuning jobs on January 6, 2027; inference on existing fine-tuned models runs until each base model deprecates.
- Mistral, Cohere, and the Gemini API are verified gone; Vertex keeps supervised tuning alive on gemini-2.5-pro, and no Gemini 3.x model is tunable yet.
- Cheapest list rate found: Tinker trains gpt-oss-20b at $0.396 per 1M tokens; Together and Fireworks start at $0.48 and $0.50.
- AWS Bedrock is the only place anywhere to fine-tune a Claude model: Claude 3 Haiku, us-west-2 only.
- A third-party benchmark from March 2026 measured Tinker cheapest but slowest, Together fastest, Nebius the best workflow.
- Before migrating, test prompt caching: one practitioner reports it "gets a lot of the same benefits as fine-tuning" for style work.
What Happened to OpenAI Fine-Tuning?
OpenAI's deprecations page sets three milestones. May 7, 2026: job creation blocked for organizations that had never used fine-tuning. July 2, 2026: job creation restricted for orgs without fine-tuned-model inference in the prior 60 days. January 6, 2027: "Active existing customers will no longer be able to create new fine-tuning jobs on this date," while "Inference on fine-tuned models will continue to be available until the base models are deprecated" (OpenAI deprecations).
What's still tunable today: the gpt-4.1 trio with SFT and DPO, gpt-4o-2024-08-06 for vision fine-tuning, and o4-mini-2025-04-16 for reinforcement fine-tuning (OpenAI docs). The next cut lands October 23, 2026, when the ft-prefixed models go. That batch covers ft-gpt-3.5-turbo, ft-gpt-4, ft-gpt-4.1-nano-2025-04-14, ft-babbage-002, ft-davinci-002, and ft-o4-mini-2025-04-16, with gpt-5.6-terra, gpt-5.6-sol, and gpt-5.6-luna listed as substitutes (OpenAI deprecations).
A staff reply on the community forum, dated 2026-04-22, tried to soften the read: the gpt-5-nano substitute column "does not currently mean that gpt-5-nano is available for fine-tuning," and "Fine-tuning itself is not being removed" (OpenAI forum). But no GPT-5-family model is a fine-tuning base, so the reassurance is thin. One user in that thread pushed back that gpt-5-nano "absolutely will not do our classification correctly without fine tuning" (OpenAI forum).
This is OpenAI's third fine-tuning shutdown cycle and the first with no successor platform: /v1/fine-tunes closed January 4, 2024, and babbage-002/davinci-002 training halted October 28, 2024 (OpenAI deprecations). The wind-down barely made news. The only Hacker News thread that drew a reply, in May, asked "what other recommendations do people have?" (HN); a second thread in April went unanswered (HN)
Who Else Has Already Left Fine-Tuning?
Three more vendors are verified gone, and one is half-out.
Mistral's fine-tuning pages now carry: "This feature is deprecated and is no longer actively supported" (Mistral docs). The old terms, a $4 minimum per job and $2 monthly model storage, still appear on the page. There is nothing left to buy.
Cohere is cleaner about it: "Cohere's fine-tuning feature was deprecated on September 15, 2025" (Cohere docs). Command R and Command R+ tuning is history.
The Gemini API has been dead longest. Google's docs state that "with the deprecation of Gemini 1.5 Flash-001 in May 2025, we no longer have a model available which supports fine-tuning," and the team has "no immediate plans" to bring it back (Google AI docs).
Vertex AI is the half-out one. Supervised tuning still runs there, and one production user operates fine-tuned gemini-2.5-pro endpoints, but a Google developer forum thread confirms "supervised fine-tuning is not yet available for any Gemini 3.x model," with no official timeline (Google Dev forum). That user's 2.5 Pro shutdown was pushed back to October 2026 by email. Tune Gemini today and you are tuning a model with an expiry date.
Which Managed Platforms Still Fine-Tune Open Models?
Six vendors still run managed fine-tuning on open models. Together's LoRA SFT starts at $0.48 per 1M training tokens for models up to 16B, with a $4 minimum per job (Together AI). Six doors, at list price:
| Platform | Cheapest listed rate | Dataset floor | What you keep |
|---|---|---|---|
| Together AI | $0.48/1M LoRA SFT, ≤16B, $4 job minimum | JSONL or Parquet upload | 33 LoRA models, 12 full-FT |
| Fireworks | $0.50/1M LoRA SFT, up to 16B | SFT: "hundreds of examples, or roughly 10M+ tokens" | "you keep the resulting weights" |
| Tinker | $0.396/1M, gpt-oss-20b | Not published | LoRA adapters export to Hugging Face |
| Nebius Token Factory | Not published | Data Lab turns production logs into training data | Deploy on Token Factory endpoints |
| Alibaba Model Studio | Qwen3-8B at ¥0.006/1K tokens | SFT: "Over 1,000 entries"; DPO: 100+ pairs | Deploy after efficient SFT |
| AWS Bedrock | Tokens × epochs, plus $1.95/month storage | No published floor | Claude 3 Haiku, Llama 3.x, Nova |
Together AI: the full rate card
Together's rate card, per 1M training tokens by model size: SFT LoRA $0.48, $1.50, and $2.90 across the ≤16B, 17-69B, and 70-100B tiers; SFT Full $0.54, $1.65, and $3.20; DPO LoRA $1.20, $3.75, and $7.25; DPO Full $1.35, $4.12, and $8.00 (Together AI). Specialized models carry their own rows: Llama 4 Scout SFT at $3, gpt-oss-120B at $5, DeepSeek-V4 Flash at $6, Kimi K2.6 at $15, GLM-5.2 at $40, each per 1M tokens. Billing counts dataset size times epochs plus any evaluation tokens, and uploads go up as JSONL or Parquet via CLI, SDK, or cURL (Together docs). Thirty-three models support LoRA and twelve support full fine-tuning (Together catalog).
Fireworks: lowest data floors
Fireworks prices by parameter count: LoRA SFT at $0.50, $3.00, $6.00, and $10.00 per 1M tokens across the up-to-16B, 16.1-80B, 80-300B, and over-300B tiers, with full fine-tuning at double (Fireworks). Its data floors are the lowest documented: SFT wants "hundreds of examples, or roughly 10M+ tokens," DPO wants "hundreds to thousands of pairs," and reinforcement fine-tuning is often satisfied by fewer than 100 prompts (Fireworks docs). Two lines matter for OpenAI migrants: "you keep the resulting weights," and "Serve fine-tuned models for the same price as base models" (Fireworks pricing).
Tinker: cheapest list rate
Tinker, from Thinking Machines Lab, prices per model: gpt-oss-20b at $0.396, Qwen3-8B at $0.44, gpt-oss-120b at $0.737, Qwen3.5-9B at $1.463, and Kimi K2.6 at $4.84 per 1M tokens, with checkpoint storage at $0.10 per GB monthly and an 80% discount on cached prefill tokens (Tinker docs). Training is LoRA and PEFT-based, adapters export to Hugging Face, and its serverless inference beta is still limited to the two Inkling models.
Nebius Token Factory: the benchmark pick
Nebius Token Factory advertises post-training workflows for open models, deploys fine-tuned checkpoints on its own endpoints, and ships a Data Lab that converts production logs into training data, but it publishes no public training prices (Nebius). Its roadmap says "Full post-training and distillation workflows will soon be available." Get a quote, or lean on the benchmark numbers below.
Alibaba Model Studio: the budget option
Alibaba Model Studio is the budget option if you'll write ChatML JSONL. It runs five training types: continual pre-training, full SFT, efficient_sft (its LoRA recommendation), dpo_full, and dpo_lora (Alibaba docs). SFT wants "Over 1,000 entries" of QA pairs, DPO wants "100+ sets," and continual pre-training wants 10 million+ tokens. Training rates match the pre-trained model's inference rate: Qwen3-14B at ¥0.03 per 1K tokens and Qwen2.5-72B at ¥0.15, which lands at ¥30 to ¥150 per 1M tokens (Alibaba billing). The tunable catalog is Qwen-heavy, and the Singapore region offers only qwen3-14b with efficient_sft.
Where Can You Fine-Tune Claude?
Exactly one place: AWS Bedrock. The live fine-tuning table in Bedrock's docs lists "Anthropic Claude 3 Haiku (anthropic.claude-3-haiku-20240307-v1:0:200k)" in us-west-2, and no other Claude model appears anywhere (AWS docs). Anthropic's direct API offers no fine-tuning. No ranking article we checked mentions this door.
Bedrock also tunes Llama 3.1 8B and 70B, Llama 3.2 1B, 3B, 11B, and 90B, Llama 3.3 70B, the Amazon Nova family, and Titan models, nearly all pinned to single regions. Billing counts tokens processed times epochs, plus model storage at $1.95 per month (AWS docs). Visible on the pricing page: reinforcement fine-tuning for gpt-oss-20b and Qwen3 32B at $80.00 per training hour (AWS pricing). Per-token training rates for the Llama 3.x and Claude rows sit behind collapsed sections, so ask your AWS account team before budgeting.
What Does Fine-Tuning Actually Cost, Measured?
List prices hide throughput, and throughput is where these platforms differ most. Vintage Data, a third-party benchmark published March 30, 2026, ran roughly 30M-token function-calling jobs through Nebius Token Factory, Tinker, and Together AI (vintagedata). Measured costs: Tinker $0.36 to $0.52 per 1M tokens for LoRA, Together $0.48 to $5.00, Nebius $0.40 to $5.00. Speeds diverge hard. Tinker took 166 to 220 minutes at roughly 2,273 to 3,012 tokens per second. Together finished in 15 to 28 minutes, one run at 33,330 tokens per second. Nebius sat between at 76 minutes and 6,646 tokens per second, and the benchmark's verdict named it "the most practical platform for our iterative use case." These are third-party measurements, not vendor list prices; your job will vary.
Why pay for training at all? The economics case for small tuned models came out of the 416-point fine-tuning debate on Hacker News in March 2026. Commenter arkmm: "typical categorical / data extraction use cases would have ~10x fewer errors at 100x lower inference cost" (HN). That's the whole argument: a tuned 8B model beating a frontier model on one narrow task, at a fraction of the per-token price.
Dead Ends: Vendors With Nothing Left to Tune
Every entry below was checked this week. Save the tab.
- xAI (Grok): docs.x.ai/docs/fine-tuning returns a 404. The docs tree is inference only, grok-4.6 models, no tuning anywhere (xAI docs).
- Groq: self-described "premier neocloud for fast inference." No training product (Groq).
- OpenRouter: the docs path for fine-tuning 404s. It routes inference; it does not train.
- DeepSeek API: no fine-tune endpoint documented, v4-flash and v4-pro chat only (DeepSeek docs). Tune DeepSeek models through Together, Fireworks, or Tinker instead.
- NVIDIA hosted (build.nvidia.com): no hosted tuning; the /explore/train path 404s while the inference pages render fine.
- Predibase: predibase.com/pricing redirects to Rubrik Agent Cloud, with no independent signup path (Predibase).
- Replicate: fine-tunes FLUX image models, not LLMs (Replicate).
- Lambda Labs: both plausible tuning service URLs 404. Nothing verifiable.
- Z.ai / Zhipu: no English fine-tuning docs. Tune GLM through Together, Fireworks, or Tinker instead.
When Should You Not Fine-Tune at All?
Before migrating anywhere, ask whether you needed fine-tuning in the first place. That same 416-point thread split the field. antirez, Redis's creator, argued modern models few-shot so well that "strong prompts plus large context windows usually win." danielhanchen, Unsloth's maintainer, countered with production evidence: Cursor's online RL lifting approval rates by 28%, Vercel's AutoFix RFT, DoorDash extracting LoRAs, Perplexity's Sonar (HN). A third commenter hit ~98% accuracy converting receipt images to structured JSON from about 1,000 SFT examples on Mistral 7B.
The cheaper alternative to test first: prompt caching. gamegoblin, on OpenAI's GPT-4o fine-tuning launch thread, made the case that a cached prompt full of examples "is a lot more developer-friendly and gets a lot of the same benefits as fine-tuning," because you can update it anytime and there's no async training job (HN). If caching hits your accuracy bar, you're done. If it doesn't, the platforms above are your move.
Run-It-Yourself Routes: $3.95 an Hour, or Free
Two routes skip the managed platforms. First, rent GPUs by the second or minute: Modal charges $0.001097 per second for an H100, roughly $3.95 an hour, with H200 at $0.001261 and B200 at $0.001736 per second (Modal). Baseten bills per minute, from a T4 at $0.01052 up to an A100-80GB at $0.06667, an H100 at $0.10833, and a B200 at $0.16633 (Baseten). Both run Unsloth or LLaMA-Factory fine and deploy the tuned model afterward.
Second route, the free one: train on hardware you already own. A 7B model trains in about 6GB of memory with 4-bit QLoRA, which fits any 16GB Mac. We walk that route end to end in our guide to training an AI to write like you, corpus ladder and overfitting guards included, and the local LLM writing stack covers the $19.99 setup those runs plug into. For the own-versus-rent call, the local LLM vs API comparison frames the decision.
Platform Fit by Situation
| Your situation | Best fit | Why | Watch out for |
|---|---|---|---|
| OpenAI refugee, small corpus | Together AI or Fireworks | $0.48/$0.50 per 1M entry rates, JSONL uploads | Together's $4 job minimum |
| Cheapest training, full control | Tinker | $0.396/1M on gpt-oss-20b, HF export | Slowest in the benchmark, 166-220 min |
| Already on AWS | Bedrock | Claude, Nova, Llama 3.3 in one console | Single-region models, $1.95/month storage |
| Must tune Claude | Bedrock, Claude 3 Haiku | The only place anywhere | us-west-2 only |
| Iterating on production logs | Nebius Token Factory | Data Lab converts logs to training data | No public prices, quote required |
| On a Mac, $0 budget | Local QLoRA | 7B trains in ~6GB at 4-bit | Corpus cleaning takes longer than training |
Related Tools & Further Reading
- The how: our guide to how you train an AI to write like you covers the local LoRA route, corpus sizes, and the overfitting guards.
- The gear: the local LLM writing stack assembles the $19.99 setup those fine-tunes plug into.
- The decision: local LLM vs API frames what to run yourself and what to rent.
- Who pays for AI writing: the best AI writing assistants roundup.
- Tools that fit the workflow: the AI tools hub, and a word counter for sizing a training corpus before you upload it anywhere.
Every platform and price above was live on September 5, 2026. Prices move, and deprecation pages move faster than anything else on a vendor's site; the January 6, 2027 cutoff arrived with months of notice, not years. Check the vendor page before you commit a corpus to any fine-tuning platform.