ToolSura Blog
ArticlesAboutContact
Search

Stay in the loop

Join thousands of developers getting weekly insights into modern web development, AI tools, and productivity.

© 2026 ToolSura Blog
AboutContactPrivacy PolicyTerms of ServiceRSS

    Table of Contents

    What Happened to OpenAI Fine-Tuning?Who Else Has Already Left Fine-Tuning?Which Managed Platforms Still Fine-Tune Open Models?Together AI: the full rate cardFireworks: lowest data floorsTinker: cheapest list rateNebius Token Factory: the benchmark pickAlibaba Model Studio: the budget optionWhere Can You Fine-Tune Claude?What Does Fine-Tuning Actually Cost, Measured?Dead Ends: Vendors With Nothing Left to TuneWhen Should You Not Fine-Tune at All?Run-It-Yourself Routes: $3.95 an Hour, or FreePlatform Fit by SituationRelated Tools & Further Reading
    HomeToolsura BlogArticle

    Where to Fine-Tune an LLM Now: Every Option, Priced (2026)

    A

    Abhay Khant

    Jan 1, 1970 • 13 min read

    You can still fine-tune an LLM in 2026, but not where you used to. OpenAI's notice reads: "OpenAI is winding down the fine-tuning platform. The platform is no longer accessible to new users, but existing users of the fine-tuning platform will be able to create training jobs for the coming months" (OpenAI docs). The hard stop for new jobs lands January 6, 2027. Most other closed-model vendors already left: Mistral and Cohere deprecated their features outright, and Google's Gemini API has had nothing tunable since May 2025 (Google AI docs).

    This post lists every door still open, with prices checked this week: Together AI, Fireworks, Tinker, Nebius, Alibaba Model Studio, AWS Bedrock, Vertex's supervised tuning on gemini-2.5-pro, and the run-it-yourself routes. The dead ends were verified dead too, so you can skip the 404 hunt.

    Key Takeaways

    • OpenAI blocks all new fine-tuning jobs on January 6, 2027; inference on existing fine-tuned models runs until each base model deprecates.
    • Mistral, Cohere, and the Gemini API are verified gone; Vertex keeps supervised tuning alive on gemini-2.5-pro, and no Gemini 3.x model is tunable yet.
    • Cheapest list rate found: Tinker trains gpt-oss-20b at $0.396 per 1M tokens; Together and Fireworks start at $0.48 and $0.50.
    • AWS Bedrock is the only place anywhere to fine-tune a Claude model: Claude 3 Haiku, us-west-2 only.
    • A third-party benchmark from March 2026 measured Tinker cheapest but slowest, Together fastest, Nebius the best workflow.
    • Before migrating, test prompt caching: one practitioner reports it "gets a lot of the same benefits as fine-tuning" for style work.

    What Happened to OpenAI Fine-Tuning?

    OpenAI's deprecations page sets three milestones. May 7, 2026: job creation blocked for organizations that had never used fine-tuning. July 2, 2026: job creation restricted for orgs without fine-tuned-model inference in the prior 60 days. January 6, 2027: "Active existing customers will no longer be able to create new fine-tuning jobs on this date," while "Inference on fine-tuned models will continue to be available until the base models are deprecated" (OpenAI deprecations).

    What's still tunable today: the gpt-4.1 trio with SFT and DPO, gpt-4o-2024-08-06 for vision fine-tuning, and o4-mini-2025-04-16 for reinforcement fine-tuning (OpenAI docs). The next cut lands October 23, 2026, when the ft-prefixed models go. That batch covers ft-gpt-3.5-turbo, ft-gpt-4, ft-gpt-4.1-nano-2025-04-14, ft-babbage-002, ft-davinci-002, and ft-o4-mini-2025-04-16, with gpt-5.6-terra, gpt-5.6-sol, and gpt-5.6-luna listed as substitutes (OpenAI deprecations).

    A staff reply on the community forum, dated 2026-04-22, tried to soften the read: the gpt-5-nano substitute column "does not currently mean that gpt-5-nano is available for fine-tuning," and "Fine-tuning itself is not being removed" (OpenAI forum). But no GPT-5-family model is a fine-tuning base, so the reassurance is thin. One user in that thread pushed back that gpt-5-nano "absolutely will not do our classification correctly without fine tuning" (OpenAI forum).

    This is OpenAI's third fine-tuning shutdown cycle and the first with no successor platform: /v1/fine-tunes closed January 4, 2024, and babbage-002/davinci-002 training halted October 28, 2024 (OpenAI deprecations). The wind-down barely made news. The only Hacker News thread that drew a reply, in May, asked "what other recommendations do people have?" (HN); a second thread in April went unanswered (HN)

    Who Else Has Already Left Fine-Tuning?

    Three more vendors are verified gone, and one is half-out.

    Mistral's fine-tuning pages now carry: "This feature is deprecated and is no longer actively supported" (Mistral docs). The old terms, a $4 minimum per job and $2 monthly model storage, still appear on the page. There is nothing left to buy.

    Cohere is cleaner about it: "Cohere's fine-tuning feature was deprecated on September 15, 2025" (Cohere docs). Command R and Command R+ tuning is history.

    The Gemini API has been dead longest. Google's docs state that "with the deprecation of Gemini 1.5 Flash-001 in May 2025, we no longer have a model available which supports fine-tuning," and the team has "no immediate plans" to bring it back (Google AI docs).

    Vertex AI is the half-out one. Supervised tuning still runs there, and one production user operates fine-tuned gemini-2.5-pro endpoints, but a Google developer forum thread confirms "supervised fine-tuning is not yet available for any Gemini 3.x model," with no official timeline (Google Dev forum). That user's 2.5 Pro shutdown was pushed back to October 2026 by email. Tune Gemini today and you are tuning a model with an expiry date.

    Which Managed Platforms Still Fine-Tune Open Models?

    Six vendors still run managed fine-tuning on open models. Together's LoRA SFT starts at $0.48 per 1M training tokens for models up to 16B, with a $4 minimum per job (Together AI). Six doors, at list price:

    PlatformCheapest listed rateDataset floorWhat you keep
    Together AI$0.48/1M LoRA SFT, ≤16B, $4 job minimumJSONL or Parquet upload33 LoRA models, 12 full-FT
    Fireworks$0.50/1M LoRA SFT, up to 16BSFT: "hundreds of examples, or roughly 10M+ tokens""you keep the resulting weights"
    Tinker$0.396/1M, gpt-oss-20bNot publishedLoRA adapters export to Hugging Face
    Nebius Token FactoryNot publishedData Lab turns production logs into training dataDeploy on Token Factory endpoints
    Alibaba Model StudioQwen3-8B at ¥0.006/1K tokensSFT: "Over 1,000 entries"; DPO: 100+ pairsDeploy after efficient SFT
    AWS BedrockTokens × epochs, plus $1.95/month storageNo published floorClaude 3 Haiku, Llama 3.x, Nova

    Together AI: the full rate card

    Together's rate card, per 1M training tokens by model size: SFT LoRA $0.48, $1.50, and $2.90 across the ≤16B, 17-69B, and 70-100B tiers; SFT Full $0.54, $1.65, and $3.20; DPO LoRA $1.20, $3.75, and $7.25; DPO Full $1.35, $4.12, and $8.00 (Together AI). Specialized models carry their own rows: Llama 4 Scout SFT at $3, gpt-oss-120B at $5, DeepSeek-V4 Flash at $6, Kimi K2.6 at $15, GLM-5.2 at $40, each per 1M tokens. Billing counts dataset size times epochs plus any evaluation tokens, and uploads go up as JSONL or Parquet via CLI, SDK, or cURL (Together docs). Thirty-three models support LoRA and twelve support full fine-tuning (Together catalog).

    Fireworks: lowest data floors

    Fireworks prices by parameter count: LoRA SFT at $0.50, $3.00, $6.00, and $10.00 per 1M tokens across the up-to-16B, 16.1-80B, 80-300B, and over-300B tiers, with full fine-tuning at double (Fireworks). Its data floors are the lowest documented: SFT wants "hundreds of examples, or roughly 10M+ tokens," DPO wants "hundreds to thousands of pairs," and reinforcement fine-tuning is often satisfied by fewer than 100 prompts (Fireworks docs). Two lines matter for OpenAI migrants: "you keep the resulting weights," and "Serve fine-tuned models for the same price as base models" (Fireworks pricing).

    Tinker: cheapest list rate

    Tinker, from Thinking Machines Lab, prices per model: gpt-oss-20b at $0.396, Qwen3-8B at $0.44, gpt-oss-120b at $0.737, Qwen3.5-9B at $1.463, and Kimi K2.6 at $4.84 per 1M tokens, with checkpoint storage at $0.10 per GB monthly and an 80% discount on cached prefill tokens (Tinker docs). Training is LoRA and PEFT-based, adapters export to Hugging Face, and its serverless inference beta is still limited to the two Inkling models.

    Nebius Token Factory: the benchmark pick

    Nebius Token Factory advertises post-training workflows for open models, deploys fine-tuned checkpoints on its own endpoints, and ships a Data Lab that converts production logs into training data, but it publishes no public training prices (Nebius). Its roadmap says "Full post-training and distillation workflows will soon be available." Get a quote, or lean on the benchmark numbers below.

    Alibaba Model Studio: the budget option

    Alibaba Model Studio is the budget option if you'll write ChatML JSONL. It runs five training types: continual pre-training, full SFT, efficient_sft (its LoRA recommendation), dpo_full, and dpo_lora (Alibaba docs). SFT wants "Over 1,000 entries" of QA pairs, DPO wants "100+ sets," and continual pre-training wants 10 million+ tokens. Training rates match the pre-trained model's inference rate: Qwen3-14B at ¥0.03 per 1K tokens and Qwen2.5-72B at ¥0.15, which lands at ¥30 to ¥150 per 1M tokens (Alibaba billing). The tunable catalog is Qwen-heavy, and the Singapore region offers only qwen3-14b with efficient_sft.

    Where Can You Fine-Tune Claude?

    Exactly one place: AWS Bedrock. The live fine-tuning table in Bedrock's docs lists "Anthropic Claude 3 Haiku (anthropic.claude-3-haiku-20240307-v1:0:200k)" in us-west-2, and no other Claude model appears anywhere (AWS docs). Anthropic's direct API offers no fine-tuning. No ranking article we checked mentions this door.

    Bedrock also tunes Llama 3.1 8B and 70B, Llama 3.2 1B, 3B, 11B, and 90B, Llama 3.3 70B, the Amazon Nova family, and Titan models, nearly all pinned to single regions. Billing counts tokens processed times epochs, plus model storage at $1.95 per month (AWS docs). Visible on the pricing page: reinforcement fine-tuning for gpt-oss-20b and Qwen3 32B at $80.00 per training hour (AWS pricing). Per-token training rates for the Llama 3.x and Claude rows sit behind collapsed sections, so ask your AWS account team before budgeting.

    What Does Fine-Tuning Actually Cost, Measured?

    List prices hide throughput, and throughput is where these platforms differ most. Vintage Data, a third-party benchmark published March 30, 2026, ran roughly 30M-token function-calling jobs through Nebius Token Factory, Tinker, and Together AI (vintagedata). Measured costs: Tinker $0.36 to $0.52 per 1M tokens for LoRA, Together $0.48 to $5.00, Nebius $0.40 to $5.00. Speeds diverge hard. Tinker took 166 to 220 minutes at roughly 2,273 to 3,012 tokens per second. Together finished in 15 to 28 minutes, one run at 33,330 tokens per second. Nebius sat between at 76 minutes and 6,646 tokens per second, and the benchmark's verdict named it "the most practical platform for our iterative use case." These are third-party measurements, not vendor list prices; your job will vary.

    Why pay for training at all? The economics case for small tuned models came out of the 416-point fine-tuning debate on Hacker News in March 2026. Commenter arkmm: "typical categorical / data extraction use cases would have ~10x fewer errors at 100x lower inference cost" (HN). That's the whole argument: a tuned 8B model beating a frontier model on one narrow task, at a fraction of the per-token price.

    Dead Ends: Vendors With Nothing Left to Tune

    Every entry below was checked this week. Save the tab.

    • xAI (Grok): docs.x.ai/docs/fine-tuning returns a 404. The docs tree is inference only, grok-4.6 models, no tuning anywhere (xAI docs).
    • Groq: self-described "premier neocloud for fast inference." No training product (Groq).
    • OpenRouter: the docs path for fine-tuning 404s. It routes inference; it does not train.
    • DeepSeek API: no fine-tune endpoint documented, v4-flash and v4-pro chat only (DeepSeek docs). Tune DeepSeek models through Together, Fireworks, or Tinker instead.
    • NVIDIA hosted (build.nvidia.com): no hosted tuning; the /explore/train path 404s while the inference pages render fine.
    • Predibase: predibase.com/pricing redirects to Rubrik Agent Cloud, with no independent signup path (Predibase).
    • Replicate: fine-tunes FLUX image models, not LLMs (Replicate).
    • Lambda Labs: both plausible tuning service URLs 404. Nothing verifiable.
    • Z.ai / Zhipu: no English fine-tuning docs. Tune GLM through Together, Fireworks, or Tinker instead.

    When Should You Not Fine-Tune at All?

    Before migrating anywhere, ask whether you needed fine-tuning in the first place. That same 416-point thread split the field. antirez, Redis's creator, argued modern models few-shot so well that "strong prompts plus large context windows usually win." danielhanchen, Unsloth's maintainer, countered with production evidence: Cursor's online RL lifting approval rates by 28%, Vercel's AutoFix RFT, DoorDash extracting LoRAs, Perplexity's Sonar (HN). A third commenter hit ~98% accuracy converting receipt images to structured JSON from about 1,000 SFT examples on Mistral 7B.

    The cheaper alternative to test first: prompt caching. gamegoblin, on OpenAI's GPT-4o fine-tuning launch thread, made the case that a cached prompt full of examples "is a lot more developer-friendly and gets a lot of the same benefits as fine-tuning," because you can update it anytime and there's no async training job (HN). If caching hits your accuracy bar, you're done. If it doesn't, the platforms above are your move.

    Run-It-Yourself Routes: $3.95 an Hour, or Free

    Two routes skip the managed platforms. First, rent GPUs by the second or minute: Modal charges $0.001097 per second for an H100, roughly $3.95 an hour, with H200 at $0.001261 and B200 at $0.001736 per second (Modal). Baseten bills per minute, from a T4 at $0.01052 up to an A100-80GB at $0.06667, an H100 at $0.10833, and a B200 at $0.16633 (Baseten). Both run Unsloth or LLaMA-Factory fine and deploy the tuned model afterward.

    Second route, the free one: train on hardware you already own. A 7B model trains in about 6GB of memory with 4-bit QLoRA, which fits any 16GB Mac. We walk that route end to end in our guide to training an AI to write like you, corpus ladder and overfitting guards included, and the local LLM writing stack covers the $19.99 setup those runs plug into. For the own-versus-rent call, the local LLM vs API comparison frames the decision.

    Platform Fit by Situation

    Your situationBest fitWhyWatch out for
    OpenAI refugee, small corpusTogether AI or Fireworks$0.48/$0.50 per 1M entry rates, JSONL uploadsTogether's $4 job minimum
    Cheapest training, full controlTinker$0.396/1M on gpt-oss-20b, HF exportSlowest in the benchmark, 166-220 min
    Already on AWSBedrockClaude, Nova, Llama 3.3 in one consoleSingle-region models, $1.95/month storage
    Must tune ClaudeBedrock, Claude 3 HaikuThe only place anywhereus-west-2 only
    Iterating on production logsNebius Token FactoryData Lab converts logs to training dataNo public prices, quote required
    On a Mac, $0 budgetLocal QLoRA7B trains in ~6GB at 4-bitCorpus cleaning takes longer than training

    Related Tools & Further Reading

    • The how: our guide to how you train an AI to write like you covers the local LoRA route, corpus sizes, and the overfitting guards.
    • The gear: the local LLM writing stack assembles the $19.99 setup those fine-tunes plug into.
    • The decision: local LLM vs API frames what to run yourself and what to rent.
    • Who pays for AI writing: the best AI writing assistants roundup.
    • Tools that fit the workflow: the AI tools hub, and a word counter for sizing a training corpus before you upload it anywhere.

    Every platform and price above was live on September 5, 2026. Prices move, and deprecation pages move faster than anything else on a vendor's site; the January 6, 2027 cutoff arrived with months of notice, not years. Check the vendor page before you commit a corpus to any fine-tuning platform.

    Frequently Asked Questions

    LLM
    AI
    comparison
    Local LLM
    open-source
    A

    About Abhay Khant

    A passionate tech enthusiast and professional developer specializing in AI, automation, and modern web development. Sharing insights and guides to help others build better software faster.

    View full profile →

    Join the Newsletter

    Get articles like this delivered to your inbox every Thursday.

    What to read next

    Jan 1, 197013 min read

    Train an AI to Write Like You: Fine-Tune a Local LLM (2026)

    Fine-tune LLM on your own writing (Sept 2026): four routes compared, real corpus numbers, MLX on a 16GB Mac, $0 local QLoRA, and the failure cases to know.

    AAbhay Khant
    Sep 3, 202613 min read

    The Local LLM Writing Stack in 2026: $19.99, $0 a Month

    Ollama drafts, LanguageTool self-hosted checks grammar, Hemingway edits: a local LLM writing stack costs $19.99 once and $0/year. Hardware numbers inside.

    AAbhay Khant
    Sep 3, 202613 min read

    Best AI Writing Assistants 2026: Free & Compared

    We compared 7 AI writing assistants for 2026, free tiers included. Grammarly Free gives 100 AI prompts, QuillBot caps at 125 words, Hemingway $19.99.

    AAbhay Khant