
Fine-tuning — the misconception that you can teach your company to AI
Contents
If you misunderstand fine-tuning as “making AI memorize your entire company manual,” you can spend a fair bit on training and still be left with a model that can’t reflect a single piece of fresh information.
Background / why this term now
In May 2026, OpenAI announced it would phase out its fine-tuning API and platform. According to the OpenAI developer docs, organizations that have never run a fine-tuning job can’t start new ones from May 7, and from July 2, organizations without inference activity on a fine-tuned model in the last 60 days are also blocked. On January 6, 2027, even existing users will no longer be able to create new jobs. OpenAI cited the fact that “modern base models follow instructions much better, and prompt-based approaches are cheaper and faster.”
Meanwhile, Google continues to support supervised fine-tuning for Gemini on Vertex AI, and Anthropic does not open fine-tuning via its own API but allows it only for Claude 3 Haiku via Amazon Bedrock. With vendors diverging like this, now is exactly the time to clarify what fine-tuning actually does.
Key data & current status
| 사업자 | 파인튜닝 지원 현황 (2026년 9월 기준) | 비고 |
|---|---|---|
| OpenAI | GPT-4.1·GPT-4.1 mini 등 기존 이용자만, 신규 불가 | 2026.05 단계적 종료 발표, 2027.01 완전 종료 |
| Vertex AI 에서 Gemini 지도 파인튜닝 계속 지원 | 텍스트·이미지·오디오·문서 데이터 지원 | |
| Anthropic | 자체 API 미지원, Bedrock 경유 Claude 3 Haiku 만 가능 | 세이프티·품질 고려로 지원 범위 제한 |
| 오픈웨이트 모델(Llama, Gemma 등) | LoRA·QLoRA 로 직접 파인튜닝 | GPU 1장·수십 분 단위로도 가능 |
Interpretation: Unlike the previous trend where “fine-tuning = a standard feature available everywhere,” in 2026 the scope of support varies by vendor. If anything, the options have broadened on the self-hosted open-weights side.
In-depth analysis
1) What exactly is fine-tuning?
Fine-tuning is the process of taking an already trained base model and further training it on a small set of example data (typically input–output pairs) to adjust part of the model’s internal weights (the numeric parameters the model holds). Unlike pretraining, which trains a model from scratch, it’s closer to layering behavior patterns onto a model that already knows language and common sense—for example, “answer in this format,” “use this tone.”
As of 2026, the most widely used methods are LoRA (Low-Rank Adaptation) and its variant QLoRA. According to a technical explainer, they freeze the original weights and only train two small low-rank matrices alongside each layer. This adjusts just 0.01–1% of the total parameters while achieving roughly 95–100% of the performance of full fine-tuning. For a 4096×4096 layer, a rank-16 setting reduces the number of trainable weights to about one hundredth.
2) A common misconception — why people read it that way
The expectation that “if we fine-tune, AI will know all our company materials” arises from two things: the impression given by the word “training,” and the failure to distinguish between what fine-tuning does well (learning format, tone, and repetitive workflow patterns) and what it does poorly (reflecting facts that change frequently).
A fine-tuned model does not know fresh facts that weren’t in its training data. If your product price changes next month, a fine-tuned model will answer with the old price until it is retrained. By contrast, retrieval-augmented generation (RAG — fetching documents relevant to the question in real time and injecting them into the model) reflects updates immediately as long as you refresh the documents. Much of what people mean by “teaching the company” is actually the job of RAG, not fine-tuning.
3) How to decide in practice
The industry’s usual order of operations is to try prompt engineering → RAG → fine-tuning. A related guide recommends considering fine-tuning only after you have at least 500+ well-curated training examples and you’ve verified a clear accuracy gap that prompt engineering and RAG cannot close. It’s time to consider fine-tuning when you can’t get the desired format via prompts alone, or when the problem isn’t fresh information lookup but enforcing repetitive behavior patterns like “always this tone,” “always answer in this order.”
You also need to consider cost structures. According to a cost guide, LoRA fine-tuning an open-weights 7B-class model can often be done for just the GPU cost on the order of hundreds of thousands of KRW, while full fine-tuning a 70B+ model can require computing costs in the tens of millions of KRW. RAG pipelines, by contrast, tend to have higher upfront build costs and steady ongoing operational costs.
4) When it’s useful and when it’s pointless
Fine-tuning is meaningful when: ① you must strictly lock down output format (specific JSON schemas, legal document templates, etc.), ② it’s cumbersome to describe recurring domain terms and tone in every prompt, ③ you want to compress a specific skill into a smaller model to reduce inference costs. Conversely, if your goal is to answer with ever-changing information—like the latest policies, prices, or inventory—fine-tuning is pointless; RAG or real-time API integration is the right approach.
In practice — four principles for reading fine-tuning correctly
- Think “fine-tuning = behavior/format learning,” not “fine-tuning = knowledge injection.” If you need fresh factual information, start with RAG.
- First verify that prompt engineering and then RAG don’t solve it; consider fine-tuning as the last step.
- Support varies by vendor. OpenAI is effectively closing new fine-tuning, while open-weights models with LoRA offer more flexibility for organizations able to self-host.
- If you’re considering fine-tuning, prepare at least hundreds of well-curated training examples and clear evidence for “why prompts/RAG aren’t enough.”
References
Contents
Related posts

API — What Does 'Plugging Someone Else's AI into Your Service' Actually Mean?
Connecting AI 'via API' does not mean bringing the model itself in-house; it means contacting someone else's server on every request to fetch an answer. This article lays out what to judge in practice, from token-based pricing and rate limits to API key security.

What is AI hallucination?
Hallucination is not the result of AI 'breaking'. It is closer to a structural side effect: today's training and evaluation methods award more points for a plausible guess than for saying 'I don't know'. Once you understand the mechanism, the way you deal with it changes too.

Training vs. Inference — When Does AI Actually Get Smarter?
It feels like AI gets smarter the more you talk to it, but it does not actually learn anything new mid-conversation. This article lays out how training and inference differ, and why test-time compute became the AI industry's new buzzword in 2026, with data.

Transformer — The Ancestor of Every AI Today
From GPT to Gemini to Claude, every major AI model in 2026 runs on the Transformer architecture Google published in 2017. This article covers how self-attention works and the line of variations from decoder-only models to MoE and Mamba hybrids.