Use retrieval to supply relevant knowledge. Use fine-tuning to adapt behaviour. Evaluate whether either is necessary before adding complexity.
Start with a baseline you can measure.
A model giving a poor answer is a symptom, not a diagnosis. It may lack the right information, misunderstand the instructions, use the wrong tool, or struggle with the task itself. Those problems call for different interventions.
Collect representative examples, decide what a good output looks like, and test a suitable model with clear instructions. Keep some examples separate for evaluation. This establishes whether the gap is knowledge, behaviour, or capability, and gives you a fair comparison for later changes.
Choose RAG when the answer depends on your knowledge.
Retrieval-augmented generation supplies relevant information to a model at request time. A retrieval pipeline finds useful passages from documents or records, and the model uses that context to form an answer. This is a natural fit for policies, product information, procedures, and other facts that change.
The retrieval system is as important as the model. Parsing, chunking, search quality, reranking, and document freshness all affect the result. Access controls must apply before information enters the prompt; a model should never be the only barrier protecting a restricted document.
- Start with authoritative, maintained sources.
- Check retrieval relevance separately from answer quality.
- Use citations so users can inspect the evidence.
- Define what happens when the source material does not contain an answer.
Choose fine-tuning when consistent behaviour is the gap.
Fine-tuning adapts an existing model using task-specific examples. It can help with specialised classification, consistent output formats, domain language, or repeated behaviours that remain difficult to achieve through prompting alone.
It needs representative data, clear usage rights, and a meaningful held-out evaluation. Training success does not establish that the model generalises. Check difficult cases, uncommon inputs, and important failure modes, as well as the average score.
Fine-tuning is not a reliable way to store a changing company knowledge base. It does not provide document-level access control, guaranteed recall, or automatic citations. Those requirements still need an application and data architecture.
Retrieval and fine-tuning can work together.
A retrieval system can supply current knowledge while an adapted model follows a specialised response format or workflow. But combining techniques also adds more things to evaluate and maintain.
Introduce one change at a time where possible. If retrieval improves factual grounding but the output format is still inconsistent, first test structured output and clearer instructions. Fine-tuning becomes useful when it produces a measurable improvement that justifies its data and operating costs.
A practical starting point for your business.
Begin with one valuable task, a small representative evaluation set, and a documented baseline. For an internal knowledge assistant, start by measuring whether the right documents are retrieved and whether the answer is supported by them. For a specialised classification task, measure performance on unseen examples.
The decision should follow the evidence: quality, privacy, latency, cost, and maintenance effort. A simpler system that meets the requirement is easier to operate and improve.
- Define the task and success criteria.
- Test prompting and retrieval before committing to training.
- Evaluate on examples the system has not been tuned against.
- Include data updates and operating costs in the decision.