Foundation models know a lot about the world but nothing about your contracts, policies, products or customers. To make generative AI useful inside a business you have to ground it in your own information. The two main techniques — retrieval-augmented generation (RAG) and fine-tuning — are often presented as alternatives. In practice they solve different problems.
What is retrieval-augmented generation (RAG)?
RAG keeps your knowledge outside the model. When a user asks a question, the system searches your documents and data, retrieves the most relevant passages and passes them to the model along with the question. The model answers using that context — and can cite its sources.
RAG is the right default when:
- Information changes often (prices, policies, inventory, tickets).
- Answers must be traceable to a source document.
- Different users are allowed to see different information.
- You want to switch models later without retraining.
What is fine-tuning?
Fine-tuning further trains a model on examples so it learns a behaviour: a format, a tone, a classification scheme or a specialised task. It changes how the model responds rather than giving it new facts to look up.
Fine-tuning is worth considering when:
- You need consistent output structure or style that prompting alone cannot achieve.
- A smaller, cheaper model must match a larger model on a narrow, repetitive task.
- You have hundreds to thousands of high-quality labelled examples.
Fine-tuning is a poor way to teach a model facts that change: the knowledge goes stale, can’t be cited and can’t be permission-filtered.
A quick comparison
| RAG | Fine-tuning | |
|---|---|---|
| Best for | Up-to-date knowledge, citations | Behaviour, format, narrow tasks |
| Data needed | Your documents and records | Curated input/output examples |
| Freshness | Update the index, instantly current | Requires retraining |
| Access control | Filter retrieval per user | Not possible inside the model |
| Main risk | Poor retrieval → poor answers | Overfitting, stale knowledge |
Making RAG work in production
Most disappointing RAG systems fail at retrieval, not generation. The fundamentals that matter most:
- Content preparation — clean extraction from PDFs, tables and slides, sensible chunking and rich metadata.
- Hybrid search — combining keyword and vector search, then re-ranking results, usually beats either alone.
- Permissions — enforce the user’s existing access rights at query time.
- Evaluation — maintain a test set of real questions with expected answers and measure retrieval accuracy, answer correctness and citation quality on every change.
- Guardrails — instruct the model to say “I don’t know” when the context doesn’t contain the answer, and filter sensitive data.
When to combine them
Mature systems often use both: RAG supplies current, permissioned knowledge, while a fine-tuned or carefully prompted model handles the house style, structured output or a specialised classification step. Start with RAG and strong prompting; add fine-tuning only when evaluation shows a specific gap that more data or better retrieval won’t close.
Choosing a platform
All three major clouds offer managed building blocks — Amazon Bedrock, Azure OpenAI Service with Azure AI Search, and Google Vertex AI — alongside open-source options such as pgvector and OpenSearch. The right choice usually follows where your data and identity already live. Our Generative AI team can help you evaluate the options against your security and cost requirements.
Frequently asked questions
Is RAG better than fine-tuning?
Neither is universally better. RAG is best for grounding answers in current, citable, permissioned business knowledge; fine-tuning is best for teaching consistent behaviour or specialised tasks. Many production systems use both.
Does RAG stop hallucinations?
It greatly reduces them by giving the model relevant source material, but it doesn't eliminate them. Good retrieval, instructions to decline when unsure, citations and ongoing evaluation are all needed.
How much data do I need to fine-tune a model?
It depends on the task, but useful fine-tuning typically needs hundreds to thousands of high-quality, representative examples.