AI Featured

RAG vs fine-tuning: when to use which

Teams reach for fine-tuning far more often than the problem requires, usually because one question was never asked out loud: does the model lack knowledge, or does it lack behaviour? Retrieval fixes the first. Fine-tuning fixes the second. Choosing wrong costs months.

The question behind the question

Take one failing example and decide which category it falls into:

Advertisement
  • Knowledge. The model does not know a fact, or knows an old version of it: a price list, an internal policy, a support ticket, a document published last week.
  • Behaviour. The model knows enough but produces the wrong shape: verbose where you need JSON, wrong tone, missing tool calls, weak classification on your taxonomy.

Most "it is wrong about our product" complaints are knowledge problems. Most "we cannot ship this output" complaints are behaviour problems.

What fine-tuning actually changes

Fine-tuning continues training on your examples and moves the weights. It is supervised learning over pairs you write and curate — a few hundred to a few thousand for a narrow task — usually through LoRA or QLoRA adapters that train a small fraction of the parameters instead of the whole model. Vendor documentation is explicit that it teaches form rather than facts.

What improves reliably: output format, tone, domain vocabulary, tool-call syntax, classification boundaries, refusal style. What does not: storing facts. A fact baked into weights cannot be cited, cannot be updated without retraining, cannot be filtered per user, and returns confidently when it goes stale.

What retrieval actually changes

Retrieval — RAG, in the paper that named it — leaves the weights alone and changes what the model sees at answer time. Every stage can fail on its own:

  • Chunking. Split on structure — headings, sections, table rows — not fixed character counts. Around 200 to 800 tokens per chunk with slight overlap suits most prose. Carry metadata with each chunk: source URL, document version, tenant and access rules.
  • Embeddings. The model used at query time must be the one used at index time. Swap embedding models and the whole corpus has to be re-embedded.
  • Search. Dense vectors miss exact identifiers such as error codes and part numbers; keyword search misses paraphrase. Run both and merge the results.
  • Reranking. A cross-encoder over the top 50 to 100 candidates, keeping the best 3 to 8, usually improves answers more than any change to the generator.
  • Grounding. Pass numbered sources, require citations, and return "I don't know" when retrieval comes back empty instead of letting the model answer from its weights.
def answer(question, tenant_id, pool=50, keep=5):
    q = embed(question, model=EMBED_MODEL)          # same model as indexing
    hits = index.search(q, k=pool, filter={"tenant_id": tenant_id})
    if not hits:
        return "I don't know based on the available documents."
    docs = rerank(question, hits)[:keep]            # cross-encoder, not cosine
    context = "\n\n".join(f"[{i+1}] {d.text}" for i, d in enumerate(docs))
    reply = llm("Answer using only the sources below and cite them as [n].\n\n"
                + context + "\n\nQ: " + question)
    return reply, [d.source_url for d in docs]

Why fine-tuning is a poor substitute for retrieval

The failure is predictable. Fine-tune on a corpus of support tickets and the model will discuss your product in a plausible voice — including the feature you deprecated last month, because that knowledge is frozen at training time. From there:

  • Changing one policy means retraining, re-validating and redeploying.
  • Deleting one record means the same, and there is no way to demonstrate the fact is gone from the weights.
  • Per-user permissions cannot be enforced: everyone gets whatever the weights contain.
  • Nothing can be cited, so reviewers cannot check an answer.
  • Out-of-scope questions degrade instead of abstaining, because the model has been pushed toward your distribution.

Retrieval turns each of those into a cheap operation: reindex one document, delete one chunk, pre-filter by tenant.

Cost, maintenance and privacy

DimensionFine-tuningRetrieval
Time to first useful resultWeeks, mostly spent curating examplesDays, mostly spent on ingestion and evals
Main cost driverLabelled data, training runs, a dedicated endpointIndexing and storage, plus extra tokens and latency per answer
FreshnessFrozen at the training runAs fresh as the last index job
Deleting one factRetrain, with no proof of removalDelete the chunk
Per-user accessNot possibleFilter inside the query
Changing base modelRe-tune the adapterUsually nothing to redo

Privacy deserves its own sentence. Retrieval still sends retrieved chunks to whichever model answers, so a hosted API sees your documents; embedding and reranking locally keeps the index in your infrastructure. Fine-tuning on sensitive data is harder to undo: weights can memorise and reproduce training text, vendors may retain training data, and individual records cannot be revoked. Do not fine-tune on anything you would refuse to leak.

The hybrid that usually wins

  • Retrieve first for anything that must be current, citable or permission-aware.
  • Fine-tune a LoRA for house style: the exact JSON envelope, the mandated greeting, the approved refusal wording, your label set. A few hundred good examples are enough.
  • Fine-tune the retriever instead of the generator when your vocabulary is unusual — statute citations, SKUs, error codes. An embedding model trained on your query and document pairs often beats a larger generator.
  • Distil rather than train from scratch: label examples with a strong model, fine-tune a small one for the narrow task, then serve it cheaply.

Evaluating, then deciding

The two halves need different scoreboards, and "it looked good in the demo" is not one.

  • Retrieval: collect 100 to 300 real questions with the document that should answer each. Track recall@k, MRR or nDCG, and how often unanswerable questions are correctly refused.
  • Generation: faithfulness — does every claim appear in the cited chunk — plus citation accuracy and task success. Judge models are useful for triage, not for certification.
  • Fine-tuning: a held-out set scored on exact match, schema validity and human ratings, alongside a general-capability regression check, because narrow tuning can quietly damage unrelated behaviour.

Then run the checklist in order:

  1. Try a better prompt and a handful of examples first. Formatting complaints often dissolve there.
  2. If answers depend on documents that change, or must be cited or access-filtered, use retrieval.
  3. If retrieval returns nothing useful, fix ingestion, chunking and reranking. Fine-tuning cannot retrieve what was never indexed.
  4. If the model knows the facts but will not hold the format, tone or labels, fine-tune.
  5. If both are true, do both and keep them separate so each can be scored on its own.
  6. If the task is small and stable, a prompted small model with a strict schema may beat both. Measure before you train.
Advertisement
khallaf

Writing about programming, AI and the tools that make engineering teams faster. Published by A1 Systems.

Last updated 19 Sep 2026

// Keep reading

Related articles

AI 5 min read

Give the Model a Role, Not a Wish

Better prompts come from context, constraints and an output shape — not from adding “please”. Here is how to test prompts, cut cost, and catch confident invention before your users do.

khallaf
Tools & Tricks 5 min read

Regex you will actually use

The small set of regex constructs that cover everyday work, the patterns worth keeping in a snippet file, and how to avoid catastrophic backtracking.

khallaf Tip