← all posts

Three ways to give a model your knowledge

A model out of the box knows a lot about the world and nothing about yours. Your product docs, your policies, your writing, last week’s numbers — none of it is in there.

There are exactly three ways to fix that, and teams argue about them constantly: retrieval-augmented generation (RAG), fine-tuning, and long context. Most “RAG vs fine-tuning” comparisons leave the third one out entirely.

In one breath: Reach for RAG when the knowledge is private or keeps changing and you want citations. Reach for fine-tuning when you need the model to behave differently — a voice, a format, a domain’s reasoning. Reach for long context when the whole corpus is small, fixed, and you want to move fast. Most production systems end up combining them.

Let’s earn that summary.

What each one actually changes

The confusion comes from treating these as three flavours of the same thing. They’re not — each changes a different part of the system.

RAG changes what the model knows right now

At query time you find the few passages relevant to the question and paste only those into the prompt. The model’s weights never move; you’ve just handed it the right page of the open book. Add a document and it’s usable instantly; every answer can point back to its source. The Writing Bed post pulls this apart from first principles — that’s RAG end to end.

Fine-tuning changes how the model behaves

You train on your own examples and the weights actually shift. This is the right tool when you need a behaviour the base model doesn’t have — a house style, a strict output format, the reasoning patterns of a specialised domain.

What it’s not good at is being a fact store. Knowledge baked into weights goes stale the day you publish something new. And it can’t tell you which document an answer came from.

Long context changes what fits in one call

Modern models take enormous prompts — hundreds of thousands to a million-plus tokens. So skip the machinery and paste everything in. No infrastructure, no index, nothing to maintain.

There are three catches. You pay for those tokens on every single request. Latency climbs with prompt size. And models genuinely lose track of things buried in the middle of a giant dump.

Pause & recall. Before reading on: which of the three would let you answer “what changed in our returns policy this morning” — and which one physically can’t, no matter how well you built it? Say why in one sentence.

Reveal the answer

RAG and long context can, because the fresh text is in front of the model at query time. Fine-tuning can’t — this morning’s change isn’t in the weights until you retrain.

Side by side

RAGFine-tuningLong context
ChangesWhat it knows nowHow it behavesWhat fits per call
Best forFresh / private factsVoice, format, domain reasoningSmall, fixed corpora
Knowledge freshnessInstant (add a file)Stale until retrainedInstant (paste it)
Can cite sourcesYesNoYes, if you ask
Setup effortMedium (index + retrieval)High (data + training)Low
Per-query costLow (a few chunks)Low (no retrieval)High (whole corpus, every call)
LatencyLowLowGrows with prompt
MaintenanceRe-index on changeRetrain on changeNone

Read down the freshness and behaviour rows and the split falls out on its own. No column wins everything. That’s why the next section is a framework rather than a verdict.

How to choose: walk these questions in order

Don’t start from the technique. Start from what your problem actually needs, in this order:

Choosing between RAG, fine-tuning and long contextStart by asking whether the answer depends on fresh or private knowledge. If it does and the corpus is large or changing, use RAG; if the corpus is small and fixed, long context is enough. Separately, if you need a specific behaviour, style, format or domain reasoning, fine-tune. If you need both fresh knowledge and special behaviour, combine fine-tuning with RAG. In all cases, measure on your own data before committing.YesYesNo — small & fixedNoYesNoYes

Does the answer depend on fresh

or private knowledge?

Is the corpus large

or always changing?

Use RAG

Long context is enough

Do you need a specific behaviour,

style, format or domain reasoning?

Fine-tune

The base model already does this

Also need special

behaviour?

Combine: fine-tune + RAG

  1. Does the answer depend on fresh or private knowledge? If yes, the knowledge must be in front of the model at answer time — that’s RAG or long context, never fine-tuning alone.
  2. Is that knowledge large or always changing? Large or changing → RAG, because you index once and updates cost little. Small and fixed → long context, because an index isn’t worth building for a page of facts.
  3. Do you need the model to behave differently — a tone, a strict format, a specialist’s reasoning? That’s a weights problem → fine-tuning. It’s a separate axis from knowledge, which is why it can stack on top of RAG.
  4. Need both fresh facts and special behaviour? Then it’s a hybrid — more on that below.
  5. Whatever you lean toward, measure it on your own data first. The lowest-cost option that passes your eval wins. “It felt better” is not an eval — here’s how to build one.

Notice steps 1–2 and step 3 are asking about different things — knowledge versus behaviour. Half of all bad architecture decisions come from mixing those two up.

  • RAG
  • Fine-tuning
  • Long context
  • Answering from a returns policy that changed this morning
  • Making every reply land in your house style
  • A two-page company handbook that never changes
  • Teaching a model to reason like a tax specialist
  • Searching ten years of engineering runbooks
  • Seeing whether the model can do the task at all

The part the vendor blogs skip: cost over time

“RAG costs less” is true on day one and not the whole story. What matters is cost as your data and traffic grow.

Interactive: twelve months of cost for all three approaches, with dials for traffic and corpus size — long context starts free and never stops climbing. Needs JavaScript.

Drag the traffic up and watch which line wins. The shape is the point: building an index is a cost that stops, and paying for the whole corpus on every call is a cost that doesn’t.

UpfrontPer queryAs data growsMaintenance cadence
RAGBuild the indexLow — a few chunksFlat — retrieval stays smallRe-index on change (low cost)
Fine-tuningExpensive (data + training)Lowest — no retrievalFlat at inferenceRetrain to update (expensive, periodic)
Long context~NoneExpensive — whole corpus every callGrows linearly, foreverNone

The trap is long context. It’s free to build, so it looks like the low-cost option — but you re-pay for the entire corpus on every request, and that bill scales with both your data size and your traffic. It’s a brilliant prototype and an expensive production default. Fine-tuning is the mirror image: painful upfront, then the lowest cost per call, as long as your knowledge doesn’t need to be current.

Scale the tool to the problem, not to the tutorial. A page of facts doesn’t need a vector database; a corpus that grows every week shouldn’t live in a prompt you pay for on every call.

When the answer is “combine them”

The real world rarely fits one column. A support assistant needs your brand voice and format (fine-tuning) and today’s product facts with citations (RAG). That pairing is common enough to have a name — RAFT, retrieval-augmented fine-tuning: fine-tune the model for the behaviour, then wrap it in retrieval for the fresh knowledge.

The rule of thumb: reach for the hybrid only once a single approach has demonstrably failed your eval — not on day one. Every layer you add is another thing to maintain.

Where this choice usually goes wrong

Same handful of mistakes, again and again:

  • Fine-tuning to “add knowledge.” The most expensive way to get a stale, un-citable fact store. Nine times out of ten they wanted RAG.
  • Building RAG for a page of facts. If the whole corpus fits comfortably in a prompt and never changes, an index is ceremony. Just paste it.
  • Treating long context as free because it “just works” in the demo. It works right up until the token bill and the latency arrive together.
  • Choosing on vibes. No eval set, no numbers — just a hunch that one “seems smarter.” Build twenty real questions with known answers before you pick.
  • Treating the decision as permanent. It isn’t. Prototype on long context, graduate to RAG when the corpus grows, add fine-tuning only when behaviour demands it. Revisit as you scale.
  • Fine-tune so the model learns this quarter's pricing
  • Prototype on long context before building an index
  • Build a vector database for a page of facts
  • Fine-tune for voice, then wrap it in RAG for the facts
  • Treat long context as free because it works in the demo
  • Write the eval before choosing the approach

A worked example: choosing for the Writing Bed

The Writing Bed answers questions from the posts on this site. Three options, one decision:

  • Long context? No. The writing grows with every new post, so every request would carry every post, and that bill only ever climbs. Fine for a weekend prototype; wrong as a default.
  • Fine-tuning? No. It would mean retraining on every new post, the model still couldn’t cite which post an answer came from, and it would confidently paraphrase things nobody wrote.
  • RAG? Yes. It updates the instant a file is added, it answers with sources so you can check it hasn’t drifted, and it scales the same at ten posts or ten thousand.

The corpus is small enough that the vector store is just an in-memory list — no Pinecone, no pgvector. That’s the same principle running the other way: RAG was right, but the heavy version of RAG would have been its own over-engineering. Match the tool to the size of the problem in front of you.

Check yourself.

  1. Which axis does fine-tuning change — what the model knows, or how it behaves?
  2. Why is long context free to build but expensive to run at scale?
  3. A teammate wants to fine-tune to teach the model this quarter’s pricing. What do you suggest instead, and why?
  4. When is reaching for the RAG-plus-fine-tuning hybrid premature?

If any answer won’t come, that’s your map of what to reread — the tables above hold all four.

Where to start

If you take one thing from this: prototype with long context, measure, then earn every layer of complexity. Paste it all in to see if the model can do the task at all. If knowledge is the gap and it’s growing, move to RAG. If behaviour is the gap, fine-tune. Add a hybrid only when a single approach has failed a real eval — never before.

The mechanics of all three are well documented and quick to pick up. Telling which gap you actually have — knowledge, behaviour, or neither — is the part that takes judgement.

Want to see the RAG end of this built from nothing? That’s the Writing Bed. Want to watch a model decide and act rather than just answer? Meet Scout — then go poke all three in the Greenhouse. 🌱

Say hello

Let's grow something together.

Consulting, teaching, speaking, or a product idea that needs an AI brain — my inbox is open.

us — Usama Shahid © 2026 Usama Shahid — reachusama.com 🌱