Canopus Software & Engineering — home

RAG or fine-tuning: which does your AI need?

Fine-tuning changes how a model behaves. Retrieval changes what it knows. Teams reach for fine-tuning when they need retrieval more often than the reverse, and the mistake is expensive because it is slow to discover.

RAG or fine-tuning — a decision guide from the Canopus AI engineering team
The short answer

Use retrieval when the answer must come from your own content. Use fine-tuning when you need a consistent format, tone or classification decision rather than new facts. Try prompting first, because it costs nothing and is often enough. Fine-tuning does not reliably teach a model facts, which is what most teams want from it.

What each one actually changes

The distinction is not a spectrum. They solve different problems and are frequently used together.

Retrieval-augmented generation

Your documents are indexed. At question time the system finds the relevant passages and hands them to the model along with the question. The model reasons over content it has just been given rather than content it memorised. Because the source passage is known, the answer can carry a citation the reader can open, and the retrieval step can filter by the permissions the user already has.

Update a policy on Friday afternoon and the next answer reflects it. Nothing is retrained.

Fine-tuning

You take a base model and continue training it on examples of the behaviour you want. It learns shape extremely well — output format, house tone, a classification taxonomy specific to your business, adherence to a rigid response structure. It learns facts unreliably, and when it gets one wrong it does so with complete confidence and no citation.

Update a policy on Friday afternoon and nothing changes until you retrain, revalidate and redeploy.

The rule of thumb

If a correct answer changes when your documents change, you need retrieval. If a correct answer changes only when your style or taxonomy changes, fine-tuning is a candidate. Wanting the model to "know our business" is almost always a retrieval requirement wearing a fine-tuning costume.

Compared on what matters in production

CriterionRetrieval (RAG)Fine-tuning
What it changesWhat the model can seeHow the model behaves
Freshness of knowledgeImmediate — reindex and it is liveFrozen until the next training run
Citations back to sourceYes, by constructionNo — nothing to cite
Per-user permission filteringYes, applied at retrievalNot possible — knowledge is baked in
Consistent output formatAchievable with a schemaIts strongest use
Cost to updateReindex the changed documentsCurate data, retrain, revalidate
Inference latencyHigher — retrieval adds a stepLower, and often a smaller model
Explaining a wrong answerInspect what was retrievedLargely opaque

Scroll the table sideways to compare both columns.

The row the table cannot carry is the monthly bill, because neither option is a one-off build cost. Retrieval pays for embeddings, a vector store and the tokens of every passage it feeds the model; a fine-tuned model pays for hosting whether anyone queries it that day or not. Both behave like every other usage-priced cloud service, which means they drift the same way — where cloud spend actually leaks is the same discipline applied to inference.

The order we try things in

  1. Prompting and output schemas. Free, instant to change, and sufficient more often than anyone expects. Constrain the response shape and force an explicit "not found" path.
  2. Retrieval. Add grounding in your own content, with citations and permission filtering. This is where the majority of business AI work lands.
  3. Fine-tuning. Only once you have evidence that prompting cannot hold the format or the classification, and you have a labelled dataset worth training on.

Skipping to step three is common and costly, because the failure is delayed — the model looks impressive in a demo and drifts in production, with no citation to explain why.

The engineering that decides success either way

The model is one component. These are the parts that determine whether the system is trustworthy.

  • Chunking that respects meaning. Splitting a clause from the condition that qualifies it produces confidently wrong answers. Semantic chunking with overlap beats fixed-size splitting on every document set we have measured.
  • Hybrid search. Vector similarity plus keyword matching, then re-ranking. Pure vector search misses exact identifiers — part numbers, clause references, error codes — which is precisely what people search for.
  • Permission-aware retrieval. Filter before the model sees anything. A support agent and a director asking the same question should get different answers from the same index.
  • An evaluation set that runs in CI. Fifty to two hundred real questions with answers your experts agree on. A prompt edit that quietly breaks accuracy then fails the pipeline instead of a customer.
  • PII redaction before inference, and audit logging after it. Under the enterprise API terms we deploy on, API traffic is contractually excluded from training by default; where data cannot leave your infrastructure at all, open-weight models on your own instances are the route.

Most of what we repair in a disappointing assistant sits upstream of the model. A nightly sync that started dropping records without raising an error, or a connector that stopped after a schema change, leaves an index quietly out of date while every dashboard still reads green — the same failure modes set out in why integrations fail silently, except the symptom here is a confident wrong answer rather than a missing row.

Where retrieval is the wrong answer too

If the rule is deterministic, write the rule. A refund threshold, a shipping band or a tax calculation belongs in a function that is right every time and costs nothing per call — not in a system that reads the policy back and paraphrases it. Retrieval earns its place where the answer lives in prose that people keep editing: policies, contracts, manuals, past tickets. Where the answer lives in a table, query the table.

Proving it before you commit

AI is the one service line where we insist on a short paid evaluation before quoting a build. Two weeks: week one builds the test set of real questions with agreed answers, week two builds a deliberately ugly prototype against real data. What comes back is measured accuracy, cost per call, latency, and the gap to what your process actually requires.

Roughly one project in five stops there. That is the point of it — a fortnight is a much better way to discover an accuracy ceiling than a quarter. You keep the test set, the measurements and the written recommendation either way.

What we will not do is quote an accuracy percentage before measuring on your data. Any number offered before that measurement is a guess with a decimal point on it. For the same reason we will not take AI work on a fixed price before that evaluation exists — you cannot fix a price against an accuracy nobody has measured yet, which makes this the clearest instance of the argument in fixed price or time and materials. Once there is a measured number, a fixed price becomes a reasonable thing to ask for.

The full approach is on AI and intelligent systems. Data handling, redaction and what we explicitly do not certify are covered in security and compliance. Since an assistant is only as useful as the systems it can read from and write to, API and system integration is usually the other half of the work. More guides are on the insights index.

Questions

Accuracy, effort and what it costs to run

What accuracy can we realistically expect?

Enough to be useful on a narrow, well-defined question set; rarely enough to remove human review in the first release. The ceiling is set by your content rather than by the model — a question answered by one clear paragraph behaves nothing like one that needs three documents reconciled and an effective date checked, and the second kind is where these systems quietly fail. So the evaluation scores each category of question separately, and scores the thing demos never show: how often the system correctly says the answer is not in your documents. A system that answers everything is not accurate, only confident.

What would fine-tuning actually require from us?

Labelled examples of the exact output you want, and more of them than teams expect — typically hundreds for a narrow classification, low thousands where the output format is intricate, plus a held-out set nobody trained on. Those examples have to come from your side; we can structure, clean and version them, but nobody here can decide what your correct output looks like. Then it repeats: a new model version, a changed taxonomy or a new document type means curating, retraining and revalidating again. That standing cost is why fine-tuning sits third in the order above, and why most projects we scope never reach it — the format problem turns out to be a prompt and an output schema.

How do we decide it is good enough to ship?

You set the bar before anything is built, and it is never a single number. Ours is usually four: accuracy per category of question, a ceiling on answers that are wrong and confident, a floor on correct refusals, and a response-time limit. Someone on your side owns those four thresholds and signs them off, because "good enough" is a judgement about what a wrong answer costs your business, not an engineering preference. And the comparison that settles it is what your own team gets right today under time pressure — not perfection, which neither a person nor a model reaches.

Which part of the bill grows when our usage doubles?

Three lines carry the cost: inference tokens, the vector store, and re-embedding when documents change. Only the first tracks usage closely, and within it you pay for the retrieved passages far more than for the user's question — so how much context you feed the model moves the invoice harder than query volume does, followed by caching the questions that repeat. Most internal assistants we have measured land between $200 and $1,500 a month in model calls, billed to your own provider account rather than to us; where your system sits in that range comes out of the evaluation, measured on your documents. A fine-tuned model inverts the shape — you pay to host it by the hour whether anybody asks it anything.

Does our knowledge base need restructuring before this works?

Cleaning, usually, rather than restructuring. Retrieval does not need your taxonomy; it needs documents that are current and unambiguous about which version is authoritative. Three PDFs of the same policy from different years is the most common cause of confidently wrong answers we find, and no amount of chunking repairs it. Worth adding to each document: an owner, an effective date and an audience, because those become retrieval filters. Scanned PDFs and spreadsheets used as source of truth need converting first — we handle that during indexing, but deciding which document wins is your team's call, not ours.

What do we do when the model is confidently wrong in production?

You design for it before launch, or you hear about it from a customer. Every answer carries citations, every response has a feedback control, and the question, the retrieved chunks and the answer are logged together. That turns a wrong answer into a debugging job with an obvious first step: read what was retrieved. Most are retrieval failures rather than model failures — a stale index, a document nobody indexed, a clause split from the condition that qualifies it. Fix in that order, then add the question to the evaluation set so the same failure cannot return quietly.

Published: Last updated: Written by Canopus delivery team

Get a Free Quote

Bring us a question your team answers a hundred times a week.

That is usually the first AI project worth doing. Tell us what it is and where the correct answer lives today — an engineer replies within one business day, including if we think a model is the wrong tool.

Scope an AI project

One business day. We will say if a model is not the right tool.
Call WhatsApp Get a Quote