AI-Ready Data Yohan Kim

RAG vs Fine-Tuning: Choose the Next Test Before the Method

Engineer comparing a contract clause with an incomplete response during a RAG vs fine-tuning evaluation.

A contract assistant misses a renewal clause. Another response identifies the clause correctly but leaves out a field the downstream application requires. Before the team invests in better retrieval or fine-tuning, it needs to understand why each response failed.

The RAG vs fine-tuning decision begins with that distinction. Retrieval-augmented generation supplies selected external information at inference time. Fine-tuning updates a model using training examples. Direct context supplies the relevant material in the request without a separate retrieval step. These approaches can work together, and none is a substitute for identifying what failed.

Start by testing whether the model can complete the task with the necessary evidence and clear instructions. That gives the team a basis for choosing its next experiment.

Separate missing evidence from an unreliable response

A knowledge gap means the model lacks information needed for the task, such as a contract clause, a current product specification, or an account’s terms. A behavior problem concerns how it uses the information it receives, including whether it extracts the required fields and follows the requested format.

These categories overlap. A model may have the right passage but misinterpret an exception. It may return valid JSON containing an unsupported answer. Calling either failure a knowledge problem or a behavior problem is a starting hypothesis, not a diagnosis.

For a failed extraction, compare the actual input with the expected output. Check whether the relevant clause and its definitions were included, and whether the instructions specified the required fields and how to handle missing information. Those checks help distinguish a training problem from one the application can address in the request.

Use retrieval when the task needs facts from an external source

RAG is worth testing when requests draw on a changing collection of documents and each request needs only part of that collection. The application selects material for the model rather than expecting the model to recall it from its weights.

In Ovadia and colleagues’ comparison, RAG outperformed unsupervised fine-tuning on the knowledge-intensive tasks they evaluated. Fine-tuning provided some improvement, but incorporating new facts through that training setup remained difficult.

Soudani and colleagues found a similar pattern in question answering about less-popular entities: fine-tuning improved performance, but retrieval performed better, particularly for the least-popular facts. Both studies support testing retrieval for specialized knowledge. Neither establishes that fine-tuning cannot teach facts or that retrieval will work equally well on a different corpus.

For a contract assistant, the relevant test is whether retrieval supplies the applicable clause and enough context to interpret it. An answer that includes a citation still needs checking: the cited passage may be irrelevant, incomplete, or inconsistent with the answer. Measure evidence selection and answer correctness separately.

Test direct context before adding retrieval machinery

Some tasks arrive with their evidence already attached. If a user asks about one contract, supplying that contract directly may be a sensible baseline. A separate search index could add work without resolving the observed failure.

The choice depends on how the application uses the material. Does each request need a small passage from a large collection, or substantial parts of one document? Can the selected model use the supplied text reliably? Does the request meet the team’s latency and cost limits?

Li and colleagues’ evaluation of long context and RAG found that results varied with the question and retrieval approach. Long context performed better overall on their question-answering benchmarks, especially Wikipedia-based questions, while RAG had advantages on dialogue-based and general queries. Summarization-based retrieval was competitive with long context. Those results concern the tested setups; they do not select an architecture for a different workload.

Compare both approaches on the same representative cases where feasible. Include questions that need a nearby definition, an exception in another section, or evidence spread across the document. A context window large enough to hold the text does not demonstrate that the model will use all of it correctly.

Consider fine-tuning for a persistent, measurable behavior gap

Consider fine-tuning when the model has the information it needs but still fails the task after the team has tested clear instructions and examples. Training requires examples of the desired behavior, along with separate evaluation cases to check whether the improvement extends beyond the training data.

Google Cloud’s supervised fine-tuning guidance describes using labeled examples to adapt model behavior, including task-specific output formats. A recurring extraction or formatting error may therefore be worth testing with supervised training.

The application still needs to validate outputs and enforce permissions. Check responses against the required schema before passing them to downstream software, and authorize actions independently of what the model says. Training does not guarantee a valid or authorized response.

Run a small experiment before committing to an architecture

Consider a hypothetical assistant that extracts renewal dates and notice periods from contracts. Reviewers have found two problems: missed exceptions and responses that omit required fields.

Begin with a reviewed set of cases and the existing application as the baseline. For each missed exception, give the model the relevant clause and definitions, selected and checked by a reviewer. Keep the model and task instructions unchanged, and repeat the comparison across the selected cases.

Same model and task instructions
Baseline input
Reviewed clause + definitions
Compare extraction
ImprovesTest evidence selection
Still failsCheck instructions and task ambiguity
Hypothetical test. Neither outcome alone proves the cause.

If extraction improves, investigate how the application selects and supplies evidence. A reviewer-selected passage can reveal an opportunity, but it does not show that an automated retriever will deliver the same improvement. Test whether retrieval or direct context can supply equally useful material under realistic conditions; other failures may remain.

If the assistant still fails with the reviewed evidence, examine the instructions, ambiguity in the task, and the model’s ability to perform the extraction. Test clearer instructions and representative examples. Fine-tuning is one possible next experiment if a consistent failure remains and suitable training data is available.

Evaluate omitted fields separately from incorrect values. A change that produces perfectly structured but incorrect answers has not solved the extraction task. Include contracts with no renewal clause so the assistant must handle missing information rather than invent a date.

If that comparison showed an improvement, a hypothetical decision record could read:

The current assistant misses renewal exceptions. Supplying reviewed clause context improves extraction on the selected cases, so the next experiment will test evidence selection. We will assess clause coverage, field accuracy, unsupported answers, latency, and cost. Format failures will be tracked separately before deciding whether training is justified.

This example is not a reported result. A real decision record should state what was observed, on which cases, and what remains untested.

Include the cost of keeping the system useful

Compare operating costs as well as the initial experiment. Retrieval requires maintaining the corpus and search process. Direct context consumes input capacity and may benefit from caching where supported. Fine-tuning requires preparing examples, evaluating the trained model, and deciding when to train again. Measure the actual setup rather than assuming one method is always cheaper.

The methods may also solve different parts of the same task. An assistant could retrieve current clauses and use a fine-tuned model to extract them into a consistent structure. Each addition should address a demonstrated gap and pass a comparison against the simpler baseline.

Record what would reopen the decision. A growing document collection, a changed output requirement, or a new model may justify another test. The original choice remains useful only while its assumptions still describe the work.

Validate the data change separately from the model change

If the investigation points to the input data, test that change on its own where practical. Correcting document structure, preserving definitions, or selecting the applicable version should be evaluated under comparable model and task conditions. Changing the data and model together may improve the system, but makes it harder to attribute the improvement.

Syntitan, CUBIG’s AI-Ready Data Platform, addresses this data-fit question. Its published approach compares original and refined data for a specified use case under fixed AI conditions. That is a different decision from choosing a retrieval architecture or training a model.

For the contract assistant, that means comparing the original document input with a version that preserves the relevant clauses and definitions, using the same extraction task. The result can show whether that data change is worth keeping, even if formatting or model errors still need separate work.

When missing or incomplete source material is part of the failure, test whether improving it helps the task. See how Syntitan compares original and refined data under the same AI conditions.

Explore Syntitan, CUBIG’s AI-Ready Data Platform

FAQ

Can RAG and fine-tuning be used together?

Yes. Retrieval can supply relevant external material while a fine-tuned model performs a task using that material. Evaluate the combination against a simpler baseline so each component has a demonstrated purpose.

What if the source documents do not contain the answer?

Neither retrieval nor a larger context window can supply a fact missing from the available sources. Define how the application should report insufficient evidence, request clarification, or route the case for review. Do not count a plausible unsupported answer as a successful completion.

Does choosing RAG settle data-access permissions?

No. Retrieval is a method for selecting information, not authorization to use it. Apply the organization's access and usage controls to the sources and the model input, regardless of the architecture selected.