Fine-tune the retrieval component first

The retrieval layer determines the evidence the generative model receives. If retrieval returns irrelevant or incomplete context, even a strong language model can produce a poor answer.

For specialist domains, general-purpose embeddings can be improved using labelled query-context pairs and contrastive or triplet-loss training so related queries and evidence are positioned closer together in vector space.

  • Create labelled query and relevant-context pairs.
  • Fine-tune embeddings on domain-specific examples.
  • Measure Precision@K and recall on a validation set.
  • Compare the fine-tuned model with the original baseline.

Use domain-specific embeddings where they add value

In biomedical decision-support experiments, domain-adapted models such as BioBERT or ClinicalBERT can provide a stronger representation of specialist terminology than completely general-purpose models. The important point is to validate retrieval performance rather than assume domain adaptation is automatically better.

Optimise FAISS for the size of the knowledge base

FAISS stores and searches vector embeddings for similarity-based retrieval. IndexFlatL2 performs exact nearest-neighbour search and can suit smaller datasets where exhaustive search is acceptable. IndexIVFFlat partitions embeddings into clusters to reduce search cost on larger collections.

  • nlist controls the number of partitions in an IVF index.
  • nprobe controls how many partitions are searched per query.
  • Increasing nprobe can improve recall but usually increases latency.
  • Evaluate index settings against both retrieval quality and response-time requirements.

Fine-tune the generative component

The generative model has a different job: synthesise the retrieved evidence into a useful response without departing from the source material. Supervised fine-tuning can use domain-specific examples containing user questions, retrieved context and expert-reviewed responses.

Instruction tuning can also teach the model to perform tasks such as explaining evidence, summarising a study, comparing guidelines or explicitly stating when the available evidence is insufficient.

Prompt for grounded answers

Prompt design should make the relationship between query, evidence and answer explicit. A useful pattern is to provide the question, retrieved evidence, an instruction to answer only from that evidence, and a requirement to state when the evidence is insufficient. In high-stakes applications, groundedness matters more than fluency.

Filter and rerank context before generation

A RAG system can retrieve more candidates than it ultimately sends to the LLM. For example, retrieve a top-10 candidate set, then reduce it to the strongest three to five passages using similarity scores, metadata, domain-term matching or a reranker. This can reduce cost and prevent loosely related passages from distracting the model.

Combine dense and keyword retrieval

Dense vector retrieval is useful for semantic similarity, but exact terminology, identifiers, medication names, legal clauses or technical codes may be better served by lexical search. Hybrid retrieval combines semantic search with keyword matching and metadata filters so the system can recover both conceptual and exact matches.

Optimise inference performance

Production performance is not only an accuracy problem. Model quantisation such as FP16 or INT8 can reduce memory use and inference time, while response-length controls such as max_new_tokens can prevent unnecessarily long generations and reduce latency and cost.

Evaluate retrieval and generation separately

A RAG response can fail because the correct evidence was never retrieved, or because the model used correctly retrieved evidence badly. Those are different failure modes and should be measured separately.

  • Retrieval: Precision@K, recall and latency.
  • Generation: grounding, relevance, accuracy and fluency.
  • Operations: token usage, failure rate and end-to-end latency.
  • Safety: unsupported claims and behaviour when evidence is insufficient.

Use human evaluation in consequential workflows

For decision-support systems, domain experts should evaluate whether retrieved evidence is relevant, whether the response accurately represents that evidence, whether the answer is actionable, and whether important claims can be traced back to sources. The AI should support professional judgement rather than replace it.

A/B test architectural changes

Compare versions of the same system using the same evaluation set. For example, generic embeddings plus vector retrieval can be compared with domain-specific embeddings plus hybrid retrieval and reranking. Measure whether the added complexity actually improves retrieval, grounding, latency, cost and expert ratings.

Build a continuous feedback loop

RAG performance changes as the knowledge base, models, prompts and user behaviour change. Production feedback should become evaluation data: user query → retrieval → generation → review → evaluation set → system improvement → regression testing.

RAG optimisation is a systems problem

A production RAG pipeline contains multiple interacting stages: query processing, embeddings, vector or hybrid retrieval, metadata filtering, reranking, context construction, generation, validation, citations, human review and evaluation. Reliability depends on the whole chain, not only the language model.