Separate private evidence from shared knowledge

In MailTrace Legal I separated private case evidence from broader legal knowledge. That matters because the system should never blur a user-owned source record with background reference material.

Retrieval needs more than embeddings

Embeddings and pgvector are useful, but good retrieval also depends on metadata, filters, source scope and deterministic signals. In investigation-style systems, who owns the document, which case it belongs to and when it occurred can matter as much as semantic similarity.

Keep answers traceable

A useful evidence assistant should return source-linked findings, not unsupported prose. Citations, event links and evidence identifiers make it possible for a user to verify what the model is saying.

  • Store stable source identifiers.
  • Preserve timestamps and participant metadata.
  • Return the evidence used for each conclusion.
  • Fail soft when retrieval confidence is weak.

Use evaluation to protect retrieval quality

RAG quality should be tested against known questions and expected evidence. Retrieval recall, citation correctness, unsupported claims, latency and cost all belong in the evaluation loop.

Best-fit business use cases

This pattern is useful for legal operations, insurance, compliance, customer complaints, procurement, due diligence, construction documentation and any environment where a model must show its working.