Keep vectors close to product data
PostgreSQL plus pgvector is attractive because ownership, case IDs, timestamps and other structured filters can live alongside embeddings. That simplifies permission-aware retrieval.
Metadata often beats another model call
Filtering by user, workspace, case, participant or date can drastically improve relevance before semantic similarity is even considered.
Chunking is a product decision
The right retrieval unit depends on what users need to verify. For evidence workflows, preserving message or document boundaries may be more important than maximizing generic semantic recall.
Evaluate retrieval separately from generation
If the wrong evidence is retrieved, a better language model will not fix the product. Measure whether the expected source appears before evaluating answer quality.
Design for permissions from day one
Private data should never be mixed into a global retrieval pool without explicit scope and authorization. Retrieval boundaries are part of the security model.