Partition by stable work units

Chunk generation and embedding calls should operate on deterministic units with stable identifiers. That makes it possible to retry failed batches without duplicating successful work.

Batch for throughput, not maximum size

Very large batches may look efficient but increase failure blast radius. Smaller bounded batches make retries cheaper and keep latency predictable.

Make jobs idempotent

A rerun should update or skip an already-processed vector rather than create a duplicate. Idempotency is what makes parallel execution safe.

Track cost and queue depth

Concurrency should be driven by rate limits, model cost and downstream database throughput. A good pipeline exposes queue depth, failed jobs, tokens processed and completion rate.