Partition by stable work units
Chunk generation and embedding calls should operate on deterministic units with stable identifiers. That makes it possible to retry failed batches without duplicating successful work.
Batch for throughput, not maximum size
Very large batches may look efficient but increase failure blast radius. Smaller bounded batches make retries cheaper and keep latency predictable.
Make jobs idempotent
A rerun should update or skip an already-processed vector rather than create a duplicate. Idempotency is what makes parallel execution safe.
Track cost and queue depth
Concurrency should be driven by rate limits, model cost and downstream database throughput. A good pipeline exposes queue depth, failed jobs, tokens processed and completion rate.