1. Count replicas
Enter replicas actively serving the same query workload.
2. Enter effective concurrency
Use the number of query slots each replica can sustain at the measured latency.
3. Add measured service time
Use end-to-end vector-search latency for the chosen index, top-k, filters, and recall target.
4. Set utilization and hours
Leave headroom below saturation and choose the number of hours represented in the volume estimate.
5. Review sustainable throughput
Use per-replica and total QPS to evaluate scaling options and planned traffic.