02 / SCALE · DISTRIBUTED INFERENCE
A wider path.
Same model. Different infrastructure. llm-d routes each request across two inference replicas.
CLIENT → LLM-D → vLLM × 2 → GPU × 2
↗
Same model.
The request finds its way through llm-d to a healthy replica.
A SMALL EXPERIMENT
NO BENCHMARK SAMPLES YET| METRIC | DIRECT | SCALED |
|---|---|---|
| Requests / sec | — | — |
| P50 latency | — | — |
| P95 latency | — | — |
| Errors | — | — |