The batch-size row is the one I would watch. End-to-end evals average it away, so a throughput dip at larger batches never shows up as a failure.
Evals tell you whether the system got better. Microbenchmarks help you understand why.
p50/p95 latency of vector search
In early October 2026 I discussed microbenchmarks as a tool for AI system reliability, describing what they are, their scope, and how they complement end‑to‑end evaluations.
A microbenchmark measures one small, isolated piece of a system in a tight, repeatable loop.
It is not the whole application.
Evals tell you whether the system got better. Microbenchmarks help you understand why.