I posted a series of X threads in early September 2026 describing how to accelerate LLM inference and cut KV‑cache cost in production.
Speculative decoding calls the target model once per block of draft tokens instead of per token, yielding 2‑3× faster generation.
- Two‑model drafter+verifier
MiniCPM5‑2B is a dense 2 B‑parameter model by OpenBMB from China that's built for reasoning, coding, and tool use on resource‑constrained hardware.
It supports SGLang, vLLM, llama.cpp, Ollama, iOS, Android, and HarmonyOS.
KV cache memory grows with the number of layers, KV heads, retained tokens, dimensions, bytes per value, and concurrent requests.
Prefix reuse lets requests with an identical prefix share already‑computed KV blocks.
Redis returned the earlier response in 0.37 seconds with zero LLM input or output tokens.
Redis reports API cost savings of up to 90% and cache‑hit responses up to 15x faster.
- EAGLE lightweight module
- Medusa multi‑head verification tree
- LayerSkip early‑layer drafter