Talk to Avi about
- LLM inference optimization
- KV‑cache reduction
- Multi‑model serving
How Avi thinks
Fine‑grained observability
Values turning trajectory comparisons into scalar rewards and adding detailed monitoring to LLM training pipelines.
For you: you should instrument your models with detailed metrics to catch issues early.
Consolidated model serving
Believes that serving several small models on one GPU via a shared inference server dramatically reduces request
For you: you can cut inference time by consolidating models behind one server.
Speculative decoding
Considers speculative decoding an effective way to make LLM inference 2, 3× faster without changing output distribution.
For you: you can boost throughput by adding a draft model that proposes multiple tokens.
Ask this first
- What are the biggest challenges in reducing KV‑cache memory?
- How does your SIE system handle multiple small models on a single GPU?
- Can you share best practices for implementing speculative decoding?
Avi's brain
Recently updated
I continued my September 2026 X thread series with a deeper dive into KV‑cache management, token‑cost reduction, positional encoding, and…
I posted a series of X threads in early September 2026 describing how to accelerate LLM inference and cut KV‑cache cost in production.
I described how I use the RULER reviewer to turn trajectory comparisons into scalar rewards for GRPO training, and how I added fine‑grained…
I compared three ways to extract decisions from LLMs: plain decoding, JSON‑masked structured output, and Jev‑style scoring, and noted when…
I captured several Jev‑style scoring use‑cases in my September 2026 X posts, each illustrating how deterministic candidate generation can…
I outlined three calibration levels for the AnyJev library, each trading label effort for confidence and latency.
The initial estimate in the capture‑recapture example was 108.
When I send a prompt to an LLM via the API, the request first enters the serving engine rather than going straight to the GPU. The serving…
I noted that the **512d binary quantized embedding outperform 3096d float32 embeddings from OpenAI v3 large**.
In late September 2026 I wrote a LinkedIn note dissecting Mixture‑of‑Experts (MoE) routing, highlighting that each token picks two experts…
I am a data‑science professional based in New Delhi, India. I co‑founded the newsletter Daily Dose of Data Science, have held…