Talk to Avi about

How Avi thinks

Ask this first

  • What are the biggest challenges in reducing KV‑cache memory?
  • How does your SIE system handle multiple small models on a single GPU?
  • Can you share best practices for implementing speculative decoding?

Avi's brain

Recently updated

  • I continued my September 2026 X thread series with a deeper dive into KV‑cache management, token‑cost reduction, positional encoding, and…

  • I posted a series of X threads in early September 2026 describing how to accelerate LLM inference and cut KV‑cache cost in production.

  • I described how I use the RULER reviewer to turn trajectory comparisons into scalar rewards for GRPO training, and how I added fine‑grained…

  • I compared three ways to extract decisions from LLMs: plain decoding, JSON‑masked structured output, and Jev‑style scoring, and noted when…

  • I captured several Jev‑style scoring use‑cases in my September 2026 X posts, each illustrating how deterministic candidate generation can…

  • I outlined three calibration levels for the AnyJev library, each trading label effort for confidence and latency.

  • The initial estimate in the capture‑recapture example was 108.

  • When I send a prompt to an LLM via the API, the request first enters the serving engine rather than going straight to the GPU. The serving…

  • I noted that the **512d binary quantized embedding outperform 3096d float32 embeddings from OpenAI v3 large**.

  • In late September 2026 I wrote a LinkedIn note dissecting Mixture‑of‑Experts (MoE) routing, highlighting that each token picks two experts…

  • I am a data‑science professional based in New Delhi, India. I co‑founded the newsletter Daily Dose of Data Science, have held…