In late September 2026 I posted a series of LinkedIn updates describing how to design, measure and operate LLM‑driven systems. I covered latency metrics, cost accounting, memory taxonomy, vector‑DB fundamentals, tool‑calling approaches, and end‑to‑end data pipelines, while repeatedly promoting the Hands‑On AI Engineering Bootcamp (https://lnkd.in/dagWE5r3).
Time to first token (TTFT): how long the user is exposed to a blank screen, the number that defines perceived latency.
Inter-token latency (ITL): how smoothly tokens stream after the first one.
End-to-end latency at p50 / p95 / p99, dominated by output length, track it per use case rather than globally.
Cost per successful task, not cost per request, a cheap request that fails is a waste.
Cache hit rate - prompt caching is often the technique that reduces cost the most.
Context window utilization - the early warning for compaction and truncation issues.
Episodic - This type of memory contains past interactions and actions performed by the agent.
Semantic - Any external information that is available to the agent and any knowledge the agent should have about itself.
Procedural - This is systemic information like the structure of the System Prompt, available tools, guardrails etc.
Approximate Nearest Neighbour (ANN) search.
Random Projection, Product Quantization, Locality-sensitive Hashing.
Cosine Similarity, Euclidean Distance, Dot Product.
I observed that practitioners are moving away from Model Context Protocol (MCP) because the extra server/client layer adds failure points, and I now prefer native function calling where flexibility is needed.
Next Cohort kicking off October 19th!
Use code YEAREND20 for 20% off - offer expiring soon.