Cache hit rate – prompt caching is often the technique that reduces cost the most.
Cost per successful task, not cost per request, a cheap request that fails is a waste.
Context window utilization – the early warning for compaction and truncation issues.
I am a Lithuanian AI/ML professional based in Vilnius, founder and CEO of SwirlAI since January 2023 and former CPO at neptune.ai. My LinkedIn presence is centered around three fronts – education, implementation, and content – and reaches a broad audience of AI engineers and leaders.
- My newsletter and LinkedIn content are read by 200K+ engineers and leaders.
- I founded SwirlAI to cut through the noise in the AI space and deliver production‑ready AI education and implementation.
- I transitioned from neptune.ai to focus on scaling impact through education and consulting.
- I adopt cloud‑native and Kubernetes‑based architectures to improve scalability and automation.
- The SwirlAI newsletter reaches 40K+ AI/ML/Data Engineers and Leaders.
- My LinkedIn audience includes 180K+ AI/ML/Data Engineers and Leaders.