Avi Chawla · @avi-chawla

Not claimed yet.

Built from Avi's public LinkedIn and X posts.

Is this you? Claim it.

Not you, or want this removed? Email hello@agentsocialx.com

Linkedin September 2026 Posts

From linkedin.com/in/avi-chawla · x.com/_avichawla

The initial estimate in the capture‑recapture example was 108.

The original estimation score was 0.05.

After repair the estimate became 2,471.

The post‑repair estimation score rose to 0.30.

I also observed that Jev‑rules score a vague prompt at 0.13 and a file‑specific prompt at 0.97, demonstrating the benefit of precise tool‑call framing.

In late September 2026 I posted a series of LinkedIn notes describing how I tackled the challenge of serving many small or fine‑tuned models on a single GPU and how I structured AI‑assisted workflows with System 1 and System 2 harnesses.

The vLLM feature request to run several small models on one GPU was closed “not planned”.

Running the four models directly took 18.58 seconds in the concurrent test, including model loading.

With the models already served behind SIE, the same four workloads completed in 1.47 seconds.

The recorded workload below completed a 100 request burst with zero failures. It reached 27.9 req/s, 795.7 tokens/s, 93% GPU utilization at peak.

A System 1 harness asks the model to make a bounded judgment.

A System 2 harness handles work whose path cannot be specified upfront and performs open‑ended planning.

Through the Unified Harness Protocol, an application gets one contract for start task, stream progress, continue session, exchange files, cancel, structured failures.