Sachin Kr. Rajput
15 years shipping Team of 8 · 6 products AWS → LLM inference

Engineering leader building LLM inference systems.

Every number I publish came from code that was executed.

Currently

Deep in LLM inference performance — benchmark harness and quantization studies on an L40S, published as I go.

15
Years building
8
Engineers led
6
Products in flight
1×L40S
Bench under test

Approach

Inference systems, and the distributed platforms under them.

Fifteen years of building systems that carry load. Today I lead a team of eight across six products, and my own work sits on inference — serving, benchmarking, quantization economics — on top of a long background in distributed AWS platforms.

serving · throughput · cost per token · AWS

Verification gates, not vibes.

Every code block is executed before it ships, and every published number traces back to printed output I can point at. If a claim cannot be reproduced from a run, it does not go out.

executed → captured → traced → published

A sequenced curriculum, one article at a time.

I work through machine learning in order, in public. The writing is the study log: what I ran, what it measured, and what it cost.

one article at a time · published as I go

Focus areas

Inference serving

Getting models to answer under real load: batching behaviour, memory headroom, and the shape of the latency curve as concurrency climbs.

Inference lab

Precision
Model size8B
Concurrency8 requests
Throughput
p50 latency
Weights in VRAM
Cost / 1M tokens

Illustrative model, not measured output. It exists to show the shape of the trade-off — quantize, and weights shrink while throughput rises; raise concurrency, and throughput and latency climb together. Real numbers appear only in the articles, next to the runs that produced them.

Writing

Published as I go.

I publish the inference work one article at a time on dev.to — harness design, benchmark results, and the cost arithmetic behind them.