shreypatel

ai + data engineer · coconut labs

I measure before I claim.
Then I publish what failed.

I build data systems and inference infrastructure, and I run a two-person research lab on the side. The through-line is not a stack, it is a habit: pre-commit the threshold, rent the hardware, run the bench, and keep the number that survived next to the caveat that limits it. Half my best results are disconfirms, and they are published like everything else.

quiet tenant vs fifo

26× tail recovery

shipped this year

6 pkgs + releases

live demos

7 in browser

disconfirms published

3 and counting

how i think

Four habits, each with a receipt

Falsify before shipping

The first kvwarden thesis was an admission cap that would smooth the latency cliff. I rented an H100, ran 16,000 requests per arm, and the cap measured 1.04x, no effect. At one operating point it made things worse. The disconfirm is in the repo next to the result that replaced it.

gate 1.5 · 16k req/arm · published as-is

Caveats travel with numbers

The hero number is a 26x tail recovery for a quiet tenant under flood. The same sentence always carries the regime: one A100 at saturation, vLLM 0.19.1, n=311. On an H100 at the same load the starvation never happens, and the page says so.

53.9 / 61.5 / 1,585 ms · A100 · n=311

Names must be true

kvwarden's cache scaffold turned out to be a ledger nothing read. Three greps proved it, so I rewrote the roadmap in public, closed my own RFC as superseded, and put "the name says KV warden but you are not touching the KV cache" in the FAQ myself.

rfc closed as superseded · faq entry shipped

Stop digging when the hole is the tool

A benchmark provider burned 512 pod spins over four days without once handing me a working machine. I stopped, wrote up the $47 of evidence that the allocator was the problem, and moved the experiment rather than the budget.

512 pods · 0 usable · finding published

timeline · 2026

Showcase, in the order it happened

Earlier work lives on the GitHub profile: an RDMA NVMe offload stack, a C++20 latency research workspace, a latent diffusion model trained from scratch, and an archived data platform kept public because the review that stopped it is worth more than pretending it lives.

working demos

Run them in your browser

demowhat it provesstatus
Silent data-regression guardrailCatches schema-quiet corruptions a type check cannot see, on a profile fitted from clean episodes.● live
Point-in-time correctnessThe one bug where offline accuracy goes up. Accuracy rewards it, schema checks miss it, this flags it.● live
Silent cache-miss guardrailFeature-store staleness that never throws, made visible.● live
Columnar scan-bytes guardrailQueries that quietly stop pruning, caught by their bytes.● live
Ingestion data-contract guardrailContract checks at the door instead of in the warehouse.● live
Agentic MLOps atlasA full lifecycle map with 25 interactive stations.● live
The proof pageEvery published number with its hardware, n, and raw artifact.● live

Every demo runs client-side and links its source on GitHub. Recorded walkthroughs are planned; until they exist this page will not pretend otherwise.

learning · academic

Study one level below the job

The practice behind the shipped work is academic on purpose: a private systems atlas that goes from the business problem down to the metal, courses run against real prototypes, and a fifteen-minute rep every day. The public edition of that practice is the masterclass.

Systems That Don't Lie

A five-chapter masterclass on honest systems: tenant fairness, single-machine inference, nanosecond signals, agents from the ground up, and observability that cannot flatter itself.

masterclass.coconutlabs.org

The Library

The private academic wing: the waterline atlas and the study shelf. What earns publication is rewritten clean and lands as a research note.

what it is · door for the two of us

notes

Written like lab records, not posts

noteone line
Tenant fairness on shared inferenceThe measurement that started the lab's public life.
A model in the roomThree jobs a model has in a creative tool, and the fence that keeps it useful.
What mixing taught me about evalsStudio monitors and eval suites fail the same way.

contact

Full-time roles, contract work, or an argument about inference fairness

All three are welcome.

shreypatel@coconutlabs.org · patelshrey77@gmail.com · github/ShreyPatel4 · github/coconut-labs