<feed xmlns="http://www.w3.org/2005/Atom"> <id>https://waqasm86.github.io/</id><title>Mohammad Waqas — AI Inference &amp; GPU Systems</title><subtitle>Technical research notes, reproducible benchmarks and engineering articles on vLLM, CUDA, NVIDIA GPUs, distributed inference, Kubernetes and AI infrastructure by Mohammad Waqas.</subtitle> <updated>2026-09-25T09:58:30+05:00</updated> <author> <name>Mohammad Waqas</name> <uri>https://waqasm86.github.io/</uri> </author><link rel="self" type="application/atom+xml" href="https://waqasm86.github.io/feed.xml"/><link rel="alternate" type="text/html" hreflang="en" href="https://waqasm86.github.io/"/> <generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator> <rights> © 2026 Mohammad Waqas </rights> <icon>/assets/img/favicons/favicon.ico</icon> <logo>/assets/img/favicons/favicon-96x96.png</logo> <entry><title>Why I Built kaggle-vllm: Reproducible vLLM Inference on Dual NVIDIA T4 GPUs</title><link href="https://waqasm86.github.io/posts/why-i-built-kaggle-vllm/" rel="alternate" type="text/html" title="Why I Built kaggle-vllm: Reproducible vLLM Inference on Dual NVIDIA T4 GPUs" /><published>2026-09-24T09:00:00+05:00</published> <updated>2026-09-24T09:00:00+05:00</updated> <id>https://waqasm86.github.io/posts/why-i-built-kaggle-vllm/</id> <content type="text/html" src="https://waqasm86.github.io/posts/why-i-built-kaggle-vllm/" /> <author> <name>Mohammad Waqas</name> </author> <category term="LLM Inference" /> <category term="kaggle-vllm" /> <summary>An evidence-first introduction to the compatibility, provenance, and experiment design behind running upstream vLLM in a managed dual-T4 notebook environment.</summary> </entry> </feed>
