About

About

Mohammad Waqas

I am a self-taught software engineer with a long-term programming background and a current focus on AI infrastructure: LLM inference, GPU systems, reproducible benchmarking, and distributed serving.

This site is an engineering notebook rather than a conventional résumé. Articles connect implementation decisions to source code, experiment environments, raw evidence, failures, and limitations. The goal is to make infrastructure conclusions auditable and reproducible.

Current technical interests include vLLM, CUDA, NVIDIA GPUs, NCCL, tensor and pipeline parallelism, Ray, Kubernetes, Prometheus, Grafana, and OpenTelemetry.