Series

Series

Article series

kaggle-vllm Engineering Notes

An evidence-first record of adapting upstream vLLM to constrained, managed GPU environments.

  1. Why I Built kaggle-vllm: Reproducible vLLM Inference on Dual NVIDIA T4 GPUs
  2. Tensor Parallelism on Dual T4 (planned)
  3. Measuring NCCL Across PCIe/PHB (planned)
  4. Concurrency Crossover: TP=1 vs TP=2 (planned)
  5. Moving the Runtime to AWS g4dn (planned)
  6. Ray Multi-Node vLLM (planned)
  7. Pipeline Parallelism Across Two Nodes (planned)

Future series

  • LLM Inference from First Principles — memory, scheduling, latency, and throughput.
  • GPU Topology and Distributed Inference — collectives, links, and placement.
  • From Kaggle T4 to Multi-Node Serving — controlled steps from notebook constraints to clustered serving.

Planned titles are a research roadmap, not claims that experiments have already run.