Series
Article series
kaggle-vllm Engineering Notes
An evidence-first record of adapting upstream vLLM to constrained, managed GPU environments.
- Why I Built kaggle-vllm: Reproducible vLLM Inference on Dual NVIDIA T4 GPUs
- Tensor Parallelism on Dual T4 (planned)
- Measuring NCCL Across PCIe/PHB (planned)
- Concurrency Crossover: TP=1 vs TP=2 (planned)
- Moving the Runtime to AWS g4dn (planned)
- Ray Multi-Node vLLM (planned)
- Pipeline Parallelism Across Two Nodes (planned)
Future series
- LLM Inference from First Principles — memory, scheduling, latency, and throughput.
- GPU Topology and Distributed Inference — collectives, links, and placement.
- From Kaggle T4 to Multi-Node Serving — controlled steps from notebook constraints to clustered serving.
Planned titles are a research roadmap, not claims that experiments have already run.