Research
Research notes
This notebook is organized around questions that can be answered with inspectable code, controlled experiments, and immutable evidence.
Inference
- vLLM engine behavior and scheduling
- KV-cache capacity and precision
- continuous batching
- time to first token (TTFT)
- time per output token (TPOT)
- request and token throughput
GPU systems
- CUDA runtime compatibility
- NVIDIA Tesla T4 behavior
- device memory and memory bandwidth
- PCIe/PHB topology
- NCCL collectives
Distributed inference
- tensor parallelism
- pipeline parallelism
- Ray orchestration
- multi-node execution
Production systems
- Kubernetes GPU serving
- Prometheus metrics
- OpenTelemetry traces
- reliability and failure analysis
Comparisons between accelerator architectures will appear only when the experiment and source evidence support them.