Research

Research

Research notes

This notebook is organized around questions that can be answered with inspectable code, controlled experiments, and immutable evidence.

Inference

  • vLLM engine behavior and scheduling
  • KV-cache capacity and precision
  • continuous batching
  • time to first token (TTFT)
  • time per output token (TPOT)
  • request and token throughput

GPU systems

  • CUDA runtime compatibility
  • NVIDIA Tesla T4 behavior
  • device memory and memory bandwidth
  • PCIe/PHB topology
  • NCCL collectives

Distributed inference

  • tensor parallelism
  • pipeline parallelism
  • Ray orchestration
  • multi-node execution

Production systems

  • Kubernetes GPU serving
  • Prometheus metrics
  • OpenTelemetry traces
  • reliability and failure analysis

Comparisons between accelerator architectures will appear only when the experiment and source evidence support them.