The GPU Performance Engineering Field Manual: Diagnose and Fix Slow, Costly AI Training Inference with CUDA, PyTorch, vLLM

Prijzen vanaf
69,99

Uitgelicht

VERGELIJK ALLE AANBIEDERS (3)

Beschrijving

Bol When a training job crawls, a GPU sits idle, or an inference bill keeps climbing, this is the manual to open. The GPU Performance Engineering Field Manual teaches a repeatable way to find out why AI workloads are slow or expensive, and how to prove that a fix worked. Every chapter starts from a symptom an engineer actually meets, explains the mechanism behind it, shows how to confirm the cause with measurements, and ends with the evidence that the problem is solved. Inside you will learn how to: - Measure before you tune, with nvidia-smi, DCGM, the PyTorch Profiler, Nsight Systems, and a benchmark harness whose numbers you can defend- Fix a GPU that waits for data, out-of-memory failures, unstable mixed precision, and host overhead- Use torch.compile, CUDA graphs, attention kernels, and custom Triton and CUDA kernels where they pay off- Scale training with DDP, FSDP, NCCL, and tensor, pipeline, and expert parallelism, and diagnose slow or hung collectives- Understand LLM inference: prefill and decode, KV cache arithmetic, TTFT and inter-token latency- Tune vLLM batching, prefix caching, and chunked prefill, and apply quantization and speculative decoding safely- Scale, route, and autoscale serving fleets, hunt tail latency, and build cost models that turn performance into money Built for ML engineers, platform and infrastructure engineers, and research engineers who train or serve models on NVIDIA GPUs with PyTorch. It assumes working Python and basic PyTorch, and no prior CUDA experience. Practical tools on every page: Triage Cards that map symptoms to first measurements and fixes Worked diagnoses with the arithmetic shown Runnable code listings 64 Triage Drills with a full answer key Fourteen end-to-end worked investigations Templates and checklists Capstone exercises A formula reference >Every slow training run and every costly inference fleet has a cause you can measure. The GPU Performance Engineering Field Manual shows you how to find it. You get Triage Cards that match symptoms to fixes, worked diagnoses with the arithmetic shown, 64 drills with a full answer key, and 14 end-to-end investigations. Order your copy today and fix the next slowdown with evidence, not guesswork.

Vergelijk aanbieders (3)

Sorteren op:

€ 69,99 Gratis verzending

€ 75,59 Gratis verzending

€ 75,59 Gratis verzending

Beschrijving (1)

When a training job crawls, a GPU sits idle, or an inference bill keeps climbing, this is the manual to open. The GPU Performance Engineering Field Manual teaches a repeatable way to find out why AI workloads are slow or expensive, and how to prove that a fix worked. Every chapter starts from a symptom an engineer actually meets, explains the mechanism behind it, shows how to confirm the cause with measurements, and ends with the evidence that the problem is solved. Inside you will learn how to: - Measure before you tune, with nvidia-smi, DCGM, the PyTorch Profiler, Nsight Systems, and a benchmark harness whose numbers you can defend- Fix a GPU that waits for data, out-of-memory failures, unstable mixed precision, and host overhead- Use torch.compile, CUDA graphs, attention kernels, and custom Triton and CUDA kernels where they pay off- Scale training with DDP, FSDP, NCCL, and tensor, pipeline, and expert parallelism, and diagnose slow or hung collectives- Understand LLM inference: prefill and decode, KV cache arithmetic, TTFT and inter-token latency- Tune vLLM batching, prefix caching, and chunked prefill, and apply quantization and speculative decoding safely- Scale, route, and autoscale serving fleets, hunt tail latency, and build cost models that turn performance into money Built for ML engineers, platform and infrastructure engineers, and research engineers who train or serve models on NVIDIA GPUs with PyTorch. It assumes working Python and basic PyTorch, and no prior CUDA experience. Practical tools on every page: Triage Cards that map symptoms to first measurements and fixes Worked diagnoses with the arithmetic shown Runnable code listings 64 Triage Drills with a full answer key Fourteen end-to-end worked investigations Templates and checklists Capstone exercises A formula reference >Every slow training run and every costly inference fleet has a cause you can measure. The GPU Performance Engineering Field Manual shows you how to find it. You get Triage Cards that match symptoms to fixes, worked diagnoses with the arithmetic shown, 64 drills with a full answer key, and 14 end-to-end investigations. Order your copy today and fix the next slowdown with evidence, not guesswork.


Productspecificaties

Merk Independently Published
EAN
  • 9798175092463
Maat


Prijshistorie

* Prijshistorie bevat geen data van Amazon, Amazon Marketplace.

Prijzen voor het laatst bijgewerkt op:

Uitgelichte Keuze
69,99
Naar shop