The GPU Performance Engineering Field Manual

Diagnose and Fix Slow, Costly AI Training and Inference with CUDA, PyTorch, and vLLM
438 Seiten, Taschenbuch
€ 97,80
-
+
Lieferung in 7-14 Werktagen

Bitte haben Sie einen Moment Geduld, wir legen Ihr Produkt in den Warenkorb.

Mehr Informationen
Reihe Independently published
ISBN 9798175092463
Sprache Englisch
Erscheinungsdatum 17.09.2026
Größe 279 x 216 mm
Verlag Independently published
LieferzeitLieferung in 7-14 Werktagen
HerstellerangabenAnzeigen
Libri GmbH
Europaallee 1 | D-36244 Bad Hersfeld
gpsr@libri.de
Unsere Prinzipien
  • ✔ kostenlose Lieferung innerhalb Österreichs ab € 35,–
  • ✔ über 1,5 Mio. Bücher, DVDs & CDs im Angebot
  • ✔ alle FALTER-Produkte und Abos, nur hier!
  • ✔ keine Weitergabe personenbezogener Daten an Dritte
  • ✔ als 100% österreichisches Unternehmen liefern wir innerhalb Österreichs mit der Österreichischen Post
Kurzbeschreibung des Verlags

When a training job crawls, a GPU sits idle, or an inference bill keeps climbing, this is the manual to open.

The GPU Performance Engineering Field Manual teaches a repeatable way to find out why AI workloads are slow or expensive, and how to prove that a fix worked. Every chapter starts from a symptom an engineer actually meets, explains the mechanism behind it, shows how to confirm the cause with measurements, and ends with the evidence that the problem is solved.

Inside you will learn how to:

    - Measure before you tune, with nvidia-smi, DCGM, the PyTorch Profiler, Nsight Systems, and a benchmark harness whose numbers you can defend- Fix a GPU that waits for data, out-of-memory failures, unstable mixed precision, and host overhead- Use torch.compile, CUDA graphs, attention kernels, and custom Triton and CUDA kernels where they pay off- Scale training with DDP, FSDP, NCCL, and tensor, pipeline, and expert parallelism, and diagnose slow or hung collectives- Understand LLM inference: prefill and decode, KV cache arithmetic, TTFT and inter-token latency- Tune vLLM batching, prefix caching, and chunked prefill, and apply quantization and speculative decoding safely- Scale, route, and autoscale serving fleets, hunt tail latency, and build cost models that turn performance into money
Built for ML engineers, platform and infrastructure engineers, and research engineers who train or serve models on NVIDIA GPUs with PyTorch. It assumes working Python and basic PyTorch, and no prior CUDA experience.

Practical tools on every page: Triage Cards that map symptoms to first measurements and fixesWorked diagnoses with the arithmetic shownRunnable code listings64 Triage Drills with a full answer keyFourteen end-to-end worked investigationsTemplates and checklistsCapstone exercisesA formula reference>Every slow training run and every costly inference fleet has a cause you can measure. The GPU Performance Engineering Field Manual shows you how to find it. You get Triage Cards that match symptoms to fixes, worked diagnoses with the arithmetic shown, 64 drills with a full answer key, and 14 end-to-end investigations.

Order your copy today and fix the next slowdown with evidence, not guesswork.

Mehr Informationen
Reihe Independently published
ISBN 9798175092463
Sprache Englisch
Erscheinungsdatum 17.09.2026
Größe 279 x 216 mm
Verlag Independently published
LieferzeitLieferung in 7-14 Werktagen
HerstellerangabenAnzeigen
Libri GmbH
Europaallee 1 | D-36244 Bad Hersfeld
gpsr@libri.de
Unsere Prinzipien
  • ✔ kostenlose Lieferung innerhalb Österreichs ab € 35,–
  • ✔ über 1,5 Mio. Bücher, DVDs & CDs im Angebot
  • ✔ alle FALTER-Produkte und Abos, nur hier!
  • ✔ keine Weitergabe personenbezogener Daten an Dritte
  • ✔ als 100% österreichisches Unternehmen liefern wir innerhalb Österreichs mit der Österreichischen Post