vLLM Serving

High¿Throughput LLM APIs with PagedAttention and KV Cache Tuning
240 Seiten, Taschenbuch
€ 50,20
-
+
Lieferung in 7-14 Werktagen

Bitte haben Sie einen Moment Geduld, wir legen Ihr Produkt in den Warenkorb.

Mehr Informationen
Themen Informatik und Informationstechnologie Computerprogrammierung und Softwareentwicklung Algorithmen und Datenstrukturen
ISBN 9798896653417
Sprache Englisch
Erscheinungsdatum 24.06.2026
Größe 229 x 152 mm
Verlag NobleTrex Press
LieferzeitLieferung in 7-14 Werktagen
HerstellerangabenAnzeigen
Libri GmbH
Europaallee 1 | D-36244 Bad Hersfeld
gpsr@libri.de
Unsere Prinzipien
  • ✔ kostenlose Lieferung innerhalb Österreichs ab € 35,–
  • ✔ über 1,5 Mio. Bücher, DVDs & CDs im Angebot
  • ✔ alle FALTER-Produkte und Abos, nur hier!
  • ✔ keine Weitergabe personenbezogener Daten an Dritte
  • ✔ als 100% österreichisches Unternehmen liefern wir innerhalb Österreichs mit der Österreichischen Post
Kurzbeschreibung des Verlags

"vLLM Serving: High¿Throughput LLM APIs with PagedAttention and KV Cache Tuning"Built for experienced ML systems engineers, platform architects, and performance-minded practitioners, this book is a deep technical guide to serving large language models with vLLM at production scale. Rather than treating inference as a black box, it explains the real control surfaces behind throughput, latency, and memory efficiency. Readers who already know LLM fundamentals but want to reason rigorously about serving behavior will find an internals-first, systems-oriented treatment.At the core of the book are the mechanisms that make vLLM distinctive: PagedAttention, continuous batching, KV cache design, and scheduler-driven execution. You will learn how request flow, cache allocation, sequence length, prefix reuse, quantized KV storage, and offloading strategies interact to determine concurrency limits and user-visible performance. The book also covers OpenAI-compatible API serving, streaming semantics, realistic benchmarking, and disciplined troubleshooting, so readers can move from conceptual understanding to evidence-based tuning and operational decisions.The emphasis throughout is on advanced mental models, trade-offs, and production diagnostics rather than introductory walkthroughs. This is a focused guide for readers comfortable with GPU inference, transformer decoding, and performance measurement who want a precise framework for designing, tuning, and operating high-throughput LLM APIs with confidence.

Mehr Informationen
Themen Informatik und Informationstechnologie Computerprogrammierung und Softwareentwicklung Algorithmen und Datenstrukturen
ISBN 9798896653417
Sprache Englisch
Erscheinungsdatum 24.06.2026
Größe 229 x 152 mm
Verlag NobleTrex Press
LieferzeitLieferung in 7-14 Werktagen
HerstellerangabenAnzeigen
Libri GmbH
Europaallee 1 | D-36244 Bad Hersfeld
gpsr@libri.de
Unsere Prinzipien
  • ✔ kostenlose Lieferung innerhalb Österreichs ab € 35,–
  • ✔ über 1,5 Mio. Bücher, DVDs & CDs im Angebot
  • ✔ alle FALTER-Produkte und Abos, nur hier!
  • ✔ keine Weitergabe personenbezogener Daten an Dritte
  • ✔ als 100% österreichisches Unternehmen liefern wir innerhalb Österreichs mit der Österreichischen Post