Practical LLM Evaluation for Production Systems: Measure, monitor, and improve AI system reliability across training and inference (Paperback)
暫譯: 實用的 LLM 評估於生產系統:衡量、監控及提升 AI 系統在訓練與推論中的可靠性 (平裝本)
Mohanna, Ammar, Kar, Indrajit, Ralte, Zonunfeli
- 出版商: Packt Publishing
- 出版日期: 2026-06-30
- 售價: $1,890
- 貴賓價: 9.5 折 $1,795
- 語言: 英文
- 頁數: 488
- 裝訂: Quality Paper - also called trade paper
- ISBN: 1807423891
- ISBN-13: 9781807423896
-
相關分類:
Large language model
海外代購書籍(需單獨結帳)
相關主題
商品描述
Build reliable Build reliable AI evaluation frameworks that measure quality, safety, grounding, and production readiness across modern LLM and SLM applications
Free with your book: DRM-free PDF version + access to Packt's next-gen Reader*
Key Features:
- Design evaluation frameworks for LLMs, SLMs, multimodal, reasoning, and agentic AI systems
- Measure quality, safety, grounding, robustness, and production readiness with practical metrics
- Apply unified evaluation methods to text, multimodal, and agentic AI systems
Book Description:
Modern AI systems are expected to do far more than generate fluent text. They should be able to retrieve information, reason through complex problems, understand images and documents, call external tools, execute workflows, and support critical business decisions. Evaluating these systems requires methods that go beyond traditional NLP benchmarks.
Taking a product-first approach, this book presents evaluation as a continuous operational capability spanning training, inference, and end-to-end system operation. You'll learn how to connect evaluation metrics directly to deployment gates, rollback criteria, monitoring systems, and production reliability objectives.
Using practical examples and real-world workflows, you'll explore evaluation strategies for text LLMs, vision-language models, multimodal conversational systems, mixture-of-experts architectures, reasoning models, agentic systems, retrieval pipelines, Text2SQL and Text2Cypher systems, embedding models, OCR workflows, and guardrail SLMs. You'll also learn how to manage non-determinism, design repeatable test suites, validate tool execution, and measure long-horizon agent behavior in production.
By the end of the book, you'll be able to design robust evaluation systems that help teams deploy reliable, safe, and economically viable LLM-powered applications with confidence.
*Email sign-up and proof of purchase required
What You Will Learn:
- Design repeatable evaluation pipelines for LLM systems
- Assess inference quality, latency, and operational cost
- Evaluate multimodal, agentic, and reasoning AI systems
- Build regression gates and deployment evaluation workflows
- Detect hallucinations and grounding failures in VLMs
- Assess routing stability in mixture-of-experts models
- Evaluate Text2SQL, OCR, and retrieval-based systems
- Translate evaluation signals into production decisions
Who this book is for:
ML engineers, GenAI engineers, AI architects, data scientists, platform engineers, and engineering managers responsible for deploying LLM-powered systems in production will benefit from this book. Applied AI researchers and technical decision-makers looking to measure reliability, safety, and operational readiness across modern AI systems will also find it valuable. Readers should have a working understanding of machine learning, Python, and modern LLM concepts.
Table of Contents
- Foundations of LLM Evaluation: Core Concepts and Primitives
- Building Reliable Text-Only LLMs Through Training-Time Evaluation
- Controlling Text-Only LLM Behavior at Inference Time
- Grounding and Reliability in Vision Language Models During Training
- Evaluating Visual Grounding and Reliability at Inference Time
- Evaluating Multimodal Conversational LLMs Across Training and Inference
- Evaluating Routing and Reliability in Mixture of Experts LLMs
- Evaluating Reliability and Control in Computer-Using Agent Systems
- Evaluating Information Extraction and Document-Understanding LLMs
- Evaluating Reasoning LLMs in Depth
- Evaluating Specialized LLM Systems
商品描述(中文翻譯)
建構可靠的 AI 評估框架,以衡量現代 LLM 和 SLM 應用程式的質量、安全性、基礎性和生產就緒性
隨書附贈:無 DRM 的 PDF 版本 + Packt 的下一代閱讀器訪問權限*
主要特點:
- 設計 LLM、SLM、多模態、推理和代理 AI 系統的評估框架
- 使用實用的指標來衡量質量、安全性、基礎性、穩健性和生產就緒性
- 將統一的評估方法應用於文本、多模態和代理 AI 系統
書籍描述:
現代 AI 系統的期望不僅僅是生成流暢的文本。它們應能檢索信息、推理複雜問題、理解圖像和文件、調用外部工具、執行工作流程,並支持關鍵的商業決策。評估這些系統需要超越傳統 NLP 基準的方法。
本書採取以產品為先的方式,將評估視為一項持續的操作能力,涵蓋訓練、推理和端到端系統操作。您將學習如何將評估指標直接連接到部署閘、回滾標準、監控系統和生產可靠性目標。
通過實用的範例和真實的工作流程,您將探索文本 LLM、視覺語言模型、多模態對話系統、專家混合架構、推理模型、代理系統、檢索管道、Text2SQL 和 Text2Cypher 系統、嵌入模型、OCR 工作流程以及防護 SLM 的評估策略。您還將學習如何管理非確定性、設計可重複的測試套件、驗證工具執行,並在生產中衡量長期代理行為。
在本書結束時,您將能夠設計穩健的評估系統,幫助團隊自信地部署可靠、安全且經濟可行的 LLM 驅動應用程式。
*需要電子郵件註冊和購買證明
您將學到的內容:
- 為 LLM 系統設計可重複的評估管道
- 評估推理質量、延遲和操作成本
- 評估多模態、代理和推理 AI 系統
- 建立回歸閘和部署評估工作流程
- 檢測 VLM 中的幻覺和基礎失敗
- 評估專家混合模型中的路由穩定性
- 評估 Text2SQL、OCR 和基於檢索的系統
- 將評估信號轉化為生產決策
本書適合對象:
負責在生產中部署 LLM 驅動系統的 ML 工程師、GenAI 工程師、AI 架構師、數據科學家、平台工程師和工程經理將從本書中受益。尋求衡量現代 AI 系統的可靠性、安全性和操作就緒性的應用 AI 研究人員和技術決策者也會覺得本書有價值。讀者應具備機器學習、Python 和現代 LLM 概念的基本理解。
目錄:
- LLM 評估的基礎:核心概念和原始元素
- 通過訓練時評估構建可靠的僅文本 LLM
- 在推理時控制僅文本 LLM 的行為
- 在訓練期間的視覺語言模型中的基礎性和可靠性
- 在推理時評估視覺基礎和可靠性
- 在訓練和推理中評估多模態對話 LLM
- 在專家混合 LLM 中評估路由和可靠性
- 在計算機使用代理系統中評估可靠性和控制
- 評估信息提取和文檔理解 LLM
- 深入評估推理 LLM
- 評估專門的 LLM 系統