Fri Frakt över 299 kr
Fri Frakt över 299 kr
Kundservice
Vision Language Models

Vision Language Models

978 kr

978 kr

Få kvar

Fre, 18 sep - tis, 22 sep

Hemleverans


Säker betalning

14-dagars öppet köp


Säljs och levereras av


Produktbeskrivning

Vision language models (VLMs) combine computer vision and natural language processing to create powerful systems that can interpret, generate, and respond in multimodal contexts. Vision Language Models is a hands-on guide to building real-world VLMs using the most up-to-date stack of machine learning tools from Hugging Face, Meta (PyTorch), NVIDIA (Cuda), and others, written by leading researchers and practitioners Merve Noyan, Miquel Farré, Andrés Marafioti, and Orr Zohar. From image captioning and document understanding to advanced zero-shot inference and retrieval-augmented generation (RAG), this book covers the full VLM application and development lifecycle. Designed for ML engineers, data scientists, and developers, this guide distills cutting-edge VLM research into practical techniques. Readers will learn how to prepare datasets, select the right architectures, fine-tune and deploy models, and apply them to real-world tasks across a range of industries. Explore core model architectures and alignment techniquesTrain and fine-tune VLMs with Hugging Face, PyTorch, and othersDeploy models for applications like image search and captioningImplement advanced inference strategies, from zero-shot to agentic systemsBuild scalable VLM systems ready for production use

Artikel.nr.

2e337774-c0b2-5429-bb4a-d49c7a815d58

Bokdetaljer

Format

Pocket

Språk

English

Antal sidor

300

Utg.datum

2026-06-23

Förlag

O'Reilly Media

Författare

Andres Marafioti and Andres Marafioti,

Dimensioner

178 × 234 × 24 mm

Vikt

702 g

ISBN

9798341624047

Ursprungsland

GB

Vision Language Models

978 kr

978 kr

Få kvar

Fre, 18 sep - tis, 22 sep

Hemleverans


Säker betalning

14-dagars öppet köp


Säljs och levereras av