Speeding Up AI Quantization Techniques for Models and Vector DBs @Weaviate
Speeding Up AI Quantization Techniques for Models and Vector DBs  @Weaviate
Uploaded March 2025 | Updated September 2026, 56 minutes ago
In this talk, Marcin Antas (linkedin.com/in/antasmarcin/), a senior Core Engineer who's been at @Weaviate for over 4 years, breaks down the essential techniques for optimizing AI models through quantization.

Learn how to significantly reduce the memory footprint of large language models and embedding models while preserving their functionality - even on constrained devices like Raspberry Pi 5!

🔑 Key Topics Covered:
- LLM quantization techniques (from FP16/FP8 to 4-bit precision)
- The GGUF format and LLAMA.cpp framework
- Why feed-forward layer parameters are more sensitive than attention layers
- Embedding model quantization using ONNX
- Vector database quantization methods (Product, Binary, and Scalar)
- Running vector databases and AI models on edge devices

This technical deep dive is perfect for developers looking to optimize AI models for memory-constrained environments or deploy vector search capabilities on edge devices.

Learn more from Weaviate at https://weaviate.io.
Speeding Up AI Quantization Techniques for Models and Vector DBsREFRAG with Xiaoqiang Lin - Weaviate Podcast #130!Window Search Tree with Josh Engels - Weaviate Podcast #98!An AI-Native Approach to App Development8. AI Agents: Tools & APIsDiversity in Search ResultsVector EmbeddingsWeaviate TECH Hands-On: Transformation Agent
Weaviate vector database |

Speeding Up AI Quantization Techniques for Models and Vector DBs

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER