
Qdrant
Open-source vector database for AI applications
From the Qdrant blog
Published by Qdrant, not by us. Every card opens the original post.
Managed Cloud Prometheus Monitoring
Monitoring Managed Cloud with Prometheus and Grafana This tutorial will guide you through the process of setting up Prometheus and Grafana to monitor Qdrant databases running in Qdrant Managed Cloud. Prerequisites This tutorial assumes that you already have a Kubernetes cluster…
Hyperbolic Embeddings in Qdrant
We choose embedding models, dimensions, and indexes. The geometry usually comes with the package. But why use a flat space, and what else could we choose? What Is a Manifold, and Where Do Our Vectors Live? A manifold is the space our embeddings live in. For embeddings, we care…
Memory Tiers in Qdrant: What to Use and When
Memory Tiers in Qdrant: What to Use and When A growing vector collection eventually outgrows the RAM it started with: Qdrant handles that by letting you assign dense vectors, the HNSW graph, quantized vectors, payloads, and payload indexes each to whichever memory tier that…
GPU-Accelerated HNSW Indexing
GPU-Accelerated HNSW Indexing in Qdrant Time: 45 min Level: Intermediate Output: GitHub Since Qdrant v1.13 , Qdrant has supported GPU-accelerated Hierarchical Navigable Small World (HNSW) indexing on self-hosted instances. Qdrant Cloud added it as a managed option more…
Configure Qdrant's Optimizer for Predictable Search Latency
A bulk load finishes, and the collection looks ready: every point is in, and the upload call has returned. Then the first queries land, and search takes hundreds of milliseconds, sometimes several seconds at a stretch, while Qdrant’s indexing, merge, and vacuum optimizers…
Hybrid Search in Qdrant
Hybrid Search in Qdrant A search result can look plausible and still be wrong. Dense retrieval can return a document on the right topic but miss an exact identifier copied into the query. Sparse retrieval can miss a relevant document when the query describes it with terms the…
When Your Collection Outgrows RAM
When Your Collection Outgrows RAM Once a collection no longer fits in RAM, the kernel evicts vector pages, and the next query waits on a disk read to get them back. Quantization buys that memory back. Qdrant keeps a compressed copy of each dense vector in RAM and moves the…
How to Tune Hybrid Search in Qdrant
How to Tune Hybrid Search in Qdrant Before you tune fusion, use the pre-tuning checks to verify index state and set a labeled baseline. Hybrid search retrieves dense and sparse candidate lists, then fuses them into one ranking. The dense prefetch finds similar meaning; the…
Prevent Unoptimized Usage
Prevent Unoptimized Usage Time: 20 min Level: Intermediate Output: GitHub After a bulk upload or a configuration change, a Qdrant collection can see higher search latency for a while. Ongoing optimizations create unindexed segments, and a query that lands on one of those…
What to Check Before Tuning a Qdrant Collection
What to Check Before Tuning a Qdrant Collection Before you change a setting, decide what better retrieval means for your workload. The right document at rank one, more candidates for a reranker, lower latency, and a smaller memory footprint each favor different settings, so…
Filtered Vector Search: What ACORN Fixes, and What Fixes ACORN
Filtered vector search breaks when metadata filters turn a healthy nearest-neighbor graph into scattered islands. HNSW’s m parameter controls how many links each point gets. At Qdrant’s default m=16 , the one-million-point collection benchmarked below averaged about…
FastEmbed: Qdrant's Efficient Python Library for Embedding Generation
Data Science and Machine Learning practitioners often find themselves navigating through a labyrinth of models, libraries, and frameworks. Which model to choose, what embedding size, and how to approach tokenizing, are just some questions you are faced with when starting your…