Using the Qdrant Python Client with FastAPI
The Qdrant Python client gives you a fully async interface that fits cleanly into FastAPI lifespan handlers, supports dense and sparse vectors for hybrid search, and exposes tuning knobs like quantization and sharding that actually matter at scale. I run this stack in production serving millions of vectors per day. Here is the practical breakdown of what works, what breaks, and the patterns I use daily.
Installation and async client setup with FastAPI lifespan
Start with the async extra. You want qdrant-client[fastapi] or just pip install "qdrant-client[async]". The sync client blocks the event loop. Don’t use it in an async app.
pip install "qdrant-client[async]" fastapi uvicornInitialize the client inside a lifespan context manager. This handles connection pooling and graceful shutdown. I point the client at a DNS name that resolves to a managed Qdrant cluster or a sidecar container. Never hardcode localhost in config.
from contextlib import asynccontextmanagerfrom qdrant_client import AsyncQdrantClientfrom qdrant_client.http import modelsfrom fastapi import FastAPIimport os
QDRANT_URL = os.getenv("QDRANT_URL", "http://localhost:6333")QDRANT_API_KEY = os.getenv("QDRANT_API_KEY") # None for local dev
_client: AsyncQdrantClient | None = None
def get_client() -> AsyncQdrantClient: if _client is None: raise RuntimeError("Qdrant client not initialized") return _client
@asynccontextmanagerasync def lifespan(app: FastAPI): global _client _client = AsyncQdrantClient( url=QDRANT_URL, api_key=QDRANT_API_KEY, timeout=10.0, # seconds prefer_grpc=True, # lower latency, higher throughput ) # Warm up connections await _client.get_collections() yield await _client.close() _client = None
app = FastAPI(lifespan=lifespan)Setting prefer_grpc=True uses the gRPC interface. It reduces serialization overhead and supports streaming. The HTTP fallback kicks in automatically if gRPC fails. I set a 10 second timeout because vector search can stall on cold segments. You need a circuit breaker upstream. I cover that pattern in Fix AsyncSession Errors in FastAPI (SQLAlchemy 2.0 Guide) for SQLAlchemy; the same retry and timeout logic applies here.
Collection creation with vector params and payload indexes
Create collections on startup or via a migration script. Do not auto-create on every request. Define your vector config and payload indexes upfront. Payload indexes are critical for filtering performance. Without them, Qdrant scans every point.
from qdrant_client.http import models
DENSE_VECTOR_NAME = "text-dense"SPARSE_VECTOR_NAME = "text-sparse"
def get_collection_config(vector_size: int = 1024) -> models.VectorParams: return models.VectorParams( size=vector_size, distance=models.Distance.COSINE, on_disk=True, # move cold vectors to disk hnsw_config=models.HnswConfigDiff( m=16, # connections per node ef_construct=100, # build time accuracy full_scan_threshold=10000, ), quantization_config=models.ScalarQuantization( scalar=models.ScalarQuantizationConfig( type=models.ScalarType.INT8, quantile=0.99, always_ram=True, # keep quantized vectors in RAM ) ), )
def get_sparse_vector_config() -> models.SparseVectorParams: return models.SparseVectorParams( index=models.SparseIndexParams( on_disk=False, ) )
PAYLOAD_INDEXES = [ models.PayloadSchemaType.KEYWORD, # for exact match filters models.PayloadSchemaType.INTEGER, # for range filters models.PayloadSchemaType.FLOAT, # for score thresholds models.PayloadSchemaType.DATETIME, # for time-based queries]The on_disk=True flag on dense vectors saves RAM. It adds ~5-10ms latency on cold reads. Worth it if your dataset exceeds memory. The ScalarQuantization config compresses float32 vectors to int8. You lose ~1-2% recall but cut memory by 4x. I run this on all collections larger than 500k vectors.
Create the collection in a startup task:
from qdrant_client.http import modelsfrom app.qdrant import get_clientfrom app.schema import ( DENSE_VECTOR_NAME, SPARSE_VECTOR_NAME, get_collection_config, get_sparse_vector_config, PAYLOAD_INDEXES,)
COLLECTION_NAME = "documents"
async def ensure_collection(vector_size: int = 1024): client = get_client() exists = await client.collection_exists(COLLECTION_NAME) if exists: return await client.create_collection( collection_name=COLLECTION_NAME, vectors_config={ DENSE_VECTOR_NAME: get_collection_config(vector_size), }, sparse_vectors_config={ SPARSE_VECTOR_NAME: get_sparse_vector_config(), }, ) # Create payload indexes for field_name, schema_type in [ ("doc_id", models.PayloadSchemaType.KEYWORD), ("tenant_id", models.PayloadSchemaType.KEYWORD), ("created_at", models.PayloadSchemaType.DATETIME), ("score", models.PayloadSchemaType.FLOAT), ]: await client.create_payload_index( collection_name=COLLECTION_NAME, field_name=field_name, field_schema=schema_type, )Upserting points with batching and payload filtering
Batch your upserts. Single-point writes kill throughput. I batch 256 points per request. Use wait=True only when you need immediate consistency. For fire-and-forget ingestion, wait=False returns immediately and Qdrant flushes in the background.
from qdrant_client.http import modelsfrom qdrant_client import AsyncQdrantClientfrom app.qdrant import get_clientfrom app.schema import COLLECTION_NAME, DENSE_VECTOR_NAME, SPARSE_VECTOR_NAMEimport uuid
BATCH_SIZE = 256
async def upsert_documents( texts: list[str], dense_vectors: list[list[float]], sparse_vectors: list[dict[int, float]], # {index: value} metadata: list[dict],): client = get_client() points = [] for i, text in enumerate(texts): point_id = uuid.uuid4().hex points.append( models.PointStruct( id=point_id, vector={ DENSE_VECTOR_NAME: dense_vectors[i], SPARSE_VECTOR_NAME: models.SparseVector( indices=list(sparse_vectors[i].keys()), values=list(sparse_vectors[i].values()), ), }, payload={ "text": text, **metadata[i], }, ) ) if len(points) >= BATCH_SIZE: await client.upsert( collection_name=COLLECTION_NAME, points=points, wait=False, ) points.clear() if points: await client.upsert( collection_name=COLLECTION_NAME, points=points, wait=False, )Filter by payload during search. The filter runs before vector scoring. This is efficient when indexes exist.
from qdrant_client.http import modelsfrom app.qdrant import get_clientfrom app.schema import COLLECTION_NAME, DENSE_VECTOR_NAME
async def search_with_filter( query_vector: list[float], tenant_id: str, date_from: str | None = None, limit: int = 10,): client = get_client() must_conditions = [ models.FieldCondition( key="tenant_id", match=models.MatchValue(value=tenant_id), ), ] if date_from: must_conditions.append( models.FieldCondition( key="created_at", range=models.DatetimeRange(gte=date_from), ) ) query_filter = models.Filter(must=must_conditions) results = await client.search( collection_name=COLLECTION_NAME, query_vector=models.NamedVector( name=DENSE_VECTOR_NAME, vector=query_vector, ), query_filter=query_filter, limit=limit, with_payload=True, score_threshold=0.7, ) return resultsHybrid search using sparse/dense vectors and fusion
Hybrid search combines dense semantic similarity with sparse keyword matching. Qdrant supports Fusion.RRF (Reciprocal Rank Fusion) and Fusion.DBSF. RRF is parameter-free and works well out of the box. DBSF lets you weight dense vs sparse. I default to RRF.
from qdrant_client.http import modelsfrom app.qdrant import get_clientfrom app.schema import COLLECTION_NAME, DENSE_VECTOR_NAME, SPARSE_VECTOR_NAME
async def hybrid_search( dense_vector: list[float], sparse_vector: dict[int, float], tenant_id: str, limit: int = 10,): client = get_client() prefetch = [ models.Prefetch( query=models.NamedVector( name=DENSE_VECTOR_NAME, vector=dense_vector, ), limit=50, filter=models.Filter( must=[ models.FieldCondition( key="tenant_id", match=models.MatchValue(value=tenant_id), ), ] ), ), models.Prefetch( query=models.NamedSparseVector( name=SPARSE_VECTOR_NAME, vector=models.SparseVector( indices=list(sparse_vector.keys()), values=list(sparse_vector.values()), ), ), limit=50, filter=models.Filter( must=[ models.FieldCondition( key="tenant_id", match=models.MatchValue(value=tenant_id), ), ] ), ), ] results = await client.query_points( collection_name=COLLECTION_NAME, prefetch=prefetch, query=models.FusionQuery( fusion=models.Fusion.RRF, ), limit=limit, with_payload=True, ) return results.pointsThe prefetch stage runs each vector search independently with a higher limit. Fusion merges the ranked lists. This is slower than single-vector search because you execute two ANN searches. Budget 50-100ms extra latency. If you need sub-50ms p99, stick to dense only and handle keywords in a separate filter.
Performance tuning: quantization, sharding, and connection pooling
Quantization is the biggest lever. I use ScalarQuantization with INT8 and always_ram=True. For collections over 10M vectors, ProductQuantization (PQ) compresses further but requires a training step. PQ adds ~20ms to search latency. Only use it when RAM is the hard constraint.
Sharding splits a collection across nodes. Configure it at creation time:
await client.create_collection( collection_name=COLLECTION_NAME, vectors_config={...}, shard_number=4, # must divide evenly across nodes replication_factor=2, # HA)You cannot change shard count later. Plan for 3-5x growth. Each shard holds its own HNSW graph. More shards means more parallelism but smaller graphs per shard, which can hurt recall.
Connection pooling: the async client maintains a pool internally. The default limit is 100 connections. For high-throughput FastAPI workers, increase it:
_client = AsyncQdrantClient( url=QDRANT_URL, api_key=QDRANT_API_KEY, timeout=10.0, prefer_grpc=True, # grpc specific options grpc_options={ "grpc.max_receive_message_length": 100 * 1024 * 1024, # 100MB },)Monitor grpc_client_msg_received_total and http_client_request_duration_seconds in Prometheus. If you see connection exhaustion, increase the pool or add a read replica.
Production patterns: retries, health checks, and monitoring
Wrap client calls in a retry policy. Transient network blips happen. Use tenacity with exponential backoff. Don’t retry on 4xx errors.
from tenacity import ( retry, stop_after_attempt, wait_exponential_jitter, retry_if_exception_type,)from qdrant_client.http.exceptions import UnexpectedResponseimport httpx
def is_retryable(exc: BaseException) -> bool: if isinstance(exc, (httpx.TimeoutException, httpx.ConnectError)): return True if isinstance(exc, UnexpectedResponse): return 500 <= exc.status_code < 600 return False
retry_policy = retry( wait=wait_exponential_jitter(initial=0.1, max=2.0), stop=stop_after_attempt(3), retry=retry_if_exception_type((httpx.TimeoutException, httpx.ConnectError, UnexpectedResponse)), retry_error_callback=lambda state: state.outcome.exception(),)Apply it to search and upsert:
@retry_policyasync def search_with_retry(...): return await client.search(...)Health checks: hit /readyz on the Qdrant HTTP port. It returns 200 when the node can serve traffic. Do not use /healthz; that only reports process liveness.
from fastapi import APIRouter, HTTPExceptionfrom app.qdrant import get_client
router = APIRouter()
@router.get("/ready")async def ready_check(): client = get_client() try: await client.get_collections() return {"status": "ok"} except Exception as e: raise HTTPException(status_code=503, detail=str(e))Monitoring: export Qdrant metrics to Prometheus. Key alerts:
collection_points_countgrowing unbounded (ingestion > deletion)search_latency_secondsp99 > 500mswal_write_latency_secondsspikes (disk IO pressure)replication_lag_seconds> 30s on replicas
I built a small wrapper that logs slow queries and emits custom metrics. It sits in the same pattern as the AI agent observability I wrote about in ai agent python code example for FastAPI and OpenAI SDK.
When not to use Qdrant
If your vector count stays under 100k and you already run Postgres, pgvector is simpler. One less moving part. No separate cluster to operate. Qdrant shines when you need multi-tenancy with strict isolation, hybrid search at scale, or quantization to fit large datasets in memory. It also handles payload filtering better than most pgvector setups.
If you need full-text search with complex linguistics (stemming, synonyms, phrase queries), pair Qdrant with Elasticsearch or Typesense. Qdrant’s sparse vectors are BM25-like but not a full search engine.
FAQ
How do I migrate a collection schema in Qdrant? You cannot alter vector params or payload indexes on an existing collection. Create a new collection with the target config, re-index data via scroll and upsert, then swap the alias. Plan for downtime or run dual-write during migration.
Does the Qdrant Python client support asyncio natively?
Yes. AsyncQdrantClient uses httpx.AsyncClient and grpc.aio under the hood. It integrates with FastAPI lifespan and asyncio.gather for concurrent searches. Do not mix sync and async clients in the same process.
What is the difference between search and query_points?
search is the legacy single-vector endpoint. query_points supports prefetch, fusion, and multi-vector queries. Use query_points for hybrid search and advanced ranking. They share the same underlying engine.
How do I handle multi-tenancy?
Option one: one collection per tenant. Simple isolation, but max collections per cluster is ~10k. Option two: single collection with tenant_id payload filter and a keyword index. This scales to millions of tenants. I use option two with a tenant_id index and filter on every query.
Key Takeaways
- Use
AsyncQdrantClientwithprefer_grpc=Trueinside FastAPI lifespan for connection management - Enable
ScalarQuantizationwithINT8andalways_ram=Trueto cut memory 4x with minimal recall loss - Create payload indexes for every field you filter on; unindexed filters scan the entire collection
- Batch upserts at 256 points with
wait=Falsefor throughput; usewait=Trueonly when consistency is required - Hybrid search via
Fusion.RRFadds latency; benchmark before adopting - Shard count is immutable; choose 4-8 shards per node for 3-5x headroom
- Wrap all client calls in
tenacityretries with exponential backoff and 5xx filtering - Monitor search latency, WAL write latency, and replication lag; alert on p99 > 500ms
Working on something similar?
If you're building backend or AI systems and want a second set of senior eyes, let's talk.