Conversation
cjnolet
reviewed
Sep 30, 2026
| } cuvsFlowannIndex; | ||
| typedef cuvsFlowannIndex* cuvsFlowannIndex_t; | ||
|
|
||
| CUVS_EXPORT cuvsError_t cuvsFlowannIndexParamsCreate(cuvsFlowannIndexParams_t* params); |
Contributor
There was a problem hiding this comment.
Rather than adding completely new API functions for flowann, let's find a way we can consolidate into existing APIs / options where possible. For e.g. if this is built on the CAGRA graph specifically, can we add an option / functions to the CAGRA APIs to enable this?
…U storage Integrate FlowANN search, graph building, serialization, benchmarks, and API documentation. Add GDRCopy and CUDA-copy queue runtimes with regression coverage, along with C and Python bindings for dense FlowANN indexes. Co-authored-by: Haoru Zhao <zhaohaoru@sjtu.edu.cn> Co-authored-by: Jingkai He <hjk020101@sjtu.edu.cn>
Use CAGRA build, search and serialization entry points for tiered dense and VPQ indexes across C++, C and Python. Own the reordered dataset and queue runtime inside the index, and retain legacy loading in the native benchmark. Internalize standalone FlowANN interfaces, update benchmark integration, parameter documentation and generated API references, and add ownership, serialization and optional-backend regression coverage. Co-authored-by: Jingkai He <hjk020101@sjtu.edu.cn>
ZhaoHaoRu
force-pushed
the
zhr/flowann-upstream
branch
from
October 4, 2026 03:12
377b285 to
b260d2e
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
This PR integrates FlowANN from our OSDI 2026 paper Disentangling Graph Dependencies for Efficient Billion-Scale GPU Vector Search into cuVS. FlowANN keeps a compressed subset of graph edges on the GPU, stores the remaining edges in CPU memory, and fetches CPU-resident edges asynchronously during search.
The PR adds the index format, graph-building pipeline, C++/C/Python interfaces, CPU-GPU request queues, serialization, benchmarks, and tests needed to run FlowANN in cuVS. FlowANN is exposed as an experimental tiered graph storage mode through
cuvs::neighbors::cagra, selected withindex_params.graph_storage = graph_storage_kind::tiered. FlowANN is excluded from the default cuVS build.co-author: Haoru Zhao zhaohaoru@sjtu.edu.cn, Jingkai He hjk020101@sjtu.edu.cn, Jingyao Zeng zengjingyao@sjtu.edu.cn, Shanghan Gao gsh20040816@gmail.com
Implementation in cuVS
Reuse the CAGRA search core
This PR changes the CAGRA single-CTA and multi-CTA JIT search loops to read neighbors through a graph accessor. Standard CAGRA continues to use the existing GPU adjacency-list accessor, preserving its execution path and graph format. FlowANN uses a new tiered graph accessor that reads resident edges from the compressed GPU graph and obtains CPU-resident edges through request queues.
Both paths share CAGRA's candidate sorting, visited-node deduplication, distance computation, and iterative search logic. FlowANN separately implements search planning, request state, and result ingestion, avoiding a duplicate CAGRA search kernel. FlowANN supports dense vectors and VPQ-F16 data, as well as single-CTA and multi-CTA search.
Add the FlowANN index and graph-building pipeline
A
cagra::indexconfigured for tiered storage owns an internal FlowANN implementation that stores:The integrated build entry point first constructs a complete CAGRA graph, then performs grouping, node reordering, edge sorting, and tiered storage. The builder derives the GPU graph row width from
device_graph_budget_bytesand writes edges that do not fit on the GPU to the CPU graph. Node reordering also updates the dataset and ID mapping, so search results continue to use the original dataset IDs.The default grouping implementation uses hierarchical k-means. Each node is divided into at most 16 child groups, with capacity repair at every level. The local-ID width and final group count can be specified explicitly or derived automatically from the dataset size and group capacity.
The graph-building implementation retains three operations:
cagra::build()constructs a CAGRA graph from vectors and produces a tiered CAGRA index. The compressed overload accepts full-precision host vectors and VPQ parameters.build_from_cagra()converts an existing CAGRA index into the tiered representation.group_graph()computes grouping results independently for an existing graph or a custom build pipeline.The public
tiered_graph_paramscarries the GPU graph budget, grouping, and seed-generation settings throughcagra::index_params::tiered. Graph conversion and grouping helpers are implementation details rather than separate public FlowANN APIs.Add the xCopier queue runtime
An internal
search_context, owned and reused by the tiered CAGRA index, creates GPU request queues and CPU polling threads. Callers usecagra::search()with optionaltiered_search_paramsfor queue, seed, and synchronization settings; they do not manage a separate context.keep_pollers_running=truekeeps CPU pollers active between calls. When expanding a node, a FlowANN kernel submits requests for CPU adjacency rows. CPU threads read the corresponding rows and write them into result buffers. The GPU consumes ready results in later search iterations and reclaims queue capacity.The default transfer path uses GDRCopy for small transfers such as adjacency rows. Building with
CUVS_FLOWANN_USE_GDRCOPY=OFFselects the CUDA-copy path instead. The queue implementation also handles batched request publication, result visibility, deferred reclamation, thread pause/resume, CUDA device binding, and exception propagation.Integrate serialization, loading, and benchmarks
This PR supports dense-vector and VPQ-F16 tiered indexes through CAGRA's
serialize()anddeserialize()entry points, while preserving historical split-index loading in the native benchmark importer. The internalindex_bundleowns both the dataset object and the FlowANN index; the publiccagra::indexholds this bundle and the queue runtime, keeping them alive across index moves and searches. Callers no longer receive or manage a separate bundle. Tiered serialization includes both the graph and its owned dataset; graph-only serialization is unsupported. Ordinary CAGRA serialization retains its existing format.The native
FLOWANN_BUILD_BENCHMARKandFLOWANN_BENCHMARKexecutables cover graph building and search, respectively. A new FlowANN backend in cuVS-bench loads an existing index, runs the native search executable throughcagra::search(), and reports timing and Recall using the number of queries actually executed.Control dependencies and build scope
FlowANN sources, tests, JIT kernels, and dependencies are controlled by CMake options. When FlowANN is disabled, standard cuVS and CAGRA builds do not require GDRCopy and do not build the FlowANN benchmarks. The public API is declared in the CAGRA headers, with generated API documentation.
flowann.hpp,flowann_build.hpp, andflowann_serialize.hppare internal implementation headers.This PR extends the existing
cuvs.neighbors.cagraPython binding, reusingIndexParams,SearchParams,Index,build(),search(),save(), andload(). Selectgraph_storage="tiered"and supplyTieredGraphParamsorTieredSearchParamsfor tiered-specific settings. The binding calls the native implementation through the CAGRA C API and supports float32, float16, int8, and uint8 dense-vector indexes. VPQ-F16 construction from host vectors is exposed throughcagra.build(..., compression=...)andcuvsCagraBuildCompressed. The cuVS-bench Python backend continues to invoke the native executable.Build and usage scope
CUVS_ENABLE_FLOWANN_SEARCH=ONCUVS_ENABLE_FLOWANN_BUILD=ONCUVS_FLOWANN_USE_GDRCOPY=ONgdrdrvmodule at runtime.OFFselects the CUDA-copy path.CUVS_BUILD_FLOWANN_BENCHMARKS=ONThe public headers are
cuvs/neighbors/cagra.hppand the C API headercuvs/neighbors/cagra.h. Tiered storage uses the CAGRA index type and entry points rather than separate public FlowANN handles. This PR covers the C++, C, and Python APIs and the cuVS-bench backend.Testing and validation
Local validation used an H100, CUDA 13.1.115, GCC 13.4, and CMake 4.4.3. The shared library, FlowANN search and build benchmarks, and related test executables were built successfully. The results below distinguish the current CAGRA API integration checks from earlier validation.
query_count=0, and requested counts exceeding the query set. These earlier checks were not rerun for this API revision.find_package(cuvs)andcuvs::cuvsagainst the build-tree export. The updated installed package was not tested.Performance regression check for the CAGRA API integration
This compares the previous direct FlowANN entry point with the unified CAGRA entry point using the same freshly built search engine, hierarchical Wiki88M index, and H800. It is separate from the two-system sweep below.
256/32/2/32256/90/2/64QPS is the median of process-level medians and includes query transfer and CPU exact reranking. Batch 1 uses 2,000 queries, multi-CTA/team32, two rerank threads, one queue and 128 seeds; one control and two API processes each have one warmup and one measured round. Batch 10,000 uses 10,000 queries, single-CTA/team8, 90 rerank threads, 44 queues and 200 seeds; three control and four API processes each have one warmup and five measured rounds. Extra large-batch runs alternate API/control. Queue statistics and persistent host pollers are enabled for both. Median search time changes are −0.16% and +0.015%, respectively. These representative checks did not show consistent API integration overhead; they are not an exhaustive performance guarantee.
Wiki88M GDRCopy-on FlowANN vs. Sharded CAGRA
These results use the corrected September 20/23, 2026 sweep, including the sharded CAGRA rerank thread-team fix. They are historical system measurements, not a new sweep of the API-integration commit. FlowANN and large-batch CAGRA were measured on September 23; small-batch CAGRA was reused from September 20 on the same machine. Earlier H100 labels were presentation labels; no separate H100 run was performed.
Data, indexes, and hardware
top-k=10; Recall@10 is computed against the first 10 IDs in each ground-truth row.CUVS_FLOWANN_USE_GDRCOPY=ON.Both methods use the same base vectors, queries, ground truth, graph degree, and VPQ specification, but their indexes are not derived by splitting the same physical graph. Their historical graphs, trained codebooks, sharding schemes, and entry-point strategies differ. This table therefore compares the two systems at specified Recall thresholds using their respective indexes; it is not a controlled experiment that changes only the search kernel.
Execution and timing methodology
batch=1, the benchmark issues 10,000 sequential single-query searches usingmulti_cta,team_size=32, and 2 CPU exact-reranking threads. Withbatch=10,000, it submits all queries in one call usingsingle_cta,team_size=8, and 90 reranking threads.batch=1, FlowANN uses 1 GDRCopy queue,empty_pause=0, and 128 seeds. Forbatch=10,000, it uses 44 queues,empty_pause=64, and 200 seeds. Shared FlowANN settings includesync_window_scale=10andsync_drop_threshold=51.hashmap_min_bitlen=14. FlowANN uses the k-means index with six groups, 24-bit local IDs, and a 7,529,758,208-byte resident graph budget. Each selected configuration has three independent process runs.itopk / max_iterations / search_width / rerank_candidates. The two methods may select parameters independently for the same Recall threshold, so configurations in one row do not necessarily match.Results at each Recall threshold
For each method, select the highest median QPS among configurations whose minimum measured Recall@10 across three runs meets the threshold. The table reports that median QPS and minimum Recall. Every entry is a measured point; no interpolation is used.
FlowANN relative QPSisFlowANN QPS / Sharded CAGRA QPS - 1, without normalizing for GPU count.512/16/2/32512/16/2/32512/20/2/32512/16/2/32512/28/2/32512/28/32/32512/28/2/1664/28/2/32512/28/4/32512/28/4/32512/28/8/32512/28/8/32At batch 1, FlowANN achieves 74.3%, 63.6%, and 75.2% of two-GPU CAGRA QPS at the three targets. At batch 10,000, its median QPS is 2.4% and 0.9% higher at Recall 0.90 and 0.95. The 0.95 QPS ranges overlap, so the median difference does not establish a stable advantage. CAGRA did not reach 0.99 within the measured configurations; this does not imply that CAGRA cannot reach that Recall with other settings.