Skip to content

Add experimental FlowANN search and graph building with tiered GPU-CPU storage & transfer - #2675

Open
ZhaoHaoRu wants to merge 2 commits into
NVIDIA:mainfrom
ZhaoHaoRu:zhr/flowann-upstream
Open

ZhaoHaoRu wants to merge 2 commits into
NVIDIA:mainfrom
ZhaoHaoRu:zhr/flowann-upstream

Conversation

@ZhaoHaoRu

@ZhaoHaoRu ZhaoHaoRu commented Sep 22, 2026 •

Copy link
Copy Markdown

Overview

This PR integrates FlowANN from our OSDI 2026 paper Disentangling Graph Dependencies for Efficient Billion-Scale GPU Vector Search into cuVS. FlowANN keeps a compressed subset of graph edges on the GPU, stores the remaining edges in CPU memory, and fetches CPU-resident edges asynchronously during search.

The PR adds the index format, graph-building pipeline, C++/C/Python interfaces, CPU-GPU request queues, serialization, benchmarks, and tests needed to run FlowANN in cuVS. FlowANN is exposed as an experimental tiered graph storage mode through cuvs::neighbors::cagra, selected with index_params.graph_storage = graph_storage_kind::tiered. FlowANN is excluded from the default cuVS build.

co-author: Haoru Zhao zhaohaoru@sjtu.edu.cn, Jingkai He hjk020101@sjtu.edu.cn, Jingyao Zeng zengjingyao@sjtu.edu.cn, Shanghan Gao gsh20040816@gmail.com

Implementation in cuVS

Reuse the CAGRA search core

This PR changes the CAGRA single-CTA and multi-CTA JIT search loops to read neighbors through a graph accessor. Standard CAGRA continues to use the existing GPU adjacency-list accessor, preserving its execution path and graph format. FlowANN uses a new tiered graph accessor that reads resident edges from the compressed GPU graph and obtains CPU-resident edges through request queues.

Both paths share CAGRA's candidate sorting, visited-node deduplication, distance computation, and iterative search logic. FlowANN separately implements search planning, request state, and result ingestion, avoiding a duplicate CAGRA search kernel. FlowANN supports dense vectors and VPQ-F16 data, as well as single-CTA and multi-CTA search.

Add the FlowANN index and graph-building pipeline

A cagra::index configured for tiered storage owns an internal FlowANN implementation that stores:

  • The compressed resident graph on the GPU.
  • Non-resident adjacency edges in CPU memory.
  • Grouping metadata, the original-ID mapping, and optional search entry points.
  • An owned dense-vector or VPQ-F16 dataset, exposed through a dataset view.

The integrated build entry point first constructs a complete CAGRA graph, then performs grouping, node reordering, edge sorting, and tiered storage. The builder derives the GPU graph row width from device_graph_budget_bytes and writes edges that do not fit on the GPU to the CPU graph. Node reordering also updates the dataset and ID mapping, so search results continue to use the original dataset IDs.

The default grouping implementation uses hierarchical k-means. Each node is divided into at most 16 child groups, with capacity repair at every level. The local-ID width and final group count can be specified explicitly or derived automatically from the dataset size and group capacity.

The graph-building implementation retains three operations:

  • Public cagra::build() constructs a CAGRA graph from vectors and produces a tiered CAGRA index. The compressed overload accepts full-precision host vectors and VPQ parameters.
  • Internal build_from_cagra() converts an existing CAGRA index into the tiered representation.
  • Internal group_graph() computes grouping results independently for an existing graph or a custom build pipeline.

The public tiered_graph_params carries the GPU graph budget, grouping, and seed-generation settings through cagra::index_params::tiered. Graph conversion and grouping helpers are implementation details rather than separate public FlowANN APIs.

Add the xCopier queue runtime

An internal search_context, owned and reused by the tiered CAGRA index, creates GPU request queues and CPU polling threads. Callers use cagra::search() with optional tiered_search_params for queue, seed, and synchronization settings; they do not manage a separate context. keep_pollers_running=true keeps CPU pollers active between calls. When expanding a node, a FlowANN kernel submits requests for CPU adjacency rows. CPU threads read the corresponding rows and write them into result buffers. The GPU consumes ready results in later search iterations and reclaims queue capacity.

The default transfer path uses GDRCopy for small transfers such as adjacency rows. Building with CUVS_FLOWANN_USE_GDRCOPY=OFF selects the CUDA-copy path instead. The queue implementation also handles batched request publication, result visibility, deferred reclamation, thread pause/resume, CUDA device binding, and exception propagation.

Integrate serialization, loading, and benchmarks

This PR supports dense-vector and VPQ-F16 tiered indexes through CAGRA's serialize() and deserialize() entry points, while preserving historical split-index loading in the native benchmark importer. The internal index_bundle owns both the dataset object and the FlowANN index; the public cagra::index holds this bundle and the queue runtime, keeping them alive across index moves and searches. Callers no longer receive or manage a separate bundle. Tiered serialization includes both the graph and its owned dataset; graph-only serialization is unsupported. Ordinary CAGRA serialization retains its existing format.

The native FLOWANN_BUILD_BENCHMARK and FLOWANN_BENCHMARK executables cover graph building and search, respectively. A new FlowANN backend in cuVS-bench loads an existing index, runs the native search executable through cagra::search(), and reports timing and Recall using the number of queries actually executed.

Control dependencies and build scope

FlowANN sources, tests, JIT kernels, and dependencies are controlled by CMake options. When FlowANN is disabled, standard cuVS and CAGRA builds do not require GDRCopy and do not build the FlowANN benchmarks. The public API is declared in the CAGRA headers, with generated API documentation. flowann.hpp, flowann_build.hpp, and flowann_serialize.hpp are internal implementation headers.

This PR extends the existing cuvs.neighbors.cagra Python binding, reusing IndexParams, SearchParams, Index, build(), search(), save(), and load(). Select graph_storage="tiered" and supply TieredGraphParams or TieredSearchParams for tiered-specific settings. The binding calls the native implementation through the CAGRA C API and supports float32, float16, int8, and uint8 dense-vector indexes. VPQ-F16 construction from host vectors is exposed through cagra.build(..., compression=...) and cuvsCagraBuildCompressed. The cuVS-bench Python backend continues to invoke the native executable.

Build and usage scope

CMake option Purpose
CUVS_ENABLE_FLOWANN_SEARCH=ON Enable FlowANN search.
CUVS_ENABLE_FLOWANN_BUILD=ON Enable FlowANN graph building. Search support must also be enabled.
CUVS_FLOWANN_USE_GDRCOPY=ON Use GDRCopy queues. This requires the GDRCopy development library and an available gdrdrv module at runtime. OFF selects the CUDA-copy path.
CUVS_BUILD_FLOWANN_BENCHMARKS=ON Build the native benchmark executables. Search support is required; the build benchmark also requires graph-building support.

The public headers are cuvs/neighbors/cagra.hpp and the C API header cuvs/neighbors/cagra.h. Tiered storage uses the CAGRA index type and entry points rather than separate public FlowANN handles. This PR covers the C++, C, and Python APIs and the cuVS-bench backend.

Testing and validation

Local validation used an H100, CUDA 13.1.115, GCC 13.4, and CMake 4.4.3. The shared library, FlowANN search and build benchmarks, and related test executables were built successfully. The results below distinguish the current CAGRA API integration checks from earlier validation.

Validation Result
FlowANN and CAGRA tiered unit/regression tests 58/58 passed with GDRCopy ON and 58/58 with GDRCopy OFF, covering graph building, serialization, compressed decoding, queue lifecycle, search, index ownership, and move behavior.
CAGRA C API 15/15 passed with FlowANN enabled; 15/15 with FlowANN disabled, including rejection of tiered builds. The existing ACE-disk test required a clean temporary directory for its fixed output path.
Existing CAGRA C++ regression cases 23 passed, 1 expected skip in the selected subset; the full matrix was not run.
cuVS-bench Python backend tests 18/18 passed.
CAGRA Python tiered binding 4/4 passed for dense/VPQ construction, repeated search, parameter checks, and serialization round trips.
Existing CAGRA Python regression cases 20/20 selected cases passed, covering data types, serialization, filtering, and parameters.
Rust FFI The FFI module and layout assertions compile; full Cargo/runtime tests were not run.
Build with FlowANN disabled Shared and C libraries built; all 15 C API cases passed. ELF inspection confirmed no GDRCopy dependency.
Earlier graph-building and cuVS-bench end-to-end validation Passed with 1,024 vectors and 64 queries for batches 1 and 64, query_count=0, and requested counts exceeding the query set. These earlier checks were not rerun for this API revision.
Earlier external CMake consumer Configured, compiled, linked, and ran using find_package(cuvs) and cuvs::cuvs against the build-tree export. The updated installed package was not tested.
Independent build with GDRCopy disabled The current native suite passes 58/58. Earlier validation reported 0 memcheck errors for 55 existing cases, 3/3 synthetic build-search checks, and successful cold starts after the pinned-host-buffer fix. The added cold-start case still timed out under memcheck, and historical Wiki88M CUDA-copy Recall was lower than GDRCopy; this API revision does not establish transport equivalence or resolve those earlier observations.
Formatting and static checks Applicable pre-commit checks, secret scanning, and generated API documentation checks passed.

Performance regression check for the CAGRA API integration

This compares the previous direct FlowANN entry point with the unified CAGRA entry point using the same freshly built search engine, hierarchical Wiki88M index, and H800. It is separate from the two-system sweep below.

Batch itopk / iterations / width / candidates Direct FlowANN QPS CAGRA API QPS Relative QPS Direct Recall range CAGRA API Recall range
1 256/32/2/32 2,018.3 2,022.7 +0.21% 0.98655 0.98565–0.98675
10,000 256/90/2/64 159,236.5 158,405.9 −0.52% 0.99078–0.99079 0.99078–0.99079

QPS is the median of process-level medians and includes query transfer and CPU exact reranking. Batch 1 uses 2,000 queries, multi-CTA/team32, two rerank threads, one queue and 128 seeds; one control and two API processes each have one warmup and one measured round. Batch 10,000 uses 10,000 queries, single-CTA/team8, 90 rerank threads, 44 queues and 200 seeds; three control and four API processes each have one warmup and five measured rounds. Extra large-batch runs alternate API/control. Queue statistics and persistent host pollers are enabled for both. Median search time changes are −0.16% and +0.015%, respectively. These representative checks did not show consistent API integration overhead; they are not an exhaustive performance guarantee.

Wiki88M GDRCopy-on FlowANN vs. Sharded CAGRA

These results use the corrected September 20/23, 2026 sweep, including the sharded CAGRA rerank thread-team fix. They are historical system measurements, not a new sweep of the API-integration commit. FlowANN and large-batch CAGRA were measured on September 23; small-batch CAGRA was reused from September 20 on the same machine. Earlier H100 labels were presentation labels; no separate H100 run was performed.

Data, indexes, and hardware

Item Configuration
Dataset Wiki88M: 87,555,327 base vectors, 10,000 queries, and 768 dimensions.
Evaluation metric top-k=10; Recall@10 is computed against the first 10 IDs in each ground-truth row.
Graph and quantization Graph degree 32; search data uses VPQ 192×8 bit with 10,000 VQ centers.
FlowANN One H800 (GPU 0), with CUVS_FLOWANN_USE_GDRCOPY=ON.
Sharded CAGRA Two H800s (GPUs 0 and 1).

Both methods use the same base vectors, queries, ground truth, graph degree, and VPQ specification, but their indexes are not derived by splitting the same physical graph. Their historical graphs, trained codebooks, sharding schemes, and entry-point strategies differ. This table therefore compares the two systems at specified Recall thresholds using their respective indexes; it is not a controlled experiment that changes only the search kernel.

Execution and timing methodology

  • Every process handles 10,000 queries. With batch=1, the benchmark issues 10,000 sequential single-query searches using multi_cta, team_size=32, and 2 CPU exact-reranking threads. With batch=10,000, it submits all queries in one call using single_cta, team_size=8, and 90 reranking threads.
  • For batch=1, FlowANN uses 1 GDRCopy queue, empty_pause=0, and 128 seeds. For batch=10,000, it uses 44 queues, empty_pause=64, and 200 seeds. Shared FlowANN settings include sync_window_scale=10 and sync_drop_threshold=51.
  • Both GPU methods use hashmap_min_bitlen=14. FlowANN uses the k-means index with six groups, 24-bit local IDs, and a 7,529,758,208-byte resident graph budget. Each selected configuration has three independent process runs.
  • QPS measures end-to-end query throughput, including query transfer, GPU search, candidate transfer/preparation, and CPU exact reranking. It excludes index loading, initial JIT compilation, and Recall computation.
  • Configurations in the table use the format itopk / max_iterations / search_width / rerank_candidates. The two methods may select parameters independently for the same Recall threshold, so configurations in one row do not necessarily match.

Results at each Recall threshold

For each method, select the highest median QPS among configurations whose minimum measured Recall@10 across three runs meets the threshold. The table reports that median QPS and minimum Recall. Every entry is a measured point; no interpolation is used. FlowANN relative QPS is FlowANN QPS / Sharded CAGRA QPS - 1, without normalizing for GPU count.

Batch Recall threshold FlowANN QPS / actual Recall / configuration Sharded CAGRA QPS / actual Recall / configuration FlowANN relative QPS
1 0.90 3,260.3 / 0.94275 / 512/16/2/32 4,390.2 / 0.95577 / 512/16/2/32 -25.7%
1 0.95 2,793.6 / 0.97446 / 512/20/2/32 4,390.2 / 0.95577 / 512/16/2/32 -36.4%
1 0.99 2,184.3 / 0.99036 / 512/28/2/32 2,905.4 / 0.99092 / 512/28/32/32 -24.8%
10,000 0.90 428,998.2 / 0.91498 / 512/28/2/16 418,951.8 / 0.90688 / 64/28/2/32 +2.4%
10,000 0.95 273,877.4 / 0.97549 / 512/28/4/32 271,524.2 / 0.95504 / 512/28/4/32 +0.9%
10,000 0.99 175,850.9 / 0.99085 / 512/28/8/32 Not reached; highest measured Recall point: 194,866.9 / 0.97614 / 512/28/8/32 N/A

At batch 1, FlowANN achieves 74.3%, 63.6%, and 75.2% of two-GPU CAGRA QPS at the three targets. At batch 10,000, its median QPS is 2.4% and 0.9% higher at Recall 0.90 and 0.95. The 0.95 QPS ranges overlap, so the median difference does not establish a stable advantage. CAGRA did not reach 0.99 within the measured configurations; this does not imply that CAGRA cannot reach that Recall with other settings.

@ZhaoHaoRu
ZhaoHaoRu requested review from a team as code owners September 22, 2026 11:02
@copy-pr-bot

copy-pr-bot Bot commented Sep 22, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@cjnolet cjnolet added improvement Improves an existing functionality non-breaking Introduces a non-breaking change labels Sep 30, 2026
@cjnolet cjnolet moved this to In Progress in Unstructured Data Processing Sep 30, 2026
Comment thread c/include/cuvs/neighbors/flowann.h Outdated
} cuvsFlowannIndex;
typedef cuvsFlowannIndex* cuvsFlowannIndex_t;

CUVS_EXPORT cuvsError_t cuvsFlowannIndexParamsCreate(cuvsFlowannIndexParams_t* params);

@cjnolet cjnolet Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Rather than adding completely new API functions for flowann, let's find a way we can consolidate into existing APIs / options where possible. For e.g. if this is built on the CAGRA graph specifically, can we add an option / functions to the CAGRA APIs to enable this?

@ZhaoHaoRu
ZhaoHaoRu requested a review from a team as a code owner October 4, 2026 00:50
ZhaoHaoRu and others added 2 commits October 4, 2026 10:51
…U storage

Integrate FlowANN search, graph building, serialization, benchmarks, and API documentation. Add GDRCopy and CUDA-copy queue runtimes with regression coverage, along with C and Python bindings for dense FlowANN indexes.

Co-authored-by: Haoru Zhao <zhaohaoru@sjtu.edu.cn>
Co-authored-by: Jingkai He <hjk020101@sjtu.edu.cn>
Use CAGRA build, search and serialization entry points for tiered dense and
VPQ indexes across C++, C and Python. Own the reordered dataset and queue
runtime inside the index, and retain legacy loading in the native benchmark.

Internalize standalone FlowANN interfaces, update benchmark integration,
parameter documentation and generated API references, and add ownership,
serialization and optional-backend regression coverage.

Co-authored-by: Jingkai He <hjk020101@sjtu.edu.cn>
@ZhaoHaoRu
ZhaoHaoRu force-pushed the zhr/flowann-upstream branch from 377b285 to b260d2e Compare October 4, 2026 03:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

improvement Improves an existing functionality non-breaking Introduces a non-breaking change

Projects

Status: In Progress

Development

Successfully merging this pull request may close these issues.

2 participants