[python] Vectorize batch raw vector scoring - #9796
Conversation
JingsongLi
left a comment
There was a problem hiding this comment.
Requirement fit: SUPPORTED. Implementation: CLEAN.
Reviewed c22880574ba9. Raw batch fallback scores every row against each query, making repeated scalar dimension loops a real CPU cost. Reusing one owned float64 block and one active query scratch buffer improves that path while preserving scalar fallback, exact accumulation and deterministic Top-K behavior. I checked Arrow offsets/layouts, null/irregular vectors and non-finite/dimension validation.
Validation: 16 tests and 51 subtests passed, including exact scalar score/byte equality and bounded batch consumption. Current head CI is green. Kernel and end-to-end throughput measurements were not rerun locally.
No actionable implementation regression found in this review.
Purpose
Batch raw-vector fallback already streams each Arrow batch once, but it expands every vector into a Python list and scores every row/query pair through the scalar dimension loop. Convert each bounded regular FLOAT vector block to one owned float64 matrix, reuse that block across all queries, and keep only one active scratch matrix while updating the existing per-query Top-K heaps.
The vectorized reductions preserve the existing scalar accumulation semantics and Top-K tie breaking. Null, irregular, unsupported, non-finite, and dimension-mismatched inputs retain the scalar validation path.
Tests
python -m pytest -q pypaimon/tests/vector_scoring_test.py pypaimon/tests/batch_vector_raw_scan_test.py: 16 passed.git diff --checkpassed.Benchmark
macOS arm64, Python 3.9.6, NumPy 2.0.2, PyArrow 19.0.1. 8,192 FLOAT vectors x 128 dimensions, 16 queries, Top-K=10. Values are medians of three runs.
The per-query-vectorized ablation converts and scores the Arrow block separately for every query. The shared-block variant is this change: it converts the Arrow block once and reuses it across queries.
All variants produced identical scores and Top-K heaps. The Top-K measurement includes Arrow-to-score conversion, scoring, and heap maintenance; it excludes table I/O, planning, and indexed search. The temporary benchmark driver is not included in the change.