Repository navigation
fix: support boolean predicates in scalar segment scans - #93
Merged
yanghua merged 3 commits intoSep 30, 2026
Merged
Conversation
There was a problem hiding this comment.
✅ Gate recommendation: approve.
Segment-scoped AND/OR/IN and ordinary NOT preserve the explicit fragment domain and recheck the complete predicate before LIMIT/OFFSET. Boolean truth tests safely retain the scoped fallback for the pinned planner, preserving matching NULL rows.
The new Substrait LabelList regressions confirm label intersections and unions, NULL/empty semantics, partial coverage, pagination, and deletion handling with both row-ID modes. The documented cross-engine boundaries match the pinned Lance semantics. This follow-up adds tests and documentation without changing production execution or the ABI.
yiguolei
pushed a commit
to apache/doris
that referenced
this pull request
Oct 8, 2026
…dex predicates (#68687) ### What problem does this PR solve? Ordinary Lance scans currently retain array label membership predicates in Doris, and scalar index planning misses usable boolean drivers. This adds Substrait pushdown for built-in `array_contains` on a direct `List<Utf8>` column with a non-NULL constant string, and for nonempty constant-string `arrays_overlap` (including the normal Nereids rewrite of three or more membership disjuncts). Positive indexable boolean conditions can select a LabelList/BTree/Bitmap segment. OR requires usable drivers on both branches of the same selected index. Literal starts_with prefixes and LIKE filters with a usable leading prefix remain eligible for BTree planning. OR branches requiring a refine filter retain fragment scans. Missing FE field/index metadata and competing logical indexes on one column preserve native automatic index selection and fragment ownership. For example, two label conditions joined by AND now reach Lance together and can reduce the index candidate set before row materialization. NULL needles, other array types, dynamic operands, UDFs, and `array_contains_all` remain residual. Doris `array_contains_all` tests a contiguous ordered subsequence, so translating it to Lance set containment would change results. Each task still selects one logical index; this change does not implement multi-column index intersection. Fragment coverage, residual evaluation, and limit handling are preserved. Complement-only predicates and expressions exceeding the pinned native node/depth budget use parallel fragment scans with scalar-index use disabled. An independent positive condition retains one search per index segment even when another conjunct is a complement. This avoids repeated segment-wide searches and prevents large overlap or unindexable OR filters from serializing fallback scans. Dependency: [lance-format/lance-c#93](lance-format/lance-c#93), merged in upstream main at `cd63420bfbe27f6f0a1edcc873b9191af7d52852`, which is now the pinned dependency. The existing Foyer payload is reapplied without implementation changes because [lance-c#73](lance-format/lance-c#73) is still open and supplies cache APIs used by Doris. Remove that remaining patch after #73 merges. The build now records and validates the installed Lance source/archive/patch fingerprint, detects incomplete publication, and refuses stale external build output. Before removing an existing installation, it verifies required rebuild inputs and rejects external definitions that differ from the checkout. The third-party CI script job runs the installation harness under the existing workflow path filters. These third-party/build changes are synchronized to #68689 for master. This PR targets `branch-4.1`. Documentation: apache/doris-website#4184 (English and Chinese; draft pending implementation). ### Validation - 70 focused FE JUnit tests passed after compiling the changed converter/planner/scan sources and tests against an existing FE dependency build. Coverage includes metadata fallbacks, competing same-column indexes in both metadata orders, balanced wide conjunctions, short IN expansion, literal/refined prefixes, nested OR guards, and existing complement/overlap limits. The new regression assertions failed before the fixes. - FE Checkstyle passed with zero violations; the updated Groovy regression suite passed syntax validation. - Executed 12 plans produced by the actual FE converter and scan planner through 30 pinned lance-c scans. Verified result rows and index counters for unavailable field IDs, omitted index metadata, unequal same-column index coverage, a 34-predicate conjunction, 62/63-label overlap combined with two-value IN, literal prefixes, refined LIKE, and direct/nested OR fallback. No segment fallback occurred; metadata/ambiguity cases retained native index loads. Previous integration also covered fully indexed, partially indexed, and unindexed datasets. - The installation harness passed on master and branch-4.1 using independent external source definitions: legacy/matching installs, missing helper/vars/patch/downloader/builder, mismatched and synchronized pin/patch updates, incomplete artifacts, stale builder output, interrupted publication, and retry. Rejected preflight cases verify that the entire installation and builder invocation count remain unchanged. The missing-helper regression failed before the fix. Shell syntax, workflow YAML, and CI path-filter checks passed. - The unchanged pinned upstream main plus Foyer dependency previously passed 75 Rust unit tests, 376 C API tests, and all 3 native C/C++/static OSS consumer tests. Third-party extraction/cache-refresh/patch-failure checks also passed on both dependency definitions. - SQL regressions assert pushed/residual EXPLAIN output, fragment versus segment grouping, actual index searches/candidates/fallbacks, partial coverage, 64/65-label boundaries, mixed complements, and selective LIMIT. Full Doris compilation and distributed SQL execution on this revision are pending CI. ### Release note Support ordinary Lance array-label membership pushdown and safe boolean scalar index planning. ### Check List (For Author) - Test - [x] Regression test added (SQL execution pending CI) - [x] Unit Test - [x] Manual integration test (FE-produced Substrait through lance-c, described above) - Behavior changed: - [x] Yes. Compatible label predicates are evaluated in Lance; eligible positive boolean drivers use scalar segments, while broad complements and unsupported or oversized expressions retain fragment parallelism. - Does this need documentation? - [x] Yes: apache/doris-website#4184 ### Check List (For Reviewer who merge this PR) - [ ] Confirm the release note - [ ] Confirm test cases - [ ] Confirm document - [ ] Confirm the upstream dependency is ready
yiguolei
pushed a commit
to apache/doris
that referenced
this pull request
Oct 8, 2026
…dex predicates (#68687) ### What problem does this PR solve? Ordinary Lance scans currently retain array label membership predicates in Doris, and scalar index planning misses usable boolean drivers. This adds Substrait pushdown for built-in `array_contains` on a direct `List<Utf8>` column with a non-NULL constant string, and for nonempty constant-string `arrays_overlap` (including the normal Nereids rewrite of three or more membership disjuncts). Positive indexable boolean conditions can select a LabelList/BTree/Bitmap segment. OR requires usable drivers on both branches of the same selected index. Literal starts_with prefixes and LIKE filters with a usable leading prefix remain eligible for BTree planning. OR branches requiring a refine filter retain fragment scans. Missing FE field/index metadata and competing logical indexes on one column preserve native automatic index selection and fragment ownership. For example, two label conditions joined by AND now reach Lance together and can reduce the index candidate set before row materialization. NULL needles, other array types, dynamic operands, UDFs, and `array_contains_all` remain residual. Doris `array_contains_all` tests a contiguous ordered subsequence, so translating it to Lance set containment would change results. Each task still selects one logical index; this change does not implement multi-column index intersection. Fragment coverage, residual evaluation, and limit handling are preserved. Complement-only predicates and expressions exceeding the pinned native node/depth budget use parallel fragment scans with scalar-index use disabled. An independent positive condition retains one search per index segment even when another conjunct is a complement. This avoids repeated segment-wide searches and prevents large overlap or unindexable OR filters from serializing fallback scans. Dependency: [lance-format/lance-c#93](lance-format/lance-c#93), merged in upstream main at `cd63420bfbe27f6f0a1edcc873b9191af7d52852`, which is now the pinned dependency. The existing Foyer payload is reapplied without implementation changes because [lance-c#73](lance-format/lance-c#73) is still open and supplies cache APIs used by Doris. Remove that remaining patch after #73 merges. The build now records and validates the installed Lance source/archive/patch fingerprint, detects incomplete publication, and refuses stale external build output. Before removing an existing installation, it verifies required rebuild inputs and rejects external definitions that differ from the checkout. The third-party CI script job runs the installation harness under the existing workflow path filters. These third-party/build changes are synchronized to #68689 for master. This PR targets `branch-4.1`. Documentation: apache/doris-website#4184 (English and Chinese; draft pending implementation). ### Validation - 70 focused FE JUnit tests passed after compiling the changed converter/planner/scan sources and tests against an existing FE dependency build. Coverage includes metadata fallbacks, competing same-column indexes in both metadata orders, balanced wide conjunctions, short IN expansion, literal/refined prefixes, nested OR guards, and existing complement/overlap limits. The new regression assertions failed before the fixes. - FE Checkstyle passed with zero violations; the updated Groovy regression suite passed syntax validation. - Executed 12 plans produced by the actual FE converter and scan planner through 30 pinned lance-c scans. Verified result rows and index counters for unavailable field IDs, omitted index metadata, unequal same-column index coverage, a 34-predicate conjunction, 62/63-label overlap combined with two-value IN, literal prefixes, refined LIKE, and direct/nested OR fallback. No segment fallback occurred; metadata/ambiguity cases retained native index loads. Previous integration also covered fully indexed, partially indexed, and unindexed datasets. - The installation harness passed on master and branch-4.1 using independent external source definitions: legacy/matching installs, missing helper/vars/patch/downloader/builder, mismatched and synchronized pin/patch updates, incomplete artifacts, stale builder output, interrupted publication, and retry. Rejected preflight cases verify that the entire installation and builder invocation count remain unchanged. The missing-helper regression failed before the fix. Shell syntax, workflow YAML, and CI path-filter checks passed. - The unchanged pinned upstream main plus Foyer dependency previously passed 75 Rust unit tests, 376 C API tests, and all 3 native C/C++/static OSS consumer tests. Third-party extraction/cache-refresh/patch-failure checks also passed on both dependency definitions. - SQL regressions assert pushed/residual EXPLAIN output, fragment versus segment grouping, actual index searches/candidates/fallbacks, partial coverage, 64/65-label boundaries, mixed complements, and selective LIMIT. Full Doris compilation and distributed SQL execution on this revision are pending CI. ### Release note Support ordinary Lance array-label membership pushdown and safe boolean scalar index planning. ### Check List (For Author) - Test - [x] Regression test added (SQL execution pending CI) - [x] Unit Test - [x] Manual integration test (FE-produced Substrait through lance-c, described above) - Behavior changed: - [x] Yes. Compatible label predicates are evaluated in Lance; eligible positive boolean drivers use scalar segments, while broad complements and unsupported or oversized expressions retain fragment parallelism. - Does this need documentation? - [x] Yes: apache/doris-website#4184 ### Check List (For Reviewer who merge this PR) - [ ] Confirm the release note - [ ] Confirm test cases - [ ] Confirm document - [ ] Confirm the upstream dependency is ready
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A scoped scalar scan currently selects only a single index leaf beneath AND. A short
INlist rewritten to OR therefore falls back to scanning the entire requested fragment domain, even when every branch is supported by the selected segment.Evaluate a scoped
ScalarIndexExprusing Lance's boolean evaluator and a loader pinned to the selected physical segment. This supports AND, OR, IN, and NULL-aware NOT through the existing C/C++ API, with no ABI or dependency changes.IS [NOT] TRUE/FALSEuse the scoped fallback before planning: the pinned planner loses their NULL semantics under negation. This guard can be removed after adopting fix(index): keep negated scalar index filters correct on NULL rows lance#9568. Ordinary equality negation remains indexed.Validation: the new IN regression was observed failing on the original implementation with
scalar_segment_fallback_no_driver=1. After the fix, the string IN regression reads four candidate rows instead of all eight rows in its fragment domain. Tests cover BTree/Bitmap, integer/string keys, NULL and empty results, stable row IDs, deleted rows, subset/partial fragment coverage, residual pagination, other indexed columns, unsafe negation, and expression budgets.cargo fmt --checkcargo check --locked --all-targetscargo clippy --locked --all-targets -- -D warningscargo test --locked: 456 tests passedcargo test --locked --test compile_and_run_test -- --ignored: 3 tests passed on implementation commit 78626b1 (before the test/documentation follow-up)Nullable-boolean regressions cover both index types and stable-row-ID modes, positive and negated truth tests, nested AND/OR, fragment scope, pagination, and continued acceleration of
NOT (key = true). The review reproducer failed with missing NULL rows before the guard was added.Array-label integration coverage: