Skip to content

fix: support boolean predicates in scalar segment scans - #93

Merged
yanghua merged 3 commits into
lance-format:mainfrom
Gabriel39:fix/scalar-segment-expressions
Sep 30, 2026
Merged

yanghua merged 3 commits into
lance-format:mainfrom
Gabriel39:fix/scalar-segment-expressions

Conversation

@Gabriel39

@Gabriel39 Gabriel39 commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

A scoped scalar scan currently selects only a single index leaf beneath AND. A short IN list rewritten to OR therefore falls back to scanning the entire requested fragment domain, even when every branch is supported by the selected segment.

Evaluate a scoped ScalarIndexExpr using Lance's boolean evaluator and a loader pinned to the selected physical segment. This supports AND, OR, IN, and NULL-aware NOT through the existing C/C++ API, with no ABI or dependency changes.

  • Retain necessary AND conditions, require candidates for both OR branches, and never negate a partially retained subtree.
  • Preserve explicit fragment boundaries, complete predicate rechecks, deletion handling, and filtering before LIMIT/OFFSET. Unsupported or inexact candidates retain the scoped scan fallback.
  • Filters containing IS [NOT] TRUE/FALSE use the scoped fallback before planning: the pinned planner loses their NULL semantics under negation. This guard can be removed after adopting fix(index): keep negated scalar index filters correct on NULL rows lance#9568. Ordinary equality negation remains indexed.
  • Bound expression selection to 128 nodes and depth 32 before evaluation. Reuse the opened segment for every leaf instead of loading a global logical index.
  • Report unknown candidate cardinality for complement masks instead of reporting zero candidates.

Validation: the new IN regression was observed failing on the original implementation with scalar_segment_fallback_no_driver=1. After the fix, the string IN regression reads four candidate rows instead of all eight rows in its fragment domain. Tests cover BTree/Bitmap, integer/string keys, NULL and empty results, stable row IDs, deleted rows, subset/partial fragment coverage, residual pagination, other indexed columns, unsafe negation, and expression budgets.

  • cargo fmt --check
  • cargo check --locked --all-targets
  • cargo clippy --locked --all-targets -- -D warnings
  • cargo test --locked: 456 tests passed
  • cargo test --locked --test compile_and_run_test -- --ignored: 3 tests passed on implementation commit 78626b1 (before the test/documentation follow-up)

Nullable-boolean regressions cover both index types and stable-row-ID modes, positive and negated truth tests, nested AND/OR, fragment scope, pagination, and continued acceleration of NOT (key = true). The review reproducer failed with missing NULL rows before the guard was added.

Array-label integration coverage:

  • Document the existing SQL/Substrait LabelList path, required list-element schema, and canonical Lance function names. No new ABI or production execution changes are needed.
  • Add a C API Substrait regression for List membership, AND/OR, all/any labels, exact candidate counts, NULL/empty/duplicate labels, residual filtering, pagination, fragment scope, partial coverage including unindexed rows, stable row IDs, and deletion handling.
  • Explain cross-engine boundaries: NULL membership and ordered-subsequence functions must not be mapped by name alone; the pinned Lance version also treats an empty all-label query as true on NULL lists.
  • The array-label follow-up passes all 456 Rust tests, cargo check, Clippy, formatting, and diff checks. Native consumer tests passed on the preceding implementation commit; this follow-up changes only tests and documentation.

lance-gatekeeper[bot]

This comment was marked as outdated.

@lance-gatekeeper lance-gatekeeper Bot added the K-changes Latest Gatekeeper recommendation requests changes. label Sep 30, 2026
@lance-gatekeeper lance-gatekeeper Bot removed the K-changes Latest Gatekeeper recommendation requests changes. label Sep 30, 2026
lance-gatekeeper[bot]

This comment was marked as outdated.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Sep 30, 2026
@lance-gatekeeper lance-gatekeeper Bot removed the K-approved Latest Gatekeeper recommendation permits acceptance. label Sep 30, 2026

@lance-gatekeeper lance-gatekeeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Gate recommendation: approve.

Segment-scoped AND/OR/IN and ordinary NOT preserve the explicit fragment domain and recheck the complete predicate before LIMIT/OFFSET. Boolean truth tests safely retain the scoped fallback for the pinned planner, preserving matching NULL rows.

The new Substrait LabelList regressions confirm label intersections and unions, NULL/empty semantics, partial coverage, pagination, and deletion handling with both row-ID modes. The documented cross-engine boundaries match the pinned Lance semantics. This follow-up adds tests and documentation without changing production execution or the ABI.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Sep 30, 2026

@yanghua yanghua left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@yanghua
yanghua merged commit cd63420 into lance-format:main Sep 30, 2026
10 checks passed
yiguolei pushed a commit to apache/doris that referenced this pull request Oct 8, 2026
…dex predicates (#68687)

### What problem does this PR solve?

Ordinary Lance scans currently retain array label membership predicates
in Doris, and scalar index planning misses usable boolean drivers. This
adds Substrait pushdown for built-in `array_contains` on a direct
`List<Utf8>` column with a non-NULL constant string, and for nonempty
constant-string `arrays_overlap` (including the normal Nereids rewrite
of three or more membership disjuncts). Positive indexable boolean
conditions can select a LabelList/BTree/Bitmap segment. OR requires
usable drivers on both branches of the same selected index. Literal
starts_with prefixes and LIKE filters with a usable leading prefix
remain eligible for BTree planning. OR branches requiring a refine
filter retain fragment scans. Missing FE field/index metadata and
competing logical indexes on one column preserve native automatic index
selection and fragment ownership. For example, two label conditions
joined by AND now reach Lance together and can reduce the index
candidate set before row materialization.

NULL needles, other array types, dynamic operands, UDFs, and
`array_contains_all` remain residual. Doris `array_contains_all` tests a
contiguous ordered subsequence, so translating it to Lance set
containment would change results. Each task still selects one logical
index; this change does not implement multi-column index intersection.
Fragment coverage, residual evaluation, and limit handling are
preserved. Complement-only predicates and expressions exceeding the
pinned native node/depth budget use parallel fragment scans with
scalar-index use disabled. An independent positive condition retains one
search per index segment even when another conjunct is a complement.
This avoids repeated segment-wide searches and prevents large overlap or
unindexable OR filters from serializing fallback scans.

Dependency:
[lance-format/lance-c#93](lance-format/lance-c#93),
merged in upstream main at `cd63420bfbe27f6f0a1edcc873b9191af7d52852`,
which is now the pinned dependency. The existing Foyer payload is
reapplied without implementation changes because
[lance-c#73](lance-format/lance-c#73) is still
open and supplies cache APIs used by Doris. Remove that remaining patch
after #73 merges. The build now records and validates the installed
Lance source/archive/patch fingerprint, detects incomplete publication,
and refuses stale external build output. Before removing an existing
installation, it verifies required rebuild inputs and rejects external
definitions that differ from the checkout. The third-party CI script job
runs the installation harness under the existing workflow path filters.
These third-party/build changes are synchronized to #68689 for master.
This PR targets `branch-4.1`.

Documentation: apache/doris-website#4184
(English and Chinese; draft pending implementation).

### Validation

- 70 focused FE JUnit tests passed after compiling the changed
converter/planner/scan sources and tests against an existing FE
dependency build. Coverage includes metadata fallbacks, competing
same-column indexes in both metadata orders, balanced wide conjunctions,
short IN expansion, literal/refined prefixes, nested OR guards, and
existing complement/overlap limits. The new regression assertions failed
before the fixes.
- FE Checkstyle passed with zero violations; the updated Groovy
regression suite passed syntax validation.
- Executed 12 plans produced by the actual FE converter and scan planner
through 30 pinned lance-c scans. Verified result rows and index counters
for unavailable field IDs, omitted index metadata, unequal same-column
index coverage, a 34-predicate conjunction, 62/63-label overlap combined
with two-value IN, literal prefixes, refined LIKE, and direct/nested OR
fallback. No segment fallback occurred; metadata/ambiguity cases
retained native index loads. Previous integration also covered fully
indexed, partially indexed, and unindexed datasets.
- The installation harness passed on master and branch-4.1 using
independent external source definitions: legacy/matching installs,
missing helper/vars/patch/downloader/builder, mismatched and
synchronized pin/patch updates, incomplete artifacts, stale builder
output, interrupted publication, and retry. Rejected preflight cases
verify that the entire installation and builder invocation count remain
unchanged. The missing-helper regression failed before the fix. Shell
syntax, workflow YAML, and CI path-filter checks passed.
- The unchanged pinned upstream main plus Foyer dependency previously
passed 75 Rust unit tests, 376 C API tests, and all 3 native
C/C++/static OSS consumer tests. Third-party
extraction/cache-refresh/patch-failure checks also passed on both
dependency definitions.
- SQL regressions assert pushed/residual EXPLAIN output, fragment versus
segment grouping, actual index searches/candidates/fallbacks, partial
coverage, 64/65-label boundaries, mixed complements, and selective
LIMIT. Full Doris compilation and distributed SQL execution on this
revision are pending CI.

### Release note

Support ordinary Lance array-label membership pushdown and safe boolean
scalar index planning.

### Check List (For Author)

- Test
    - [x] Regression test added (SQL execution pending CI)
    - [x] Unit Test
- [x] Manual integration test (FE-produced Substrait through lance-c,
described above)
- Behavior changed:
- [x] Yes. Compatible label predicates are evaluated in Lance; eligible
positive boolean drivers use scalar segments, while broad complements
and unsupported or oversized expressions retain fragment parallelism.
- Does this need documentation?
    - [x] Yes: apache/doris-website#4184

### Check List (For Reviewer who merge this PR)

- [ ] Confirm the release note
- [ ] Confirm test cases
- [ ] Confirm document
- [ ] Confirm the upstream dependency is ready
yiguolei pushed a commit to apache/doris that referenced this pull request Oct 8, 2026
…dex predicates (#68687)

### What problem does this PR solve?

Ordinary Lance scans currently retain array label membership predicates
in Doris, and scalar index planning misses usable boolean drivers. This
adds Substrait pushdown for built-in `array_contains` on a direct
`List<Utf8>` column with a non-NULL constant string, and for nonempty
constant-string `arrays_overlap` (including the normal Nereids rewrite
of three or more membership disjuncts). Positive indexable boolean
conditions can select a LabelList/BTree/Bitmap segment. OR requires
usable drivers on both branches of the same selected index. Literal
starts_with prefixes and LIKE filters with a usable leading prefix
remain eligible for BTree planning. OR branches requiring a refine
filter retain fragment scans. Missing FE field/index metadata and
competing logical indexes on one column preserve native automatic index
selection and fragment ownership. For example, two label conditions
joined by AND now reach Lance together and can reduce the index
candidate set before row materialization.

NULL needles, other array types, dynamic operands, UDFs, and
`array_contains_all` remain residual. Doris `array_contains_all` tests a
contiguous ordered subsequence, so translating it to Lance set
containment would change results. Each task still selects one logical
index; this change does not implement multi-column index intersection.
Fragment coverage, residual evaluation, and limit handling are
preserved. Complement-only predicates and expressions exceeding the
pinned native node/depth budget use parallel fragment scans with
scalar-index use disabled. An independent positive condition retains one
search per index segment even when another conjunct is a complement.
This avoids repeated segment-wide searches and prevents large overlap or
unindexable OR filters from serializing fallback scans.

Dependency:
[lance-format/lance-c#93](lance-format/lance-c#93),
merged in upstream main at `cd63420bfbe27f6f0a1edcc873b9191af7d52852`,
which is now the pinned dependency. The existing Foyer payload is
reapplied without implementation changes because
[lance-c#73](lance-format/lance-c#73) is still
open and supplies cache APIs used by Doris. Remove that remaining patch
after #73 merges. The build now records and validates the installed
Lance source/archive/patch fingerprint, detects incomplete publication,
and refuses stale external build output. Before removing an existing
installation, it verifies required rebuild inputs and rejects external
definitions that differ from the checkout. The third-party CI script job
runs the installation harness under the existing workflow path filters.
These third-party/build changes are synchronized to #68689 for master.
This PR targets `branch-4.1`.

Documentation: apache/doris-website#4184
(English and Chinese; draft pending implementation).

### Validation

- 70 focused FE JUnit tests passed after compiling the changed
converter/planner/scan sources and tests against an existing FE
dependency build. Coverage includes metadata fallbacks, competing
same-column indexes in both metadata orders, balanced wide conjunctions,
short IN expansion, literal/refined prefixes, nested OR guards, and
existing complement/overlap limits. The new regression assertions failed
before the fixes.
- FE Checkstyle passed with zero violations; the updated Groovy
regression suite passed syntax validation.
- Executed 12 plans produced by the actual FE converter and scan planner
through 30 pinned lance-c scans. Verified result rows and index counters
for unavailable field IDs, omitted index metadata, unequal same-column
index coverage, a 34-predicate conjunction, 62/63-label overlap combined
with two-value IN, literal prefixes, refined LIKE, and direct/nested OR
fallback. No segment fallback occurred; metadata/ambiguity cases
retained native index loads. Previous integration also covered fully
indexed, partially indexed, and unindexed datasets.
- The installation harness passed on master and branch-4.1 using
independent external source definitions: legacy/matching installs,
missing helper/vars/patch/downloader/builder, mismatched and
synchronized pin/patch updates, incomplete artifacts, stale builder
output, interrupted publication, and retry. Rejected preflight cases
verify that the entire installation and builder invocation count remain
unchanged. The missing-helper regression failed before the fix. Shell
syntax, workflow YAML, and CI path-filter checks passed.
- The unchanged pinned upstream main plus Foyer dependency previously
passed 75 Rust unit tests, 376 C API tests, and all 3 native
C/C++/static OSS consumer tests. Third-party
extraction/cache-refresh/patch-failure checks also passed on both
dependency definitions.
- SQL regressions assert pushed/residual EXPLAIN output, fragment versus
segment grouping, actual index searches/candidates/fallbacks, partial
coverage, 64/65-label boundaries, mixed complements, and selective
LIMIT. Full Doris compilation and distributed SQL execution on this
revision are pending CI.

### Release note

Support ordinary Lance array-label membership pushdown and safe boolean
scalar index planning.

### Check List (For Author)

- Test
    - [x] Regression test added (SQL execution pending CI)
    - [x] Unit Test
- [x] Manual integration test (FE-produced Substrait through lance-c,
described above)
- Behavior changed:
- [x] Yes. Compatible label predicates are evaluated in Lance; eligible
positive boolean drivers use scalar segments, while broad complements
and unsupported or oversized expressions retain fragment parallelism.
- Does this need documentation?
    - [x] Yes: apache/doris-website#4184

### Check List (For Reviewer who merge this PR)

- [ ] Confirm the release note
- [ ] Confirm test cases
- [ ] Confirm document
- [ ] Confirm the upstream dependency is ready
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

K-approved Latest Gatekeeper recommendation permits acceptance.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants