Skip to content

[fix](function) Split UTF-8 strings correctly with empty regexp - #68020

Open
Mryange wants to merge 1 commit into
apache:masterfrom
Mryange:fix-split-regexp-utf8
Open

Mryange wants to merge 1 commit into
apache:masterfrom
Mryange:fix-split-regexp-utf8

Conversation

@Mryange

@Mryange Mryange commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

split_by_regexp advanced one byte at a time when the pattern was empty, which split multibyte UTF-8 characters into invalid array elements and could truncate a character when a limit was specified. Root cause: the empty-pattern path bypassed RE2 and treated each byte as a complete character. The implementation now advances by the current UTF-8 code point length while preserving the existing byte-based behavior for ASCII input. Regression coverage verifies Chinese and emoji input with and without a split limit.

Release note

None

Check List (For Author)

  • Test

    • Regression test
    • Unit Test
    • Manual test (add detailed scripts or steps below)
    • No need to test or manual test. Explain why:
      • This is a refactor/code format and no logic has been changed.
      • Previous test can cover this change.
      • No code files have been changed.
      • Other reason
  • Behavior changed:

    • No.
    • Yes.
  • Does this need documentation?

    • No.
    • Yes.

Check List (For Reviewer who merge this PR)

  • Confirm the release note
  • Confirm test cases
  • Confirm document
  • Add branch pick label

### What problem does this PR solve?

Issue Number: N/A

Problem Summary: split_by_regexp advanced one byte at a time for an empty pattern, so multibyte UTF-8 characters were split into invalid array elements. Advance by the current UTF-8 code point length and add regression coverage with and without a split limit.

### Release note

Fix split_by_regexp with an empty pattern to split UTF-8 strings by character instead of by byte.

### Check List (For Author)

- Test: Regression test
    - RELEASE BE build and targeted test_split_by_regexp regression
- Behavior changed: Yes; empty-pattern splitting now preserves UTF-8 character boundaries.
- Does this need documentation: No
@hello-stephen

Copy link
Copy Markdown
Contributor

Thank you for your contribution to Apache Doris.
Don't know what should be done next? See How to process your PR.

Please clearly describe your PR:

  1. What problem was fixed (it's best to include specific error reporting information). How it was fixed.
  2. Which behaviors were modified. What was the previous behavior, what is it now, why was it modified, and what possible impacts might there be.
  3. What features were added. Why was this function added?
  4. Which code was refactored and why was this part of the code refactored?
  5. Which functions were optimized and what is the difference before and after the optimization?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants