Skip to content

Migrate GeometricWidthDiscretiser.fit() to narwhals, add polars support - #1041

Open
solegalli wants to merge 2 commits into
narwhals-migrationfrom
narwhals-geometric-width-discretiser
Open

Migrate GeometricWidthDiscretiser.fit() to narwhals, add polars support#1041
solegalli wants to merge 2 commits into
narwhals-migrationfrom
narwhals-geometric-width-discretiser

Conversation

@solegalli

Copy link
Copy Markdown
Collaborator

Migrates GeometricWidthDiscretiser.fit() to narwhals with polars support.

fit()'s only pandas dependency was X[var].min()/.max() to compute the geometric progression's anchors — everything downstream (np.power/np.r_/np.sort bin-edge math) was already plain numpy. Replaced the pandas indexing with nw.from_native(X, eager_only=True).get_column(var).min()/.max(), which returns a numpy/python float scalar on both backends and feeds np.power identically.

Merge vs split: benchmarked old pandas-native fit() vs narwhals-on-pandas / narwhals-on-polars at 10k/50k/100k rows × 1/2/10 cols (min/max dominate cost; bin-edge math is O(bins) not O(n)):

  • narwhals-on-pandas: 1.0–1.3x at realistic sizes (the 1.8x at 10k/1-col is sub-ms fixed overhead) — minimal loss, single narwhals path, no is_pandas split.
  • narwhals-on-polars: ~0.35–0.7x (1.4–2.8x faster).

Verified new fit() bin edges against the old pandas implementation across edge cases (skewed/normal/mixed-sign distributions, two-point range, min == max degenerate case) on both backends — exact equality. Cross-checked full fit_transform() (both return_object and return_boundaries) between pandas and polars — identical.

Tests: dataframe-touching tests in test_geometric_width_discretiser.py parametrized over pd.DataFrame/pl.DataFrame, replacing pandas-only fixtures with local dicts. Init-only param-validation tests unchanged.

Verified: tests/test_discretisation — 114 passed, same 5 pre-existing check_estimator failures. flake8 / mypy clean, sphinx -W clean. User-guide worked example re-verified against real output; "With polars" section added.


Stacked on narwhals-discretisation-base (its own PR). Until that merges this PR's diff also contains the shared BaseDiscretiser commit; review that one first.

solegalli and others added 2 commits August 25, 2026 17:08
Shared base for ArbitraryDiscretiser, EqualFrequencyDiscretiser,
EqualWidthDiscretiser and GeometricWidthDiscretiser (not
DecisionTreeDiscretiser, which extends a different base). Only
transform() needed migrating - _fit_setup(), _get_feature_names_in()
and _check_transform_input_and_state() are inherited unchanged from
BaseNumericalTransformer, already fully narwhals-migrated.

transform()'s only pandas dependency was pd.cut, applied per column to
sort values into the bins already fixed by fit() (binner_dict_).
Replaced it with a plain numpy implementation: pandas.cut is itself
built on bins.searchsorted() internally (verified against pandas 3.0's
_bins_to_cuts source), so np.searchsorted + the same include_lowest
index-1 special case reproduces its bin-index logic exactly, with no
per-backend branch needed - values come from
nw_X.get_column(feature).to_numpy() regardless of backend, and results
are re-attached via nw.new_series()/with_columns(), so the same code
path runs for pandas and polars.

Benchmarked old pd.cut vs the new numpy+narwhals path at 10k/50k/100k
rows x 1/2/10 columns:
- return_boundaries=False (bin codes): narwhals-on-pandas lands at
  ~1.0-1.2x of pandas-native at realistic sizes (50k-100k rows, the
  ~1.9x seen only at the smallest 10k-row/1-col case is fixed
  per-call overhead, sub-millisecond either way) - minimal loss,
  merged into a single path, no is_pandas split. narwhals-on-polars is
  ~1.0-1.3x *faster* than pandas-native at every size tested.
- return_boundaries=True (interval-label strings): the numpy path is
  12-20x faster than pd.cut on pandas itself (e.g. 100k rows x 10
  cols: 647ms old vs 40ms new) - pd.cut's Categorical/IntervalIndex
  machinery has heavy per-call overhead that np.searchsorted plus
  plain string formatting avoids entirely. polars is ~1.2x faster
  still than the new pandas path.
Given both branches favour or are at parity with a single numpy-driven
path, there was no case for a pandas fast-path split here.

return_boundaries=True's interval-label formatting
("(lower, upper]" text, e.g. "(-0.001, 20.0]") replicates pandas.cut's
_round_frac/_infer_precision/lowest-edge-adjustment algorithm in pure
numpy so it works identically on both backends - verified against real
pd.cut(...).astype(str) output across positive/negative/duplicate-
inducing/inf-edge bins, and against the California housing dataset
used in the existing test. return_object=True now builds a nw.Object
column (narwhals' cross-backend equivalent of pandas' "O" dtype,
already used by variable_handling for categorical-column detection)
instead of a pandas-only astype("O") call.

Verified: tests/test_discretisation full suite unchanged (109 passed,
5 pre-existing failures in test_check_estimator_discretisers.py -
sklearn's check_estimator feeds raw numpy arrays, which check_X() has
rejected since the narwhals migration's dataframe-only contract;
reproduced identically on the unmodified file). Manually diffed
transform() output against real pd.cut() across ~10 edge cases (NaN,
out-of-range values on both ends, negative bins, exact-edge values,
precision auto-widening, single bin) plus the three sibling
discretisers' documented doctest examples (EqualWidthDiscretiser,
ArbitraryDiscretiser, EqualFrequencyDiscretiser value_counts()) -
all numerically identical to old pd.cut output; the "Name: x" vs
"Name: count" and bare-fit()-repr mismatches those doctests already
show are a pre-existing pandas-3.0 doc-staleness issue unrelated to
this migration (reproduced on the unmodified files too). flake8 and
mypy clean. Module imports with pandas blocked (loaded standalone,
since sibling discretiser files in this package are not yet migrated
and still import pandas at their own module level). sphinx -W build
clean (only the pre-existing unrelated linkcode_resolve warning).

test_base_discretizer.py's test_transform is now parametrized over
pd.DataFrame/pl.DataFrame per AGENTS.md - its MockClassFit hard-codes
binner_dict_ rather than actually fitting, so it needed no pandas-only
logic to begin with. The other four discretisers' own test files stay
pandas-only for now: their fit() methods still call pd.cut/pd.qcut
directly and aren't migrated by this branch.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
fit()'s only pandas dependency was X[var].min()/.max() to compute the
geometric progression's min/max anchors - everything downstream (the
np.power/np.r_/np.sort bin-edge math) was already plain numpy and
needed no changes. Replaced the pandas indexing with
nw.from_native(X, eager_only=True).get_column(var).min()/.max(),
which returns a numpy/python float scalar on both backends and feeds
np.power identically either way.

Benchmarked old pandas-native fit() vs the new narwhals-on-pandas and
narwhals-on-polars paths at 10k/50k/100k rows x 1/2/10 columns (200
iterations each, min/max dominate cost either way since bin-edge math
is O(bins) not O(n)):
- narwhals-on-pandas: 1.0-1.3x of pandas-native at realistic sizes
  (50k-100k rows); the 1.8x seen only at the smallest 10k-row/1-col
  case is sub-millisecond fixed per-call overhead. Minimal loss -
  merged into a single narwhals path, no is_pandas split.
- narwhals-on-polars: ~0.35-0.7x of pandas-native (i.e. 1.4-2.8x
  *faster*), consistent with the sibling BaseDiscretiser.transform()
  migration finding polars faster at every size tested.

Verified: diffed new fit() bin edges against the old pandas
implementation across edge cases (skewed/normal/negative-and-positive
distributions, two-point range, and the min==max degenerate case) on
both backends - numerically identical (exact equality, not just
close). Cross-checked full fit_transform() (both return_object and
return_boundaries combinations) between pandas and polars inputs -
identical output values. Manually reran the GeometricWidthDiscretiser
user guide's house_prices worked example (binner_dict_ and interval
width numbers) against real output to confirm the docs still match
current behaviour (the precision example there was already fixed in
#986, prior to this branch) before adding a new "With polars" section
with verified output.

tests/test_discretisation/test_geometric_width_discretiser.py: the
dataframe-touching tests are now parametrized over
pd.DataFrame/pl.DataFrame per AGENTS.md, replacing the pandas-only
df_normal_dist/df_na/df_vartypes fixtures with local dicts so the same
input produces and asserts the same output on both backends (bin
edges, transform values via narwhals-agnostic extraction, dtype
checks, and NA-error cases). Init-only param-validation tests are
unchanged since they never touch a dataframe.

flake8 and mypy clean. Module imports with pandas blocked (loaded
standalone, since sibling discretiser files in this package aren't
migrated yet and still import pandas at their own module level).
sphinx -W build clean (only the pre-existing unrelated
linkcode_resolve warning). Full tests/test_discretisation suite: 114
passed, same 5 pre-existing failures as the unmodified base branch
(test_check_estimator_discretisers.py - sklearn's check_estimator
feeds raw numpy arrays, rejected by check_X()'s dataframe-only
contract since the narwhals migration; unrelated to this change).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant