Skip to content

[SPARK-58732][ML] Avoid MLlib vector conversion in Normalizer - #57955

Closed
zhengruifeng wants to merge 4 commits into
apache:masterfrom
zhengruifeng:SPARK-58732-normalizer-native-vector
Closed

[SPARK-58732][ML] Avoid MLlib vector conversion in Normalizer#57955
zhengruifeng wants to merge 4 commits into
apache:masterfrom
zhengruifeng:SPARK-58732-normalizer-native-vector

Conversation

@zhengruifeng

@zhengruifeng zhengruifeng commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

What changes were proposed in this pull request?

Update ML Normalizer to calculate the norm and normalize values with ML vectors directly. The normalization implementation follows the existing MLlib Normalizer transform method, including cloning dense and sparse value arrays and reusing sparse indices.

Why are the changes needed?

ML Normalizer currently converts every input vector from ML to MLlib and converts the normalized result back to ML. Avoiding those conversions reduces allocation and transform overhead.

Does this PR introduce any user-facing change?

No.

How was this patch tested?

The existing NormalizerSuite covers dense, sparse, zero-vector, and parameterized normalization behavior.

build/sbt 'mllib/testOnly org.apache.spark.ml.feature.NormalizerSuite'

Was this patch authored or co-authored using generative AI tooling?

Generated-by: Codex (GPT-5)

@zhengruifeng
zhengruifeng marked this pull request as draft August 12, 2026 08:31
@zhengruifeng
zhengruifeng marked this pull request as ready for review August 12, 2026 13:53
@uros-b

uros-b commented Aug 12, 2026

Copy link
Copy Markdown
Member

Thank you @zhengruifeng!

zhengruifeng added a commit that referenced this pull request Aug 12, 2026
### What changes were proposed in this pull request?

Update ML Normalizer to calculate the norm and normalize values with ML vectors directly. The normalization implementation follows the existing MLlib Normalizer transform method, including cloning dense and sparse value arrays and reusing sparse indices.

### Why are the changes needed?

ML Normalizer currently converts every input vector from ML to MLlib and converts the normalized result back to ML. Avoiding those conversions reduces allocation and transform overhead.

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

The existing NormalizerSuite covers dense, sparse, zero-vector, and parameterized normalization behavior.

build/sbt 'mllib/testOnly org.apache.spark.ml.feature.NormalizerSuite'

### Was this patch authored or co-authored using generative AI tooling?

Generated-by: Codex (GPT-5)

Closes #57955 from zhengruifeng/SPARK-58732-normalizer-native-vector.

Authored-by: Ruifeng Zheng <ruifengz@apache.org>
Signed-off-by: Ruifeng Zheng <ruifengz@foxmail.com>
(cherry picked from commit cddd9ae)
Signed-off-by: Ruifeng Zheng <ruifengz@foxmail.com>
@zhengruifeng

Copy link
Copy Markdown
Contributor Author

Merge Summary:

Posted by merge_spark_pr.py

@zhengruifeng
zhengruifeng deleted the SPARK-58732-normalizer-native-vector branch August 12, 2026 23:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants