Skip to content

perf: avoid redundant permute+cont in non-interleaved RoPE - #2102

Merged
leejet merged 3 commits into
leejet:masterfrom
daniandtheweb:non-interleaved-rope
Oct 6, 2026
Merged

leejet merged 3 commits into
leejet:masterfrom
daniandtheweb:non-interleaved-rope

Conversation

@daniandtheweb

@daniandtheweb daniandtheweb commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This PR reduces the number of full ggml_cont passes in apply rope for the non-interleaved RoPE layout from 3 copies to 1.

After the single required H/L reorder, the two contiguous halves of the head dim (i, i + d_head/2) are read as views and the rotation is assembled with one ggml_concat. This eliminates 2 of the 3 full copies. The interleaved path is unchanged.

The output is identical to the previous implementation.

As for performance goes, on my RX 7800XT an Anima generation at 1024x1024, cfg 5 goes from 2.21 s/it to 2.15s/it.

This performance optimization possibility was found by a Pi agent running Qwen 3.8 27B locally.
The optimization was then made by me and the AI helped with the review.

Checklist

@leejet
leejet merged commit 9db0db6 into leejet:master Oct 6, 2026
@leejet

leejet commented Oct 6, 2026

Copy link
Copy Markdown
Owner

Thank you for your contribution.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants