Conversation
…] in free-threaded builds
`ins1()`(used by `list.insert()`) and `list_ass_item_lock_held()` shift the list with one atomic release store per element in free-threaded builds. Those atomic stores prevent the compiler from vectorizing the loops, making them much slower than `memmove()` for large shifts. I propose using the existing `ptr_wise_atomic_memmove()` helper here. It uses `memmove()` as a fast path when the list is owned by the current thread and is not shared, and keeps the atomic element-wise stores for lists that other threads can observe.
The default (GIL) build keeps the current loops because the compiler can optimize them directly in the calling function, and they perform better in benchmarks.
pyperf, free-threaded release build (--disable-gil, -O3):
del l[0] 2.26x faster
del l[len(l)//2] 2.14x faster
l.insert(0, x) 1.95x faster
l.insert(len(l)//2, x) 1.86x faster
insert(0)+del[0], n=1000 2.86x faster
sliding window, n=1000 2.46x faster
Compiling both versions in GIL mode produces the same code for `ins1()` and `list_ass_item_lock_held()`, so the default build is not affected.
This continues pythongh-129069, which introduced `ptr_wise_atomic_memmove()`.
|
You are now making the new fragmented logic for the free-threaded build. How worth to maintain this? How many performance portion from this logic in general workload. |
We could add a macro for that pattern, or a small function which does it but i'm not sure the compiler would inline the call AND the loop anymore. |
Yeah in that case, would be fine, but I prefer to not adding free-threading only logic as possible unless it's really worth to do. |
Use Sphinx roles in the NEWS entry (:meth:`list.insert`, :keyword:`del`, :manpage:`memmove(3)`) and drop a comment in ins1() that repeated the rationale already given in the commit message.
|
@corona10 I agree that the branching does introduce some maintenance cost, but I think it's worths it. The point is that I think the best trade-off is to keep the existing compiler-optimized loop for the GIL build, while using |
Move the free-threaded/GIL split out of ins1() and list_ass_item_lock_held() into two small direction-specific list_shift_items_*() helpers, so the call sites stay platform-independent and each #ifdef lives in one place. The helpers are static inline and keep the original loop shapes; the generated code is unchanged in both build configurations.
No it will not be allowed. |
|
My concern about this kind of optimization is that we can not add all kinds of free-threaded specific operations because of the uncommon workload.(I am not saying this is uncommon workload) So that is why we need to prove this is worth doing. |
Benchmark resultsI benchmarked three variants:
All benchmarks were run with release builds ( 1. Free-threaded builds
The optimization consistently improves the affected operations, with larger gains for larger shifts. The benefit increases with list size. For 2. Realistic workloadsI also tested several patterns that perform list insertion/deletion in realistic ways.
3. GIL buildThe current implementation has no measurable performance regression in the GIL build. The GIL branch retains the existing implementation, so this PR does not introduce a new code path for GIL builds. I also tested the alternative of using This is not a strict improvement:
Using a separate implementation for free-threaded builds avoids regressing short-list workloads in the default GIL build while still providing the substantial benefit in free-threaded builds. 4. General-purpose benchmarksI ran a subset of For the free-threaded build, the geometric mean was 1.00x, with no significant overall regression. The GIL build with the unconditional ConclusionThe optimization provides
The additional maintenance cost is limited to two small bench_list_patterns.py |
Functions
ins1()(used bylist.insert()) andlist_ass_item_lock_held()shift the list with one atomic release store per element in free-threaded builds. Those atomic stores prevent the compiler from vectorizing the loops, making them much slower thanmemmove()for large shifts. I propose using the existingptr_wise_atomic_memmove()helper here. It usesmemmove()as a fast path when the list is owned by the current thread and is not shared, and keeps the atomic element-wise stores for lists that other threads can observe.The default (GIL) build keeps the current loops because the compiler can optimize them directly in the calling function, and they perform better in benchmarks.
pyperf, free-threaded release build (--disable-gil, -O3):
Compiling both versions in GIL mode produces the same code for
ins1()andlist_ass_item_lock_held(), so the default build is not affected.The benchmark following is produced by AI and verified by me.