Skip to content

gh-158790: Improve performance for list.insert() and del list[i] free-threaded builds - #158791

Open
hetaozdh wants to merge 3 commits into
python:mainfrom
hetaozdh:list-shift-memmove
Open

hetaozdh wants to merge 3 commits into
python:mainfrom
hetaozdh:list-shift-memmove

Conversation

@hetaozdh

@hetaozdh hetaozdh commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Functions ins1()(used by list.insert()) and list_ass_item_lock_held() shift the list with one atomic release store per element in free-threaded builds. Those atomic stores prevent the compiler from vectorizing the loops, making them much slower than memmove() for large shifts. I propose using the existing ptr_wise_atomic_memmove() helper here. It uses memmove() as a fast path when the list is owned by the current thread and is not shared, and keeps the atomic element-wise stores for lists that other threads can observe.

The default (GIL) build keeps the current loops because the compiler can optimize them directly in the calling function, and they perform better in benchmarks.

pyperf, free-threaded release build (--disable-gil, -O3):

del l[0]                 2.26x faster
del l[len(l)//2]         2.14x faster
l.insert(0, x)           1.95x faster
l.insert(len(l)//2, x)   1.86x faster
insert(0)+del[0], n=1000 2.86x faster
sliding window, n=1000   2.46x faster

Geometric mean           2.23x faster

Compiling both versions in GIL mode produces the same code for ins1() and list_ass_item_lock_held(), so the default build is not affected.

The benchmark following is produced by AI and verified by me.

#!/usr/bin/env python3
"""pyperf benchmark for list.insert() / del list[i] element shifting.

Usgae:
    PYTHONPATH=... <interpreter> bench_list_shift.py -o result.json -p 1 -n 10
    python -m pyperf compare_to base.json patched.json --table --table-format md

Each benchmark function performs one workload iteration; pyperf calibrates
the number of loops per value.  Results are compared with
``pyperf compare_to`` in the accompanying report.
"""

import pyperf


def bench_del_front(n=100_000, k=2_000):
    lst = list(range(n))
    for _ in range(k):
        del lst[0]


def bench_del_mid(n=100_000, k=2_000):
    lst = list(range(n))
    for _ in range(k):
        del lst[len(lst) // 2]


def bench_insert_front(n=100_000, k=2_000):
    lst = list(range(n))
    for _ in range(k):
        lst.insert(0, None)


def bench_insert_mid(n=100_000, k=2_000):
    lst = list(range(n))
    for _ in range(k):
        lst.insert(len(lst) // 2, None)


def _small(n, k):
    # insert(0) and del[0] both shift ~n elements; the size stays n
    lst = list(range(n))
    for _ in range(k):
        lst.insert(0, None)
        del lst[0]


def bench_small_n8(n=8, k=200_000):
    _small(n, k)


def bench_small_n16(n=16, k=200_000):
    _small(n, k)


def bench_small_n32(n=32, k=200_000):
    _small(n, k)


def bench_small_n1000(n=1000, k=20_000):
    _small(n, k)


def bench_sliding_window(n=1000, k=50_000):
    # deque-like pattern: push to the front, pop from the back
    lst = list(range(n))
    for i in range(k):
        lst.insert(0, i)
        lst.pop()


def bench_control_del_last(n=100_000):
    # no shift: exercise the unchanged "delete last element" guard path
    lst = list(range(n))
    for _ in range(n):
        del lst[-1]


def bench_control_append(n=1_000_000):
    # no shift: unrelated fast path, should stay flat
    lst = []
    append = lst.append
    for i in range(n):
        append(i)


BENCHMARKS = [
    ("del_front_n100k_k2k", bench_del_front),
    ("del_mid_n100k_k2k", bench_del_mid),
    ("insert_front_n100k_k2k", bench_insert_front),
    ("insert_mid_n100k_k2k", bench_insert_mid),
    ("small_n8_k200k", bench_small_n8),
    ("small_n16_k200k", bench_small_n16),
    ("small_n32_k200k", bench_small_n32),
    ("small_n1000_k20k", bench_small_n1000),
    ("sliding_window_n1000_k50k", bench_sliding_window),
    ("control_del_last_n100k", bench_control_del_last),
    ("control_append_n1M", bench_control_append),
]


if __name__ == "__main__":
    runner = pyperf.Runner()
    for name, func in BENCHMARKS:
        runner.bench_func(name, func)

…] in free-threaded builds

`ins1()`(used by `list.insert()`) and `list_ass_item_lock_held()` shift the list with one atomic release store per element in free-threaded builds.  Those atomic stores prevent the compiler from vectorizing the loops, making them much slower than `memmove()` for large shifts.  I propose using the existing `ptr_wise_atomic_memmove()` helper here. It uses `memmove()` as a fast path when the list is owned by the current thread and is not shared, and keeps the atomic element-wise stores for lists that other threads can observe.

The default (GIL) build keeps the current loops because the compiler can optimize them directly in the calling function, and they perform better in benchmarks.

pyperf, free-threaded release build (--disable-gil, -O3):

    del l[0]                 2.26x faster
    del l[len(l)//2]         2.14x faster
    l.insert(0, x)           1.95x faster
    l.insert(len(l)//2, x)   1.86x faster
    insert(0)+del[0], n=1000 2.86x faster
    sliding window, n=1000   2.46x faster

Compiling both versions in GIL mode produces the same code for `ins1()` and `list_ass_item_lock_held()`, so the default build is not affected.

This continues pythongh-129069, which introduced `ptr_wise_atomic_memmove()`.
@corona10

corona10 commented Oct 4, 2026

Copy link
Copy Markdown
Member

You are now making the new fragmented logic for the free-threaded build. How worth to maintain this? How many performance portion from this logic in general workload.

Comment thread Misc/NEWS.d/next/Core_and_Builtins/2026-10-04-01-40-31.gh-issue-158790.F3kQ7z.rst Outdated
Comment thread Objects/listobject.c Outdated
@picnixz

picnixz commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

You are now making the new fragmented logic for the free-threaded build. How worth to maintain this? How many performance portion from this logic in general workload.

We could add a macro for that pattern, or a small function which does it but i'm not sure the compiler would inline the call AND the loop anymore.

@corona10

corona10 commented Oct 4, 2026

Copy link
Copy Markdown
Member

We could add a macro for that pattern, or a small function which does it but i'm not sure the compiler would inline the call AND the loop anymore.

Yeah in that case, would be fine, but I prefer to not adding free-threading only logic as possible unless it's really worth to do.

Use Sphinx roles in the NEWS entry (:meth:`list.insert`,
:keyword:`del`, :manpage:`memmove(3)`) and drop a comment in ins1()
that repeated the rationale already given in the commit message.
@hetaozdh

hetaozdh commented Oct 4, 2026

Copy link
Copy Markdown
Contributor Author

@corona10 I agree that the branching does introduce some maintenance cost, but I think it's worths it.

The point is that memmove() is already a better implementation for large shifts in general. In the GIL build, replacing the existing loop with memmove() makes these benchmarks only about 10% slower overall (geometric mean). In the free-threaded build, it makes them 2.23x faster overall.

I think the best trade-off is to keep the existing compiler-optimized loop for the GIL build, while using ptr_wise_atomic_memmove() for the free-threaded build. The helper also retains the atomic element-wise path for shared lists and uses memmove() only when the list is owned by the current thread and is not shared.

Move the free-threaded/GIL split out of ins1() and
list_ass_item_lock_held() into two small direction-specific
list_shift_items_*() helpers, so the call sites stay
platform-independent and each #ifdef lives in one place.

The helpers are static inline and keep the original loop shapes; the
generated code is unchanged in both build configurations.
@corona10

corona10 commented Oct 5, 2026

Copy link
Copy Markdown
Member

In the GIL build, replacing the existing loop with memmove() makes these benchmarks only about 10% slower overall (geometric mean)

No it will not be allowed.

@corona10

corona10 commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

My concern about this kind of optimization is that we can not add all kinds of free-threaded specific operations because of the uncommon workload.(I am not saying this is uncommon workload) So that is why we need to prove this is worth doing.

@hetaozdh

hetaozdh commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Benchmark results

I benchmarked three variants:

  • baseline: current implementation on the merge base
  • N: this PR — ptr_wise_atomic_memmove() for free-threaded builds, original loop for GIL builds
  • A: ptr_wise_atomic_memmove() used for both GIL and free-threaded builds

All benchmarks were run with release builds (-O3 -DNDEBUG) on Apple Silicon arm64.

1. Free-threaded builds

Operation Speedup
del l[0], n=100k 2.17x
del l[len//2], n=100k 2.19x
insert(0, x), n=100k 1.94x
insert(mid, x), n=100k 1.85x
insert(0) + del[0], n=1000 2.67x
sliding window, n=1000 2.29x

The optimization consistently improves the affected operations, with larger gains for larger shifts.

The benefit increases with list size. For insert(0) + del[0], the free-threaded build reaches 2.85x at n=4096, with the crossover occurring around n=16.

2. Realistic workloads

I also tested several patterns that perform list insertion/deletion in realistic ways.

Workload Speedup
bisect.insort() into sorted list 2.04x
top-K sorted window 1.18x
list-based MRU window 2.42x
queue drained from the front 2.91x
prepending output lines 2.55x

3. GIL build

The current implementation has no measurable performance regression in the GIL build. The GIL branch retains the existing implementation, so this PR does not introduce a new code path for GIL builds.

I also tested the alternative of using memmove in the GIL build as well.

This is not a strict improvement:

  • it is up to ~30% slower for short lists;
  • the benefit only appears for sufficiently large lists;
    • for example, insert(0) + del[0] is 0.87x at n=32, but 1.20x at n=4096.

Using a separate implementation for free-threaded builds avoids regressing short-list workloads in the default GIL build while still providing the substantial benefit in free-threaded builds.

4. General-purpose benchmarks

I ran a subset of pyperformance benchmarks covering general-purpose and stdlib workloads.

For the free-threaded build, the geometric mean was 1.00x, with no significant overall regression.

The GIL build with the unconditional memmove variant also showed no significant regression in the general-purpose benchmark subset. The trade-off is primarily visible in the short-list cases described above.

Conclusion

The optimization provides

  • 1.85x–2.67x speedups on the directly affected operations;
  • 1.18x–2.91x speedups on the tested realistic workloads;
  • no measurable cost in the default GIL implementation;
  • using memmove unconditionally would introduce regressions for short lists;
  • no regression in the general-purpose pyperformance subset;
  • all relevant tests pass.

The additional maintenance cost is limited to two small static inline helpers and the #ifdef, while reusing the existing atomic bulk-move primitive.

bench_list_patterns.py
bench_list_shift_sweep.py
bench_list_shift.py

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting review type-feature A feature request or enhancement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants