Skip to content

LazyImportType.resolve() does not replace the module global and can import twice #152298

Description

@1sgtpepper

Bug report

Bug description:

Calling resolve() on a lazy import proxy forces the import and returns the resolved object, but it does not replace the original lazy object in the module globals. A later normal global access reifies the same lazy object again.

import builtins
import types

real_import = builtins.__import__
calls = []

lazy import target_module as target

def custom_import(name, globals=None, locals=None, fromlist=None, level=0):
    if name == "target_module":
        index = len(calls) + 1
        calls.append((index, name, globals.get("__name__"), fromlist))
        module = types.ModuleType(name)
        module.VALUE = f"value-{index}"
        return module
    return real_import(name, globals, locals, fromlist, level)

builtins.__import__ = custom_import
try:
    def resolve_from_dict():
        lazy_obj = globals()["target"]
        resolved = lazy_obj.resolve()
        print("resolved:", resolved.VALUE)
        print("global after resolve:", type(globals()["target"]).__name__)

    resolve_from_dict()

    print("direct:", target.VALUE)
    print("global after direct:", type(globals()["target"]).__name__)
    print("calls:", calls)
finally:
    builtins.__import__ = real_import

Expected: one import, with the module global replaced after resolve().

Actual on current main:

resolved: value-1
global after resolve: lazy_import
direct: value-2
global after direct: module
calls: [(1, 'target_module', '__main__', None), (2, 'target_module', '__main__', None)]

I confirmed this across tier1, tier2, and tier2 without uop optimization:
https://github.com/Kuhai9801/cpython/actions/runs/28243825007

CPython versions tested on:

CPython main branch

Operating systems tested on:

Linux

Linked PRs

Activity

  1. 1sgtpepper commented on Jun 26, 2026

    @1sgtpepper
    Author

    I'm working on a fix for this.

  2. added 12 commits that reference this issue on Jun 26, 2026
  3. 17 remaining items

  4. added 9 commits that reference this issue on Jul 4, 2026
  5. brittanyrey commented on Sep 18, 2026

    @brittanyrey
    Contributor

    Putting up a new PR to try to address this. #157730

  6. encukou commented on Sep 28, 2026

    @encukou
    Member

    Hmm, is this a bug? I thought that since we have a cache in sys.modules, repeated imports are (intended to be) fast, and so resolve() doing a re-import is not a problem.
    In the same way, it's not a problem that with lazy from module import a, b, c we can eventually import module three times (unlike the non-lazy variant). Is it?

  7. brittanyrey commented on Oct 7, 2026

    @brittanyrey
    Contributor

    @encukou I did a little digging to try and quantify if addressing this change actually provides value and it looks like the case where it helps the most is if resolve() is called repeatedly (which may be a scenario that is most likely hit during a stress test).

    benchmark numbers

    Benchmark Base Average PR Average Verdict
    reify_import — first access of 2,000 lazy import bindings 0.795 ms 0.773 ms within noise (+0.1–0.3% instructions)
    reify_from — first access of 2,000 lazy from bindings 0.270 ms 0.271 ms within noise (+0.1–0.2% instructions)
    steady_global — 1M reads of a resolved lazy global 17.654 ms 17.374 ms within noise
    resolve_repeat — 100k proxy.resolve() calls on one proxy 26.390 ms 7.502 ms 3.5× faster
    startup_argparse — import argparse + parse (subprocess) 23.902 ms 23.846 ms within noise

    plausible-ish scenario where the improvement could come into play

    import annotationlib
    lazy import json
    
    # A handler signature where one annotation names a type that only exists for type checkers.
    def handler(request: TypeCheckingOnly, decoder: json.JSONDecoder) -> None: ...
    
    # e.g. a web framework, CLI library or DI container inspecting handlers repeatedly
    for _ in range(1000):
        annotationlib.get_annotations(handler, format=annotationlib.Format.FORWARDREF)

    Without the PR, we materialize json via resolve 1000x and with the change, it happens once and is cached.
    Granted for this codepath to even trigger, we need a function with mix of annotations that can and cannot be accessed at runtime to trigger the ForwardRef codepath, we need json to not be loaded prior to being used as an annotation, and we need the codepath (handler) in this case to be hit again and again...and even in that case, the raw time savings isn't huge...it just feels "more correct" to load it once than on every single hit.

    With all that said, I'm glad to drop the PR + we can close the issue if it seems mostly moot and just added complexity for not enough benefit. Just wanted to take a holistic pass on making sense around if this is worth addressing

  8. encukou commented on Oct 8, 2026

    @encukou
    Member

    Thank you!

    I read that as “it's possible to deliberately bypass caches”. Adding a third one on top of sys.modules & _PyLazyImport_Reify/_PyDict_ReplaceItemIf still seems unnecessary to me.

    In your scenario, the get_annotations caller should add a cache. The function's docs describe an expensive uncached operation, even if they aren't explicit about it.

  9. brittanyrey commented on Oct 9, 2026

    @brittanyrey
    Contributor

    That makes sense, we can handle this corner case at the application level.
    Want to close the issue + PR?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions