Skip to content

Qwen Image #758

Description

@yggdrasil75

A request to add Qwen Image support

Activity

  1. Seas0 commented on Aug 9, 2025

    @Seas0
    Contributor

    I'll investigate if I have free time. A LOT OF CODE COPYING from the Qwen2.5-vl model seems to be necessary due to the complexity of the visual & textual encoder, maybe we shall copy these things or import them from the llama.cpp main repo and integrate better with them?

  2. stduhpf commented on Aug 9, 2025

    @stduhpf
    Contributor

    Yes this kind of things have been discussed before, it seems that including llama.cpp as a submodule in the project would be the way to do this, or maybe including this project into llama.cpp.

  3. Seas0 commented on Aug 9, 2025

    @Seas0
    Contributor

    Yes this kind of things have been discussed before, it seems that including llama.cpp as a submodule in the project would be the way to do this, or maybe including this project into llama.cpp.

    Any reference for discussion context?

  4. stduhpf commented on Aug 9, 2025

    @stduhpf
    Contributor

    Yes this kind of things have been discussed before, it seems that including llama.cpp as a submodule in the project would be the way to do this, or maybe including this project into llama.cpp.

    Any reference for discussion context?

    #653

  5. henk717 commented on Aug 13, 2025

    @henk717

    @LostRuins considering you have experience in integrating them might be userful to chime in.

  6. LostRuins commented on Aug 13, 2025

    @LostRuins
    Contributor

    I didn't do the implementation though, I only have a minimal understanding of how it works, mostly using the existing clip api. ngxson is the one you want.

    having said that, a submodule is often very messy as it becomes hard to track changes to files. I would recommend extracting out the files you need and modifying them instead.

  7. stduhpf commented on Aug 13, 2025

    @stduhpf
    Contributor

    a submodule is often very messy as it becomes hard to track changes to files. I would recommend extracting out the files you need and modifying them instead.

    Doesn't that make tracking upstream changes even harder?

  8. LostRuins commented on Aug 13, 2025

    @LostRuins
    Contributor

    Doesn't that make tracking upstream changes even harder?

    In some ways yes, but the idea is that only the necessary code for the vision/text encoders is ported out (in like one or two files) and handled independently - it essentially becomes part of sd.cpp's code base and frozen in it's current functional state for the purposes of qwen image.

    I think that the overlap between the rest of llama.cpp's code and stable-diffusion.cpp's code is sufficiently small that you don't really want to grab all the other stuff just for the purposes of a single function or two. But that is just my 2c.

  9. rujialiu commented on Aug 13, 2025

    @rujialiu

    I've done similar things before. If the overlap is small enough and you sometimes want to modify the copied code and your modification is not general enough to make a PR for the upstream, then this approach is much better than submodule. Sometimes modification is unavoidable if you want to remove codes that you don't need, but is considered a major functionality in the upstream repo. If you don't need to update too frequently, it's actually quite manageable

  10. triple-mu commented on Aug 21, 2025

    @triple-mu

    same request!

  11. rujialiu commented on Aug 22, 2025

    @rujialiu

    When attemping to fix an issue with qwen2.5-vl in llama.cpp, I debugging through a lot of codes and becomes a little bit familiar with them, so I would like to have a try.

    However, I have zero knowledge about stable-diffusion.cpp's codebase. Anyone can give me some guidance? Let's get things work first, then decide the best way (copy code? submodules? other ways?) later. Looking at the actual code may help us decide.

  12. leejet commented on Aug 29, 2025

    @leejet
    Owner

    Once this PR #778 that adds Wan support is merged, I will try to add support for Qwen Image.

  13. LostRuins commented on Sep 13, 2025

    @LostRuins
    Contributor

    Sounds exciting. Qwen image edit is basically SOTA for image editing.

  14. marked Qwen img support #848 as a duplicate of this issue on Sep 22, 2025
  15. leejet commented on Sep 22, 2025

    @leejet
    Owner

    Support for Qwen image has been added #851.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions