Repository navigation
Qwen Image #758
Description
Activity
I'll investigate if I have free time. A LOT OF CODE COPYING from the
Qwen2.5-vlmodel seems to be necessary due to the complexity of the visual & textual encoder, maybe we shall copy these things or import them from thellama.cppmain repo and integrate better with them?Yes this kind of things have been discussed before, it seems that including llama.cpp as a submodule in the project would be the way to do this, or maybe including this project into llama.cpp.
Reacted by Gilad S.Reacted by LostRuins ConcedoYes this kind of things have been discussed before, it seems that including llama.cpp as a submodule in the project would be the way to do this, or maybe including this project into llama.cpp.
Any reference for discussion context?
Yes this kind of things have been discussed before, it seems that including llama.cpp as a submodule in the project would be the way to do this, or maybe including this project into llama.cpp.
Any reference for discussion context?
Reacted by Seas0@LostRuins considering you have experience in integrating them might be userful to chime in.
I didn't do the implementation though, I only have a minimal understanding of how it works, mostly using the existing clip api. ngxson is the one you want.
having said that, a submodule is often very messy as it becomes hard to track changes to files. I would recommend extracting out the files you need and modifying them instead.
a submodule is often very messy as it becomes hard to track changes to files. I would recommend extracting out the files you need and modifying them instead.
Doesn't that make tracking upstream changes even harder?
Doesn't that make tracking upstream changes even harder?
In some ways yes, but the idea is that only the necessary code for the vision/text encoders is ported out (in like one or two files) and handled independently - it essentially becomes part of sd.cpp's code base and frozen in it's current functional state for the purposes of qwen image.
I think that the overlap between the rest of llama.cpp's code and stable-diffusion.cpp's code is sufficiently small that you don't really want to grab all the other stuff just for the purposes of a single function or two. But that is just my 2c.
Reacted by Rujia Liu and Piotr Wilkin (ilintar)I've done similar things before. If the overlap is small enough and you sometimes want to modify the copied code and your modification is not general enough to make a PR for the upstream, then this approach is much better than submodule. Sometimes modification is unavoidable if you want to remove codes that you don't need, but is considered a major functionality in the upstream repo. If you don't need to update too frequently, it's actually quite manageable
same request!
Reacted by LostRuins ConcedoWhen attemping to fix an issue with qwen2.5-vl in llama.cpp, I debugging through a lot of codes and becomes a little bit familiar with them, so I would like to have a try.
However, I have zero knowledge about
stable-diffusion.cpp's codebase. Anyone can give me some guidance? Let's get things work first, then decide the best way (copy code? submodules? other ways?) later. Looking at the actual code may help us decide.Reacted by Gilad S. and NobodyOnce this PR #778 that adds Wan support is merged, I will try to add support for Qwen Image.
Reacted by Seas0, LostRuins Concedo, Rujia Liu, lin72h, Lubos Lenco, 0xF, Victor Dvornikov, Rupesh Sreeraman and Rubén SantosReacted by lin72h, Lubos Lenco, Victor Dvornikov and Sean GallagherReacted by Gilad S., Rujia Liu, lin72h, Piotr Wilkin (ilintar), Lubos Lenco, Victor Dvornikov, Rubén Santos and fszontaghReacted by stuartjingSounds exciting. Qwen image edit is basically SOTA for image editing.
Reacted by Sean GallagherSupport for Qwen image has been added #851.
Reacted by LostRuins Concedo, Rubén Santos, Shareef, Rujia Liu, iwr-redmond, Lubos Lenco, Raphaël Thériault, sjxx and Erik Scholz
A request to add Qwen Image support