Replying to @⁨Chee_Koala@lemmy.world⁩

For your hardware, the VRam is not enough to run 27b but, I’d recommend Qwen 3.5 9b for image / text to text.

And I’m planning to experiment with Qwen 3.8 9b for text to text.

4_k_m quantization is the sweet spot for performance and ram usage.

Also, I find Llama cpp is better than Ollama in terms of performance.

huggingface.coDavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science.
Edited ⁨⁨Aug⁩ ⁨31⁩, ⁨2026⁩, ⁨11:50⁩⁩en