← Back to post

Edit history

Most recent

For your hardware, the VRam is not enough to run 27b but, I’d recommend Qwen 3.5 9b for image / text to text.

And I’m planning to experiment with Qwen 3.8 9b for text to text.

4_k_m quantization is the sweet spot for performance and ram usage.

Also, I find Llama cpp is better than Ollama in terms of performance.

Original

I’d recommend Qwen 3.5 9b for image / text to text.

And I’m planning to experiment with Qwen 3.8 9b for text to text.

4_k_m quantization is the sweet spot for performance and ram usage.

Also, I find Llama cpp is better than Ollama in terms of performance.