My Local Models Gemma4 and Qwen3.8
So I have 3 GPUs:
- AMD AI Pro R9700 (32gb)
- Nvidia 5070 (12gb)
- Nvidia 1080 (8gb, no tensorcores)
I honestly have no idea how the local model community and Huggingface keep up with so many different models. There is honestly a choice paralysis with so many distilations and fine tunes and model names that sound like gibberish. As such I've mostly limited myself to the models that have been released by Unsloth since they seem to be the authoritative open-source model enabler. While the frontier labs release open source models, it ends up being Unsloth who makes sure those models work on local software. So i've been using their Quantized Aware Trained (QAT) versions of Google's Gemma 4 (12B, 26-A4B and 31B models). I was using Unsloth's version of Qwen3.8-27B but honestly the overthinking was getting to be too much so I switched to a fine-tuned version that has some capped thinking.
Some quick thoughts
- Gemma 4 12B:
- Thank you google for throwing a bone to the folks with 12gb GPUs as the 12B model at QAT4 fits like a glove.
- Certainly not a coding model but in terms of tool calling quality, large context window, and speed (120 tokens/second), it is a useful model
- Gemma 4 26B-A4B:
- This model is how i discovered why RAM memory prices have quadrupled.
- Even at QAT4, this model does not entirely fit on the 12GB Nvidia 5070 card... except it is a Mixture of Experts (MoE) model which means I can unload multiple layers onto system RAM
- Even with some layers spilling over to RAM, im getting 80 tokens/second which is still very fast
- It was interesting how when asking 12B and 26B the same question ("Who is Keith David"), the 12B model just pulled from its training data while 26B did a proper web search
- Gemma 4 31B:
- Smarter yet slower than 26B (40 tokens per second)
- Currently occupying the Architect and reviewer role in my software factory
- Not a good coder and when asked to write python files, does not consider the rest of the code base
- Functionally a slower, dumber version of Gemini 3.8 but still has its uses.
- Qwen3.8-27B:
- I'm still working with it, but I honestly don't get the hype?
- It certainly thinks alot and does produce somewhat decent code. It is not bad but this model makes me question some of those LLM benchmarks
- I might be expecting more from it than I should though
- It is currently acting as my Developer, TDD agent