upvote
There is a Gemma 4 model with 12B parameters which might be worth trying. e.g. https://huggingface.co/mlx-community/gemma-4-12B-it-qat-4bit

That said, your computer will still get hot!

reply
I just run Gemma4 12B MLX via Ollama and it's been doing fantastic work!
reply
If you go small enough it should be no problem. For example Gemma 4 E4B in Q6 or Q4 quantization should run well on your laptop. It shouldn't be too taxing, but would still want to eat 7-9 GB of VRAM or so.

Now that model is mostly useful for writing or chatting.

reply
You need more RAM, plus the models take up a lot of space. 32gb min but I’d recommend 48/64gb, you won’t get close to frontier but it’s still fun to play with, images are very good
reply
heating up is normal, that cannot be avoided. it should become laggy, but you just have very little RAM (I assume 16 GB?) So most models are too big with other stuff running.
reply
Tbh, the only reason to run a model locally is when you want to be completely safe.

From productivity point of view, it doesn’t make sense to have any notebook running a local LLM.

We have one life. We should spend it wisely.

reply