upvote
Fascinating, thank you. I am trying to switch to pi from opencode (my own thinking loops and burnout are a challenge lately).

It had not occurred to me that you could nudge it to stop thinking with a proxy. Nice idea.

Will favourite your comment and come back to it.

ETA: Incidentally you've helped me put into words the difference between the way Muse Glimmer thinks to the way Qwen thinks. There is a clear sense of urgency in Glimmer's thinking traces.

reply
Having spent a good part of the day with it, glimmer reminds me of Rorschach from The Watchmen. No unessential parts of speech, action oriented, brief and to the point. From a token perspective anyway it’s great, and it seems to hold its own well against more verbose models.

I really do feel like it’s effective tok / s is way higher because it doesn’t waste them.

reply
I am very struck by the way open weights LLMs seem to reflect a culture.

I don't really enjoy the way Qwen writes prose, and I find its thinking a bit exhausting, though it clearly writes very good code.

I like the neutral, clear way the Gemma models write, which I sometimes use to get myself a "getting started" document on something I want to understand; it also summarises well. It is neutral, sensible, un-showy. It writes in a way that is fairly close to what I would use for documentation. The 12B and 26B models are also very good for talking about art and photography. Analysing my own photographic work has helped me more than I expected it to.

This model, honestly, has made me smile. It also feels like it is more creative at a given temperature than Gemma. I am trying to motivate myself to do something quite open-ended so I asked it about what other people's considerations might be in my situation, and at the risk of anthropomorphising, the things it has come up with feel like the work of a more curious mind, somehow. More eclectic. I have enjoyed testing it and I really want to test it more, which might help me get over a motivation hump there, too.

(I am also exploring its hard-wired policies by asking it to analyse some studio art nude work I have done; it definitely thinks out loud about its policies in a way I have not seen Gemma do.)

reply
I think we’re going to see a lot more “product“ focus in the future with deliberate attention paid to these kind of properties. Historically though there are some obvious differences, the focus has been on benchmark maximizing. As that saturates, I expect more interesting choices about writing and thinking style designed to be differentiators instead of a side effect. Kudos to the PM here for taking it in a different directions, there’s obviously been thought put into it.
reply
I ran a 9B over my like 100k photo library — it was very good at it. And extracting any text. All local.
reply
It might be an indication that analysis of photography is something of an ideal discipline for an LLM since so much content online involves discussion of pictures.

A lot of what I am trying to do with my photography is sort of meta-photography. I am really interested in early photographic history, pictorialism and its opponents etc., but I try to avoid reproduction, so I try not to simulate processes too closely or to use vintage tropes in props and settings, but I use simple, undercorrected lenses and some vintage lenses, to gain some of the visual language.

Finding out that LLMs (including Gemma-4 12B with its built-in image encoder) understands a lot of my references and influences, could recommend me my (still semi-obscure) favourite historical photographer and other photographers who clearly engage in the same work, is amazing. And sometimes it says stuff I had not thought of, which is what I am looking for, since my photographic journey is somewhat lonely.

And that is just sort of brushing past the fact that these things can describe the contents of photographs with an accuracy that you can almost take for granted.

reply
You really have to get the models to end their thinking. Almost any commercial model serving has safe guards like this to tune how much they think.
reply