upvote
Pretty much matches my experience.

> 'sleepy time' means sleeping → start_vacuum with room 'bedroom' to start cleaning

The "DeepSeek 4 Flash grade" claim seems far fetched.

reply
thanks for these haha, you can actually edit the tools and/or their descriptions, the demo is just a "get started" preset. But still we do have room for reasoning improvement!
reply
What kinds of things do you expect to work?

Edit - I’m struggling to get anything useful. Reasoning is often utter nonsense and the actions are very often very wrong. To the point of seemingly needing very precise sentences to work at which point you may as well do regexes. Very simple things like clean one room then another with the vac fails.

reply
Thanks for the feedback! Implications and relations are hard for the model to understand (things like go to the living room, then the kitchen, and back), so yes the cleanest use cases involve direct language. Reasoning isn't true reasoning in the way general LLMs do it, it is more like grounding for the model that it generates itself. This can often become nonsensical specifically when the model gets things wrong, providing signal to the confidence.
reply
Can you share an actual example of where it works please?
reply
Maybe with fine tuning?
reply
those are all expecting far too much for models this size
reply
Given the title of the post says 'can match deepseek v4 flash' I think it's fair to call out these sort of dumb mistakes.
reply