upvote
>We run all our Personal Assistants now on flash

are you worried about sending all your data to third parties, especially if they're in different countries?

reply
Not shilling for them but Ollama cloud hosts domestically with ZDR afaik. I run 95% of my open weight inference through them. The rest goes through Opencode Go $10 plan (which is enough to run 3 hermes agents on DSF 4.1 and leave plenty of left to experiment with when new models drop).
reply
Just a reminder that any API using a Cloudflare TLS certificate isn't ZDR.

The model engine provider might be ZDR, but the service as a whole isn't.

reply
Cloudflare doesn't retain request bodies by default, and if you don't trust that then you shouldn't trust the third party AI provider either. Cloudflare does cache responses, but that doesn't typically apply on API endpoints and is trivially disableable.
reply
did not know. ty! have my updoot as thanks
reply
We use hosted providers in the EU. They have a cost markup but its worth it. I enjoy the idea that I do t talk to big tech.

I cant ofc be fully sure because i don’t own the chain end 2 end.

reply
Could you please share which ones? I need serious ones with generic DPA but without Enterprise plan costs.
reply
I use it through OpenRouter, which has ZDR enforcement.
reply
should you not be worried more about sending data to providers of your own country? I think they are equally bad personally, but if you think one of them is acceptable, why do you prefer it to be your own government that has more ties to your life?
reply
I have asked OP that question but I think there are providers who are not in China and they just host the model/inference.
reply
Just use a provider hosting it in your country especially if your country has major data centers then its the same as using Anthropic or GPT of GCP Model Garden or AWS Bedrock

nobody here is talking about running frontier level intelligence locally so if you’re Chinaphobic and prefer layers of corporations siphoning your data in between you and the party there are plenty of options instead of directly to the party

reply
what would be the concern here?
reply
What is the cost of access like for DeepSeek-v4.1-flash, compared to GLM-5.3-flash via ZAI's Coding Plan? Because that's what I use; and often hit the "wait". I wouldn't mind trying a new model subscription or even API access which hits around glm-5.3-flash level weight class (which seem to be enough for me; with quite some human suprvision and nudging) but gives muuuuuuch moooore tokens for the same price.

(And what are the preferred providers?)

reply
Have you compared it against actual SOTA models like latest Fable or Astra?
reply
The author explains this very well tbh:

> Today's models are now good enough for high-quality unattended tasks. Chasing the latest and greatest is silly. It is fun to see the new Fable capabilities, but the tasks we throw at them are usually ridiculous (maybe even insulting) if you believe in LLM sentience. It's like asking a math PhD to organize the files on your desktop.

I'm using DS V4.1 Flash as my main model since their release and it works great for all my coding tasks. My setup is OpenCode Go subscription and obra/superpowers skill.

The only times I try to change models are on general planning tasks (like research this codebase for tech debt mitigation opportunities) or if I need deep research which would benefit from searching the web, in which I still think Gemini is still the best because of the speed and access to google search index. But these are not even 20% of my daily tasks.

reply
I have had good luck rotating between the three tiers of GPT 5.6 with occasional jumps up to Astra or Fable. Most of my work is with GPT-5.6-Sol. Simple tasks like data extraction or trivial refactors (rename this variable etc) I often push down to Terra or Luna, or even to self-hosted Gemma4:31B. Very tricky stuff, like planning a new feature, design review, or code review of a complex change across multiple repositories is where I leverage Astra.

I've had middling success with models like DS V4.1 Flash and free Gemini. They tend to be pretty good at very easy stuff ... but they're more likely to go down rabbit holes, confidently assert falsehoods, or fix bugs with changes to my test harness rather than my code.

I asked about comparison to the well-known SotA models specifically because I use either Astra or Sol for ~85% of my daily tasks. When I try to use smaller/cheaper models, I have had very mixed success. Sometimes it's perfect, while other times it fails in subtle and hard to catch ways.

reply
Of course there's still a huge performance gap

but DS 4.1 Flash is good enough for most tasks

reply
What personal assistants are you using?
reply
Hi. We are building our own, but the core is openClaw. Hoever, that will change soon
reply