upvote
A VLM (a vision-language model) is already being used by Waymo[1]. It's useful for scenarios that require reasoning and general knowledge.

[1] https://waymo.com/blog/2025/12/demonstrably-safe-ai-for-auto...

reply
Fast forward to me sitting at a Green Light waiting for my usage to reset for the week so I can get to where I’m going
reply
In a sense, they already are. The giant leap in self driving cars we've seen in the last handful of few comes from using transformer models.
reply
I remain worried about prompt injection style attacks against self-driving cars.

Imagine if someone finds a weird image pattern that gets misinterpreted as instructions and hangs that off a bridge over a freeway.

reply
It'll get rooted over the uplink/WiFi/BT long before that. Probably even more likely for non-SDVs.
reply
deleted
reply
It's not a coincidence that Waymo started becoming viable after GPT-3.
reply
News to me. Where did you learn about that?
reply
GPT-6 is multimodal, LLMs alone have no vision capability
reply
"LLM" is now in practice a superset of "LMM"
reply
To my understanding gemma is shipped with waymo, in a highly modified fashion.

I'm skeptical, tho. Cost will push for right sizing, much like we have right sized a lot of things about modern cars.

reply