The main question I tend to get from hacker-types is "what are the tech specs?". The demos run on a Mac with custom software I built for projection mapping/rendering/etc. (the hand tracking uses Apple's built-in framework, which is very good), a small consumer 4k laser projector, and a basic webcam for vision.
But really, this is primarily a research/design prototype - I think it's much more important to figure out what the user experience should ideally feel like, if the interactions are even plausible & desirable in the first place; and then work backwards to the technology.
Possible use case: projecting menus, ads, and bills on restaurant tables. With a camera and enough automated layout to work around what's on the table.
It might sell to the people who buy refrigerators with displays on the door. Could be useful as a cooking training aid, like cooking videos but present where you need them.
The light bulb form factor is cute but probably not that useful. Most of the light will be aimed in non-useful directions. Look into ultra short throw projectors.
How many watts does the projector use for how many lumens? I would be curious if these numbers plus size would be a fundamental limit to the tech.
I wasn’t sure if any of this was real and that example looked extremely fake to me because I had a hard time believing the device being able to read text from that distance.
(I think I overestimated how much of an angled shot it was the first time I saw it, on top of how easy or hard turning your finger into “actually capture this text”… but you’re totally right that in principle a camera at a non agressive angle can capture that text)
(I like this overall concept a lot! Cooking is the right example, voice interfacing’s 96% use case is basically “my hands are dirty but I must use computing”)
It's not mentioned in the piece, but for the latter I don't see any technical obstacle to having several of these in different rooms in the house, or a couple in the same room to augment coverage.
One small quality-of-life improvement that could be added: have the software timestamp when you say "here" and then (using a short low-FPS video buffer?) analyze the corresponding frame of video. This would eliminate the pause when you're pointing at something, waiting for the computer to respond.
Alternately a low-latency speech analyzer could look for the word "here" and quickly snap a screenshot. Perhaps project a temporary dot on the surface, so the user knows they can put their hand down now?
Quite the impressive demo! There's huge potential here.
I mean, merge one of these with a modern anti-consumer "Smart TV", and we're uncomfortably close to 1984's telescreen.
One big difference in direction is that Dynamicland + Folk Computer use projector+computer vision as substrate to try and reinvent/rethink the interactive/dynamic medium, and what end-user programming could be (which is very cool).
In that way I am much less ambitious; this exploration is more about how there is an entire class of existing interactions (the ones showed in the piece) that we perform today on screens, that could be much more interesting + collaborative if overlaid in the real world. And I am also terribly uninspired by the eyewear-based computing vision all major companies have pushed lately.
Youtube videos - https://www.youtube.com/results?search_query=sixth+sense+tec...
Sure, it's a little more hands free, and projecting around definitly is nicer than screens. But projecting has a lot of problems too. Shadows, projection surface, ambiant light, heat are difficult to handle, but to me the worst is that it's not persistant. You can't just project 24/7 your family planner on the wall.
I'd really rather have something that prints, or better draws.
You want a recipe => print it, don't beam it.
You want a map of kyoto => tiny robot walks around the page and draws it for you. And you can draw over it if you want.
And it doesn't have to be exclusive, you can do both beam and print. But I think protable printing / drawing has a lot of unexplored potential.
Printing otoh requires hands and prep. That's greater effort. You wouldn't print a recipe if you were just a little uncertain about step 3, you would ask for some kind of just in time modality instead. Voice might be fine for short info but sometimes you'd prefer to have it up and there's beam.
But more importantly I ultimately think that one of the unstated goals is to reduce screen time. And setups like these are only one voice command away from beaming youtube videos all around your house.
Reminds me of the projection bracelet scam that Captain Disillusion debunked years ago https://youtu.be/Cw7g8ixbomU?t=75
Most of the videos are shot in real life, during summer daylight (only exception are the 2 person around a table ones, which are shot at night in a room with regular ceiling lighting).
And the projector I used is only 1100 ISO lumens; fancier models go way higher.
If you project mostly white content, it's remarkably readable + visible. Outdoors in direct sunlight would be another story, of course.
Holographic computing, I think, will be the future. At least for interfaces.