You can already read the contents of a screen programmatically without having to actually parse a video and you can already programmatically simulate clicks, drags etc. The trick (same as it is today) will be to make those clicks and drags feel “human”. Not too fast, not too slow, etc etc. But all those challenges exist today.