When we got a form of such search, it searched all of the internet alongside with the functions of your software on your computer. You had to tell it manually what software you have and when it found something you had to fish out the description of what to click to get the functionality you are seeking and then click through some dumb ui to actually invoke it.
Agents are the first thing that can lead you directly from "I know what I want my computer to do." to "Actually doing it."
Success of agents is founded on the profound and sustained failure of all of the software designers and developers ever.
apropos ?
Obviously I'd prefer to ask specific questions to my expert human friend sitting at the next desk over, but sadly they don't exist.
I just hope there is a future where this tech exists without such big downsides.
Even the data centre using this probably uses less net energy, and costs less, to build one of these dumb arrows than the aggregate sum of the energy used to keep you alive while you Googled for the answer. (Including heating/cooling the building, powering the equipment growing the food you eat, etc.)
People have been saying "we don't need docs", "nobody reads docs", "they'll figure it out". They experience the pain of a new employee onboarding but never treat this as a signal for the user experience. The user can't just walk over to the developer's desk. But luckily we got stack overflow, blogs, and Google. Somewhere someone explains it! The problem became finding it!
But hey, why do things the "hard" way when you can just throw money at the problem? Or better yet, dismiss the problem by calling the user an idiot.
You're right. What a time to be alive. We've created a world where this product is useful. And not even to because dark patterns exist, but simply because we've always wanted to avoid taking a step back and thinking or getting an outside view. Because it isn't obvious if a robot has to teach you what button to press. If it does, you probably should be ashamed (there are, as always, exceptions)
That said, I do think there's still a lot of utility you this. Those exceptions aren't uncommon. This'll help people learn programs like FreeCAD or Blender, or whatever. Where the complexity is naturally high. But also I do think we should recognize the silliness of many problems that this does solve that shouldn't be problems in the first place.
Garage sale signs are a good example of this. The person making the sign knows what it says, so they can "read" it from farther away than someone who doesn't already know what it says.
Indeed, what a time to be able.
/s
https://github.com/seb3773/ntfs-repair-rfc
https://github.com/Kotivskyi/screenshot-tool
I think it was more popular around the beginning of the year to mid-year
Whenever I see these vibecoded apps, my move is to just do what you've described: point an agent at it and tell it to reproduce it (usually with my own little customizations). It's a bit absurd, but I don't trust that they vibecoded the app as well as "I" could lol.
The open models that are chasing the frontier labs will stop being open once things slow down and there is less incentive to undercut the front runners. Time will tell if GPU compute gets cheap enough to run stuff locally.
Again, I think it heavily depends on if people are writing comprehensive specs with sufficient detail.
I think the next generation of exploit will be putting up code all over the internet that does x, y and z and then when asked to one shot whatever that code does the agents bake in the exploit that was indirectly included in their training corpus.
Everything your building is now streaming through a third party who could steal it, the same way Amazon steals open source software and repackages it.
All you need to do is look at palentir's forward deployed engineers. Once you let a third party observe your processed, you're cloneable and replaceable _as an entity_.
I think we can look at our own human psychology as it developed into social groups. If our brains were obviously deterministic, it would be easy for people to prey upon us via structured inputs. Some level of randomness in our thought process is almost evolutionarily required to avoid becoming part of the larger super structure, similar to an ant colony.
So these AI companies are definitely more dangerous than just the "AI is super intelligence". They're basically a intelligence platform everyone is wiring into.
That's today. Tomorrow, they'll be run by an MBA which is the enshittification.
Not really. Ants are superorganisms, they share more than half of their DNA with each other. You could think of each ant as being like a cell in a human.
If you can make a facebook clone it doesn't mean you can steal facebook's lunch. there are exceptions of course.
It makes sense to not trust AI models with the ability to read and alter your screen for many many many reasons, but the developer of the tool knows that as well. A tool that lets your ai do something doesn't also let it do anything.
So, as always, it's a question of if you trust the software producer. You're giving them the ability to draw on your screen too. Would you have the same fear if there were no AI involved?
For stuff like this I imagine AI misbehavior as a novel failure trigger not a novel failure type. It can only do the harm that the software was able to do on its own already.
I don't know why anyone would ever willingly want this.
Brilliant. This is exactly the kind of content I seek when I visit Hacker News.
Truly art.
It's just the epitome of living in an era where code is so cheap to generate that we don't invest any effort into even fully fleshing out the idea. It's just "Claude, make me a program that draws arrows on screen". And HN, I guess, is dominated by two groups: one still stuck in the mentality of "code = effort" and trying to recognize that; and another stuck in the mentality of "AI = cool".
A big arrow to an interface element which ... has a label that explains what it does?
So...a label for a label?
Helping people deal with bad UX.
Everybody who has worked in a traditional corporate environment has encountered poorly tutorialized, poorly labeled, counterintuitive user interfaces where the only way to know how to get something done is to have somebody sit over your shoulder and tell you which buttons to click.
Oh, you want to figure out which OIM entitlements your coworker has that you don't to debug an access issue? You just need to know that you have to click the "compliance" button to get there. Why is it called "compliance" when what it actually does is let you export comparisons from the user tables? Because fuck you that's why. Memorize it.
An AI assistant that can read the procedures for you, or sometimes even read the source code, or the website code, can show you how to accomplish something that the UI on its own fails to do.
For the first experiment, Claude stepped me through how to edit its settings.json to remove warnings from its configuration: a little clunky as I had to respond to Claude's CLI prompts to get it to auto run.
In the second experiment, Claude walked me through how to use Apple's Compressor.app to speed up a video. It walked me through dragging the file into compressor and clicking on the appropriate UI elements to do this.
So, one use case for the bigarrow skill enables Claude to step users through a relatively complex UI to accomplish a task.
There are likely more efficient ways to accomplish the tasks at hand (e.g. ffmpeg for the second experiment), but this use case is interesting if using a GUI tool is required.
She doesn't have to call you anymore.
Do you remember when PCs used to come with a completely soup-to-nuts tutorial that would talk to you like you had never seen a PC before? Things like this: https://www.youtube.com/watch?v=3ScS4OYDfHE
This kind of baked-in interactivity could really help in a training or disability context.
- rainbow dripping arrows
- angrily pointing arrows
- flame-surrounded text boxes with particle effects
- the ability for the agent to play airhorn.wav at max volume, overriding existing volume or mute settings
EDIT: ayy lmao https://github.com/franzenzenhofer/big-arrow-on-the-screen/p...
https://hyperties.org/cabinet/symelec/
PIXIE and FORTH on PDP-7 Cabinet Emulator with Type 340 Vector Graphics Display:
It's already got the power to be obnoxious at you
Maybe isoprophlex will add that to his PR!
That would be a godsent for those of us with aging parents.
Right yeah? I mean I have freagin remote desktop installed on all my parent's devices. I am the agent for them every time they call frantically for help.
It's Time to redesign the apple genius bar for the modern aging boomer using this.
For web apps, WebMCP is a way where you can make a frontend way more discoverable to an LLM rather than that it is going to read all your code. I sometimes create this in my apps at home and sometimes in my apps at work and it makes LLM assistants way more usable. It also helps if they can highlight certain things of a visual.
We humans can do this too on paper. We can point at something, we can highlight something with a marker. Why not give an LLM these capabilities as well?
We built something to help with that [1].
Whenever I use web tools that don't have "deep linking" [0], I love to throw out this quote:
"Have we thought about using hyperlinks? They are these SUPER useful things that were invented in Switzerland back in the 1990s."
What terminal / agentic workflows spawn the demonstrated dialogue boxes that require the User to click/reject the action? Aren't most such flows actually inline UIs? And finally, if the core issue is that these confirmation boxes are tied to the terminal that triggered them, which does not autofocus, what's the point of said "big arrow" that is also lurking behind without focus?
If the said "big arrow" automatically gains focus, isn't the real fix here to just make the dialogue boxes themselves gain focus automatically? Both are similarly disruptive anyway.
"Teach me to do this" kinda stuff. "Guiding agent" rather than "doing agent".
Honestly, @franze, if you're reading this, maybe update your screenshots to show an agent in tutorial mode on some complicated app?
(OP, nice project! Sorry for my rant.)
But it's a subpar experience compared to just opening github on the browser.
Why does this need to be a skill? Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows” and you get the same thing without burning tens of thousands of tokens in context.
> burning tens of thousands of tokens in context.
~1400?
> Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows”
Is there a macos native thing it uses? Or are you talking about screenshots? This draws on your screen instead.
Does one need 4 programming languages to draw something on a mac?
> "It's this window, not that one." You have 14 Chrome windows. The agent knows which one it means: --window "Google Chrome:Pull request". It even picks the right tab: --app "Google Chrome:Pull request".
> Remote help. "No, the other gear icon." Point at it instead of describing it.
To me this feels like it takes away from what the human is supposed to do (read, understand the consequences of the action, then.. consent or abort)
There is a reason your AI Agent won't automate these clicks for you
Grim. If you're just there to click sudo buttons for the bot, you might as well give it root access and be done.
--norms-taboos-offI remember taking your SEO course many moons ago.
https://laughingsquid.com/yahoo-neon-billboard-in-san-franci...
I absolutely loathe this phrasing now. I don’t even know what part of it is good or notable.
% @(#)handy.ps
%
% Handy Pointer
% Copyright (C) 1989.
% By Don Hopkins. (don@brillig.umd.edu)
% All rights reserved.
https://donhopkins.com/home/archive/psiber/cyber/pointer.psPSIBER Space Deck and Pseudo Scientific Visualizer Demo:
https://youtu.be/_fqCeuue5Ac?t=213
The Shape of PSIBER Space: PostScript Interactive Bug Eradication Routines — October 1989:
https://medium.com/@donhopkins/the-shape-of-psiber-space-oct...
But I feel dumber just by looking at it. If this is how this will look like in 10 years, why to make desktop at all. Just connect mic and speaker to your PC or talk to the phone: 'I need new pair of socks. Order 10 for me. In your favourige color it will be 21.37. Should I charge your credit card?'.
It is not like most of the people enjoy computers. I am pretty sure they do not. They just need them to operate systems they need: government websites, banks, maps, restaurant menus etc. If some agent will do that for them, why bother looking at screen at all? Rich people have it with their own personal assistants.
I found this awesome vibe-coded Pixel News Network project recently. How fucking cool is this?? Would never have been created otherwise. https://pnn.watch/
I actually somewhat enjoyed the times when vibed projects were a broken mess, lots of accidental comedy in that era.
~guywithnopowertodisallowit