Yesterday I received a new thermometer for my aquarium to replace an old broken one. Both were bluetooth, but different models. I just told claude "I'm going to set up up my new bluetooth thermometer for my fish tank in a few minutes, keep an eye out for it and replace the old broken one with it in Home Assistant" and then walked away and put a battery in it and put it in my aquarium.
When I came back it had found it, replaced all my existing entities for the broken one with the new one, and verified it was all working with my existing graphs and automations.
The agent can get to the resource through the MCP server or using API key. I personally do not see the benefit MCP is providing here. Sure you can reduce the exposed surface at MCP layer, but I do that at the API layer. I don't need to add another layer here.
I can kinda understand if you do not have control of the API layer and/or you have to expose the API layer to the public Internet as well. Most of the time that is not the case for me.
I find it way better to be able to confidently tell agents to use CLIs than worrying about partially implemented mcps that need configuration and are often yolo’d with npx @latest anyway
Not only giving an API key to an agent can leak to the model because of harness issues or too broad reading rights, but also most providers don't give the ability to apply principle of least privilege to an API key.
I don't want to give an agent full R/W access to any of my services/accounts.
Recently I wanted to set up a slightly complex routine involving some lights and a couple of motion sensors. It feels like magic to be able to describe the behavior I want, briefly discuss the implementation, and walk past the sensor and see it in action.
User will interact or build apps with simply text like: "Give me all the issues that are X, context: https://somedomain.com/llm.txt"
llm.txt will have all the API instructions
Having an agent keep an eye on stuff and fix things proactively should make the experience much better. Plus I can no longer write yamls at last.
Definitely helps that the home assistant api is documented online most likely in the training data.
It is more of a curse than a blessing. MCP pollutes agent context even when you are not using it. Use a manually invoked skill instead if you don't want to be wasting tokens on every turn and bloating up agent context making it dumber in the process.
MCP in some harnesses bloats context. MCP in some harnesses doesn't bloat context.
They're so powerful and yet get out of your way when you're not using them. I couldn't imagine being without them. Thanks y'all!
But in my case, a CLI was not enough.
Like, to the MCP I might say:
Set Clop to make every video copied in ~/shots smaller, 2x and silent
Then the MCP can use elicitation and say: By smaller, you mean re-encode to compress file size (factor can be 0 to 100 max compression) or downscale resolution (100% same size, 50% half size)?
And does 2x mean faster speed? In which case do you want to keep frames so the video plays smoother or drop frames for size? Or does 2x mean upscale?
And the agent will present those as nice choice menus I can decide schematically on.With the CLI I have to first read, learn and memorize the requests and commands needed for each app, the accepted values and formats and the steps to reach a specific result.
There's only so much space in my head I can leave for implementation details of arbitrary apps. I'd rather have an agent care about that.
And yes I get the irony, those are my apps, I coded them by hand for years, I should know their implementation details, yet even I forget if I should pass 50% or 0.5 for half size.
Btw Clop is a media file compressor for context: https://lowtechguys.com/clop
Plus I can add some complex commands like in the rcmd Stages [1] case where the agent can create a 4 monitor layout with apps and windows placed where you want, with every window opening the document/folder/project/URL you want and running the terminal commands you need. Sure you can do that with the CLI, but it's hard enough to get right because of shell quoting issues, that even an agent can get it wrong.
For simple tasks though, sure, the CLI is just enough and the agent can use it without needing to install yet another MCP. You'll know when you need it.
[1] https://lowtechguys.com/rcmd
---
EDIT: I just remembered, you can even hook the Claude/Codex/Gemini desktop app to the MCP, while you can't get it to use the CLI. so there's that for users that still don't feel comfortable at a terminal, which is a number higher than you might estimate.
This seems like exactly the sort of thing I've done with shell scripts or even makefiles.
But this is for people that already use the app, researchers, writers, students, people that aren't necessarily comfortable with a terminal. And given Clop already implements the basics: an efficient file events watcher, tuned encoders for the Mac silicon, fail safe backups and UI for seeing the result and interacting with it in real time, it has advantages over trying to do it yourself.
I will point out that the shell script way uses less resources than an LLM making a tool call. But I understand that these scenarios are not necessarily meant for the same user.
Oh for sure, I would prefer to have the automations as invisible things running at the system level, doing exactly what I want and nothing else, not wasting resources on UIs and event watchers I might not need. I would get rid of my own apps if that was easy to do.
But it seems we need to waste some resources to get some usability in return.
Definitely saving it for further use, I sometimes need to have small invisible watchers and I don't want a full fledged app or shell scripts for that.
Why is AI involved for any other reason than building the original test implementation?
Not sure if you got the right context, your question doesn't really make sense to me.
We're back onto the original use cases for natural language processing. This is where all the value always was, and now the market has proven to itself what anyone with even a bachelor's in computer science already knew.
I'm working on a sideproject called Rowbly[1]. It acts as a sharable data store where LLM's can dump research rather than keeping it in their memory or throwing it into a spreadsheet.
At first I thought "I don't need an MCP, I'll just expose a CLI" but that carries a pretty big limitation in that it only works with agents on your computer (Codex, Claude Code, Pi, etc). For the folks on here this is not an issue and is often times preferable but I'm also targeting the average LLM user that primarily interfaces with it via "consumer AI" and with those tools the stuff you can do is very very limited.
I still think MCP's have a long way to go maturity wise and hey maybe in a few years we will figure out a better way to do things but for now, if you want to interact with consumer AI apps, there's just no way around using them.
[1]: https://rowbly.com
There were some teething problems on Golden Gate: keystrokes didn't seem to make it to the rcmd popup, but your recent remediations seem to have solved it. It's a beta OS version too, of course :-)
Still working on finding all the edge cases, so sorry if you still encounter problems there. I've been using it since June and still find problems in input handling.
Like there's this thing, where if an app has Accessibility Permissions and listens to key events (like rcmd does) and then you revoke that permission while the app is running, then your whole system will stop responding to keys and clicks. Until you kill the app in question, but how are you going to do that without a keyboard?
All apps have this problem, even established ones like BTT, because it's a recently introduced macOS behavior in how the internals of CGEventTap work.
Is that not how it works out of the box?
It's a very specific thing for me really, I connect the HDD specifically for doing backups as fast as possible then I want to disconnect and store it back so I can keep using my laptop. I don't have a desk anymore where I can keep these things connected all the time.
Everyone on this forum has an absolute paucity of imagination when it comes to applying LLMs to any use case that doesnt involve coding.
Of course MCP has its use case like if you want auth, or session based actions.
NO IT CANT!!! why dont you understand that not all agents have access to a terminal!
Recent example: The AWS CLI has a convenient "s3 sync" command that does a one-way directory sync, only downloadiong new files if existing files with matching sizes and timestamps don't exist, but the closest thing the .NET SDK has unconditionally overwrites destination files.
Another example: one of the nicest things about C# is how the compiler exposes itself as extensive library functionality, so, not only can I compile and run code at runtime, I can generate this code from programmatically created ASTs instead of text, parse expressions into ASTs, etc.
These are things devs often take for granted as being part of ‘how chat agents operate’ but they are specific to how claude code/opencode/pi/codex operate.
Giving agents a Unix computer account they can play with is definitely a powerful tool that makes them capable of doing a lot more (see: meta muse, OpenAI dots), in much the same way that giving a human a computer they are trained to use makes them way more capable… but it’s surely not the only way we can run these things.
They pay tons of money for hardware, only to use it the same way I was using those DG/UX terminals at the university.
Naturally there are no coders in other operating systems as well.