I don't know how it just happens that sending text messages to people can manage to result in specs that are painful to implement.
Can you provide more details here? I've never seen an issue with the protocol at all, more issues with feature difference between servers depending on what they have decided to implement or not.
For instance, it works as a constant, uninterrupted stream. You never close the document until the disconnection. Which means you absolutely have to parse it with a stream parser.
It could have been something sensible, like a document per message: "Here's 300 characters of XML document: <xml>....". But nope.
And what's up with this? https://xmpp.org/extensions/xep-0394.html
That made me remember IBM JSONx... https://www.ibm.com/docs/en/datapower-gateway/10.6.x?topic=2...
Are there more elegant or natural ways to do it? Probably. But when you say ‘extremely complex’, saying ‘codepoints 7 to 15 should be bold’ is not what comes to my mind.
* It's not standard. You can't plug in some normal formatting library there and be done with it.
* You'll probably end up having to write conversions back and forth.
* It's fragile -- if anything gets misaligned it'll break
* It's tooling unfriendly. Think things like Nagios, Grafana, etc sending formatted notifications. Everything can do this or <b>this</b>, but practically nothing is set up to accommodate this format.
* It leaks into other subsystems. You can't eg, just search/replace/insert/delete words if you need to for any reason without breaking this.
Is it the worst thing ever? No, but it seems to go with the pattern that nothing in XMPP is normal or comfortable. Everything has this weird particular flavor to it and more complexity than necessary.
If it were up to me, my option would be to support 2 things: Markdown for simple cases, and embedded HTML for when you really have to get fancy. Each of those can be handed out to existing code and you don't ever have to keep adding extensions because what if people want colors or something now.
And you could use markdown. Just get the rendered spanned text result and transpile it. This would be true for any editor as well, just get convert it to spanned text, edit, convert back.
And with HTML you'll still have to specify what subset you support, and many tools won't support that either. Markdown suffers from this too, by supporting embedded HTML, though at least commonmark is pretty baseline and well supported.
<message>
<body>This XEP supports many things:
* inline markup
* code blocks
* lists
* and possibly more!</body>
<markup xmlns="urn:xmpp:markup:0">
<list start="31" end="89" ordered="false">
<li start="31"/>
<li start="47"/>
<li start="61"/>
<li start="69"/>
</list>
</markup>
</message>My own preference goes to NOSTR, just a set of public/private keys as identity and nothing else needed.
Thank you for the reference to that project called anproto, never had heard of it and I'm still without understanding what problem it really solves. Seems to encrypt texts but then contains zero outside metadata to route it somewhere. The website is very scarce on details for someone unfamiliar the scuttle-something they mention.
https://news.ycombinator.com/item?id=9772968
https://news.ycombinator.com/item?id=31133082 (article and discussion)
But TL;DR:
* It's hard to even parse, XMPP uses an uninterrupted XML stream. * The contents are often baroque and complex * Standards are a mess, and stuff that should be in core isn't * Data loss is possible * Protocol wasn't made for mobile devices * Multiple devices are terribly supported * Data loss is possible
From my attempts long ago, and other testimonials, writing an XMPP client is a full time job of solving weird problems that shouldn't exist in something better designed.
I can't take this write-up seriously if it starts like that. I still read the whole thing though and I couldn't find any solid argument as to why someone would prefer XML to JSON for XMPP.
> This is especially true in browser environments, where XMPP streams run over WebSockets, which naturally frames the XMPP protocol. That’s why you are never actually working with XML trees consuming large chunks of memory. Modern implementations like XMPP.js go further and use LTX—a lightweight parser built specifically for XMPP’s streaming model—rather than the browser’s DOM parser. The result: developers work with JSON-like objects anyway. The wire format becomes invisible to your application code.
That to me, is an argument for using JSON, not XML. XML is strong when you have elements referencing other elements in your document structure or DOCTYPE for grammar defs. I might be missing something but I don't get how streaming XML is to be preferred over JSON for XMPP.
>> XML remains the best format for representing trees—deep hierarchies of nested data. JSON handles flatter structures well, but good messaging protocols are extensible: extensions can be embedded at different levels and composed together, like Lego bricks. That’s where XML shines.
Anything reasonable you'd be doing with messaging protocol is pretty flat.
Good protocols are opinionated and make concrete choices.
Extensibility typically leads to compatibility issues like the ones you are describing.
I briefly entertained building a modern chat experience on top of IRC when I came to this realization.
Considering I'd never hosted an XMPP daemon and didn't know anything about the protocol (I'd used XMPP clients a little bit, but had never looked at the protocol) and got a server (that part, I didn't write), an auth connector for our website's authentication system (so the daemon would authenticate against that instead), prod-ready and the features I wanted all working smoothly and reliably in maybe three weeks of very part-time work (this'd be, like, 3-4 part-time days with LLMs now, tops, from the same starting point)... seems decent to me? I mean I did direct work with the protocol, didn't just glue together libraries, and it was pretty damn good. Also (and I know browsers seem to be retreating on this front, which sucks) being XML made it very nice to work with in a Web context, since you can just ask the browser to turn ~any XML into a DOM for you, and get a bunch of functionality for free.
What's wrong with it?
This is what boggles my mind. We're talking: Text. Over the Internet. It should not be a complex, difficult problem! We've been sending text over the Internet from the moment the Internet went online. So how is it that 10 companies have managed to find 30 different ways to do it, which are all incompatible with each other? You have to TRY to fail this badly.
Come on, try! Start listing some of the core features of the different protocols/clients and answer your own question as to why it isn't simple! Presence. Chat history while offline. Threading. Mobile devices. Intermittent/mobile/cellular/wifi/CGNAT connectivity. Multiple clients on the same identity and presence. Identity establishing and confirming. Encryption. End-to-end encryption. Backing up and restoring chats, encrypted chats, group chats. Federation. Friends/groups/permissions/trust boundaries. Markup, markdown, formatting, images, emoji, unicode, code blocks. Audio, video, broadcast, multicast. Discovery. Extensibility. Federation. Store-and-forward.
> "We've been sending text over the Internet from the moment the Internet went online"
And IRC has RFC 1459, RFC 2810, RFC 2811, RFC 2812, RFC 2813, RFC 7194 and it does almost none of those things.
And you can still use IRC. But IRC is not good. At this point it's either wilfull ignorance of what people use computers for, or it's indistinguishable from deliberate trolling.
Ultimately the tech is only as good as the humans that use it. If you're chatting with folks that can't understand the difference between a thread or not, that's a meat-bag issue and not anything an RFC can fix.
- opening threads in a new window is in some hard to find and hard to remember location
- threads can’t be inlined (like you’ve described). I actually like how they’re in a pane to the right but sometimes that causes issues (as you’ve described). Sometimes it’s just that the view is too narrow.
- there’s no way to shift non-threaded comments into the thread. I’ve lost count of the number of times conversations have happened over multiple threads and/or non-threaded replies.
The way they implemented it is a complete clusterfuck.
In this case, it's just a waste of tokens as something not very unique was generated that doesn't have a real use case, or solves anyone's problem. As with many AI generated projects, I'm willing to bet that OP themselves will not using it any more in a month.
How you went from seeing somebody's post to deciding they don't care is a pretty big jump, one that isn't warranted. This kind of post seems like the old but now more elaborated form of hating on something that you don't even know what it is.
It just isn't fair to the work that has been put in.
If you tell it to "build me an application X to do Y using programming language Z", it will comply.
If you interactively ask "I need X, can you suggest some options? use an existing product ? or build something using a library ? or build entirely myself ? can you suggest alternatives with pro's and con's ? Anything I should think of before deciding ?" then you will get an entirely different answer.
Maybe we need additional modes, aside from thinking mode, agentic mode etc. we need e.g. "sparring mode" ? or does such a thing already exist ?
The repository does not even mention this software is LLM-generated.
I wonder when/if the collective realization will land that LLM's failure to retain provenance during training means the code it generates exists in an indeterminate state of public domain / IP trap.
It's valid that someone just wanted to build something like this for their own purposes.
In this case, someone shipped something that works for their use case, which is a great deal more than talking.
There's lots of approaches out there that say if you don't ship something embarrassing, it's too late.
I think this is the main point I was trying to make and our exchange helped me bring it forward, thanks.
That’s because it’s a paradigm shift, not meant to be planned or organized by humans. The prior paradigm was that there were subject-matter experts who were human. The new paradigm is that the only author of all code is to be a machine. This new paradigm requires removing all prior authors (who happen to be human), and removing them simply requires overcrowding their output such that it’s not locatable or trustworthy.
Sit back, grab some popcorn.
The system has broken under the noise.