[1] https://archive.ph/FhG5L (the original either got deleted or login-walled, here is the archived version)
[2] https://hn.algolia.com/?q=always+bet+on+text
[3] https://news.ycombinator.com/item?id=26164001
[4] https://news.ycombinator.com/item?id=8451271
For one, there are a few different ways to terminate lines. All major operating systems now tend to use just \n, but I have older files that use \r\n (Microsoft), \r (Macintosh) or \n\r (RiscOS).
There are also different opinions on how text files end in different operating systems. In POSIX, all lines are terminated by \n, even the last one. Microsoft software still tends to insist that the last line of a file is special case that doesn't need to be terminated even now that they have otherwise adopted POSIX style line endings. In Microsoft's view, it seems that the line ending sequence separates lines rather than terminate them. Files created according to this view don't play well with tools like cat(1) if your intent is to concatenate the lines of two files, but it seems other Unix clone tools have adapted to the possibility that the last line isn't terminated properly.
Finally there's the encoding problem. I don't know of a good tool that determines the original encoding based on heuristics and re-encodes to UTF-8 but if someone does I'd love to know. If I know that the input language is English for example it shouldn't be too hard to determine what encoding the funny byte used in contractions or the funny bytes used in quotes belong to. Still, in English most of the files that use 8-bit encodings remain quite readable if you just box out the invalid bytes.
Granted, I knew that all files are UTF-8, so I didn't need to have it "guess" what the encoding was without the BOM.
[0] Windows newlines not withstanding.
As it is programs have to guess the following;
A) encoding. If ANSI which code page? If unicode which encoding? In the case of utf-16 big or little endian?
B) line endings? CR? LF? CRLF? I guess no-one uses LFCR... right?
C) number formats? 100.000 or 100,000? Date formats? is that mm-dd-yyyy or dd-mm-yyyy?
D) what human-language is it in?
E) CSV? Don't get me started...
Yes. Text is a lot easier to load, parse, make guesses about than say XLS. Yes all of the above things can be "guessed" to a greater or lesser extent.
Yes BOM exists to at least try solving the encoding question. Of course most files don't use them. Lots of tools don't support them. And they don't solve any of the other issues.
But sure, ASCII files in English with US date formats, and Windows line endings....no problems at all...
The line endings problem really isn't a problem. Pretty much every text editor out there can handle different line endings.
I don’t think number nor date formats are relevant here. For example you could have that same problem entering text into MS Word. That’s really more of an issue if you want to use text as a database rather than a document format, which isn’t really something that even plain text advocates would generally recommend.
As for human language detection, that’s a much easier problem to solve than decoding a proprietary binary blob.
CSV definitely has its warts. But it’s not like that’s the only plain text option for serialising data. (JSON, jsonlines, YAML, XML, etc). Or you could use the actual ASCII codes reserved for records, if you really wanted something that didn’t require quoting and escaping in plain text. It’s actually a pity nobody does this.
Basically it's a text-based language for embedding semantic metadata into some kind of underlying text stream. Is this interesting to you?
“Windows” newlines are also the standard in many communication protocols like HTTP and SMTP. That’s not because of Windows or DOS, it’s because it was the standard for teletypes which needed bot CR and LF. It’s arguably systems like Unix that deviated from that standard.
I agree that beyond ASCII there is a slope from well-supported to less-supported and quirky to problematic areas of Unicode. For example, HN filters many Unicode text elements like combining characters (Zalgo text) and emojis, and there is no specification to point at what it supports.
Even within ASCII, most control characters don’t have a portable meaning. So it’s really just the printable subset of ASCII, and strictly speaking not even that, given that there are regional variants of ASCII, such as the Japanese one where backslash becomes the Yen sign.
Fortunately they’re easy to test for and most languages have standard libraries that make this painless.
And in practice, still a rather large number of them, going back to the 1970's - https://en.wikipedia.org/wiki/Extended_ASCII
Symbols on paper is best.
Plain text in a file that can be opened by any computer comes close, but needs a tool.
Fancy file formats that need not only hardware but also special software are the worst.
However, there's an important tradeoff, which is the fancier formats can present information in ways mere symbols might struggle with, and can also use interactivity to improve understanding.
Don't ask me what my point is
With plaintext you can hand the file off to anyone on any device--the caveat being absolutely no one wants to be handed a plaintext script. The software I used ten years ago to write plays is long since deprecated and those files are basically unopenable. Plaintext however remains.
# INT. JOE'S BEDROOM, DAY.
JOE
This is my dialogue, yeehaw.
JOE exits, and we know JOE is exiting because this is a stage direction, and we know this is a stage direction because it is free text and doesn't satisfy any other formatting requirement.
Sharing aside, I imagine a plain text format would also be helpful for version control.
I built sharing and (sort of) versioning into the web app. Everything is saved in localStorage and on a MySQL server. You can Save or Save As, and every time you Save As it inserts a new database row and gives you a new, sharable URL.
I wanted an experience where you didn't have to login to use it, but ultimately, having a way to keep track of all those URLs would actually be a pretty decent versioning system--add diff checking between script versions and you end up with something pretty handy.
It's not baked-in to the weights of any model, but agents can write their own tools to work with arbitrary binary formats and get many of the benefits of off-the-shelf unix utilities.
I love that the Godot game engine has a textual scene description. That's a stark contrast to Unreal Engine's binary format for blueprints (visual scripting).
Oh and now you could have LLMs write the Markdown, but chances are no human reads that, or even if you do, the formatting is going to be insane. I have to keep telling Claude to give me a .txt instead of .md when asking for a context dump. Maybe .md is popular for agent readmes because they optimize around its headings.
I'm assuming any tool in this regard would expect the user to write an EBNF grammar.
I found ANTLR to be nice but it's way too convoluted to use as a tool with Java dependencies and seems to be stagnating. And tree sitter just is too convoluted and requires the added C/C++ overhead to understand how to use it.
I created Brashtag [1]. It is simpler than markdown.
> 'this is a bag named code with a blob, therefore it is code."
This would make #code{} a special bag. Right now no bag is special. Also it would be hard to find the closing } if somebody wrote unbalanced paren code in the bag like #code{ func main() { }.
Imagine you could just say #footnote{This will become footnote} and have a little program that would produce html that will show it at the bottom of the page.
See how easy it is to write programs to process brashtag [1].
I can chat with llm, interact with jupyter kernel and do literate programming from any text editor all in the same document [3].
It took me like 10 minutes to build this mermaid like thing [2].
List of frustrations:
- One day Anki crashed. --- You can just write #card{....} in a text file. Some program would read it and show flashcards in browser.
- Another day JabRef crashed. --- You can just write #bib{....} put link to a paper and its citation info there and have a program download the paper and copy the citation to bib file.
- Markdown parser processed mathjax wrong. Why keep trying different markdown parsers? Just write #h1{}, #b{} #i{} and just convert them to html tags.
- Literate programming tool I was using failed at some corner case.
- My notes were scattered across files and tools. I couldn't find anything. I tried to have a text file where sections were separated by four dashes. But that didn't allow nesting. Also, the parser I wrote for that hit corner cases. Why not just put #note{} #JIRA-334{} etc. in text file.
[1] https://github.com/PratikDeoghare/brashtag#some-programs [2] https://github.com/PratikDeoghare/brashtag/tree/master/cmd/m... [3] https://www.youtube.com/watch?v=IMXgIE0Vljg
That's how Unicode started but nowadays Unicode expresses things that are only used because Unicode itself introduced them.
And this while many important things in the "have used text in history" category have been unfinished or not tackled at all.
The fact that unicode maps the lower 7 bits to its own character set is a nice touch but none of the unicode sets are plain text.
Unicode are multibyte characters with variable byte length and endianess at play. If you read it wrong or guess the length wrong your results might be anything but useful.
> The fact that unicode maps the lower 7 bits to its own character set is a nice touch but none of the unicode sets are plain text.
That is only true for the mapping of Unicode character (code point to be exact) to UTF-8, which is an encoding of Unicode characters.
> Unicode are multibyte characters with variable byte length and endianess at play.
That is only true of the UTF-16 encoding. UTF-8 does not have endianess, UTF-32 does not have variable length per code point.
None of that is true for Unicode, because it is abstracted away from any byte representation. Furthermore, getting back to the start:
> ASCII (7-Bit) is the only widely understood charset there is. Everything beyond this point depends on the loaded charset.
ASCII also depends on how you try to decode your text. If you interpret two bytes as one character, ASCII will never be correct. There is nothing magical about one byte mapping to one character. I'd even argue that the only reason ASCII support is so universal internationally is because of UTF-8. Otherwise many countries would default to encodings where ASCII is not a subset (as they did before UTF-8 became common). So IMHO, UTF-8 and Unicode are the only widely understood encoding and character set.
I wish he'd spent a little more time on the fundamental issue of CR vs CR/LF vs LF. That's been a "plain" text nightmare since before Unicode and many of the other complexities existed.
All of these use slightly incompatible formats, the irony of which will not escape the educated reader.
Nope. Skintones are part of Unicode. It also has 33 control characters from ASCII including one that rings a bell... There are also numerous characters added by Unicode that are literally called "layout controls".
If you want a format that provides purely semantic information, then "plaintext" doesn't fit the bill.
imagine the world wide web but markdown (and decentralized). gno.land is that.
Just the other day I had opus read a 10 year old proprietary file format. The idea that we will lose the ability to use file formats is not consistent with reality.
My understanding is that HN has started incorporating AI tooling in the back-end to scan for LLM-generated submissions and source content, in an effort to discourage it and encourage human-made content (with human discussions, one would hope). Why do we keep having these kinds of articles every day? Entire "apps" are generated - see the lighthouse one - and submitted, and are plain-as-day LLM-generated nonsense.
Anyway, flagged.
* All the posts are walls of text. In many of the posts, all of the paragraphs are just about exactly the same length.
* None of the posts contain personal stories, experiences, or anecdotes.
* All of the posts are hyper-focused around advocacy of "old tech" and decentralization. Which is great, but real blogs with this many posts have at least _some_ variation in subject matter and quality. This one's surprisingly uniform.
* The blog is relatively new, posts began about two months ago and there are 17 posts, that's a little over 2 posts a week. Sure, there are bloggers who post that often, sometimes daily, but they are usually much shorter posts, tend to be shorter, or do it for their job.
* The picture of the author on the About page is very obviously AI generated.
* The author has a GitHub account that had virtually no activity until March of this year, and basically all of the commits were co-authored by Claude.
To be clear, I'm sure a real human is behind the site/posts (as opposed to a fully autonomous agent), but I'd bet dollars to donuts that each article started out its life as a very short prompt.
Worse the tools for detecting it have an insane false detection rate. ESL = FP. Good writer = FP. Garbage human slop writer = A'ok.
And, well, neither of us really knows the ground truth. Maybe the guy who has a repo filled with articles co-authored by Claude is actually not posting slop this time and HN users are wrong and Pangram flagging it is just a false positive. But I can't blame anyone for being tired of giving this the benefit of the doubt.
I'm ESL. The mistakes I make are part of what makes me human. If HN'ers are having an allergic reaction at the first hint of slop, it's because we've been drowning in it for months.
The author's picture clearly has an LLM-generated background, (and is a little tacky, IMO.) That being said, I wouldn't dismiss an article merely because someone used an LLM while setting up their blog.
Even has a simple an/a mistake that I doubt a chatbot is going to make.