upvote
> Why didn't they simly replace the original bad one?

Starting with version 2.1 they made it so that code points can't ever be unassigned. People were getting really upset with the removals, since they can result in old data being misinterpreted. (Er, nothing actually was removed in 2.1, but it probably didn't become official policy until 3.0… ?)

Besides, it's useful to have the bad characters in there for people who want to write posts like this one discussing them.

reply
But it's even more useful for people who want to use the actual character to be able to use it! And you can discuss using other means, not like drawings or old standard data disappears Weird absolutism re error preservation instead of striving for correctness
reply
> Not that hard to imagine, OCR existed back then?

How do you train OCR on a character that doesn't exist though? Even these days, OCR often confuses the 52 Latin characters, so I wouldn't expect as good of results with older technology and some 20k CJK characters.

reply
It does exist? It's part of Unicode!
reply
It's listed in the Unicode character database, but I doubt that it was in (m)any fonts 20 years ago, it would have been in zero printed books [0], and nobody knew the meaning of the character, so I'd argue that it's more an artifact of the Unicode compilation process than a "real" character.

[0]: Aside from the "Overview of National Administrative Districts" book mentioned in the article

reply
You don't need (m)any fonts, only need one font used by the book and you can get it the same way it was originally done - gluing some parts of other characters together. Or you could just draw it if the Unicode version is too dissimilar from the book version for a reliable OCR. (though if the character came from this single book, likely the Unicode reference is its replica?)
reply
(Talking about Japanese here)

I don't know if this is how OCR would work here, but they're made of components called radicals. It's kind of like how letters make words, but in a two dimensional layout, and one letter was wrong and a nonexistent word was made.

For example 彁 is made of parts of 引 and 歌 (the first that popped into my mind).

IIRC there's only like 200-something radicals, though some can be hard to spot due to how they overlap or nest inside each other.

reply
I find 鸚哥(いんこ) to be the easiest way to type the right hand side of 彁.
reply
OCR was slow and unreliable and was for a very long time.
reply
"slow" - wasn't like they were pressed for time. It took them almost 20years to even start the investigation! Also not certain you needed that much reliability compared to what was available to match one pic to a quality scan to get a reasonable number of candidates
reply