So far as I can work out, literally all of the plans are bad. The "why" varies, but they're all bad.
I'm too tired to double check your maths, so I will assume correct: one likely difference even for this plan is a terminator following sun-synchronous orbit, which means you'll only see it twice a day despite the orbital period being about 90-100 minutes, and when you see it will be specifically at sunrise and sunset.
Visibility is also a question of reflection, not just size. Terminator following orbits are worse than normal satellites, because one of the tricks for reducing e.g. Starlink visibility is to tilt them as they cross the terminator and you can't do that if they're always on the terminator.
The SpaceX plans (a million small ones) becomes a glitter band in some parts of the sky and will appear visually contiguous in other parts, though I need to double check my maths and assumptions about visibility given this happens during sunrise and sunset so the sky itself is pretty bright.
1) Don’t the radiators need to have more surface area than the solar? (Unless the chips run very hot.)
2) The datacenter will have the pesky earth between it and the sun some fraction of the time, and have to either shut down or run off batteries. ~50%, assuming LEO, right? If you leave LEO, then the latency sucks, so they’re training-only clusters. At 50% the solar doubles and you need 370 megawatt hours per hour of darkness, or you run the machines 50% of the time, rebooting for each orbit. If you make the orbit shorter (so you can have smaller batteries), then they wear out faster. The batteries also emit heat. Plus, you need to double the solar so they charge while the workload is running.
3) How do they cope with cosmic rays? The standard approach is still to duplicate or triplicate all computation, or use larger/slower processes, right?
The obvious answer to each question makes the engineering design at least twice as dumb, and they stack. There are many other problems like these.
I think the big problem with the idea is that GPUs have a failure rate and even if they didn't, they become obsolete. Most orbital DC plans are really more like flying server racks with no servicing in orbit. So when the GPUs die the satellite is a flying brick. All the power equipment, all the thermal equipment, all the comms equipment now depreciates at the same rate as GPUs. Very different economics from terrestrial DCs. And that's all assuming that you can launch everything up there quite efficiently. And the satellites take time to engineer but DCs are a more known quantity.
my personal understanding is that in-orbit compute is perfectly practical up to some obvious limits like the ones you describe. a few reasonably sized clusters up there (tens of kilowatts) doing high priority processing jobs paid by the flop is a great idea. localized compute on existing satellites already does some of this but some earth observation company being able to rapidly scale up image processing for an hour is a great option to have. the really big stuff is just a fantasy.