upvote
Doing stuff isn't free. For instance, Go compiles relatively quickly for a modern language, but the biggest reason for that is that it does less stuff than most compilers... less optimization, less checking, and some stuff built into the language to avoid some of the problems with having to read lots of headers just to compile a file and other ways of doing less stuff, but mostly the key is it does less stuff, in both the good and bad senses of that.

If you want something like Rust that offers guarantees and checks and cross-checks by the boatload, it adds up. Macros, monomorphization, implicit code generation with traits and all those other things add up too. And you can't always get O(n) or O(n log n) code to implement those checks. Maybe it can be sped up and maybe there's tricks here or there, but at the Pareto frontier, a language that has more checks will be slower to compile than one that has fewer.

And that's not a bad thing or a deficit in Rust, it's just the nature of the beast.

reply
Nah, the only reason is lack of tooling.

Instead of Go, you could have reached out to complex languages with fast compilation times like D, OCaml, Haskell, Ada, Delphi, C++.

All of them have alternative implementations with fast compilation times.

D, use dmd for fast development workflows, gdc or ldc for the ultimate performance at the expense of compilation times.

OCaml, use the REPL or bytecode interpreter for fast development times, the full blow compiler for ultimate performance.

Haskell, use the REPL, GHCi for the fast development cycles, GHC for the release build.

Ada and Delphi, have had fast implementations since forever, although Ada/SPARK is indeed somehow expensive.

C++, yes it isn't a mistake. Use Live++, VS hot reload, coupled with binary libraries, or a REPL like CINT (nee ROOT), binary libraries for dependencies, incremental compilation and incremental linking for the development workflow.

The problem with Rust isn't the language itself, rather the ecosystem currently lacking such kind of options being available.

reply
deleted
reply
Dlang is my weapon of choice and I prefer it to rust.
reply
including c++ with all of those qualifiers would be like me saying rust is fast with sccache, subsecond, cranelift, and making every single module a separate crate.
reply
The difference being the adoption culture across the ecosystem.

C++ game engines also don't need Bevy like tutorials, because most studios aren't compiling them from scratch, and tools like Live++ or VC++ hot reload are relatively easy to use.

reply
unreal engine is compiled from scratch (depending on your definition of "from scratch") and has notoriously long compile times. heck you have to recompile the editor to make certain gameplay code changes. and it's by far the most popular AAA game engine on the planet. if it was so easy to fix with your suggestions they would have done it already.

but hey, just add subsecond and switch to cranelift. problem solved :)

reply
Unreal Engine has an installer, and they have done my suggestions, as they are the main customers of Live++ [0], advocates of tooling like Blueprints and now Verse for doing full games.

Many studios have delivered games with little to no changes to the underlying C++ code.

[0] - https://dev.epicgames.com/documentation/unreal-engine/using-...

reply
so... c++ compile times are fast, because you can simply not compile c++ and instead use blueprints? seems odd since we're talking about c++/rust compile times. turns out you can make rust compile times infinitely faster by not compiling it too...

but when you do have to compile c++, it can be real slow. maybe not as slow as rust, but it's not in the same league as eg. go, c#, zig, etc.

ue5's hot reload is by no means perfect either - many gameplay code changes require recompiling the editor. exposing c++ properties on a blueprint? recompile. modify a constructor? recompile. changing parameters, return types, or adding/removing UFUNCTION or UPROPERTY macros? recompile. you can simply search the web to read the experiences of thousands of ue devs complaining about slow compile times and workflows. it's the same with unity btw, search "reloading domain unity".

also verse is not in unreal engine. no studio is using verse for ue games. not sure why you threw that in.

i get you love c++ and dislike rust, but you aren't really making a good case for c++ when you group it in with other languages that have fast compile times ootb, make arguments that involve not using c++ (like blueprints, lol), and rely on brittle third-party hot-reloading solutions.

reply
Rust wants badly to have its cake and eat it too. I get it, but it's not the sensibilities I have. The rustc compiler, if it really can achieve everything all at once, will be a beast like none other, making C++ compilers look simple by comparison.

I'd love something halfway. Go is maybe a bit radical in some regards, but also, with the news of the new SIMD package for Go, it has occured to me just how little I missed having things like, say, autovectorization.

(I know also that some people have tried halfway, but the big thing is figuring out how to keep a relatively simple type system that can still support a borrow checker. Even if there is some middleground, is it truly worth it? As nice as it sounds, I've been more skeptical. Go seems to exist in a very narrow space where its simplifications barely can be made to work.)

reply
The middle ground is pretty much nim.
reply
Yeah doing optimizations in a compiler on things like loops, addition, or string layouts is never free.
reply
deleted
reply
And since you bring up Go and contrast the compile time with Rust, it's been experienced at Google (and also Volvo and other places) that Rust and Go teams are as productive whereas C++ is less than half as productive: https://www.youtube.com/watch?t=27012&v=6mZRWFQRvmw&feature=...

So the focus on Rust compile time is misplaced. It's not a big deal in terms of overall productivity.

reply
I don't think that follows - just that compile times aren't the only thing that matters. Maybe rust could be twice as productive as go if compile times went to zero for instance (or maybe not - just disputing that the evidence proves the claim here).
reply
But it can't go to zero, that was the point -- at some level you have to pay for what you get. I'm not saying the compile times couldn't be better, but Rust can't be Go in terms of compile times if it wants to offer the guarantees and options it does. Meanwhile if Go wants to offer more options and stronger guarantees, it won't be able to preserve the short compile times everyone points to.
reply
The Cranelift backend is extremely fast. Most of the slowness is all the crazy optimizations that LLVM does + other things (generating debug info etc).
reply
For symbol heavy projects, linking is a surprising bottleneck.

Some ideas to speed up compilation by not evaluating items that are not used might pan out significantly for big crates in your dep tree (that behavior might never be stable because that would allow items with compile errors in a crate that would still let your application compile, which is against the Rust approach). The same work to do that would also allow overlapping of crate evaluation between different rustc instances called by cargo, as it would require partial evaluation of crates (to do name res only and gather the symbols needed from its deps).

Another thing is that stable rust doesn't treat macros as idempotent (because that wasn't a requirement from the start, there are crates that do dynamic IO to generate types), but if they are then incr comp can be faster by not evaluating them unnecessarily.

I know people are working on a bunch of different strategies to improve both first and incremental compile times, and I'm looking forward to the fruit of their labor.

Compile times are rarely the bottleneck for me, but that doesn't mean I won't welcome any improvements on that front.

reply
> For symbol heavy projects, linking is a surprising bottleneck.

Is that still true with modern linkers like `wild`? Seems like it can link Chromium in 1-2s.

> Another thing is that stable rust doesn't treat macros as idempotent

This seems absolutely insane to me given how pervasive macros are in Rust (I wonder how much of incremental compilation time is just repeated serde derives?). Obviously we can't just blindly treat all macros as idempotent, but an opt-in attribute on a macros that pinky-promises that it is seems like it ought to be pretty easy to implement?

Maybe somebody's looked into it and it doesn't help much? But I've heard (unverified) rumours of the opposite.

reply
> Is that still true with modern linkers like `wild`? Seems like it can link Chromium in 1-2s.

Wild improves things significantly, but it really depends on the kind of project. For some a third of the time can be linking. And not everyone configures their environment to use a different linker. Another thing that can easily improve performance is changing the global allocator, but I've seen teams that measured 10% perf improvements in their own metrics decide against going with anything other than the "default".

I fully expect that if wild delivers an effective incremental linking architecture, then rustc will be able to produce the patches directly (instead of wild having to produce patches from two versions of the object files), which would mean both that rustc is producing less LLVM bytecode and that wild gets to do what it (will) do best and make linking as fast as mechanically possible.

> This seems absolutely insane to me given how pervasive macros are in Rust

There was some work done on this front, but I haven't kept up to date on the current status of that. There was a PR showing promise https://github.com/rust-lang/rust/pull/129102 (later landed as https://github.com/rust-lang/rust/pull/145354, 10% on a specific serde-heavy crate, reports of 32% improvements in the original PR). The tracking issue doesn't have any updates https://github.com/rust-lang/rust/issues/151364, but you can try out nightly with -Zcache-proc-macros to see what the effect could be on your projects.

Part of the problem I see is that crates will have to opt-in (and we might be able to change the default over an edition boundary) to get the perf benefit, and it might require a granularity lower than "a crate" (which would complicate implementation). I haven't seen additional discussions around these design considerations (needed to stabilize), which will need to happen before people can see progress.

Sadly, Rust is, like many open source projects, a show-up-o-cracy: you need a motivated (group of?) individual(s) to deliver a feature to fruition, and if the person driving a feature is fine with using nightly for their purposes, and only needs a subset of a feature, they will drive the feature to that state (lets say 80% completion), and then the feature will linger (until someone else with enough motivation to see the other 80% through shows up). This is exacerbated because the project is unwilling to have 80% solutions on stable unless the state is very clearly not going to preclude other future work, or the path to completion is 100% visible (but if that were the case, then it would have been completed already). So we end up with situations like the Allocator APIs.

reply
(I did a quick test, and I'm seeing between 15-20% for https://github.com/servo/stylo which is a macro-heavy crate. I guess it only matter for crates you're actually editing though, macro will already be cached as part of "entire crate is unchanged")

> Another thing that can easily improve performance is changing the global allocator, but I've seen teams that measured 10% perf improvements in their own metrics decide against going with anything other than the "default".

Yeah, although I think there's more of a trade-off there. Alternative allocators can add significant amounts of compile time. Whereas, modulo maturity, I think a faster linker is more of a pure win. I would imagine we will make wild the default linker at some point if development continues on the trajectory it seems to be on.

reply
Cranelift's biggest blocker for me is that it completely breaks debuggers:

https://github.com/rust-lang/rustc_codegen_cranelift/issues/...

reply
I’m not sure if Cranelift support will ever become stable. It’s really complex, it has tradeoffs like you mentioned, and it requires its own backend to be maintained. TPDE looks like something that could more realistically be stabilized as it just slots into the existing LLVM support.
reply
Zero-cost abstractions aren't zero cost in compilation time. High-level abstractions translate to a lot of boilerplate that the compiler has to optimize out.

In unoptimized builds often the linker is the bottleneck. Rust/Cargo can parallelize most of the build, generating tons of code and debug info, but then the poor linker has to consume all of it at once. The object/exe formats were designed in ancient times, so they're hard to build incrementally or in parallel (some linkers are trying).

reply
Is there any experimentation with new ABIs? I mean certain linker flags like --f-lto literally hijack the linker protocol to dump an AST into the backend.

At least for the fully static binary part of rust, there should be some optimizations there w.r.t. compilation. Sure you're not going to interface with shared libraries well but maybe a small experimental feature for fully owned projects? Idk.

reply
I once heard (and don't know if it's still true) that the biggest compile-time sinks are macros and codegen.

Both are kind of outside the Rust compiler's influence. Macros can be almost arbitrarily complex: you pay for what you order. Codegen is LLVM, and that's a fixed choice. You can use Cranelift to get around it, but then you pay elsewhere.

Also, generics and monomorphization regularly come up in these discussion, while common wisdom seems to be that cost for the additional static analysis over other languages like C++ is no a major contributor.

Regardless, it's nice to see performance improvements in the compiler, even if you have to cooperate to benefit from them (e.g. by keeping your macros light and use less generics).

reply
Contrast this with zig, where compiler performance is so important that it has influenced the language design.

Zig and Go are similar on that front.

reply
Ultimately it's the fact that compile time was not a first-class consideration during the design of Rust's important features. There's only so much you can do to mitigate the consequences.
reply
Compilation times were a consideration, but they were a subservient consideration to the prime considerations of 1) being as memory-safe as GC'd languages, 2) exploring the limits of statically-guaranteed thread safety, and 3) being as fast and memory-efficient as C and C++. If Rust had been willing to compromise on those goals and instead had just, for example, used a virtual machine with pervasive garbage collection and been designed for dynamically-dispatched generics then you'd get a lot of compilation time reductions for free, but the world didn't need another Java, it needed a more secure systems language to stand up against C++ where every previous challenger had failed.
reply
No, Rust is not at the paretto frontier for your mentioned 1, 2, 3 and also 4 compilation speed. It could have all of those things and also just have faster compilation. Rust devs have talked about how they have some regrets about not optimizing more for compilation, but instead another attribute got optimized in terms of paretto efficiency - the language's development time. Arguably I think that is the single worst trait you can optimize for in language development, because by saving a little time developing the language you cost humanity, let's say hundreds of millions of hours of development time downstream from you, with millions of users being less productive.
reply
First of all, the fact that the article talks about "speeding up" the rust compiler doesn't automatically mean that the compiler is "slow"[0].

Now, is rustc slower than e.g. clang? by how much? why?

Those are different (and complicated) questions. It really depends on what you're compiling, but I'd say rustc can be 1-5x slower (maybe more at times?).

The reasons are many and varied, but in general rust compilation is slower because the compiler is doing way more things compared to C (monomorphization, complex trait resolution + type inference, borrow checker..)

[0]: Also I'd argue that "slow" without a concrete point of reference is a meaningless term in this context.

reply
> the fact that the article talks about "speeding up" the rust compiler doesn't automatically mean that the compiler is "slow"

Well, if it weren't "slow" for some definition of "slow", nobody would bother speeding it up, would they?

reply
The go compiler is considered fast, yet Google spends effort speeding up the go compiler. Not to the same extent, but Google has a lot of go code, so it does make a difference.
reply
Its similar to C++ compilers, there is no one reason. It is a language that tries to optimize a lot, its a big language, it does safety checks, it uses llvm which is a bit slow, its a language that makes use of generics which generate extra code etc etc.
reply
C++ compilers have it better, despite the fame, because on the C++ world, people like their binary libraries (not compiling the world from scratch on each git clone), there better support for incremental compilation, and incremental linking, there are interpreters and hot code reload tooling available.
reply
> there better support for incremental compilation

What do you mean by this?

reply
Some C++ compilers and linkers have better support for incremental compilation and linking, and do it more fine grained than Rust, e.g. VC++.

Some tooling ideas go all the way back to when C++ vendors started adopting ideas from Smalltalk and Lisp, e.g. Energize C++ or Visual Age for C++ v4.

Or build systems like ClearMake from ClearCase, where the object files and binary libraries are shared across everyone on the cluster with the same views (ClearMake speak for what files/branches are selected).

reply
It's doing static analysis that many other languages don't do at compile time.
reply
This is not actually the main reason, most of the time.

Generics/monomorphization and how iterators work results in a lot of compiler bytecode that has to be churned through. More bytecode = longer compilation. It increases the size of the (debug) binaries, the debuginfo in general, causes performance issues with debug binaries in some situations unless you bump the optimization level, causes more IO, etc.

reply
If you don't use generics and monomorphization, then what explains it?
reply
It's unlikely that many people are using Rust without using Option<T> or Result<T, U> a fair bit. Idiomatic Rust fundamentally uses a lot of generics. And a basic for loop expands into quite a lot of intermediate representation due to Iterator.
reply
Other languages support generics, monomorphization, and iterators (e.g. Zig or D), but they're not as slow as Rust to compile, what's the reason for that?
reply
It's likely easier to answer for specific languages

All three of those words can mean "just like Rust" but equally "Not at all like Rust" for different languages.

Two examples to contrast: In C++ the iterators are basically a pointer analog (in some cases they're just literally pointers) and that's a very difference "feature" but it's still definitely iterators. In Ginger Bill's Odin, the iterators are a function, possibly generic, which returns a pair, the next item and a boolean telling you whether the iterator was exhausted.

reply
Zig has their own backend, monomorphization is afaik slow because LLVM itself.

On top the borrow checker and other features are non-existent in these languages.

reply
AFAIK it's because of trait solving. And also not every language does monomorphization.
reply
Trait solving isn't a bottleneck for Rust compilation unless you're doing some extremely cursed things like trying to implement Doom in the type system. At the end of the day it's still mostly that Rust just generates a ton of IR for LLVM to chew on (because of monomorphization), then LLVM takes a while to process all of it (because LLVM is designed primarily to produce high-performance code, often resulting in reasonable tradeoffs against compilation speed), and then the linker takes a while to connect it all together. With an alternative backend to LLVM (e.g. Cranelift) you could choose to design it with a greater emphasis on compilation speed (likely trading off generated code quality in the process). And with a more radical vertically-integrated model where Rust controls the linker you could do in-place linking and nearly completely skip the final step (at least for debug builds), but that's a big change compared to the classic C-style build model.
reply
Not sure why you are being downvoted here.
reply