upvote
They're a bit old and missing some details, but I like Agner Fog's manuals.

https://www.agner.org/optimize/

reply
I've not read Fog's Optimizing software in C++ but I see it's freely available there as a PDF (182 pages). Looks like a great resource on these topics.

https://www.agner.org/optimize/optimizing_cpp.pdf

reply
Learn which instructions SIMD nicely (sqrt / fabs, etc). Use ternaries in loops for masking. Use trig identities and lookup tables (don't recompute sin(3t) when you can use two vector multiples using a table of sin(t) eg. sin(t) * sin(t) * sin(t)). Use divisible constexpr constants in loops to eliminate the SIMD tail. Be careful with type casts and floats. `float x; x += 0.5` will introduce *cvt instructions even if the compiler statically knew better otherwise (use 0.5f). Compile with --fast-math and friends so errno doesn't invalidate your SIMD pipeline.
reply
Most applications (including most applications that care about numerical performance) should not use -ffast-math.
reply
That has a similar problem to the article, it's trying to fit far too much into too small a format.

What you've written mostly makes sense to someone who already has a solid understanding of SIMD and of C++ (although I can't say I follow all of it), but the target audience is people who don't. For them, each point needs a much lengthier explanation.

reply
Likely the best tip would to `objdump -d` and inspect the assembly then checking performance counters. Prepending (__attribute__((used)) will allow you to inspect your functions.

A quick restrict example:

    #define fn __attribute__((used))

    fn void copy1(int* to, const int* from, const int size)
    {
        for(int i = 0; i < size; i++)
            to[i] = from[i];
    }

    fn void copy2(int* to, const int* from)
    {   
        constexpr int size = 1024;
        for(int i = 0; i < size; i++) 
            to[i] = from[i];
    }

    fn void copy3(int* restrict to, const int* restrict from)
    {
        constexpr int size = 1024;
        for(int i = 0; i < size; i++) 
            to[i] = from[i];
    }

    gcc test.c -c -O3 && objdump -d ./test.o
copy1 is 52 lines, copy2 is 28 lines, copy3 is 2 lines (just a call to memcpy).

This is a good starting point for self teaching. The impact of your TLB, L1, and overall instruction count (with IPC) can further be measured with `./perf stat -d -d -d ./a.out`. If you want a quick rule of thumb, no instructions are fast instructions.

reply
This seems like domain specific advice.
reply
Do you have any recommended essential reading for this?
reply
I'm no expert in this stuff but:

creata's comment [0] mentions the works of Agner Fog, which seem very good, and are freely available.

I haven't read C++ High Performance [1] but it looks like it covers the sorts of topics you'd expect, although it looks like it doesn't cover computer architecture in detail e.g. branch prediction. There are books on that too, of course.

[0] https://news.ycombinator.com/item?id=49868657

[1] https://www.packtpub.com/en-us/product/c-high-performance-97...

reply