Rendered at 09:15:34 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
asveikau 15 hours ago [-]
This article reminds me of performance advice I was starting to see in the 2000s decade. Basically it was to not introduce a bunch of pointer heavy data structures to get lower algorithmic complexity. Stuff it all into a vector. You will use some algorithms that the computer science textbook will say it's slower, but if it fits all in cache it doesn't matter. The cache misses following pointers all over town hurts you more.
tialaramex 9 hours ago [-]
This is partly because the C++ stdlib looks like they've given you all the basic tools you need - unlike the C standard library - yet in fact many of these tools are hopelessly obsolete. It's not quite PHP's "fractal of bad design", these are all reasonable tools... if it's 1985. The linked lists make sense on hardware where five pointer fetches and five consecutive memory reads cost roughly the same - 1985 hardware.
The growable array type std::vector<T> is least impacted by these archaic choices out of the tools in the box you're likely to reach for. So it will make sense very often to choose this type first.
adrianN 7 hours ago [-]
Linked lists are great data structures for the use cases where you need their properties. It’s just that you don’t encounter those scenarios very often in most kinds of software.
someonebaggy 2 hours ago [-]
You say not very often, but the true usefulness is almost never. Extremely little. One in a trillion times. Even when you think linked lists would be faster, they usually aren't.
socalgal2 2 hours ago [-]
vs what? What languages make these choices better?
tialaramex 2 hours ago [-]
So lets take C++ versus Rust
The C++ 23 containers are: array, vector, deque, forward_list, list, set, map, multiset, multimap, unordered_set, unordered_map, unordered_multiset, unordered_multimap
Firstly, Rust doesn't consider "array" a library type here, in C++ the language has built-in arrays but they're very poor because they are the C arrays - so you use the library feature to get good arrays. In Rust they... just fixed the language, because duh.
Next thing you'll notice is that C++ has lots more of these types, about twice as many. I stopped at C++ 23 because in C++ 26 they added even more. These are a significant maintenance burden and of course having more means in practice maintenance gets worse. But this could be good if these types were all high quality and kept that way.
All of the C++ unordered containers are the same crap hash table design but with slightly different parameters. The Rust HashMap and HashSet are Swiss Tables though they do not promise that and if a better design comes along they will probably switch. C++ can't change the design because the API welds them to a very specific shape for this data structure, a shape which delivers bad performance on any vaguely modern hardware.
std::deque is the most horrible surprise. A modern programmer who has thought about it at all is expecting a type like Rust's VecDeque. Generalise the amortized growable array from the language to use it as a ring buffer. Cheap push & pop at both ends, canonically use it as a FIFO but also practical in lots of other situations. But that's not what std::deque is at all, instead inside it's an array of links to small arrays. On MSVC it's effectively a linked list again because those inner arrays contain only one item due to ABI considerations.
std::set and std::map are very principled red-black trees. I say principled because in practice this is too expensive on modern hardware because (say it with me) it spends too long chasing pointers up and down your tree. Rust's choice here in BTreeMap and BTreeSet packs more data in each "node" on the tree, which makes the big-O worse but the practical performance better. Figuring out how to best do this for the general case is an active area of research but "I bet a pure red-black tree will be fast" is not a good guess for the past several decades.
Finally std::forward_list and std::list are the singly and doubly extrusive linked list types. The thing you most likely have seen in some high performance software is an intrusive linked list, and C++ doesn't provide those. In an intrusive linked list each item in the list itself links to where the next (and for simple double links also the previous) item is, so the item needs to know it's in a list [in some systems more than one list, thus more than one set of links]. C++ provides extrusive linked lists where those links live in a separate object and so the items in the list don't know about this at all. Rust provides only a doubly-linked extrusive list exactly like C++ std::list, but again, this almost certainly isn't what you wanted, you most likely do not need a linked list and if you do have a good reason for a linked list you probably want an intrusive linked list.
smj-edison 8 hours ago [-]
It's interesting because I tried to follow this advice when I wrote my own interpreter, but either 1. I just had a bad intuition and it's gotten better, or 2. It's trickier with interpreters when you have thousands of objects.
For example, since I allowed for objects to be shared between threads, I decided to use struct of arrays so the reference count, metadata, and value would be stored in separate cache lines. This ended up hurting me because object initialization touched three separate cache lines (obvious in hindsight, but the advice of using SoA failed me here). I also heard that you want to pack your values as tight as possible, so I used a packed string index, but then I ended up with integer division to unpack the string (also a mistake, but again the advice failed me). I used a custom allocator to avoid indirection with lists (list items were allocated directly after the list head), but then I had heap fragmentation and the implementation complexity exploded.
Anyways, I am now happily using two to three levels of indirection in my data structures, large structs, and malloc for individual objects, and it's still been faster in my end to end testing. So maybe this is unique to interpreters, and maybe I could have done it better, but the suggestions don't automatically apply in my experience.
someonebaggy 2 hours ago [-]
"Good advice tends to come with a rationale so you can tell when it becomes bad advice" - Raymond Chen
I'd say the rule was followed in this case - the rationale of SoA is to reduce cache misses when iterating all objects and only using some of the attributes, which is something games do all the time, but it's bad if you are always accessing one object at a time. Maybe an array language interpreter would have luck with SoA.
carlmr 2 hours ago [-]
>This ended up hurting me because object initialization touched three separate cache lines (obvious in hindsight, but the advice of using SoA failed me here).
I mean this is kind of what happens with any advice that has nuance to it, that's not carried with the advice.
E.g. if you have a point in 3D space with x, y, z coordinates. Array points as SoA of individual dimensions makes sense only if you do a lot of averaging and such on the individual dimensions.
If you mostly use the 3 coordinates together, SoA will have bad caching behavior.
So the better advice would be to try to keep things that are used together in the same cache line, whether it's on dimension or all 3. Usage makes the difference.
spider-mario 14 hours ago [-]
Or even, stuff it into several parallel vectors (structure of arrays instead of array of structures).
stackghost 14 hours ago [-]
I too am in the "premature optimization bad" camp.
Beyond the low-hanging fruit like ensuring you aren't creating O(n^2) complexity by accident, I think C++ is fast enough/has mature-enough compilers that by the time you're worrying about cache hits materially affecting performance, you're probably also sufficiently staffed and capitalized to pay people to A/B test that performance.
Pannoniae 10 hours ago [-]
1. Compilers barely do even basic optimisations such as interprocedural register allocation when faced with non-trivial code.
You often also need the most aggressive optimisation settings, LTO or even PGO enabled for many of these.
2. Virtuals are, with the exception of PGO, mostly a black box i.e. you get a hard optimisation boundary, no inlining at all.
3. The C++ standard library is usually comically slow (yes, even compared to Java/C#/the likes) so if your project uses std::vector and the such instead of specialised libraries, you've already lost at the beginning.
4. If you don't pay attention to performance from the get-go, the approximate amount of autovectorisation you'll get is close to zero. Some compilers are better than others (Clang>MSVC for example) but I've seen codebases with 8 figures of LoC where the number of vectorised divides/multiplys was like less than ten when you dumped the object listing. In the whole program.
5. Since aliasing and other optimisation barriers (you didn't use restrict or manually hoist, did ya?), it's not uncommon for large C++ programs to spend a third of their runtime doing atomic increments because shared_ptr is supposedly cheap and who cares about lifetimes anyway.
6. If you're targeting Windows, the default new operator / malloc is also comically slow. Luckily that one is fairly easy to fix with installing mimalloc and deploying the hijack dll, but the negative effects on cache by the fragmented allocations is also significant.
someonebaggy 2 hours ago [-]
Regarding 3, you're probably using MSVC in debug mode. Switch to release mode and rerun your benchmarks.
Pannoniae 2 hours ago [-]
No I'm not. And it's not just MSVC-specific either, they're just not very good.
std::vector doesn't have trivial relocation so any type with a destructor ends up doing elementwise destruct+construct instead of a memcpy.
std::map and std::list are memes and if you use them you're giving your CPU the 1995 treatment with all that pointer chasing.
You thought std::unordered_map is better? Well, actually not because node stability, so it's still chained-bucket, you almost always want to use a flat map like boost::unordered_flat_map or the abseil/eastl version.
<random> is hard-to-use and isn't very performant, std::regex is "you might as well write it in Python and it'd be faster", <iostreams> is virtual calls galore, both the formatting and the stdio functionality are slow.
The conveniently-named std::function is a very general device resulting in a heap allocation and usually a virtual call, there's specific optimisations but don't rely on it.
The STL string manipulation functions are also usually slow, they check the locale for string manipulation rules.
The floating-point functions set errno preventing vectorisation and emitting branches in your straight-line float code unless you use fastmath (the thing people tell you never to do) or one of the more fine-grained compiler-specific switches to turn it off.
std::shared_ptr is Arc<T>, not Rc<T> and eating the cost of atomics can add up in many situations especially with all the other memory traffic going on.
std::variant and std::visit are also not very fast either.
std::filesystem as a whole also has several pain points like iteration which is like a magnitude slower than the native APIs, std::chrono isn't much better either
std::error_code sounds like a simple integer or even a struct.... lol no guess what, more virtual calls
ahartmetz 9 hours ago [-]
Regarding 5., I have fortunately never seen a program overusing shared_ptr like that, but when I recently had a performance-sensitive use case for shared_ptr, I found boost::local_shared_ptr with non-atomic reference counting.
stackghost 4 hours ago [-]
>If you're targeting Windows, the default new operator / malloc is also comically slow. Luckily that one is fairly easy to fix with installing mimalloc and deploying the hijack dll, but the negative effects on cache by the fragmented allocations is also significant.
Surely nobody outside microsoft is doing serious work targeting Windows any more are they? Isn't that a dead platform? I read somewhere a while back they're now below 60% market share.
Agentlien 4 hours ago [-]
There's an enormous amount of serious work targeting Windows. For example, the video game industry still has a very strong PC user base and is very concerned with performance.
nnevatie 2 hours ago [-]
> C++ is fast enough
Yeah, it's really not.
There are multiple areas of work, where C++ can be considered a glue language. The high-performance work is then done in explicit SIMD (intrinsics, ISPC, etc.) and/or GPU-targeting languages such as CUDA or Vulkan.
In these areas of work, high performance is part of the design and not something that can be easily added as after-thought.
Also, relying on optimization features such as compiler auto-vectorization is way too finicky - your hot-loop performance may completely break without anyone noticing by someone changing a trivial-looking part of a loop.
Agentlien 4 hours ago [-]
This depends so much on what your work and industry is. I hear A/B and immediately think this is alien and inapplicable to me.
I work in game development and for the last six years I've spent most of my time specifically on optimization. A lot of that effort has been focused on cache behaviors. Not because it's fun, but because it's often the difference between being able to ship the game on weaker hardware (e.g. Nintendo Switch) or not.
djmips 7 hours ago [-]
You really aren't in the "premature optimization bad" camp you just don't realize you optimize all the time but justify it as obvious. The main thing to know is that what is 'obvious' isn't unless you are profiling.
stackghost 4 hours ago [-]
"Good design" is not what I would call optimization.
As a contrived example: there are specific cases when a particular non-quicksort algorithm is optimal. In almost all real world scenarios, though, you're just going to say fuck it and use quicksort until profiling determines that the sort is the bottleneck.
Unless you already have specific knowledge that your data comes in a particular shape, defaulting to quicksort is good design (IMHO). Worrying about pathological sorting before you've seen benchmarks is premature optimization.
djmips 3 hours ago [-]
Semantics - not worth me arguing about
stackghost 3 hours ago [-]
And yet you posted an argumentative reply to my comment.
asveikau 12 hours ago [-]
I don't know if this advice is strictly advocating to avoid premature optimization. Many problems are modeled intuitively with lots of tiny allocations and pointer heavy structures, and this advice is saying to avoid that.
I think it's more like: prioritize cache locality over big O compexity.
srean 13 hours ago [-]
And how would that staff have learned it ?
stackghost 12 hours ago [-]
I don’t understand the question. Are you implying someone cannot know how to do something in a particular codebase unless they’ve already done it on that same codebase?
jeffbee 12 hours ago [-]
I don't really agree because it's so hard to reform a full application that's been written without regard to performance, after it's been written. You really need to pay attention from the beginning.
bluGill 13 hours ago [-]
While this advice isn't wrong, it is misleading. In my benchmarks std::map beats vector after 9 elements. Less than that and linear search is better but branch prediction and cache loading is very good.
Run your own benchmarks on your own data of course. Also map is not considered the best key value store.
mandarax8 13 hours ago [-]
Because in your benchmark all std::map nodes were allocated in succession, most likely being placed in adjacent memory locations...
This likely won't be true in a real application with a non-trivial allocation pattern.
bluGill 13 hours ago [-]
That might or might not be true in the real world. Often in my applications I'm creating at startup and then referencing later.
Still a custom map that allocated a bunch of nodes would be a useful optimization.
einpoklum 13 hours ago [-]
Actually it really depends, because allocators can also be kind of smart (and you don't have to use the default allocator).
And then, on the other hand - I really doubt GP's map beats a vector, with all of those pointers bins and stuff, in a non-contrived benchmark with 10 elements.
Finally - it's not either-or: There are better hash maps whose memory is sequentially allocated and/or are otherwise cache-aware. And there are data structures geared towards parallel execution on multiple threads; and towards SIMD; etc. etc.
someonebaggy 2 hours ago [-]
That is not a plausible result, sorry. Perhaps you're using the painfully slow MSVC debug mode vector?
danbolt 10 hours ago [-]
I’ve definitely been on teams where they ran the numbers, and found that they were mostly working with smaller containers, and std::vector was the way to go.
If you’re down to that sort of decision-making, you have to measure.
Jeaye 13 hours ago [-]
While we're here, has anyone seen any resources related to data-oriented design when GCs are involved? So much of data-oriented design is arena-focused, but that's not always possible, when the lifetime model of the code requires a GC (for whatever reason).
I feel like the DoD movement is a slow-moving, but big, change through how systems programming is done, but that there's still insufficient material for how to do this in different scenarios. I would really like to apply this more to my areas of work, which are also in C++, but there seems to be a gap between what they're presenting and how it can be applied.
More specifically, I'm using C++ to build a dynamic programming language runtime for a Clojure dialect. That runtime is required to be garbage collected, type-erased, and highly polymorphic. So I surely can't just SoA or AoS everything. Yes, I can pack my data, and I can avoid the GC whenever possible, both in compiler/runtime code and in generated code via escape analysis. But what about everything else, which is the 80% or more of the system? It could be that this runtime is too far at odds with DoD, but I generally see things as a gradient rather than black and white.
smj-edison 6 hours ago [-]
Not sure if this is relevant, but I was in a discussion about heap layout a while ago, and one of the commenters was talking about how they were able to have classic lisp-style linked lists with decent performance by using a copying garbage collector: https://ziggit.dev/t/memory-layout-suggestions-for-tcl-inter...
Or were you referring more to all the intermediate allocations that aren't the object heap? V8's zones are interesting in this area, because they're like an arena, except that they're only partially reset when a zone ends, so zones can nest inside each other.
chombier 51 minutes ago [-]
A few weeks back there was this nice little database query language as a language library [1] posted on HN, with the constraint that everything had to be 6NF.
At the time I thought this would naturally fit in a data-oriented design/ECS system to run complex queries. I wonder whether anyone has tried this before and whether this actually works in practice?
There's no mention of branch prediction, or context switching, or synchronisation. Depending on what you're doing, they could be very consequential. There's only very brief mention of parallelisation with threads and with SIMD.
High-performance programming is a big topic. The scope is far too broad for a single blog post, which naturally gives only cursory discussion of C++ and computer architecture. The article isn't bad considering, but I do think it's the wrong format. A blog series, or even a book, would be more fitting.
creata 16 hours ago [-]
They're a bit old and missing some details, but I like Agner Fog's manuals.
I've not read Fog's Optimizing software in C++ but I see it's freely available there as a PDF (182 pages). Looks like a great resource on these topics.
Learn which instructions SIMD nicely (sqrt / fabs, etc). Use ternaries in loops for masking. Use trig identities and lookup tables (don't recompute sin(3t) when you can use two vector multiples using a table of sin(t) eg. sin(t) * sin(t) * sin(t)). Use divisible constexpr constants in loops to eliminate the SIMD tail. Be careful with type casts and floats. `float x; x += 0.5` will introduce *cvt instructions even if the compiler statically knew better otherwise (use 0.5f). Compile with --fast-math and friends so errno doesn't invalidate your SIMD pipeline.
creata 16 hours ago [-]
Most applications (including most applications that care about numerical performance) should not use -ffast-math.
someonebaggy 2 hours ago [-]
Why not?
MaxBarraclough 16 hours ago [-]
That has a similar problem to the article, it's trying to fit far too much into too small a format.
What you've written mostly makes sense to someone who already has a solid understanding of SIMD and of C++ (although I can't say I follow all of it), but the target audience is people who don't. For them, each point needs a much lengthier explanation.
glouwbug 16 hours ago [-]
Likely the best tip would to `objdump -d` and inspect the assembly then checking performance counters. Prepending (__attribute__((used)) will allow you to inspect your functions.
A quick restrict example:
#define fn __attribute__((used))
fn void copy1(int* to, const int* from, const int size)
{
for(int i = 0; i < size; i++)
to[i] = from[i];
}
fn void copy2(int* to, const int* from)
{
constexpr int size = 1024;
for(int i = 0; i < size; i++)
to[i] = from[i];
}
fn void copy3(int* restrict to, const int* restrict from)
{
constexpr int size = 1024;
for(int i = 0; i < size; i++)
to[i] = from[i];
}
gcc test.c -c -O3 && objdump -d ./test.o
copy1 is 52 lines, copy2 is 28 lines, copy3 is 2 lines (just a call to memcpy).
This is a good starting point for self teaching. The impact of your TLB, L1, and overall instruction count (with IPC) can further be measured with `./perf stat -d -d -d ./a.out`. If you want a quick rule of thumb, no instructions are fast instructions.
asveikau 15 hours ago [-]
This seems like domain specific advice.
Jeaye 15 hours ago [-]
Do you have any recommended essential reading for this?
MaxBarraclough 14 hours ago [-]
I'm no expert in this stuff but:
creata's comment [0] mentions the works of Agner Fog, which seem very good, and are freely available.
I haven't read C++ High Performance [1] but it looks like it covers the sorts of topics you'd expect, although it looks like it doesn't cover computer architecture in detail e.g. branch prediction. There are books on that too, of course.
I write in C++ almost every day but never have the need to optimize for speed. Even when you write straightforward code it's already blazingly fast.
gbin 17 hours ago [-]
It is probably very domain specific. In robotics for example everything is a zero sum game: CPU, memory bandwidth, GPU, battery life etc ... So it is really a topic, probably true for anything embedded actually. Some other offline applications: HFT, Telco etc..
I wish the GUI apps devs respect more the laptop resources they are running on, don't get me started on the 4 instances of chrome I need to run just for discord, signal etc ...
serbuvlad 13 hours ago [-]
GUI engine developers need to trade EVERYTHING for execution time, otherwise JavaScript would simply not be fast enough to handle modern applications.
If your device has enough resources to power V8, modern GUIs are certainly very pleasant and snappier than a more minimal GUI like HN. Otherwise they are horrendous and very laggy.
gbin 13 hours ago [-]
I don't know if you got my point. Starting a multi gigabyte machinery for the web just for a chatting app is pure insanity sorry. It will be slow to start, slow to react and a battery hog vs a comparable quality QT app. The worse part is usually people use the web stack for desktop app because they don't want to bother giving a good experience to the people on their own native platform.
rubymamis 2 hours ago [-]
True, I wrote my own chat app in Qt, and it is substantially faster and more efficient (lower battery usage) than anything made with web technologies.[1]
I do wish more things were written in Qt, but Qt Widgets is very unpleasant to work in for a modern dynamic app.
QtQuick/QML/JS is very pleasant and I do wish more people would use it, but from I've seen it's 25-50% the resource use of electron, not some multi-order-of-maginute improvement, do I understand why many people still prefer electron for portability in this case.
14 hours ago [-]
flowerbreeze 16 hours ago [-]
When writing code for end-user applications, I think it's mostly true. When it's writing code for a database engine, a game engine, a 3d renderer, or anything else that involves heavy data processing, optimization is the core "thing" often and it might not even be a good enough solution without it. Although, a lot of time even then C++ is good enough even then when picking reasonable data structures to represent the data.
hn_submit 13 hours ago [-]
These are what I like to call "infinity applications" where the need for speed is essentially infinite.
Even if you write them in hand-optimized assembly they would still clamor for more speed.
8n4vidtmkvmk 15 hours ago [-]
Definitely need to optimize a bit for games and huge scale web apps. We've been finding big optimizations in our app recently. App works without them because we can scale horizontally but cutting CPU usage by 30% by eliminating redundant work and reducing copies of big objects? Why wouldn't we want to do that? This isn't even fancy algorithm stuff, mostly just shoddy initial implementations by 100s of eng working on a codebase over 7 years (not even that old). Stuff like that creeps in.
Agentlien 3 hours ago [-]
I work in game development with graphics, porting, and performance. The stuff mentioned in this article is absolutely essential to ship games with sufficient performance
senderista 14 hours ago [-]
Then why are you using C++? Java/C#/Go are already fast enough for general application development. Why would you accept the footguns if not for performance?
bluGill 13 hours ago [-]
In my case, about 5% of our code needs the power of C++. Mixing C++ with any other language is a huge pain. Even if we were using C, mixing C with anything else is a pain, and that's despite being the most supported FFI.
Note that we started our project before Rust was an option. These days I would certainly look at rust to see if that would cover our 5% of the needs but now we have a lot of C++ and mixing rust with C++ is a pain.
palata 12 hours ago [-]
I agree that mixing languages adds complexity. And I say that as someone who routinely does it, because many times it's better to reuse a mature/audited component than rewrite it from scratch.
einpoklum 12 hours ago [-]
> Mixing C++ with any other language is a huge pain.
It is less pain than for most other languages, except for C. The pain is in exposing a C API for your C++ code. Then you build a library and you're set - because basically every language has the ability to call C code. Python, Rust, Java, etc. etc.
palata 12 hours ago [-]
Well once you have a C API, it is relatively easy to call it from any other language.
The painful part is to have to go through a C API (modern languages can express much richer APIs and of course there are different constraints on the different runtimes, e.g. GC).
The annoying part is that each language adds overhead (its runtime). I wouldn't call it painful (I don't have much to do about it), I say "annoying" just because I would rather minimise the amount of code I ship.
palata 13 hours ago [-]
Sometimes it's about the libraries. E.g. writing Computer Vision is nicer in C++ right now (IMHO) because most CV libraries are in C++.
Similarly I like to do video stuff in C just because I call gstreamer/ffmpeg directly in C, rather than having to bridge everything.
glouwbug 16 hours ago [-]
True, but moving from a list of unique polymorphic pointers to a std::variant gains you at least a 2-3x speed up in terms of TLB and cacheline locality. From there, swapping to SOA will net you another 4-8x, so you're looking at nearly 25x improvement by going data first. That may not matter in the unique case of say, games, where rendering a million entities will dwarf the cost of SIMD processing a million entities, but in something like numerical simulations (fluids) or quant it will be warmly welcomed
someonebaggy 2 hours ago [-]
You don't render a million entities, only the ones that are actually on-screen. But when you do render them, you also want to render them in SoA style, culling them by comparing all the bounding boxes against the view frustum, generating a single big list of draw commands and passing it to draw-indirect one time.
cenamus 16 hours ago [-]
Is the improvement from using std:variant vs polymorphism just due to the indirection you save on?
glouwbug 15 hours ago [-]
That, and it frees the compiler from reasoning about virtual inlining, and that the std::variant approach can pack potentially more than one object into a single cacheline. TLBs also work with 4096 byte pages, so 32 polymorphic 128 byte entities may (at the absolute worst case) use 32 distinct pages which requires 32 TLB virtual translations, while the std::variant one uses 1.
The next step of going SOA benefits from all of the above, it just further unlocks you packed quad and oct instructions (AVX256 and 512 depending if you buy AMD or not).
It's odd reading an AI generated article about something that happened in 2022. Anachronistic.
creata 15 hours ago [-]
If you take the linked benchmark and use the latest compiler version, the std::variant version is faster. The annoying thing about std::variant (and with some other features of modern C++) is that it generates a bunch of code that the compiler has to optimize away.
vlovich123 15 hours ago [-]
Somehow and for some reason rust enums don’t have this problem and are far more ergonomic and easier to work with (not to mention compile times are amazing). I don’t know exactly why it’s better to have it as a first class language primitive and why the compiler has such a problem with std::visit, but clearly c++ meta programming slows down things in a super linear way such that the compiler has problems both from code gen and then optimization.
hackrmn 16 hours ago [-]
I started writing a [CPU-only] 3-D rendering library in C++ recently, after having written the equivalent in C as a proof-of-concept and an experiment. The reason I decided to write it in C++ after C, is not only because I wanted to tap into meta-programming which is facilitated much better with C++, or that I wanted niceties like procedure overloading, but because some things with C or C++ aren't automagically optimised -- like if you want to leverage struct-of-array (SoA) memory layouts because it allows fewer SIMD (AVX in my case) instructions in the rendering pipeline. You do _not_ get that "for free" just writing a single procedure in C++, much less with C. Both languages are layout-sensitive, I mean this is in part what gives you the speed -- optimising with memory layout for cache locality etc. But you have to do it yourself. Meaning that if you need array-of-struct (AoS) or in fact don't know which path the CPU would prefer, there's no other way than roll up your sleeves and one way or another implement both.
The kicker is, in my case I chose C++ because templates allow me to reuse most of the code in the rendering pipeline _regardless_ of whether I go for AoS or SoA layout. I leverage operator overloading to do vector by matrix multiplication which is implemented in both variants. I do have to specify the desired variant during building, but I've profiled and for Intel x86 AVX in my case SoA is something like twice as efficient because I process ("shade") 8 vertices with 4-5 instructions instead of 1 vertex at a time (still shaded with vectorisation -- just "rotated", i.e in the pipeline axis and not vertex buffer axis).
TL;DR; C++ gives you plenty fast by default, but it's not always enough. The difference between 5 and 15 frames per second, well, makes all the difference -- our eyes are only fooled once the frames-per-second rate goes sufficiently up, anything below an acceptable threshold and it's completely different experience. You then either sacrifice resolution or level of detail etc, or decide to squeeze more from the language by helping the compiler.
smallstepforman 4 hours ago [-]
To do something like this, you really need to look at both cache locality and CPU core access patterns:
40’000 NPC in game with collision avoidance, steering, on 8yo hardware.
loeg 15 hours ago [-]
This is a sign you might be more productive in a higher-level language.
creata 15 hours ago [-]
And the code might be faster, too, like in that 2005 series of articles by Raymond Chen and Rico Mariani (in which one of them wrote a program in C++ and the other wrote the same program in C#).
In (soft) realtime audio programming, your audio callback might only have a time budget of 1.3 milliseconds. Everytime you exceed that limit, you'll hear a dropout. That's when you'll start to optimize the hell out of your program :)
cjbgkagh 17 hours ago [-]
I rarely use C++ but when I do it is for speed. It’s not uncommon that carefully crafted intrinsics can 10x the straightforward naive implementation.
aldanor 13 hours ago [-]
Depends on your field. When one microsecond is considered "hellishly slow", you might reconsider
12 hours ago [-]
wat10000 14 hours ago [-]
There’s code where there exists a concept of “fast enough,” and code where there is no such thing.
mathisfun123 16 hours ago [-]
Then you don't work on a product that has any scale <shrug>.
11 hours ago [-]
112233 18 hours ago [-]
"This article was originally published in Polish in issue 4/2013" — a lot of excellent advice. Sad to see C++ have moved in last decade in a direction that makes writing efficient, simple low level code harder and harder :(
jandrewrogers 16 hours ago [-]
Writing clear, concise, and efficient code in C++ has never been simpler or easier. The improvements in C++ over the last 15 years have been qualitative.
So many complex, esoteric, and difficult to maintain incantations that used to be required for efficient code generation are no longer necessary.
112233 14 hours ago [-]
How do you process read-only mmaped data in C++, in accordance with the language rules? As an example.
jandrewrogers 8 hours ago [-]
That was addressed circa C++17 IIRC, same with all of the technically UB surrounding DMA memory. You needed non-obvious incantations for some use cases but they were expressible. C++23 eliminated most of those incantations. Before all of this you had to use “blessed” incantations that were technically UB but which compilers needed to allow because the code was expressing a valid use case.
I find it crazy that popular systems languages didn’t have an explicitly valid way to deal with all of these ambiguous ownership and lifetime issues around memory until relatively recently.
Agentlien 4 hours ago [-]
Can you give some specifics? I feel like this is adjacent to my work and I'm not sure what you're referring to which makes me feel I've missed something important.
senderista 13 hours ago [-]
I think there are lots of Unix APIs that are impossible to use without UB, e.g. SCM_RIGHTS (maybe io_uring as well?).
someonebaggy 2 hours ago [-]
Implementations are free to define UB.
mdspan 10 hours ago [-]
That's not something that's unique to C++ though. For example, Rust has the same issue.
jll29 17 hours ago [-]
I think it has become EASIER: for instance, since C++23 Rust-like move semantics can be used, which provides the compiler with extra information that can be leveraged for the generation of better code.
Or take constexpr - it permits to move computations to compile time that are complex and in older versions either had to be done at runtime, or an ugly workaround had to be used (e.g. assigning a mysterious literal pre-computed in another run or by hand).
creata 17 hours ago [-]
> C++23 Rust-like move semantics can be used
What C++23 feature allows that?
aw1621107 15 hours ago [-]
The closest thing I can think of is trivial relocation [0] which matches the bitwise copy + no destructor on moved-from bits of Rust's moves, but that was only added to the draft for C++26 and was removed late in the process anyways [1].
I think that's basically clarifying the conditions under which C++11-style moves can be performed.
fooblaster 18 hours ago [-]
How? you can write exactly the same low level code today.
beached_whale 17 hours ago [-]
My thought too.
There are so many things that are expressible in C++ now that could not be without writing much more code or using per-compilation tools back then. The ability to run code at compile time that is not run at runtime is huge, #embed lets us make other tools output available without linker scripts or compiler specific tools that.
Also, most of the code from the past still works(from 10 years ago definitely works)
AlotOfReading 17 hours ago [-]
Shot in the dark, but maybe the OP is referring to the fact that these code conventions are explicitly discouraged by the C++ core guidelines. The SoA example falls afoul of the rule requiring T* to be used only for singular object pointers, for example.
someonebaggy 2 hours ago [-]
Core guidelines, and any other pattern document, should be downstream of working code, not upstream.
AlotOfReading 55 minutes ago [-]
Okay? The point of the guidelines is to document principles the committee's thinks "good" C++ should follow, within the much larger universe of possible C++ code.
someonebaggy 52 minutes ago [-]
So if you find good code that works a different way, do you update your beliefs on what good code is, or on whether that code is good?
cjbgkagh 17 hours ago [-]
Not a regular C++ programmer but wouldn’t you use std::span here instead? Sure it’ll carry a few redundant lengths but it makes using functions that take spans easier. When I do write C++ it’s usually for speed so I’m often working at the intrinsics level, though AI has gotten good enough at it that I now generally delegate this work to an agent.
jandrewrogers 16 hours ago [-]
Depending on the specific code, the compiler may even eliminate the redundant lengths.
112233 14 hours ago [-]
eh, yes-ish, but only thanks to compiler writers. at one point committee went all-out enforcing their lifetime model, making bit_cast not an option. If your code accesses same data using different types, you are spelunking ruins with snake pits and lava. more and more stuff needs magic support code in std::, making no-lib code less and less possible (it used to be that with no-rtti and no-exceptions, you could use all c++ features and only needed cxa_at_exit, operator delete, and few other little things. NOT ANY MORE).
On one hand, you have consteval and stuff, letting you FINALLY initialize data at compile time (hey, 20 years late but still!)
on other hand, it is done in most non-debuggable way possible. try setting breakpoint or adding print to constexpr function that causes your requires clause to fail...
so no, newer C++ the language is not possible to use for low level work. The dialects that compiler makers support are. We will see for how long
add2 10 hours ago [-]
In most cases, the performance improvement was negligible, and the code merely became more complex.
"Premature optimization is the root of all evil"
FpUser 16 hours ago [-]
My latest C++ project is assessment engine covering various actuarial type things like calculates risk for insurance etc. Typical performance for bulk calculation reaches millions to 10s of millions assessments per second on 16 core server. Well there is a trick there that inside it JIT compiles rules from a DSL to an executable code. interpreter mode (used mainly for audit mode) is about 3-5 times slower which is still insanely fast
einpoklum 11 hours ago [-]
> We know that the C++ keyword volatile is used to prevent a value from being optimized in a way that would allow the compiler to keep it in a processor register rather than fetching it from its original memory location each time.
Actually, we don't know that. The meaning of volatile is rather subtle
someonebaggy 2 hours ago [-]
It does mean the compiler needs to treat loading and storing the value as a side effect - actually write a load instruction every time it's loaded in the source code. It's as if *X was a call to get_X() the compiler can't optimize around. However the next sentence which says you can use it for thread synchronization is wrong.
jocelyner 4 hours ago [-]
[dead]
tug2024 17 hours ago [-]
[dead]
uwagar 15 hours ago [-]
i heard a lot of AI and LLM is in python?
pjmlp 2 hours ago [-]
As glue language for the actual work implemented in a mix of C, C++, Fortran and more recently Rust.
Just that Python culture has a strange way to call bindings to native code, "Python libraries".
bee_rider 14 hours ago [-]
Python is often used as (very useful!) glue for calling CUDA, C, C++, Fortran, etc… codes.
So, if you are thinking about the sort of “business logic” that’s often Python, but the performance comes from the parts that are usually not.
uwagar 4 hours ago [-]
well you'd be surprised how much of that glue becomes logic and eats up the ground water.
cosmicradiance 4 hours ago [-]
Bumblebees shouldn’t be able to fly; they do anyways.
hydrocephalitic 15 hours ago [-]
Yes, but the numerical backend of the python libraries are written in lower-level languages. For example, pytorch uses a C++ backend.
tom_ 15 hours ago [-]
You heard correctly. This discussion is about C++ though.
The growable array type std::vector<T> is least impacted by these archaic choices out of the tools in the box you're likely to reach for. So it will make sense very often to choose this type first.
The C++ 23 containers are: array, vector, deque, forward_list, list, set, map, multiset, multimap, unordered_set, unordered_map, unordered_multiset, unordered_multimap
Rust's collections are: BTreeMap, BTreeSet, BinaryHeap, HashMap, HashSet, Vec, VecDeque
Firstly, Rust doesn't consider "array" a library type here, in C++ the language has built-in arrays but they're very poor because they are the C arrays - so you use the library feature to get good arrays. In Rust they... just fixed the language, because duh.
Next thing you'll notice is that C++ has lots more of these types, about twice as many. I stopped at C++ 23 because in C++ 26 they added even more. These are a significant maintenance burden and of course having more means in practice maintenance gets worse. But this could be good if these types were all high quality and kept that way.
All of the C++ unordered containers are the same crap hash table design but with slightly different parameters. The Rust HashMap and HashSet are Swiss Tables though they do not promise that and if a better design comes along they will probably switch. C++ can't change the design because the API welds them to a very specific shape for this data structure, a shape which delivers bad performance on any vaguely modern hardware.
std::deque is the most horrible surprise. A modern programmer who has thought about it at all is expecting a type like Rust's VecDeque. Generalise the amortized growable array from the language to use it as a ring buffer. Cheap push & pop at both ends, canonically use it as a FIFO but also practical in lots of other situations. But that's not what std::deque is at all, instead inside it's an array of links to small arrays. On MSVC it's effectively a linked list again because those inner arrays contain only one item due to ABI considerations.
std::set and std::map are very principled red-black trees. I say principled because in practice this is too expensive on modern hardware because (say it with me) it spends too long chasing pointers up and down your tree. Rust's choice here in BTreeMap and BTreeSet packs more data in each "node" on the tree, which makes the big-O worse but the practical performance better. Figuring out how to best do this for the general case is an active area of research but "I bet a pure red-black tree will be fast" is not a good guess for the past several decades.
Finally std::forward_list and std::list are the singly and doubly extrusive linked list types. The thing you most likely have seen in some high performance software is an intrusive linked list, and C++ doesn't provide those. In an intrusive linked list each item in the list itself links to where the next (and for simple double links also the previous) item is, so the item needs to know it's in a list [in some systems more than one list, thus more than one set of links]. C++ provides extrusive linked lists where those links live in a separate object and so the items in the list don't know about this at all. Rust provides only a doubly-linked extrusive list exactly like C++ std::list, but again, this almost certainly isn't what you wanted, you most likely do not need a linked list and if you do have a good reason for a linked list you probably want an intrusive linked list.
For example, since I allowed for objects to be shared between threads, I decided to use struct of arrays so the reference count, metadata, and value would be stored in separate cache lines. This ended up hurting me because object initialization touched three separate cache lines (obvious in hindsight, but the advice of using SoA failed me here). I also heard that you want to pack your values as tight as possible, so I used a packed string index, but then I ended up with integer division to unpack the string (also a mistake, but again the advice failed me). I used a custom allocator to avoid indirection with lists (list items were allocated directly after the list head), but then I had heap fragmentation and the implementation complexity exploded.
Anyways, I am now happily using two to three levels of indirection in my data structures, large structs, and malloc for individual objects, and it's still been faster in my end to end testing. So maybe this is unique to interpreters, and maybe I could have done it better, but the suggestions don't automatically apply in my experience.
I'd say the rule was followed in this case - the rationale of SoA is to reduce cache misses when iterating all objects and only using some of the attributes, which is something games do all the time, but it's bad if you are always accessing one object at a time. Maybe an array language interpreter would have luck with SoA.
I mean this is kind of what happens with any advice that has nuance to it, that's not carried with the advice.
E.g. if you have a point in 3D space with x, y, z coordinates. Array points as SoA of individual dimensions makes sense only if you do a lot of averaging and such on the individual dimensions.
If you mostly use the 3 coordinates together, SoA will have bad caching behavior.
So the better advice would be to try to keep things that are used together in the same cache line, whether it's on dimension or all 3. Usage makes the difference.
Beyond the low-hanging fruit like ensuring you aren't creating O(n^2) complexity by accident, I think C++ is fast enough/has mature-enough compilers that by the time you're worrying about cache hits materially affecting performance, you're probably also sufficiently staffed and capitalized to pay people to A/B test that performance.
2. Virtuals are, with the exception of PGO, mostly a black box i.e. you get a hard optimisation boundary, no inlining at all.
3. The C++ standard library is usually comically slow (yes, even compared to Java/C#/the likes) so if your project uses std::vector and the such instead of specialised libraries, you've already lost at the beginning.
4. If you don't pay attention to performance from the get-go, the approximate amount of autovectorisation you'll get is close to zero. Some compilers are better than others (Clang>MSVC for example) but I've seen codebases with 8 figures of LoC where the number of vectorised divides/multiplys was like less than ten when you dumped the object listing. In the whole program.
5. Since aliasing and other optimisation barriers (you didn't use restrict or manually hoist, did ya?), it's not uncommon for large C++ programs to spend a third of their runtime doing atomic increments because shared_ptr is supposedly cheap and who cares about lifetimes anyway.
6. If you're targeting Windows, the default new operator / malloc is also comically slow. Luckily that one is fairly easy to fix with installing mimalloc and deploying the hijack dll, but the negative effects on cache by the fragmented allocations is also significant.
std::vector doesn't have trivial relocation so any type with a destructor ends up doing elementwise destruct+construct instead of a memcpy.
std::map and std::list are memes and if you use them you're giving your CPU the 1995 treatment with all that pointer chasing.
You thought std::unordered_map is better? Well, actually not because node stability, so it's still chained-bucket, you almost always want to use a flat map like boost::unordered_flat_map or the abseil/eastl version.
<random> is hard-to-use and isn't very performant, std::regex is "you might as well write it in Python and it'd be faster", <iostreams> is virtual calls galore, both the formatting and the stdio functionality are slow.
The conveniently-named std::function is a very general device resulting in a heap allocation and usually a virtual call, there's specific optimisations but don't rely on it.
The STL string manipulation functions are also usually slow, they check the locale for string manipulation rules.
The floating-point functions set errno preventing vectorisation and emitting branches in your straight-line float code unless you use fastmath (the thing people tell you never to do) or one of the more fine-grained compiler-specific switches to turn it off.
std::shared_ptr is Arc<T>, not Rc<T> and eating the cost of atomics can add up in many situations especially with all the other memory traffic going on.
std::variant and std::visit are also not very fast either.
std::filesystem as a whole also has several pain points like iteration which is like a magnitude slower than the native APIs, std::chrono isn't much better either
std::error_code sounds like a simple integer or even a struct.... lol no guess what, more virtual calls
Surely nobody outside microsoft is doing serious work targeting Windows any more are they? Isn't that a dead platform? I read somewhere a while back they're now below 60% market share.
Yeah, it's really not.
There are multiple areas of work, where C++ can be considered a glue language. The high-performance work is then done in explicit SIMD (intrinsics, ISPC, etc.) and/or GPU-targeting languages such as CUDA or Vulkan.
In these areas of work, high performance is part of the design and not something that can be easily added as after-thought.
Also, relying on optimization features such as compiler auto-vectorization is way too finicky - your hot-loop performance may completely break without anyone noticing by someone changing a trivial-looking part of a loop.
I work in game development and for the last six years I've spent most of my time specifically on optimization. A lot of that effort has been focused on cache behaviors. Not because it's fun, but because it's often the difference between being able to ship the game on weaker hardware (e.g. Nintendo Switch) or not.
As a contrived example: there are specific cases when a particular non-quicksort algorithm is optimal. In almost all real world scenarios, though, you're just going to say fuck it and use quicksort until profiling determines that the sort is the bottleneck.
Unless you already have specific knowledge that your data comes in a particular shape, defaulting to quicksort is good design (IMHO). Worrying about pathological sorting before you've seen benchmarks is premature optimization.
I think it's more like: prioritize cache locality over big O compexity.
Run your own benchmarks on your own data of course. Also map is not considered the best key value store.
This likely won't be true in a real application with a non-trivial allocation pattern.
Still a custom map that allocated a bunch of nodes would be a useful optimization.
And then, on the other hand - I really doubt GP's map beats a vector, with all of those pointers bins and stuff, in a non-contrived benchmark with 10 elements.
Finally - it's not either-or: There are better hash maps whose memory is sequentially allocated and/or are otherwise cache-aware. And there are data structures geared towards parallel execution on multiple threads; and towards SIMD; etc. etc.
If you’re down to that sort of decision-making, you have to measure.
I feel like the DoD movement is a slow-moving, but big, change through how systems programming is done, but that there's still insufficient material for how to do this in different scenarios. I would really like to apply this more to my areas of work, which are also in C++, but there seems to be a gap between what they're presenting and how it can be applied.
More specifically, I'm using C++ to build a dynamic programming language runtime for a Clojure dialect. That runtime is required to be garbage collected, type-erased, and highly polymorphic. So I surely can't just SoA or AoS everything. Yes, I can pack my data, and I can avoid the GC whenever possible, both in compiler/runtime code and in generated code via escape analysis. But what about everything else, which is the 80% or more of the system? It could be that this runtime is too far at odds with DoD, but I generally see things as a gradient rather than black and white.
Or were you referring more to all the intermediate allocations that aren't the object heap? V8's zones are interesting in this area, because they're like an arena, except that they're only partially reset when a zone ends, so zones can nest inside each other.
At the time I thought this would naturally fit in a data-oriented design/ECS system to run complex queries. I wonder whether anyone has tried this before and whether this actually works in practice?
[1] https://prela-lang.org/tutorial/
High-performance programming is a big topic. The scope is far too broad for a single blog post, which naturally gives only cursory discussion of C++ and computer architecture. The article isn't bad considering, but I do think it's the wrong format. A blog series, or even a book, would be more fitting.
https://www.agner.org/optimize/
https://www.agner.org/optimize/optimizing_cpp.pdf
What you've written mostly makes sense to someone who already has a solid understanding of SIMD and of C++ (although I can't say I follow all of it), but the target audience is people who don't. For them, each point needs a much lengthier explanation.
A quick restrict example:
copy1 is 52 lines, copy2 is 28 lines, copy3 is 2 lines (just a call to memcpy).This is a good starting point for self teaching. The impact of your TLB, L1, and overall instruction count (with IPC) can further be measured with `./perf stat -d -d -d ./a.out`. If you want a quick rule of thumb, no instructions are fast instructions.
creata's comment [0] mentions the works of Agner Fog, which seem very good, and are freely available.
I haven't read C++ High Performance [1] but it looks like it covers the sorts of topics you'd expect, although it looks like it doesn't cover computer architecture in detail e.g. branch prediction. There are books on that too, of course.
[0] https://news.ycombinator.com/item?id=49868657
[1] https://www.packtpub.com/en-us/product/c-high-performance-97...
If your device has enough resources to power V8, modern GUIs are certainly very pleasant and snappier than a more minimal GUI like HN. Otherwise they are horrendous and very laggy.
[1] https://www.get-vox.com/
QtQuick/QML/JS is very pleasant and I do wish more people would use it, but from I've seen it's 25-50% the resource use of electron, not some multi-order-of-maginute improvement, do I understand why many people still prefer electron for portability in this case.
Even if you write them in hand-optimized assembly they would still clamor for more speed.
Note that we started our project before Rust was an option. These days I would certainly look at rust to see if that would cover our 5% of the needs but now we have a lot of C++ and mixing rust with C++ is a pain.
It is less pain than for most other languages, except for C. The pain is in exposing a C API for your C++ code. Then you build a library and you're set - because basically every language has the ability to call C code. Python, Rust, Java, etc. etc.
The painful part is to have to go through a C API (modern languages can express much richer APIs and of course there are different constraints on the different runtimes, e.g. GC).
The annoying part is that each language adds overhead (its runtime). I wouldn't call it painful (I don't have much to do about it), I say "annoying" just because I would rather minimise the amount of code I ship.
Similarly I like to do video stuff in C just because I call gstreamer/ffmpeg directly in C, rather than having to bridge everything.
The next step of going SOA benefits from all of the above, it just further unlocks you packed quad and oct instructions (AVX256 and 512 depending if you buy AMD or not).
https://stackoverflow.com/questions/69444641/c17-stdvariant-...
GCC 11 (2021) std::visit was slower than virtual dispatch.
GCC 12 (2022) optimized std::visit so it can be faster than virtual dispatch.
https://shubhankar-gambhir.github.io/posts/your-stdlib-imple...
The kicker is, in my case I chose C++ because templates allow me to reuse most of the code in the rendering pipeline _regardless_ of whether I go for AoS or SoA layout. I leverage operator overloading to do vector by matrix multiplication which is implemented in both variants. I do have to specify the desired variant during building, but I've profiled and for Intel x86 AVX in my case SoA is something like twice as efficient because I process ("shade") 8 vertices with 4-5 instructions instead of 1 vertex at a time (still shaded with vectorisation -- just "rotated", i.e in the pipeline axis and not vertex buffer axis).
TL;DR; C++ gives you plenty fast by default, but it's not always enough. The difference between 5 and 15 frames per second, well, makes all the difference -- our eyes are only fooled once the frames-per-second rate goes sufficiently up, anything below an acceptable threshold and it's completely different experience. You then either sacrifice resolution or level of detail etc, or decide to squeeze more from the language by helping the compiler.
https://youtu.be/jsdwRf3JvZM?si=0tysoCsvaWl0R_PZ
40’000 NPC in game with collision avoidance, steering, on 8yo hardware.
https://devblogs.microsoft.com/oldnewthing/20060731-15/?p=30...
https://learn.microsoft.com/en-us/archive/blogs/ricom/perfor...
"Just because I don't write about .NET doesn't mean that I don't like it"
"Performance Quiz #6 -- Chinese/English Dictionary reader"
How 100Gbit NICs could your filter through your stateful firewall at line speed? And with 64 byte packets?
https://web.archive.org/web/20250201145327/https://users.ece...
In (soft) realtime audio programming, your audio callback might only have a time budget of 1.3 milliseconds. Everytime you exceed that limit, you'll hear a dropout. That's when you'll start to optimize the hell out of your program :)
So many complex, esoteric, and difficult to maintain incantations that used to be required for efficient code generation are no longer necessary.
I find it crazy that popular systems languages didn’t have an explicitly valid way to deal with all of these ambiguous ownership and lifetime issues around memory until relatively recently.
Or take constexpr - it permits to move computations to compile time that are complex and in older versions either had to be done at runtime, or an ugly workaround had to be used (e.g. assigning a mysterious literal pre-computed in another run or by hand).
What C++23 feature allows that?
[0]: https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2025/p27...
[1]: https://herbsutter.com/2025/11/10/trip-report-november-2025-...
https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2022/p22...
There are so many things that are expressible in C++ now that could not be without writing much more code or using per-compilation tools back then. The ability to run code at compile time that is not run at runtime is huge, #embed lets us make other tools output available without linker scripts or compiler specific tools that.
Also, most of the code from the past still works(from 10 years ago definitely works)
On one hand, you have consteval and stuff, letting you FINALLY initialize data at compile time (hey, 20 years late but still!)
on other hand, it is done in most non-debuggable way possible. try setting breakpoint or adding print to constexpr function that causes your requires clause to fail...
so no, newer C++ the language is not possible to use for low level work. The dialects that compiler makers support are. We will see for how long
"Premature optimization is the root of all evil"
Actually, we don't know that. The meaning of volatile is rather subtle
Just that Python culture has a strange way to call bindings to native code, "Python libraries".
So, if you are thinking about the sort of “business logic” that’s often Python, but the performance comes from the parts that are usually not.