• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Intel Nova Lake increasing core counts

Lakados

[H]F Junkie
2FA
Joined
Feb 3, 2014
Messages
13,912
https://www.techpowerup.com/331392/intel-nova-lake-test-cpu-appears-targeting-2026-launch

https://overclock3d.net/news/cpu_mainboard/intel-nova-lake-leaks-hint-a-huge-core-count-increase/

Leaks show a dual chiplet variant where each chiplet is an 8+16 configuration, so 16 performance cores and 32 efficient cores

Supposedly Intel has already begun shipping variants of the CPUs to OEMs and AIBs for validation testing, rumors say the chips are primarily on the Intel 18A node.
https://www.tomshardware.com/pc-com...-samples-intended-for-validation-and-research
https://www.msn.com/en-us/lifestyle...ut-things-are-more-complicated-on-the-desktop

It's worth noting that the Intel 18A node seems to have better power delivery than TSMC's current 2N process, but TSMC is doing better with SRAM cell size.
TSMC uses a power workaround they named "NanoFlex", but backside power delivery isn't coming back until the TSMC A16 node.
https://www.tomshardware.com/tech-i...et-nanoflex-n2p-loses-backside-power-delivery

Here's a little info on Backside Power Delivery and why it is crucial for GAA designs, as Intel learned early on and TSMC is learning now.
https://semianalysis.com/2024/10/01/clash-of-the-foundries/

Without backside power delivery, the smaller the chip, the better, so I suspect that Intel continues to use TSMC for its GPU functionality with the CPU being on the Intel 18A node... that is speculation, but given that moving a GPU design down the TSMC stack is probably easier than moving it to a completely different tech it seems reasonable to me that Intel leaves the AI and GPU stuff at TSMC until they do their next major redesign.
 
Not that core count was a big issue for Intel (against the competition of apple-amd at least), so by itself maybe an average news (what percentage of users will use more than the 20-24 cores of the previous gen).

But what it must mean in power consumption for a p-core if they can fit 16 of them on a single cpu (that has 32 e-core on it), that must be a really good sign for that part. And the fact all of that would fit, must be good sign for 18A node density/power efficacy.
 
Its interesting how CPU companies who are behind on per thread performance tend to increase core counts where it doesn't matter to try to make up for it.

It was silly when AMD did it and now it is silly when Intel is doing it.
 
I'll keep an open mind. I think next upgrade for me is going to be 2026 so will see what Zen 6 and Nova Lake look like and go from there.
 
Its interesting how CPU companies who are behind on per thread performance tend to increase core counts where it doesn't matter to try to make up for it.

It was silly when AMD did it and now it is silly when Intel is doing it.
It just goes to show that when someone can't get the single-core crown, they compensate with more cores.

AMD with Zen 1 and 2, as soon as they hit IPC and ST Supremacy with Zen 3, they stopped increasing cores. Now it's Intel's turn.
 
I've got some 144 core Sierra forest single socket servers, they tend to work pretty well for telecom/MSO workloads, so I don't necessarily mind increasing core counts on consumer products.
 
It just goes to show that when someone can't get the single-core crown, they compensate with more cores.

AMD with Zen 1 and 2, as soon as they hit IPC and ST Supremacy with Zen 3, they stopped increasing cores. Now it's Intel's turn.

Just imagine the uselfulness of all those cores for your average PC user using nothing but Office, Outlook, a web browser and occasionally watching a youtube video :p
 
Apple does not use core count much as a selling point and imo, they seem to optimise it "perfectly" for the average most common personnal computer user.

Base M4, 8-10 core, 4 p-core, 6 e-core, 8-10 core that perfectly fine for most people.
M4 pro: 12-14 core, perfectly fine for pro users
Max: 14-16, make sense to use almost all your extra silicon for GPU instead of CPU core past the 14 core mark...

But the phrasing of it:
  • NVL-S: 2x(8+16)
  • NVL-HX: 1x(8+16)
  • NVL-S/NVL-H: 4+8
  • NVL-U: 4+0
Maybe they just gained the ability to have 2 compute tile at the same time, like AMD chiplet, in that case, no foul, no harm, it will be just for the subset of people that need lot of cores with nothing loss for everyone else maybe. Could mean very little about intel18, power reduction of p-core and so on if that the case. The tile would still go up 8+16 exactly like before....
 
I still don't understand why they don't just add 3D V-Cache when they obviously can. Instead, they just keep upping the amount of NPU's. I get it, the "consumer market" isn't a big enough market, but if it isn't a big enough market, why not just dump consumer-grade CPU's and GPU's altogether and just focus on the more "lucrative" server market?

I'm just surprised Intel or AMD hasn't yet implemented 3D V-Cache in any of their mobile CPU's.

https://www.techpowerup.com/328853/...3d-v-cache-tech-in-2025-just-not-for-desktops
"... the company is working on augmenting its processors with large shared L3 caches, however, it will begin doing so only with its server processors."
 
Good luck to Intel. Really hope these don't suck. If they are smart they learn lessons from Ultra... and improve. I fear they have more then one design team and its possible any lessons to learn will be ignored, making 2 crap generations in a row.

I am sure the top end skus will do well, the catch will how they cut and position all the stuff down from there. I think AMD has proven their are sweet spots to be found down from the big full 2 dies of cores. Maybe Intel can find their own sweet spot gaming chip with more cache per core if they plan their design right.
 
making 2 crap generations
Would be more the third one, meteor lake was the first generation of that model I feel like, it just did not work as a desktop CPU, 2xx ultra was the second generation that learned from meteor lake major failure and seem to have made an quite good laptop version at least (they would have learned more in that segment having had an actual field release).
 
Last edited:
  • Like
Reactions: ChadD
like this
I still don't understand why they don't just add 3D V-Cache when they obviously can. Instead, they just keep upping the amount of NPU's. I get it, the "consumer market" isn't a big enough market, but if it isn't a big enough market, why not just dump consumer-grade CPU's and GPU's altogether and just focus on the more "lucrative" server market?

I'm just surprised Intel or AMD hasn't yet implemented 3D V-Cache in any of their mobile CPU's.

https://www.techpowerup.com/328853/...3d-v-cache-tech-in-2025-just-not-for-desktops
It all comes down to price and manufacturing capacity.
TSMC's CoWoS packaging facilities are at capacity, TSMC has made it clear they are not expanding on it, and they are instead focusing on newer packaging technologies which are also at capacity but focus on the AI and data center markets.
To make the 3D chips, they have to make the first chip, then make the cache chip, and finally, they line them up perfectly and glue them together, it is not a fast process and adds significant cost. The only reason AMD can currently justify it is because, in the grand scheme of things, they sell relatively few of them, and the few they are selling is maxing out TSMC's capacity to assemble them.
Offering them across the board isn't something that TSMC could facilitate, and in the mobile segment, it isn't needed as much due to the shared memory configuration for the CPU and GPU that most devices use and because of the nature of LPDDR which is being used more frequently in laptops and other mobile devices.
The extra cache works as a buffer, reducing the number of times the CPU needs to utilize the memory channels, which leaves them more available for the GPU which deals with large data transfers. In a shared memory environment, there aren't any large data transfers from SRAM to VRAM so both are then using small short calls, and the need for the extra cache is reduced. Most mobile devices are now also using LPDDR modules which are faster with lower latency and are closer to the CPU and GPU which further speed things up which further reduces the need for the stacked cache.
In a modern laptop, especially the ones that AMD primarily sells, the extra cache would not likely bring anything noticeable to the table except a larger price tag.

Intel does something different with its stacked cache; they don't attach it to the chip for it to function as a memory buffer, they build it into the packaging process to work as a shared data layer between the CPU and other on-chip accelerators, AI, GPU, etc...
That way, if the CPU and those devices need to share data directly, it is not sending them down to system RAM, and back they can share a cache pool directly on the packaging fabric itself. What they are demonstrating is the evolution of the Intel Adamantine Cache which is technically an L4 cache, not L3 like what AMD does, it transforms the interposer from a passive chip that only exists to facilitate data and power to an active processor that works to accelerate the chips connected to it.
Active Interposers are expensive, really expensive, and it would not be an exaggeration to say that the interposer can cost as much or more than the chips being connected to it, which is why it is currently only found on Datacenter components.
 
Good luck to Intel. Really hope these don't suck. If they are smart they learn lessons from Ultra... and improve. I fear they have more then one design team and its possible any lessons to learn will be ignored, making 2 crap generations in a row.

I am sure the top end skus will do well, the catch will how they cut and position all the stuff down from there. I think AMD has proven their are sweet spots to be found down from the big full 2 dies of cores. Maybe Intel can find their own sweet spot gaming chip with more cache per core if they plan their design right.
It's a full new P and E core design, and they scrapped one that was supposed to come between them, which was for Intel 2A, so I am somewhat hopeful for this because it would be externally the 3'rd generation, but the 4th internally as they decided that one of them wasn't even worth releasing.
These should be based on the cores used in the current Xeon lineup, which hold up well, I would say they are the most competitive Xeons in more than a few generations.
 
Its interesting how CPU companies who are behind on per thread performance tend to increase core counts where it doesn't matter to try to make up for it.

It was silly when AMD did it and now it is silly when Intel is doing it.

Easy way to compensate. However parallelism is increasing across the board so it's not all for naught in terms of performance we can use.

I still don't understand why they don't just add 3D V-Cache when they obviously can. Instead, they just keep upping the amount of NPU's. I get it, the "consumer market" isn't a big enough market, but if it isn't a big enough market, why not just dump consumer-grade CPU's and GPU's altogether and just focus on the more "lucrative" server market?

I'm just surprised Intel or AMD hasn't yet implemented 3D V-Cache in any of their mobile CPU's.

https://www.techpowerup.com/328853/...3d-v-cache-tech-in-2025-just-not-for-desktops

3D vcache doesn't work well with all workloads, it also adds cost and complexity as well. Especially on very small form factor laptop chips which is the bulk of the consumer space.
 
I can't wait to be disappointed by this processor when Intel finally releases it.
 
I had no idea what Intel's next Desktop chip was called but 2026 is a long way away. Hopefully I can get a GPU by then.
 
IMO, the big news here is 18A may actually work and reach consumer hands. Honestly, pretty surprised.

Its interesting how CPU companies who are behind on per thread performance tend to increase core counts where it doesn't matter to try to make up for it.

It was silly when AMD did it and now it is silly when Intel is doing it.
It got Intel to shrink its' roadmaps and release consumer 6 core and 8 core, a lot sooner than originally planned.

I still don't understand why they don't just add 3D V-Cache when they obviously can. Instead, they just keep upping the amount of NPU's. I get it, the "consumer market" isn't a big enough market, but if it isn't a big enough market, why not just dump consumer-grade CPU's and GPU's altogether and just focus on the more "lucrative" server market?

I'm just surprised Intel or AMD hasn't yet implemented 3D V-Cache in any of their mobile CPU's.

https://www.techpowerup.com/328853/...3d-v-cache-tech-in-2025-just-not-for-desktops

3D vcache doesn't work well with all workloads, it also adds cost and complexity as well. Especially on very small form factor laptop chips which is the bulk of the consumer space.

AMD is on its 2nd Vcache laptop chip.
 
AMD is on its 2nd Vcache laptop chip.
I know. I made a comment about it, but it got deleted and I got a warning for going off topic, lol. The whole conversation is about things Intel are capable of doing, but don't. See, I mentioned Intel, it's not off topic!
 
Not that core count was a big issue for Intel (against the competition of apple-amd at least), so by itself maybe an average news (what percentage of users will use more than the 20-24 cores of the previous gen).

But what it must mean in power consumption for a p-core if they can fit 16 of them on a single cpu (that has 32 e-core on it), that must be a really good sign for that part. And the fact all of that would fit, must be good sign for 18A node density/power efficacy.

Well if it doesn't mean Intel has solved it's power efficiency problem, we'll finally be getting the hotter than the core of a nuclear reactor to hotter than the surface of the sun chips that people were predicting ~25 years ago from just continuing to scale clock speeds and power levels from one generation of pentium to the next.
 
Good luck to Intel. Really hope these don't suck. If they are smart they learn lessons from Ultra... and improve. I fear they have more then one design team and its possible any lessons to learn will be ignored, making 2 crap generations in a row.

I am sure the top end skus will do well, the catch will how they cut and position all the stuff down from there. I think AMD has proven their are sweet spots to be found down from the big full 2 dies of cores. Maybe Intel can find their own sweet spot gaming chip with more cache per core if they plan their design right.

I don't know if they still do; but before the 10nm manufacturing debacle turned all their long term plans upside down Intel did have 2 independent CPU design teams working on offset 4 year cycles. That occasionally resulted in features moving onto - off of - and then back onto the CPU in successive generations because one team either didn't think it was worth it, or it was too late in their design cycle to add it once it was shown to be good.
 
  • Like
Reactions: ChadD
like this
3D vcache doesn't work well with all workloads, it also adds cost and complexity as well. Especially on very small form factor laptop chips which is the bulk of the consumer space.

Not really true, the Vcache works just fine with all workloads: it either greatly boosts it, boosts it a little, or does nothing.

The issue was the heavy reduction in clock speeds associated with the first two generations of Vcache chips, which the most recent gen almost entirely eliminates.
 
I don't know if they still do; but before the 10nm manufacturing debacle turned all their long term plans upside down Intel did have 2 independent CPU design teams working on offset 4 year cycles. That occasionally resulted in features moving onto - off of - and then back onto the CPU in successive generations because one team either didn't think it was worth it, or it was too late in their design cycle to add it once it was shown to be good.
Pretty sure they now have a team that is working on Intel fab versions of their tech, and the one designing to TSMC. Probably hasn't ever really changed.
 
Its interesting how CPU companies who are behind on per thread performance tend to increase core counts where it doesn't matter to try to make up for it.

It was silly when AMD did it and now it is silly when Intel is doing it.
Era dependant. AMD did it after a lifetime of Intel quad-core bullshit. Worked out amazing for them so not sure why you think it was silly?
Care to clarify the silliness?
 
Its interesting how CPU companies who are behind on per thread performance tend to increase core counts where it doesn't matter to try to make up for it.

It was silly when AMD did it and now it is silly when Intel is doing it.
I don't think it was silly when AMD was doing it. In fact, even with the per thread performance increases they're able to achieve now AMD is still actively increasing cores. We all know now Intel had the performance crown because of all the shortcuts they were taking. Once Intel could no longer rely on those shortcuts, they ran into the same issue that AMD did, that increasing performance came in smaller increments and the way to compensate is to increase cores.
 
AMD is still actively increasing cores.

In the market tier that matter, regular consumer market their offering have been 4/6 to 16 since the 3000 series launched in 2019, 5+ years ago. (if I have not missed a release or maybe something in the Laptop space ? did they started to push high count Zen-C cores there?)

, that increasing performance came in smaller increments and the way to compensate is to increase cores.

That work for certain workload, but past 5-6 cores it can get quite hard for most of them

apple-m-series-1.jpg


AMD 2600x->5600x went from a ST passmark of 2,382 to 3,362 and to a 9600x up to 4,580... that nearly doubling from a 2600x (2018) to a 9600x (2024)

Really not that bad, Foundry Node stagnation was happening at Intel in a worse way that the physic laws made it possible or impossible, I feel like.
 
Last edited:
I don't think it was silly when AMD was doing it. In fact, even with the per thread performance increases they're able to achieve now AMD is still actively increasing cores. We all know now Intel had the performance crown because of all the shortcuts they were taking. Once Intel could no longer rely on those shortcuts, they ran into the same issue that AMD did, that increasing performance came in smaller increments and the way to compensate is to increase cores.

It is still true that adding cores improves performance in a minority of workloads while increasing per core performance improves performance in all workloads.

AMD benefitted little to nothing from being the first to 8-corez with bulldozer, due to the per thread performance being utter trash.

There were some workloads they were quite nice in though. I - for one - used one (first an 8120, and later upgraded to a 8350 after they were dirt cheap) as one of my first consumer hardware based VM servers. In that application it was held back by only being able to support 32GB of RAM though.

For most consumer workloads including games they were terrible.

Games eventually caught up to where they can benefit from 6-8 cores but it took a long time, and that is likely the ceiling of how threaded most games will be. (There may be some limited exceptions, like turn based strategy titles with AI computations between turns, but most times won't be able to support much more parallelism. There is simply a lot of code that cannot under any circumstances be multi threaded without serious side effects like thread locking and performance hits due to threads waiting for input threads to complete.

A lot of the multi threading in games we are today isn't even true multi threading where the total workload is chopped into small bits, and evenly distributed between the available threads. . It is "fake" multi threading by splitting out things to individual threads. Main game logic gets one thread, audio gets another thread, physics gets yet another thread, etc. and then they sync with each other. . This is not multi threading, but rather multiple instances of single threaded code running at the same time.

Even as it stands, most of the benefits beyond ~6C/12T come solely from the fact that people are running more shit in the background than we used to (Silly screen recording/streaming, RGB LED software, management software for mice and every other damn accessory or cooler, constant Windows background tasks, etc. etc.) not from that the workload itself actually benefits much of at all from multithreading beyond 6C/12T.

I have been using my 24C/48T Threadripper 3960x since 2019, and in real use, the times where I benefit from more than 6 of those cores are pretty damn rare. I bought it primarily for the 64 PCIe lanes, not for the cores.

Either way, the beauty of HEDT was that it did everything well. Top per core performance in it's generation that supported high end consumer workloads and games, while at the same time having many workstation-like features for other things. It allowed us to build one badass system to do it all.

Unfortunately HEDT is now dead, what with all of the new HEDT-like platforms being more true Workstation products that perform poorly on consumer workloads.

Because of this I have begrudgingly made the decision to switch to two systems. One workstation, and one game machine.

And guess what? The game machine will only be getting an 8 core CPU, because I don't stream and never will, and I run a streamlined system with as little as possible running in the background. Absolutely nothing it will ever do will benefit but having more cores than that.

And if that ever changes, well, it will probably be time for the biannual upgrade anyway.
 
Well if it doesn't mean Intel has solved it's power efficiency problem, we'll finally be getting the hotter than the core of a nuclear reactor to hotter than the surface of the sun chips that people were predicting ~25 years ago from just continuing to scale clock speeds and power levels from one generation of pentium to the next.
The A18 process compares favourably to the TSMC 3N process and approaches some of the things TSMC 2N is supposed to do.

I am optimistic for the process, especially if they can get it operating with very high yields.

The last report I read on the topic was favourable, small chips and tiles were going to have few issues, larger monolithic designs… well fortunately most have moved away from those.
Either way this chip will be the poster child for A18 and if it’s good Intel can hopefully scrounge up some customers.
 
Even as it stands, most of the benefits beyond ~6C/12T come solely from the fact that people are running more shit
A lot of the time is also how they structure their SKUs line, Intel for example you get more cache and boost clock up the line has well has more core.

And a lot of people mistook the game performance gain to be caused by the presence of those extra core and not the bigger enabled L3 cache and boost clock, for a lot of game the 14900k>14600k performance difference stay even if you disable cores, sometime even get better.

Almost all benchmark we see are exceptionally clean machine, fresh install, only installed drivers-steam-the game and 1 tools, doing exceptionally little when the game is running, putting a lot of effort to remove all the noise that did not let higher core cpu shine in those regards.
 
It is still true that adding cores improves performance in a minority of workloads while increasing per core performance improves performance in all workloads.

AMD benefitted little to nothing from being the first to 8-corez with bulldozer, due to the per thread performance being utter trash.

There were some workloads they were quite nice in though. I - for one - used one (first an 8120, and later upgraded to a 8350 after they were dirt cheap) as one of my first consumer hardware based VM servers. In that application it was held back by only being able to support 32GB of RAM though.

For most consumer workloads including games they were terrible.

Games eventually caught up to where they can benefit from 6-8 cores but it took a long time, and that is likely the ceiling of how threaded most games will be. (There may be some limited exceptions, like turn based strategy titles with AI computations between turns, but most times won't be able to support much more parallelism. There is simply a lot of code that cannot under any circumstances be multi threaded without serious side effects like thread locking and performance hits due to threads waiting for input threads to complete.

A lot of the multi threading in games we are today isn't even true multi threading where the total workload is chopped into small bits, and evenly distributed between the available threads. . It is "fake" multi threading by splitting out things to individual threads. Main game logic gets one thread, audio gets another thread, physics gets yet another thread, etc. and then they sync with each other. . This is not multi threading, but rather multiple instances of single threaded code running at the same time.

Even as it stands, most of the benefits beyond ~6C/12T come solely from the fact that people are running more shit in the background than we used to (Silly screen recording/streaming, RGB LED software, management software for mice and every other damn accessory or cooler, constant Windows background tasks, etc. etc.) not from that the workload itself actually benefits much of at all from multithreading beyond 6C/12T.

I have been using my 24C/48T Threadripper 3960x since 2019, and in real use, the times where I benefit from more than 6 of those cores are pretty damn rare. I bought it primarily for the 64 PCIe lanes, not for the cores.

Either way, the beauty of HEDT was that it did everything well. Top per core performance in it's generation that supported high end consumer workloads and games, while at the same time having many workstation-like features for other things. It allowed us to build one badass system to do it all.

Unfortunately HEDT is now dead, what with all of the new HEDT-like platforms being more true Workstation products that perform poorly on consumer workloads.

Because of this I have begrudgingly made the decision to switch to two systems. One workstation, and one game machine.

And guess what? The game machine will only be getting an 8 core CPU, because I don't stream and never will, and I run a streamlined system with as little as possible running in the background. Absolutely nothing it will ever do will benefit but having more cores than that.

And if that ever changes, well, it will probably be time for the biannual upgrade anyway.
And let’s not forget the performance increases gained by increased core count are incredibly dependent on the system scheduler, of which Microsoft is …. Not great at tweaking or optimizing in a timely fashion.

I would really love to see some CPU benchmarks some time comparing server to desktop. Especially in high core count systems.
 
A lot of the time is also how they structure their SKUs line, Intel for example you get more cache and boost clock up the line has well has more core.

And a lot of people mistook the game performance gain to be caused by the presence of those extra core and not the bigger enabled L3 cache and boost clock, for a lot of game the 14900k>14600k performance difference stay even if you disable cores, sometime even get better.

Almost all benchmark we see are exceptionally clean machine, fresh install, only installed drivers-steam-the game and 1 tools, doing exceptionally little when the game is running, putting a lot of effort to remove all the noise that did not let higher core cpu shine in those regards.

Yep. That has been frustrating me for some time. It used to always be tradeoff. If you got more cores, they were often clocked lower (like in server parts) to reduce heat and power use. Because of this fewer core parts used to be clocked higher, as they could clock higher without overheating.

These days as dies have kept shrinking, the CPU makers seem to keep their best CPU bins only for the highest core count parts which is really annoying for those of us who have no need for many core parts, and just want the best possible clocks in fewer cores.
 
I just hope it's good. We had a great run with AMD and Intel being so close, we don't need to be stuck on whatever AMD decides we deserve for 10 years.
 
And let’s not forget the performance increases gained by increased core count are incredibly dependent on the system scheduler, of which Microsoft is …. Not great at tweaking or optimizing in a timely fashion.

I would really love to see some CPU benchmarks some time comparing server to desktop. Especially in high core count systems.

That is true. Microsoft does a terrible job with their scheduler, but that impacts low threaded and multithreaded code alike.

It's actually funny. I have seen some tests where CPU performance is actually improved when running windows inside a VM on a Linux-based KVM system, because the Linux kernel is so much better at picking which thread to run a particular virtual CPU on than the native Microsoft scheduler.

Like - for instance - if you have a traditional 8-core hyperthreaded CPU, and you are only loading up 4 threads, the Linux kernel is so much better at selecting the threads that need the performance the most, and spreading them out, one thread per physical core, to maximize both turbo clocks, and avoid interruptions from the second logical core.

I suspected this was the case, but the results actually floored me when I saw them, as I was expecting the Linux kernel to do a better job here, but it to still be slower than bare metal due to the inefficiencies involved in virtual CPU's. It was truly surprising to see the VMs perform better than bare metal in these cases.

There are similar benefits with modern big/little P/E core optimizations in the linux kernel that the microsoft scheduler fails at.

The microsoft scheduler really sucks at this. Not to mention the whole unnecessary AMD cache flushing thing. Microsoft should be ashamed of themselves.
 
That is true. Microsoft does a terrible job with their scheduler, but that impacts low threaded and multithreaded code alike.

It's actually funny. I have seen some tests where CPU performance is actually improved when running windows inside a VM on a Linux-based KVM system, because the Linux kernel is so much better at picking which thread to run a particular virtual CPU on than the native Microsoft scheduler.

Like - for instance - if you have a traditional 8-core hyperthreaded CPU, and you are only loading up 4 threads, the Linux kernel is so much better at selecting the threads that need the performance the most, and spreading them out, one thread per physical core, to maximize both turbo clocks, and avoid interruptions from the second logical core.

I suspected this was the case, but the results actually floored me when I saw them, as I was expecting the Linux kernel to do a better job here, but it to still be slower than bare metal due to the inefficiencies involved in virtual CPU's. It was truly surprising to see the VMs perform better than bare metal in these cases.

There are similar benefits with modern big/little P/E core optimizations in the linux kernel that the microsoft scheduler fails at.

The microsoft scheduler really sucks at this. Not to mention the whole unnecessary AMD cache flushing thing. Microsoft should be ashamed of themselves.
I have just finished flashing a USB key with a Server 2025 Standard image to install on my old gaming rig to try out something similar.
Windows server does not use the same CPU scheduler as Windows Desktop and I am curious to know if it works better.
It's not too hard to remove the few server monitoring tasks in Windows Server that make it as barebones as it gets, the Image I just installed on the VM for my new print server is using a whopping 1.3GB in RAM at idle, and the stripped-down Win 11 image on that same host I use for testing deployments idles at twice that.
Win Server also doesn't seem to have any issues with high core counts or multiple CPUs like the desktop variants, so it will mostly come down to how well Nvidia's drivers and the various anti-cheats from the game launchers hate Windows Server.
But yeah unless something pulls me away from the house that's my weekend... Civ7 gets to be the first test game on the box.
 
It is still true that adding cores improves performance in a minority of workloads while increasing per core performance improves performance in all workloads.
It is not 2010, and increasing cores improves performance in nearly all workloads at this point.

AMD benefitted little to nothing from being the first to 8-corez with bulldozer, due to the per thread performance being utter trash.
AMD was about 5 years too early back in 2011 with the debut of Bulldozer and their CMT architecture.
At the time the integer performance was about 55% clock-for-clock what Sandy Bridge was, and the FX 8-core CPUs were the only ones with 4 FPUs.

The threading would become more important once the PS4 and XBone were both released by 2013 with 8 cores/threads, and games and software started becoming more optimized with more than 4 cores/threads.
AMD was too early with CMT's larger thread counts for software to be optimized for it at the time, and Meltdown/Spectre/Foreshadow/etc./etc./etc. hadn't become public knowledge yet giving Intel an artificial boost with their now-known garbage CPU security implementations.

Had nearly all software been optimized for 6-8 threads (like it has been for the last decade) and had Intel actually implemented the proper security designs in their CPU architectures, the performance would have been much closer.
I'm not going to white knight for AMD, though at the same time lets not rewrite history.
 
Last edited:
Back
Top