• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Intel Meteor Lake Technical Deep Dive

erek

8=D
2FA
Joined
Dec 19, 2005
Messages
18,292
Nice

"Intel provided us with a close look at its upcoming "Meteor Lake" microarchitecture powering the next generation of Intel Core processors for PCs. "Meteor Lake" is Intel's very first processor to fully realize Intel's IDM 2.0 Strategy of redesigning its future processors such that it can maximize the utilization of its latest in-house semiconductor foundry capacity for the specific components on the processor that need it the most; and then carefully disaggregating the various components onto slightly older foundry nodes—some even from external foundries—and building an integrated device that's greater than the sum of its parts.

Intel attaches great importance to the success of "Meteor Lake," not just because it is Intel's first disaggregated chiplet-based processor, but also because it powers Intel's biggest push for consumer AI hardware acceleration, with Intel AI Boost. With this, Intel hopes to mainstream AI in the client space, and get the PC ecosystem to embrace on-device accelerated AI for consumer software applications spanning from everything between limited function apps to complex productivity suites. "Meteor Lake" is hardly the first device to accelerate AI for the client—GPUs with AI acceleration hardware have been around for close to 5 years now—but Intel still holds the reins to the PC ecosystem with an over 80% market-share in PC processors, and the fact that a majority of PCs don't use discrete GPUs. With Intel taking the lead in consumer AI, software vendors would finally have the confidence to invest big in new AI-accelerated features since they know there's a sizable userbase.

title.jpg



2023 is a very different point in time from 2008. Back then, Intel had just perfected its 45 nm HKMG foundry node, and had a credible path toward 14 nm, and could execute the famous "Tick-Tock" product development strategy, where the company would introduce a new foundry node every two years, and a new microarchitecture every intervening two years. Since the company had industrial leadership over foundry nodes, it could build processors on monolithic dies. The 1st Gen Core "Nehalem" was pivotal, in that it saw the first aggregation effort by Intel in a decade, with the PC's northbridge being integrated with the processor package. Over subsequent generations, the aggregation trend would continue, driven by clear performance incentives—first the memory controller, followed by the integrated graphics, and then platform I/O. The company's last monolithic processor die, "Raptor Lake," integrates nearly every device of the modern PC onto a single die built on the Intel 7 process, except the platform I/O.

Fast forward to 2023, and Intel has just emerged from a very slow jump from 14 nm to Intel 7 (10 nm Enhanced SuperFin), the prices of wafers on the latest EUV-based nodes are significantly higher than what they used to be, and so Intel is incentivized to disaggregate—to identify the specific components on the processor that don't benefit as much from being built on the latest foundry node, and spin them off into chiplets, or tiles as Intel likes to call them, on older foundry nodes. Alongside "Meteor Lake," Intel is introducing the new Intel 4 node, the company's first to utilize EUV lithography, and offer both transistor densities and energy efficiency rivaling 4 nm-class nodes by TSMC. Besides Intel 4, the company has its mature Intel 7 node, and even third party nodes for specific tiles.

The secret sauce here is the Foveros chip packaging technology, a combination of high-density inter-die and via-substrate electronic connections between the various tiles, which make them work as if they were a single monolithic die. The effort here is to provide both bandwidth and latencies close to those of on-die connections, or else the processor is essentially an MCM and disaggregated as things were before Nehalem.

In this article, we dive deep into the fascinating world of "Meteor Lake" and how Intel plans to make this chiplet-based processor greater than the sum of its parts and realize CEO Pat Gelsinger's vision of IDM 2.0 manufacturing, which will redefine the way chips will be built for the foreseeable future. This is purely an architecture-focused article, Intel has not detailed any specific processor models based on "Meteor Lake," nor what its next-gen Core lineup will look like; and so obviously we have no performance numbers in this review besides anything Intel would have claimed in its presentations to us."

Source: https://www.techpowerup.com/review/intel-meteor-lake-technical-deep-dive/
 
Nice

"Intel provided us with a close look at its upcoming "Meteor Lake" microarchitecture powering the next generation of Intel Core processors for PCs. "Meteor Lake" is Intel's very first processor to fully realize Intel's IDM 2.0 Strategy of redesigning its future processors such that it can maximize the utilization of its latest in-house semiconductor foundry capacity for the specific components on the processor that need it the most; and then carefully disaggregating the various components onto slightly older foundry nodes—some even from external foundries—and building an integrated device that's greater than the sum of its parts.

Intel attaches great importance to the success of "Meteor Lake," not just because it is Intel's first disaggregated chiplet-based processor, but also because it powers Intel's biggest push for consumer AI hardware acceleration, with Intel AI Boost. With this, Intel hopes to mainstream AI in the client space, and get the PC ecosystem to embrace on-device accelerated AI for consumer software applications spanning from everything between limited function apps to complex productivity suites. "Meteor Lake" is hardly the first device to accelerate AI for the client—GPUs with AI acceleration hardware have been around for close to 5 years now—but Intel still holds the reins to the PC ecosystem with an over 80% market-share in PC processors, and the fact that a majority of PCs don't use discrete GPUs. With Intel taking the lead in consumer AI, software vendors would finally have the confidence to invest big in new AI-accelerated features since they know there's a sizable userbase.

View attachment 599924


2023 is a very different point in time from 2008. Back then, Intel had just perfected its 45 nm HKMG foundry node, and had a credible path toward 14 nm, and could execute the famous "Tick-Tock" product development strategy, where the company would introduce a new foundry node every two years, and a new microarchitecture every intervening two years. Since the company had industrial leadership over foundry nodes, it could build processors on monolithic dies. The 1st Gen Core "Nehalem" was pivotal, in that it saw the first aggregation effort by Intel in a decade, with the PC's northbridge being integrated with the processor package. Over subsequent generations, the aggregation trend would continue, driven by clear performance incentives—first the memory controller, followed by the integrated graphics, and then platform I/O. The company's last monolithic processor die, "Raptor Lake," integrates nearly every device of the modern PC onto a single die built on the Intel 7 process, except the platform I/O.

Fast forward to 2023, and Intel has just emerged from a very slow jump from 14 nm to Intel 7 (10 nm Enhanced SuperFin), the prices of wafers on the latest EUV-based nodes are significantly higher than what they used to be, and so Intel is incentivized to disaggregate—to identify the specific components on the processor that don't benefit as much from being built on the latest foundry node, and spin them off into chiplets, or tiles as Intel likes to call them, on older foundry nodes. Alongside "Meteor Lake," Intel is introducing the new Intel 4 node, the company's first to utilize EUV lithography, and offer both transistor densities and energy efficiency rivaling 4 nm-class nodes by TSMC. Besides Intel 4, the company has its mature Intel 7 node, and even third party nodes for specific tiles.

The secret sauce here is the Foveros chip packaging technology, a combination of high-density inter-die and via-substrate electronic connections between the various tiles, which make them work as if they were a single monolithic die. The effort here is to provide both bandwidth and latencies close to those of on-die connections, or else the processor is essentially an MCM and disaggregated as things were before Nehalem.

In this article, we dive deep into the fascinating world of "Meteor Lake" and how Intel plans to make this chiplet-based processor greater than the sum of its parts and realize CEO Pat Gelsinger's vision of IDM 2.0 manufacturing, which will redefine the way chips will be built for the foreseeable future. This is purely an architecture-focused article, Intel has not detailed any specific processor models based on "Meteor Lake," nor what its next-gen Core lineup will look like; and so obviously we have no performance numbers in this review besides anything Intel would have claimed in its presentations to us."

Source: https://www.techpowerup.com/review/intel-meteor-lake-technical-deep-dive/
"Intel CEO Pat Gelsinger, in the Q&A session of InnovatiON 2023 Day 1, confirmed that the company is developing 3D-stacked cache technology for its processors. The technology involves expanding the on-die last-level cache (L3 cache) of a processor with an additional SRAM die physically stacked on top, and bonded with the cache's high-bandwidth data fabric. The stacked cache operates at the same speed as the on-die cache, and so the combined cache size is visible to software as a single contiguous addressable block of cache memory.

AMD has used 3D-stacked cache to good effect on its processors. On client processors such as the Ryzen X3D series, the cache provides significant gaming performance uplifts as the larger L3 cache makes more of the game's rendering data immediately accessible to the CPU cores; while on server processors such as EPYC "Milan-X" and "Genoa-X," the added cache provides significant uplifts to memory intensive compute workloads. Intel's approach to 3D-stacked cache will be different at the hardware level compared to AMD's, Gelsinger stated in his response. AMD's tech has been collaboratively developed with TSMC, and hinges on a TSMC-made SoIC packaging tech that facilitates high-density die-to-die wiring between the CCD and cache chiplet. Intel uses its own fabs for processor dies, and will have to use its own IP.
gJlvss8G5ONmmtt1_thm.jpg
"When you reference V-Cache, you're talking about a very specific technology that TSMC does with some of its customers as well. Obviously, we're doing that differently in our composition, right? And that particular type of technology isn't something that's part of Meteor Lake, but in our roadmap, you're seeing the idea of 3D silicon where we'll have cache on one die, and we'll have CPU compute on the stacked die on top of it, and obviously using EMIB that Foveros we'll be able to compose different capabilities," Gelsinger said.

"We feel very good that we have advanced capabilities for next-generation memory architectures, advantages for 3D stacking, for both little die, as well as for very big packages for AI and high-performance servers as well. So we have a full breadth of those technologies. We'll be using those for our products, as well as presenting it to the Foundry (IFS) customers as well," he added.

Intel recently provided an architecture deep-dive into its upcoming "Meteor Lake" client processor, in which its Foveros packaging tech and tile-to-tile interconnects allow the various tiles (chiplets) to work like a cohesive silicon. In particular, Intel appears to have solved the latency issues of having a the iGPU, CPU cores, and memory controllers on separate tiles."

Source: https://www.techpowerup.com/313818/pat-gelsinger-says-3d-stacked-cache-tech-coming-to-intel
 
"When you reference V-Cache, you're talking about a very specific technology that TSMC does with some of its customers as well. Obviously, we're doing that differently in our composition, right? And that particular type of technology isn't something that's part of Meteor Lake, but in our roadmap, you're seeing the idea of 3D silicon where we'll have cache on one die, and we'll have CPU compute on the stacked die on top of it, and obviously using EMIB that Foveros we'll be able to compose different capabilities," Gelsinger said.
So, stacked cache won't be in meteor lake? Am I reading that correctly?
 
So, stacked cache won't be in meteor lake? Am I reading that correctly?
I think he's saying "we don't call it that because that's what AMD calls it, but we'll still have something like it" probably because "V-Cache" is a trademark. He says they'll stack a CPU tile on top of a cache tile (instead of the other way 'round).
 
So, stacked cache won't be in meteor lake? Am I reading that correctly?
Not stacked just more of it, they will have the 128MB cache for the interposer (L4), but L2 cache for the P cores is going up to 2MB from 1.25, and each e core cluster (grouping of 4) is going up to 4MB from 2, and L3 cache is going up from 30MB to 36, so there is just a lot more onboard to play with.

But Intel has been working with Samsung on this issue and Samsung calls it their CacheDRAM, kac77 posted an article about it not too long ago.
https://hardforum.com/threads/intel...50-faster-60-more-efficient-than-hbm.2030321/

Intel doesn't expect to have the process commercially viable until 2025, but they have shown demos of it to both Nvidia and AMD, and Jensen is on record stating that he hopes to integrate it into their future generations of data center hardware.

It should be noted though that it would allow upwards of 64GB to be attached directly to the Processor as a stacked cache when launched, and upwards of 1TB on a single SoC.
 
Last edited:
That thing looks pretty massive, unless that's a really blown up picture. Wonder how the heat will be and what sort of heatsink shape you'll need to cool it with. Also a little confused as to why Intel feels the need to go hard into NPUs on the desktop market. I guess that's mostly going to be a thing on laptops, because most desktop users with any interest in AI are going to know you kind of need a dedicated card... probably from Nvidia. I guess assuming there are going to be more and more programs that just casually leverage AI, it's definitely going to be good for that.

I'm kind of curious if any games could take advantage of it. Maybe somehow accelerating physics workloads by leveraging the NPU to give a different way of calculating it...? But unless AMD supports it, I don't know how likely that would fly.
 
That thing looks pretty massive, unless that's a really blown up picture. Wonder how the heat will be and what sort of heatsink shape you'll need to cool it with. Also a little confused as to why Intel feels the need to go hard into NPUs on the desktop market. I guess that's mostly going to be a thing on laptops, because most desktop users with any interest in AI are going to know you kind of need a dedicated card... probably from Nvidia. I guess assuming there are going to be more and more programs that just casually leverage AI, it's definitely going to be good for that.

I'm kind of curious if any games could take advantage of it. Maybe somehow accelerating physics workloads by leveraging the NPU to give a different way of calculating it...? But unless AMD supports it, I don't know how likely that would fly.
It's a really blown-up picture, the chip is designed for ultrabooks and general business laptops, it's a 15-65w part.
 
It's a really blown-up picture, the chip is designed for ultrabooks and general business laptops, it's a 15-65w part.

Oh both of those were laptop parts, including the last topic. I guess I misread what rick said in the last topic and thought this was a desktop variant.

Well, I guess it's nice if you want to do some (very basic) stable diffusion on the go... Otherwise I'm not sure why they're so hyperfocused on including AI accelerators with the CPU. It kind of has a place in mobile phones because of how much postprocessing they need to do on images, or the fact that they need to recognize your face to unlock (which I think sucks and is a privacy invasion and no one should use but that's an aside). Or that they want to scan your images for things various entities might find disagreeable (we hope that's not happening outside of cloud).

But for AI work, a lot of common stuff already revolves around Nvidia and its ecosystem. For Stable Diffusion, xformers by itself is basically a game changer. I guess this is just everyone doing the gold rush now.
 
Oh both of those were laptop parts, including the last topic. I guess I misread what rick said in the last topic and thought this was a desktop variant.

Well, I guess it's nice if you want to do some (very basic) stable diffusion on the go... Otherwise I'm not sure why they're so hyperfocused on including AI accelerators with the CPU. It kind of has a place in mobile phones because of how much postprocessing they need to do on images, or the fact that they need to recognize your face to unlock (which I think sucks and is a privacy invasion and no one should use but that's an aside). Or that they want to scan your images for things various entities might find disagreeable (we hope that's not happening outside of cloud).

But for AI work, a lot of common stuff already revolves around Nvidia and its ecosystem. For Stable Diffusion, xformers by itself is basically a game changer. I guess this is just everyone doing the gold rush now.
It's more for the other AI capable things that people take for granted that are present in Android phones and basically any Mac product.

AI noise filters for picking out the speaker in a noisy environment, image enhancement for not-so-good webcams, sound enhancements for virtual meetings, text prediction, speech-to-text, text-to-speech, network packet optimization, real-time language translation (both spoken or written), numerous photo editors, etc ...

Lots of day to day things use AI, or fancy Machine Learning, stuff that Intel and AMD have traditionally just done with x86, but Apple is kinda of whooping their buts in the mobile space, especially in regards to efficiency, and that is mostly because of ML and AI acceleration, which these deliver.
Apple Neural Engine, Ryzen AI, Intel AI Boost, and ARM just calls them their NPUs, All the same thing, x86 mobile is just getting caught up to the ARM mobile space for features they have had for a good while now.

Intel is adding their AI accelerator to the whole stack, AMD has chosen to just include it at the top and it is supposed to debut in their 7040u series.
 
AI noise filters for picking out the speaker in a noisy environment, image enhancement for not-so-good webcams, sound enhancements for virtual meetings, text prediction, speech-to-text, text-to-speech, network packet optimization, real-time language translation (both spoken or written), numerous photo editors, etc ...

Those are good examples. I didn't think of that, likely because I don't really use those AI features very much. Well, and I don't have a Mac product (although I have used an iPhone for quite some time in the past). Although I don't know about real time language translation. I think that's mostly still done by online services? Or has Apple actually decided to implement a fully local solution? I wouldn't be surprised either way. They definitely have some bright minds over there. Although some annoying minds as well, behind their more "courageous" moves. AI packet optimization is not something I have heard about at all before.
 
Those are good examples. I didn't think of that, likely because I don't really use those AI features very much. Well, and I don't have a Mac product (although I have used an iPhone for quite some time in the past). Although I don't know about real time language translation. I think that's mostly still done by online services? Or has Apple actually decided to implement a fully local solution? I wouldn't be surprised either way. They definitely have some bright minds over there. Although some annoying minds as well, behind their more "courageous" moves. AI packet optimization is not something I have heard about at all before.
Lots of those online services do the crunching and send it back. Grammerly is the big stand out there, you send them everything you type so they can process it and train the algorithms and such.
But what if you work for a company or an environment where that sort of privacy impact is unacceptable. Would having all the capabilities of Grammerly built into Word running 100% local not be a pretty big game changer? Same with virtual assistants like Alexa, Seri, etc, can’t use them offline really but what if you could?

Apple is moving more locally and that’s where their Apple silicon Neural Engine machine learning stuff comes in. They distinctively call it machine learning and not AI more stuff sent is weaker privacy and higher data usage and Apple is pushing hard to keep that stuff internal. So they are working to take all the big services that the majority of their users are using and building an internal and offline equivalent.
Siri is currently too large to take offline but they are hopeful they can at least keep the simple things local soon.

But the accelerators are mostly about taking jobs normally dispatched to a GPU or inefficient on a CPU and building a simple bit of silicon to deal with. It’s mostly about improved feel to a system and improved battery.
 
Lots of those online services do the crunching and send it back. Grammerly is the big stand out there, you send them everything you type so they can process it and train the algorithms and such.
But what if you work for a company or an environment where that sort of privacy impact is unacceptable. Would having all the capabilities of Grammerly built into Word running 100% local not be a pretty big game changer? Same with virtual assistants like Alexa, Seri, etc, can’t use them offline really but what if you could?

Apple is moving more locally and that’s where their Apple silicon Neural Engine machine learning stuff comes in. They distinctively call it machine learning and not AI more stuff sent is weaker privacy and higher data usage and Apple is pushing hard to keep that stuff internal. So they are working to take all the big services that the majority of their users are using and building an internal and offline equivalent.
Siri is currently too large to take offline but they are hopeful they can at least keep the simple things local soon.

But the accelerators are mostly about taking jobs normally dispatched to a GPU or inefficient on a CPU and building a simple bit of silicon to deal with. It’s mostly about improved feel to a system and improved battery.
New stuff


View: https://youtu.be/_CMGdhtDlqI?si=53sP3geQqiikKxRL
 

I like what I hear about their new testing and binning process that’s cool. The idea of the NPU’s being available as an M.2 is also cool.

Overall dope and the LPe cores are fun sounding, combined with the changes to the thread director that’s just cool.

And if anybody wants to know why 16 frames for the decode buffer that’s because 30 fps streamed content from Netflix, Disney, etc, uses some alternate frame rendering magic and runs far less than 30 ((last-first)+key frames)/2 so ((30-1)+2)/2=15.5 fps.
 
Back
Top