• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Intel thoughs

Duke3d87

Limp Gawd
Joined
Jul 6, 2004
Messages
430
I posted this on XFN also. I'm not claiming to be an expert or anything, just a nerd. :D

So, lately there has been a lot of news about Intel canceling Whitefield and basically shooting themselves in the foot. From an outsiders standpoint, it sounds like a tragic loss and it is. Whitefield was supposed to be magic quad core with low power and an integrated memory controller. When I read google news about Intel, it's like "grr...i want to strangle Paul Ontellini", but from a closeup viewpoint I personally don't think it's going to be that bad.

1. The CEO is a business man and certainly knows how to manage a company - i think he knows what he's doing. He was most likely the man that canceled the 4 GHz processor and the Tejas core.

2. Intel claims that the Tigerton will be a better chip. That has to count for something. They are also replacing the FSB with a point to point connection between the external chipset and the processor. While it has more latency then the AMD series, it's a step in the right direction. What i am wondering is how Intel plans on getting enough bandwith to feed 8 bandwith hungry cores.

3. Development has been sent to the Israeli team. When i read that, i was like "All is good!" Why? Becuase they basically specialize in low power chips. They were the people behind the Banias, Dothan and Yonah cores. They actually made a processor with an integrated memory and video controller before Banias, but Intel killed that processor before it was launched. So, to give the team something to do, Intel funded what we call the Pentium-M. (I don't remember where i read that) So, i don't think it will be that bad. They know what they are doing.

4. Also, Intel is said to have the upper hand when it comes to the 45nm process. If they have problems competing with AMD, they can just add more cache or cores if needed.
 
The answer to #2 is FB-DIMM. It seems FB-DIMM will basically allow for a quad channel memory controller due to the lowered stress on the memory controller.

Makes sense that they killed the IMC chip. Wouldn't sell as many chipsets with the CPU if the MC on was the CPU.
 
robberbaron said:
The answer to #2 is FB-DIMM. It seems FB-DIMM will basically allow for a quad channel memory controller due to the lowered stress on the memory controller.

Makes sense that they killed the IMC chip. Wouldn't sell as many chipsets with the CPU if the MC on was the CPU.
but the problem is with FBDIMM, you would still have a bottleneck

if you have 533 MHz FBDIMM

that's 533*8*4 = 17.056 GB/sec. That's for 8 processors.

That's 2.132 GB/sec. That's not that much. If they can scale FBDIMM to DDR2 speeds, then that might be a bit better...

however, AMD has plans of using DDR4 and maybe Intel will use that too? 2.8 GHz RAM sounds really good.
 
For some reason I'm not convinced that DDR2 FB-DIMMs are confined to speeds of 533MHz.

As for DDR4, that's years away. DDR3 is barely standardized, DDR4 isn't relevant enough to speculate on right now.
 
1. He had no choice but to cancel the future processors, Intel could not continue the megahertz race. They were producing to much heat and had to much leakage.

2. Of course Intel claims there next processor is going to be a better Chip.

3. Yep there searching high and low throughout there company for a good solid processor design that will allow them to compete with Amd.


4 They don't have a upper hand. Reguardless of the manufacturing process they use it still costs die space to add cache and it's expensive. In order for this strategy to be cost effective they would have to be 2-3 processes ahead of Amd. They will also have to have an effective and efficient core to make use of that cache. As you can see even though intel is boosting its cache per core by upwards of 2 meg it isint making a dent. However it does make a dent in there profit per chip.

Also keep in mind that Amd can also boost cache on there chips to up there performance. And if they did they would still consume less power then Intels chips.

5. Intels only chance is to get a better processor then Amd has. They cant win in a price war because to many people are starting to want amd's product.

Amd is gaining to much market share in the desktop and the server markets. They are now targetting the mobile sector aggressively as well.

Current estimates put Intel up to 2007 before they have processors that can compete with Amds offering. And this estimate is based upon the fact that no one has a clue about how good amds k10 chips will be.
 
eblislyge said:
4 They don't have a upper hand. Reguardless of the manufacturing process they use it still costs die space to add cache and it's expensive. In order for this strategy to be cost effective they would have to be 2-3 processes ahead of Amd. They will also have to have an effective and efficient core to make use of that cache. As you can see even though intel is boosting its cache per core by upwards of 2 meg it isint making a dent. However it does make a dent in there profit per chip.
Ok, we're talking XeonMP here. The processor that tops at around $4,000. So, i doubt that adding a few megs of cache will really make a big difference. I don't remember which processor it is, but i think it was supposed to have like 16 MB of L2 cache for a Xeon which is pretty good. And, it costs like $40 for Intel to make a Pentium 4. Since the XeonMP is very similiar to the Pentium 4, it won't increase manufacturing costs that much. And since the quad cores released in 2007 are expected to be made off of the 45nm process, manufacturing won't be that big of an issue. And BTW, there have been times when Intel almost randomly added cache to boost performance.

eblislyge said:
Also keep in mind that Amd can also boost cache on there chips to up there performance. And if they did they would still consume less power then Intels chips.
That's not necessarly true either. Intel has tweaked the 45nm process and is said to have killed all leakage (sounds too good to be true by the way). Once they switch to the Conroe type architecture, power consumption will plummet. A Conroe dual core processor with 4 MB of shared L2 will only consume around 65 watts. the Server end around 80ish.

robberbaron said:
For some reason I'm not convinced that DDR2 FB-DIMMs are confined to speeds of 533MHz.

As for DDR4, that's years away. DDR3 is barely standardized, DDR4 isn't relevant enough to speculate on right now.
well i got that from the fact that on AMD's roadmap, they have plans for DDR4. You might be right about FB-DIMM. It gets really hot, but i think it is still in its infancy
 
Josh_B said:
Dear tool,

AMD engineers hail from the Alpha engineering teams formerly at DEC. They already had a direct connect architecture with both Alpha and the K7. This direct connect architecture was the EV6 bus. (S2K) At that time, Intel continued to use the outdated FSB, while AMD went down the direct connect path. You tell me who's brighter.

Furthermore, Intel does not own productivity, and you completely neglect that the Pentium M does quite well for "gamer fags". Obviously this goes to show the quality of your commentary.

I will gladly buy an ATI card as soon as their Linux drivers even install on any Linux machine without three patches to the source code. Oh yes, and the performance sucks my bag.
They had an integrated controller type thing a long time ago, but randomly killed it. And, while the AthlonMPs did use a point to point bus, it still used the FSB. The reason why Intel's FSB was so bad was becuase it was a shared FSB. There was one connection to the MCH, and it was the clock speed of the slowest FSB in the system. Dual independant FSB should help alleviate some of that problem.
 
Duke3d87 said:
They had an integrated controller type thing a long time ago, but randomly killed it. And, while the AthlonMPs did use a point to point bus, it still used the FSB. The reason why Intel's FSB was so bad was becuase it was a shared FSB. There was one connection to the MCH, and it was the clock speed of the slowest FSB in the system. Dual independant FSB should help alleviate some of that problem.

Sorry, hadn't had my coffee yet this morning. You are quite right. As you say, I was mainly pointing out that there was a dedicated circuit for each CPU, even with the Athlon MP. Shared FSBs are ghey.
 
Josh_B said:
Sorry, hadn't had my coffee yet this morning. You are quite right. As you say, I was mainly pointing out that there was a dedicated circuit for each CPU, even with the Athlon MP. Shared FSBs are ghey.
which is why they are getting rid of it. One good thing came out of a lack of bandwith and that's L3 cache. While i'm not sure if it was invented for that purpose, but it helps performance of the bandwith starved XeonMP...the earlier ones were only getting (800 MB/sec)/processor and that was maximum (i.e if it was 100% efficient). Now, the situation is a bit better, but not ideal. Intel could no longer use the shared bus as once you get to quad qual core Xeons or quad quad core Xeons, you lose a lot of performance due to bottlenecks.
 
Duke3d87 said:
but the problem is with FBDIMM, you would still have a bottleneck

if you have 533 MHz FBDIMM

that's 533*8*4 = 17.056 GB/sec. That's for 8 processors.

That's 2.132 GB/sec. That's not that much. If they can scale FBDIMM to DDR2 speeds, then that might be a bit better...

however, AMD has plans of using DDR4 and maybe Intel will use that too? 2.8 GHz RAM sounds really good.
FB-DIMMs are meant to enable upto a 4X increase in bandwidth and a 24X increase in capacity while using fewer pins versus a standard DDR2 implementation.

The current Xeon MP chipsets, Truland and IBM's Hurricane already use 1 FSB per 2 sockets and have 21.3 GB/s of total memory bandwidth, though only half of it to the CPUs. The Hurricane also multiple high-speed ports to connect to other 4 socket nodes to scale up to 32 sockets efficiently.

http://realworldtech.com/page.cfm?ArticleID=RWT042405213553
 
I've deleted some posts and re-opened the thread. Please keep it on topic...... :cool:

Thanks - B.B.S.
 
SLee said:
FB-DIMMs are meant to enable upto a 4X increase in bandwidth and a 24X increase in capacity while using fewer pins versus a standard DDR2 implementation.

The current Xeon MP chipsets, Truland and IBM's Hurricane already use 1 FSB per 2 sockets and have 21.3 GB/s of total memory bandwidth, though only half of it to the CPUs. The Hurricane also multiple high-speed ports to connect to other 4 socket nodes to scale up to 32 sockets efficiently.

http://realworldtech.com/page.cfm?ArticleID=RWT042405213553
I read that and printed that. The only thing that i don't understand is where they get 21.3 GB/sec

They are using DDR2 400 RAM, and four memory controllers

so that's

400*8 = 3.2 GB/sec

If it's in single channel mode: 3.2*4 = 12.8 GB/sec

if it's in dual channel mode: 3.2*8 = 25.6 GB/sec

Maybe this is going to be the kind of design that Intel will use with the Tigerton type setup? It sounds like the most logical way of solving the problem although it would just be using a point to point connection between the CPU and MCH. I think it would be a lot better then having a high speed FSB b/c if you increase the memory bandwith, you don't really get much benefit for the CPU b/c it is fixed by a fixed speed.

And another thing is they never specified how many memory controllers that you can add. I think it would be interesting if IBM modified it for FB-DIMM and had in 6 channel mode (The technology supports up to six channels of eight double-sided FB-DIMM modules each, or a maximum 192GB of RAM. - PC Stats). or something like that so it could maximise the amount of bandwith through to the CPU...

logically:

dual channel FB-DIMM 800 = 12.8 GB/sec
quad channel FB-DIMM 800 = 25.6 GB/sec
six channel FB-DIMM 800 = 38.4 GB/sec...

and if you maxed out all four controllers, you could have 115.2 GB/sec of bandwith, but that would require a lot of RAM.
i might be wrong though. But, it sounds pretty good.
 
Xeons, P4's or celerons no matter what flavor of chip it is it costs big time to add cache. If it did not then every processor you see would have 512 meg of cache on board.

And yeah it may cost Intel $40 to make some chips but that figure is without and excessive amounts of cache.

It is not possible to remove leakage on a silicon based processor it is inherit to the design and simply by shrinking the processor it creates more leakage. Also leakeage is the main reason Intel couldnt scale clock speeds. There chips even future ones are promising slower clock speeds not faster. So Ill call Bullshit on that one.

Intel may lower the leakage and thats grand, but they can't remove it. If they could we would be seeing 5 ghz+ processors by now.

Ill go back to main point. Unless Intel pulls something really secretive out of its perverbial magic hat we won't see anything competitive from them until the end of 2006 till the middle of 2007.

And this is providing that Amd doesnt change there cache levels on the current chips or introduce some nextgen chip tech.

And as far as Amd not being able to more cache to there current chips why not? There using 90nm production same as Intel, only reason they havent added more cache is they dont need to.

Also Amd is not developing in a vacume, There new 65 nanometer process is coming along and Ibm has signed another contract with them so they will be scaling down to newer production tech also.

People talk about Amd versus Intel and rave about what Intel is going to be coming out with and how its going to kill what Amd offers. But all of these assumptions have Amd keeping the exact same chips they have now with no new revisions and no new tech. Ive never figured that one out.

]
 
"Think happy thoughts here people, from what several sources have told the INQ, the leakage problem is solved, and I mean solved, not lessened. This will be a massive gain for Intel, and unless AMD and IBM can match it, it will pretty much hand it the mobile space, not to mention anything else where power matters."

http://theinquirer.net/?article=25512


Intel Researchers Develop Breakthrough Transistor Technologies To Fight Power, Heat Issues In Future Processors

SANTA CLARA, Calif., Nov. 5, 2003 - Intel Corporation today announced it has identified new materials to replace those that have been used to manufacture chips for more than 30 years. The breakthrough is a significant accomplishment as the industry races to reduce electrical current leakage in transistors -- a growing problem for chip manufacturers as more and more transistors are packed onto tiny pieces of silicon.

Intel researchers have developed record-setting, high-performance transistors using a new material, called high-k, for the "gate dielectric" and new metal materials for the transistor "gate." Transistors are the microscopic, silicon-based switches that process the ones and zeros of the digital world. The gate turns the transistor on and off and the gate dielectric is an insulator underneath it that controls the flow of electric current. Together, the new gate and gate dielectric materials help drastically reduce current leakage that leads to reduced battery power and generates unwanted heat. Intel said the new high-k material reduces leakage by more than 100 times over the silicon dioxide used for the past three decades.
http://www.intel.com/pressroom/archive/releases/20031105tech.htm

The cost of the Pentium 4 was probably the 6xx core
http://theinquirer.net/?article=26141

And one of the reasons why there is not 512 MB is becuase although Intel adds cache relatively easily, they do not add that much cache. I was talking in the ballpark of 2 MB which is technically a lot in terms of memory operating at 3 GHz.

Competition wise, Intel will have dual dual core Xeons with dual 1066 Mhz FSBs. They will be connected to the memory in dual/quad channel FB-DIMM. This is a huge boost for Intel as there will be a combined bandwith of 17 GB/sec which is several times better then the current Intel Xeons. It will operate at up to 3.46 GHz and have 4 MB of L2 cache (2 MB/core). Intel is also positioning the Yonah core as a server which is a good move. They are fast, efficient and dissipate little heat. Most of these processors are transitions to Intel's Merom, Conroe and Woodcrest chips. BTW, the Presler which requires the 975 series chipset will be compatable with the Conroe processor.

5 Ghz processors are not necessary now. If can only handle one thread. And the reason why Intel stumbled with the lastest Pentium 4's is becuase no one really expected the 90nm process to be so hard. Look at IBM. They watercool the G5 and with AMD, they had to turn to IBM for help on SOI. IBM was willing to help AMD for money, but word has it that IBM wanted access to Intel's profitible chipset market. Also, the 65nm process was supposed to be really bad at first with more leakage then the 90nm process, but Intel cut the leakage and used other power saving techniques that enables them to lower the TDP. Look at Merom, Conroe. Conroe will have a TDP of around 65 watts and that's a high speed core with 4 MB of L2 cache.
 
As I posted above they cannot remove leakage, they may lessen it but it cannot be removed from silicon as your inq article reinforces. Reducing leakge does not equate to removing. Furthermore with each scaledown in size leakage increases.

They are already producing processors with 2 meg of cache, Hell the Itanium has 16 meg of it. But what does it matter if you the processor has poor performance to begin with. Cache is simply a crutch for Intel right now.

Duel core zeons they already have, and the 1066mhz bus hasent done anything to boost performance of there current extreme processors. What its going to do for xeons wont be any different. They WILL NOT see any noticeable gains till they go to a on chip memory controller. And you can take that to the bank.

Bandwidth is irrelevant Since Intel cannot max the pipes of there current chips as it stands. If they said they will have 10000 mbs of bandwidth It wouldnt matter.

To many people who dont know what there reading get hung up by the big numbers pr people like to spin for companies like Intel. They have been spinning bs since Amd took the performance crown with the release of the Athlon 64. Some people by into it. Thats why we have Benchmarks. You see Intel boast about 2 meg of cache but they conveniently neglect to mention its a last ditch effort to crutch up performance.

5 Ghz processors are very nessasary If Intel had one they wouldnt be in second place. This very site was started from a guys desire to achieve better performance from overclocking his gaming PC.

And one thread is fine for the majority of pc users since the number of single threaded apps outnumber multi threaded 99-1. Itl bel a couple years before we start seeing mainstream multithreaded apps. Only professional pc users can make an immediate benefit. Or the people who claim that the entire time there on there pc there burning DVds. Anyone whos pirating that much data needs some help anyways.

But again tha'ts irrelevant since Intel can't compete on the multicore front either.

Im not against Intel in anyway, they just simply dont have anything to offer for a while, and the way they keep scrapping chip designs just enforces that point. I really hope they do come up with something or were just going to see another large chipmaker dictating market price like Intel has all these years.

Competition is the onlything that will keep the chip industry accelerating and prices low.
 
eblislyge said:
As I posted above they cannot remove leakage, they may lessen it but it cannot be removed from silicon as your inq article reinforces. Reducing leakge does not equate to removing. Furthermore with each scaledown in size leakage increases.

At 65nm Intel says it managed to reduce some sources of leakage 1000x in its P1265 process allowing extremply cool products.

[quoter]
They are already producing processors with 2 meg of cache, Hell the Itanium has 16 meg of it. But what does it matter if you the processor has poor performance to begin with. Cache is simply a crutch for Intel right now. [/quote]

Hmm..they had 2 MB since 1999 , 6MB since 2003 , 9MB 2004 and will top 12MB next year ( talking /core , not per cpu ).

And the proc doesn't have poor performance to begin with , cache is needed to help scaling above 4 way&commercial aplications like ERP and databases, not to mention HPC.Power 5 uses 36MB L3/ 2 cores and 144MB / MCM.PA-8900 has a 64MB L2. So your agument above cache is kinda of BS....

Duel core zeons they already have, and the 1066mhz bus hasent done anything to boost performance of there current extreme processors. What its going to do for xeons wont be any different. They WILL NOT see any noticeable gains till they go to a on chip memory controller. And you can take that to the bank.

BS , sorry. 1066 Mhz didn't provide significant performance / core because BW was enough already.But for Xeons going from a 800Mhz shared FSB ( 6.4GBs/ 4 cores ) to DIB 1066Mhz FSBs ( 17GBs/ 4 cores ) will provide a massive increase in performance for DP Xeons.

Bandwidth is irrelevant Since Intel cannot max the pipes of there current chips as it stands. If they said they will have 10000 mbs of bandwidth It wouldnt matter.

Depends on the application.For example Power 5 has 25.6GBs / cpu which allows it to rock in commercial benchmarks and altough the Madison core is more powerfull it is BW starved ( 4.2GBs/ cpu at best ).

But again tha'ts irrelevant since Intel can't compete on the multicore front either.

:D Oh..it will...it will compete so hard by this time next year that AMD will be working overtime to match it...As a sidenote Intel currently has 16 dual core projects and 10 quad-core projects for the 2005-2009 timeframe.Some are stop-gap measures but some are truly awesome.

Im not against Intel in anyway, they just simply dont have anything to offer for a while, and the way they keep scrapping chip designs just enforces that point. I really hope they do come up with something or were just going to see another large chipmaker dictating market price like Intel has all these years.

Yes , you are.Otherwise you woulnd't have wrote the part I had to remove....
 
Duke3d87 said:
And the reason why Intel stumbled with the lastest Pentium 4's is becuase no one really expected the 90nm process to be so hard.

Umhh...no.Intel's 90nm worked flawlessly.The problem was with Prescott which had the monster number of 70million logic transistors , compared to K8s less than 45 million.Prescott leaked far more as a result and used a lot more power...
 
At 65nm Intel says it managed to reduce some sources of leakage 1000x in its P1265 process allowing extremply cool products.

Where are these Cool products, show me one. Show me something, anything, from a tangible review site. I wanna see where they have made gains in something other then a press release.

Lest we forget Intel is a publicly traded company with investors to keep happy and they have been known to stretch the truth.

((Hmm..they had 2 MB since 1999 , 6MB since 2003 , 9MB 2004 and will top 12MB next year ( talking /core , not per cpu )))
I have no idea what that has to do with anything..

And the proc doesn't have poor performance to begin with , cache is needed to help scaling above 4 way&commercial aplications like ERP and databases, not to mention HPC.Power 5 uses 36MB L3/ 2 cores and 144MB / MCM.PA-8900 has a 64MB L2. So your agument above cache is kinda of BS....

If it doesnt have poor performance compared to Athlon 64's then why is this post here? Am I confused? Is this post about how Intel is dominating Amd right now or what?

Power 5 processors have nothing to do with anything being discussed these are multi thousand dollar processors of a completely different architecture.

Furthermore Intel increasing Cache to compete is exactly what they have tried to do. If they hadent been looking for a performance crutch they would have never done it in the first place. But since they couldnt scale Mhz anymore there only choice was cache. It is and has been a crutch.

BS , sorry. 1066 Mhz didn't provide significant performance / core because BW was enough already.But for Xeons going from a 800Mhz shared FSB ( 6.4GBs/ 4 cores ) to DIB 1066Mhz FSBs ( 17GBs/ 4 cores ) will provide a massive increase in performance for DP Xeons.

By massive you mean what exactly? Are we going to be reading benchmarks and be in awe? Or are you sure it wont be like it was before where the gains are so negligible no one notices?

The bandwidth is not the problem its the processor. Xeons are almost identical to P4's the only way to scale performance is to increase the mhz on that processor and add lots of cache. Im talking 500 mhz boosts and multi meg cache implements.

:D Oh..it will...it will compete so hard by this time next year that AMD will be working overtime to match it...As a sidenote Intel currently has 16 dual core projects and 10 quad-core projects for the 2005-2009 timeframe.Some are stop-gap measures but some are truly awesome.

I Have already said time multiple times in this thread that it will be next year before intel can compete. Maybe you should read the thread and not just one post before commenting on what people are discussing.

However there is so little known about what Amd is working on that Intel may not do as good as some people think they may.
Yes , you are.Otherwise you woulnd't have wrote the part I had to remove....

I dont know what it is your trying to say here, I have reread my post's ending just to make sure I didnt break Hards rules on posting. I think you should to, nothing I posted is vulgur in any way.

Unless me saying that Intel had controlled the processor prices for the last few years hurt your fellings. :confused:
 
eblislyge said:
Where are these Cool products, show me one. Show me something, anything, from a tangible review site. I wanna see where they have made gains in something other then a press release.

Huh?? That's 65nm tech which started mass production last quarter in the form of Presler and Cedar Mill cpus ( using P1264 ).

P1265 will be used for the Xscale familly and flash products ( dunno if general purpose cpus are going to use it ) and as a logical assumption , products based on it will appear in 2006.And those gains are fully explained in IDF pdfs regarding manufacturing....

Lest we forget Intel is a publicly traded company with investors to keep happy and they have been known to stretch the truth.

And Intel , btw , is a company that invented the microprocessor , the x86 ISA and several thousand other things ranging from math coprocessors to busses like PCI and standards like Plug'n'Play just to name a "few"

You are talking as if a 90000 men company suddenly turned flat stupid overnight and isn't able to do anything...


If it doesnt have poor performance compared to Athlon 64's then why is this post here? Am I confused? Is this post about how Intel is dominating Amd right now or what?

Power 5 processors have nothing to do with anything being discussed these are multi thousand dollar processors of a completely different architecture.

Furthermore Intel increasing Cache to compete is exactly what they have tried to do. If they hadent been looking for a performance crutch they would have never done it in the first place. But since they couldnt scale Mhz anymore there only choice was cache. It is and has been a crutch.

Guess what , it depends on what you define poor performance.Cause if you remove gaming benchmarks , a P4 is within +/- 5% of an A64.I hardly see that as "poor".

About the cache thing , the easiest way to increase performance without increasing thermals is to add cache.And 2MB L2 which is now standard on P4s is hardly a crutch...depending on aplications its performance benefit can vary.Secondly , since it is offered virtually for free why complain ??

As for server processors , had you bothered to check , that's where companies uses mamooth caches in order to improve performance. That's teh point after all.Should we condamn them for this as you do with Intel ??


By massive you mean what exactly? Are we going to be reading benchmarks and be in awe? Or are you sure it wont be like it was before where the gains are so negligible no one notices?

Massive means giving dual core Opteron a real run for their money...

The bandwidth is not the problem its the processor. Xeons are almost identical to P4's the only way to scale performance is to increase the mhz on that processor and add lots of cache. Im talking 500 mhz boosts and multi meg cache implements.

:D You're wrong. P4/Xeons use a 128bit L2 cache line size.Even if you transfer 1 bit or 128 you'll use 128bit of BW.K8 uses 64 coupled with very low latency makes it BW-free.

Intel used such an aproach ( Netburst in fact ) because in multimedia apps with constant data flows it shines...That's why in multimedia/content creation P4s rule...

At just 1.6GBs of BW/core todays Xeons are simply BW starved even more than the original P4 on a 400Mhz FSB.Getting the BWW increase to 4.3GBs/core brings those Xeon to the P4B level , where if you remember P4Bs crushed the AXP.Of course the K8 is not the K7 but Blackford is kinda of a savior for Xeon DP.
 
Savantu, I don't hold Intel using cache to boost performance as something negative, im all for them getting all the performance they can.

In your zest to try to prove me wrong you have entrirely missed my point with the cache in teh first place.

The typical Athlon 64 will have 512 kilo to 1 megabyte of cache depending on its designated pr rating.

Todays P4's has 1 meg or more and sometimes 2 meg. The reason Intel is using so much cache is because they cannot boost megahertz enough to get reasonable gains. Unlike the Athlon 64's the P4 architecture has to many stages and is not efficient. A 100 megahertz clock increase for intels chips is not worth the bother. However a 100 mhz increase on a athlon is a considerable performance gain when comparing the two.

Now since Intel cannot boost megahertz they boosted cache to gain performance. Amd however still has room left to increase megahertz and they have yet to play the cache card.

This is and was my point Intels cache increases are a knee jerk reaction to hold ground. Amd has not had to do this since its performance is currently superior. The downside for Intel is this does eat up wafer space which results in a loss of money earned per wafer.

However if your selling as much chip as intel does it doesnt hurt to terribly bad. But you can bet your bippy Intel would love to be able to beat Amd's performance with 512 to 1 meg caches on chip.

As far as xeons go I dont see Intel competing with Amd in the server market till the 2nd half of 2007 maybe as far away as 2008.

You keep forgetting that Amd is working on newer tech, and they still have some megahertz and cache cards to play. Have you noticed the difference 512 k of cache makes on a Amd chip? Care to guess how a duel core 2 meg cache per core Opteron would run? How about if it's scaled over to the newer lower latency DDr2 memory thats coming to market right now.?

Amd can scale performance easily with there current chips, Intel has not been able to. And industry analysts dont expect anything from them to compete till next year. Check the news on Hard ocps front page for an article on that. Hell google Amd versus Intel and see what you get.

As far as intel spinning PR and maybe stretching the truth, well thats true. You think any tech company is going to say hey we just poured millions into something to try to compete and it didnt turn out like we hoped. Hardly they play there cards same as anyone else.

I know what Intel has and hasent created( at least everything that has been marketed) And no I don't think they dont know what there doing. But they have the strongest competition they have ever had. There whole company was based on megahertz. Now that route of development has hit a wall there having to try to turn and refocus on multi core processors shorter pipelined more efficent processors. Higher performance less power consuming less heat building chips.

Literally having to try to execute a move like this overnight, and wanting to do it faster. That's why Intel isint competing and why they cannot compete until next year. They dont have the nessasary pieces in place yet.
 
savantu said:
Huh?? That's 65nm tech which started mass production last quarter in the form of Presler and Cedar Mill cpus ( using P1264 ).

Guess what , it depends on what you define poor performance.Cause if you remove gaming benchmarks , a P4 is within +/- 5% of an A64.I hardly see that as "poor".

BTW, anandtech's review of the 65nm Pentium D's stated that this is preproduction and engineering samples. There is a really good chance that the retail 65nm processors will use less power and overclock better.

savantu said:
:D You're wrong. P4/Xeons use a 128bit L2 cache line size.Even if you transfer 1 bit or 128 you'll use 128bit of BW.K8 uses 64 coupled with very low latency makes it BW-free.
The Pentium 4's and Xeons actually have a 256 bit L2 cache. the Athlon64 also uses a different kind of cache that Intel uses which makes a difference. Thus, the Pentium 4/D/Xeons all thrive on large L2 caches. That is why if you have a large L1, you have a higher latency and stuff like that.

savantu said:
Intel used such an aproach ( Netburst in fact ) because in multimedia apps with constant data flows it shines...That's why in multimedia/content creation P4s rule...
I think the fact that the Pentium 4/D/Xeon does so well with multimedia is becuase of the high core clock speed and it's highly optmized SSE2 FPU. If you read articles around when the Pentium 4 first came out, they will say that Intel wanted to market SSE2 as the new standard in FPU (multimedia and stuff like that) opertions.


savantu said:
At just 1.6GBs of BW/core todays Xeons are simply BW starved even more than the original P4 on a 400Mhz FSB.Getting the BWW increase to 4.3GBs/core brings those Xeon to the P4B level , where if you remember P4Bs crushed the AXP.Of course the K8 is not the K7 but Blackford is kinda of a savior for Xeon DP.
You have to keep in mind that the Pentium 4B also had a shorter and more efficient pipeline. I still wonder why Intel lengthened it. If you look at the Paxville DP processors, you see a really big problem. 800 MHz feeding four cores. That brings us back to the XeonMP processors a year or two ago. And they had L3 cache to help with bandwith. Blackford will be nice, but if the INQ is correct, Intel will struggle for a while with bandwith issues as the INQ reports that Intel will have the FSB structure for some time. I also remember hearing the INQ say that Intel might have used a higher latency, higher clock speed FSB for the Pentium Ds. And thus, given that the bus has latency which i'm sure it does, it would be wise for Intel to forgo the bus and use the IBM hurricane type chipset (that SLee posted about) with point to point busses. I'm not a computer engineer...yet, so i don't know the indept stuff about the memory controllers other then the fact that if Intel was to do that, they would be spending a lot of money. IBM spend over $100 million researching the hurricane type chipset.
 
Duke3d87 said:
BTW, anandtech's review of the 65nm Pentium D's stated that this is preproduction and engineering samples. There is a really good chance that the retail 65nm processors will use less power and overclock better.

I second that.

The Pentium 4's and Xeons actually have a 256 bit L2 cache. the Athlon64 also uses a different kind of cache that Intel uses which makes a difference. Thus, the Pentium 4/D/Xeons all thrive on large L2 caches. That is why if you have a large L1, you have a higher latency and stuff like that.

I'm talking about cache line size which is 128bytes as P4 technical description states.Not talking about the cache hieracrchy being inclusive or exclusive. K8s have a 64bytes cache line size , thus they aren't so affected by BW starvation ( look at socket 754 vs. 939 ).Put a P4 on a 400Mhz bus and it crawls...

I think the fact that the Pentium 4/D/Xeon does so well with multimedia is becuase of the high core clock speed and it's highly optmized SSE2 FPU. If you read articles around when the Pentium 4 first came out, they will say that Intel wanted to market SSE2 as the new standard in FPU (multimedia and stuff like that) opertions.

Indeed , but Netburst as a whole was designed to be very multimedia friendly.Intel thought 10 years ago that by now the vast majority of apps will be multimedia friendl and partly they were right.

You have to keep in mind that the Pentium 4B also had a shorter and more efficient pipeline.

Research done by Intel suggested the pipeline length should be around 40+ stages while IBM considered a 60+ pipeline stage.

I still wonder why Intel lengthened it. If you look at the Paxville DP processors, you see a really big problem. 800 MHz feeding four cores. That brings us back to the XeonMP processors a year or two ago. And they had L3 cache to help with bandwith. Blackford will be nice, but if the INQ is correct, Intel will struggle for a while with bandwith issues as the INQ reports that Intel will have the FSB structure for some time. I also remember hearing the INQ say that Intel might have used a higher latency, higher clock speed FSB for the Pentium Ds. And thus, given that the bus has latency which i'm sure it does, it would be wise for Intel to forgo the bus and use the IBM hurricane type chipset (that SLee posted about) with point to point busses. I'm not a computer engineer...yet, so i don't know the indept stuff about the memory controllers other then the fact that if Intel was to do that, they would be spending a lot of money. IBM spend over $100 million researching the hurricane type chipset.

Blackford and the current E8500 ( for Xeons MPs ) have DIB , which are point to point busses.Paxville wasn't designed for the current E7x20 chipsets and as a result performs poorly due to BW starvation.The core itself is an old revision and burns a lot of power.

IMO , Blackford and Dempsey are going to rock.There's no way for Dempsey to be BW starved , in fact it will have 30% more BW / core than dual core Opterons which is a first for an Intel server.

Hurricane is truly a marvel of engineering , and its next revision which will support the 800Mhz FSB Paxville MPs due for launch in a week will be a solid performer.
It is ironical that most people think Intel will equal the Opteron based servers in 2007 when in fact the X366 and x460 ( Xeon based Hurricane servers ) are the best performing x86 servers today from 4 way and up ( to a monster 32 way / 512GB memory x86 beast ) beating all other Opteron servers from HP,SUN,etc.
 
I just saw benchmarks by IBM. A quad 3.66 XeonMP outperformed a quad 2.6 Opteron by 8% and that's with the slower FSB. The next revsion as noted by you should be quite intersting. What i don't get is how Intel severs outperform AMDs servers. If you look at the desktop line, the Pentium D series are usually the most cost effective procesors, but their performance is a bit lower then AMDs. While i'll relent on the fact that the NetBurst core was never designed to be put into dual core configuration, it seems strange. Maybe it's the fact that the XeonMPs have up to 8 MB of L3 cache?

And why would Intel and IBM want a 40-60 stage pipeline? It makes it a lot harder to make an efficient processor. I often thought that if i was the Intel roadmap guy, i would tellt he engineers that we wanted a long pipeline processor, have them design it so that it is really efficient and be like "BTW, I was kidding about the long pipeline." and have really efficient processor.

With Intel's next generation servers, 1066 MHz is really good, but i don't really like how all of them won't be 1066 Mhz. You will have Xeons with slower busses.

When i meant point to point, i mean like have a hypertransport type thing and not a FSB...

for example..

Have a dual core Xeon with NO FSB and just a hypertransport connection to a Hurricane type chipset. And then have your quad channel FBDIMMs in your four memory controllers. That way, you are not limited by the FSB, but by how man sticks of ram you have. You also should have less latency from the lack of FSB. Becuase the problem that i think IBM has to deal with is that although the hurricane chipset can crank out 21ish GB/sec, the bus can only do about 12.8 GB/sec which leaves a good chunk of bandwith to other parts rather then the CPU.
 
Duke3d87 said:
I just saw benchmarks by IBM. A quad 3.66 XeonMP outperformed a quad 2.6 Opteron by 8% and that's with the slower FSB. The next revsion as noted by you should be quite intersting. What i don't get is how Intel severs outperform AMDs servers. If you look at the desktop line, the Pentium D series are usually the most cost effective procesors, but their performance is a bit lower then AMDs. While i'll relent on the fact that the NetBurst core was never designed to be put into dual core configuration, it seems strange. Maybe it's the fact that the XeonMPs have up to 8 MB of L3 cache?

The 3.66 Xeon MP has only a 1MB L2 cache and costs less than 900$ while a 2.6Ghz Opteron 852 is over 1500$ ( list price ).The 3.66Ghz / 1MB L2 part , known as "poor man's Xeon MP" is only available in 4P configurations ( aka 366 ).From what I've seen , over 4P the 3Ghz dual core Paxville and single core 3.33Ghz/8MB L3 part is used. ( for scaling reasons ).

If you want to see real muscle , look for IBM's 8P Hurricane TPC-C submission with 3.33Ghz /8MB L3 Xeon MPs and then compare it to the 4 socket dual core 2.4Ghz HP 585 ( Opteron 8P ).Look at the memory quantity for both.

And why would Intel and IBM want a 40-60 stage pipeline? It makes it a lot harder to make an efficient processor. I often thought that if i was the Intel roadmap guy, i would tellt he engineers that we wanted a long pipeline processor, have them design it so that it is really efficient and be like "BTW, I was kidding about the long pipeline." and have really efficient processor.

Probably the engineers where asked : "make a study and see what's the ideal pipeline lenght leaving aside any process restrictions ( heat,etc )".

With Intel's next generation servers, 1066 MHz is really good, but i don't really like how all of them won't be 1066 Mhz. You will have Xeons with slower busses.

From what I've seen :

H1 2k6

Xeon DP 800Mhz shared FSB ( 6.4GBs ) replaced by 1066Mhz DIB ( 17GBs )
Xeon MP 667Mhz DIB ( 10.6GBs ) coexists with 800Mhz DIB ( 12.8GBs )

H2 2k6

Xeon DP 1066 DIB ( 17GBs ) coexists with 1333Mhz DIB (21GBs )
Xeon MP 800Mhz DIB ( 12.8GBs )

Not too bad I'd say....

When i meant point to point, i mean like have a hypertransport type thing and not a FSB...

for example..

Have a dual core Xeon with NO FSB and just a hypertransport connection to a Hurricane type chipset. And then have your quad channel FBDIMMs in your four memory controllers. That way, you are not limited by the FSB, but by how man sticks of ram you have. You also should have less latency from the lack of FSB. Becuase the problem that i think IBM has to deal with is that although the hurricane chipset can crank out 21ish GB/sec, the bus can only do about 12.8 GB/sec which leaves a good chunk of bandwith to other parts rather then the CPU.

Depends.The FSB aproach has some advantages , some of which are very important for enterprise level servers.

First of all , having a chipset handling memory means you can have extensive RAS like hot swap , memory RAID , ChipKill , etc etc
By giving each cpu its RAM ( opteron ) a cpu failure will crash the system no matter what you do.A server with chipset integrated memory controller survives.it is simply more fault tolerant.

Look at the RAS features of high end Xeon/Itanium boxes and compare those to Opteron based ones.They are a world apart.And some are exactly there due to not having an OMC.

Intel's aproach with Tigerton , using dedicated PoP links per socket is an excellent compromise , but it requires a complicated chipset.Fortunately Intel is the best at building state-of-the art chipsets.
 
What i meant was not have an integrated memory controller, but have all the memory controllers on the motherboard and have point to point busses to it. Anyway, what is memory raid? I've heard that a lot recently and never really found out what that is. Do you think that the upcomming DIB chipsets is Intel's version of practicing for Tigerton? I know that they are really good at building complicated chipsets, but they are also known to have a lot of memory bottlenecks. And is that 1.33 GHz FSB that you mentioned for Q2 of '06 the Woodcrest one?
 
Can we talk more about what the original poster said, that being intel as a revenue driven company vs. pure technology for technology's sake. Anyone have any information closer to the inner business workings behind the technology decisions?
 
Personally i think some are forgetting just how BIG intel is compared to AMD. Intel is a giant with incredibly marketing and financial power. Just like Microsoft they are driven by the cold hard cash which they dont need as much as AMD. Intel can go trailing behind no problem and rest on their fat bellys until the sun doesnt raise anymore. Never count Intel out of the game, they always had a greater plan and they will stay the giant in this game.

:D
 
Duke3d87 said:
What i meant was not have an integrated memory controller, but have all the memory controllers on the motherboard and have point to point busses to it.

Intel did this with partly with the E8500 for the Xeon MP ( altough there are still 2 scokets/bus ) and the upcoming Blackford for Xeon DP ( 1 socket / bus ).Tigerton will feature a similar 1 socket/bus aproach but for Xeon MPs.

Anyway, what is memory raid? I've heard that a lot recently and never really found out what that is.

For HPs server , DL7xx , similar to RAID Level 4 for HDDs.

"Servers with HP Hot Plug RAID Memory use five memory controllers to control five memory cartridges. Each cartridge can hold up to eight industry-standard DIMMs When the memory controllers need to write data to memory, they split the data into four blocks and write them to four of the memory cartridges. A RAID engine calculates parity information, which is stored on the fifth cartridge. With the four data cartridges and the parity cartridge, the data subsystem is completely redundant such that if the data from any DIMM is incorrect or any cartridge is removed, the data can be recreated from the remaining four cartridges.The redundancy in HP Hot Plug RAID Memory allows customers to hot-replace, hot-add, and hot-upgrade DIMMs without shutting down the server. Hot replace is replacing a failed DIMM while the system continues to operate."

Another option is Mirrored Memory mode which is "a fault-tolerant memory option that will provide a higher level of availability than Online Spare Memory. Online Spare Memory mode protects against single-bit errors and entire DRAM failure, but Mirrored Memory mode provides full protection against single-bit and multi-bit errors."

It is in fact very similar to RAID 1 for HDDs.You divide the memory into 2 sections : system memory and mirrored memory.The system writes data to both but only reads from system memory.Any error detected in the system memory the sections change role , systems becomes mirrrored and viceversa , still keeping the write/read policies.
Thus you are virtually immune to just about every type of memory error. ( it fails only if the same error happens in the same place in both system and mirrored which is highly unlikely - but then you have memory RAID to get you out of this :D )

Add to this advanced ECC , Online Spare , Hot Plug , Hot Add ....

Do you think that the upcomming DIB chipsets is Intel's version of practicing for Tigerton? [/quote ]

Well , E8500 is 6 months old by now and it has DIB. :) Tigerton is different because it will use 1 socket/bus vs. the current 2 sockets/bus. Tigerton's aproach should be called MIB ( multiple independent bus ) or smt like that. :p

I know that they are really good at building complicated chipsets, but they are also known to have a lot of memory bottlenecks.

What memory bottlenecks ?

For example the F8 chipset designed in 2001 for Xeon MPs by Compaq had a 8.5GBs memory BW while the CPUs used 3.2GBs.The remainng BW allowed excellent peripheral performance and thus it competed well with Opteron until this year when it was replaced by the E8500 ( which is a masterpiece btw )

E8500 has 4 XMBs ( eXternal Memory Bridge) , each with 2 channels , 4 dimms /channel of DDR 2 memory ( 32DIMMs in all ).If you use 4GB dimms you end up with 128GB of RAM.

So you have 4 Independent Memory Interface (IMI) ports, each with up to 5.33 GB/s read bandwidth and 2.67 GB/s write bandwidth simultaneously , 40-bit addressing supporting up to one terabyte (2 x 102 bytes) addressing (this is in excess of maximum physical memory supported by the Intel® E8500 chipset platform)

So altough the CPUs only get 10.6GBs , the chipset has 21.3GBs available.This allows it to use peripherals to the max ( PCI-E , PCI-X , GB LAN , 10GB LAN ).

Are this memory bottlenescks ???

And is that 1.33 GHz FSB that you mentioned for Q2 of '06 the Woodcrest one?

Yes.
 
Ultra Wide said:
Can we talk more about what the original poster said, that being intel as a revenue driven company vs. pure technology for technology's sake. Anyone have any information closer to the inner business workings behind the technology decisions?

Wow , tell me a company which is not revenue driven ??

Who , for God's sake , makes some technology just for technology's sake ??

Whitefield was cancelled because it apparently had a serious problem with its CSI implementation.The cores , were basically 2 Woodcrest with a large shared L3 cache ( 2.93+ GHz , 2/4MB L2 , 16MB shared L3 ) are problem free.Who knows...
 
savantu said:
Intel did this with partly with the E8500 for the Xeon MP ( altough there are still 2 scokets/bus ) and the upcoming Blackford for Xeon DP ( 1 socket / bus ).Tigerton will feature a similar 1 socket/bus aproach but for Xeon MPs.

Well , E8500 is 6 months old by now and it has DIB. :) Tigerton is different because it will use 1 socket/bus vs. the current 2 sockets/bus. Tigerton's aproach should be called MIB ( multiple independent bus ) or smt like that. :p[/QUOTE]

How do you know that is it 1 socket/bus approach? From what i hear, it will be the same FSB as the previous generation Xeons, just modified.

savantu said:
What memory bottlenecks ?

For example the F8 chipset designed in 2001 for Xeon MPs by Compaq had a 8.5GBs memory BW while the CPUs used 3.2GBs.The remainng BW allowed excellent peripheral performance and thus it competed well with Opteron until this year when it was replaced by the E8500 ( which is a masterpiece btw )

E8500 has 4 XMBs ( eXternal Memory Bridge) , each with 2 channels , 4 dimms /channel of DDR 2 memory ( 32DIMMs in all ).If you use 4GB dimms you end up with 128GB of RAM.

So you have 4 Independent Memory Interface (IMI) ports, each with up to 5.33 GB/s read bandwidth and 2.67 GB/s write bandwidth simultaneously , 40-bit addressing supporting up to one terabyte (2 x 102 bytes) addressing (this is in excess of maximum physical memory supported by the Intel® E8500 chipset platform)

So altough the CPUs only get 10.6GBs , the chipset has 21.3GBs available.This allows it to use peripherals to the max ( PCI-E , PCI-X , GB LAN , 10GB LAN ).

Are this memory bottlenescks ???
a. I'm not that familiar with XeonMP chipset so ya.
b. I mean to the CPU. IMO, if you can crank out 21 GB/sec and only 10 GB/sec goes to four CPUs, i consider it a bottleneck b/c the FSB is preventing some bandwith to the CPU.
 
Back
Top