• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Hyperthreading: Please explain

tubedogg

n00b
Joined
Mar 8, 2009
Messages
40
My (most likely incorrect) understanding is that hyperthreading works by fooling the O/S into thinking that there are twice the number of cores, with each running at half the speed.

I guess the theory is IF the typical task the PC is used for can be split up into even chunks (ie: video editing), then it makes sense to enable hyperthreading. However, why does this make things faster?

example, before:
core 1, 3GHZ
core 2, 3GHZ
core 3, 3GHZ
core 4, 3GHZ

after:
core 1, 1.5GHZ
core 2, 1.5GHZ
core 3, 1.5GHZ
core 4, 1.5GHZ
core 5, 1.5GHZ
core 6, 1.5GHZ
core 7, 1.5GHZ
core 8, 1.5GHZ

Seems like net processing power should be the same?
 
no. no. no.

each core still runs at 3ghz, but it can handle 2 requests at the same time. it's a slight speed improvement. when pentium 4's came out with HT, it wasn't anything groundbreaking, but it did offer some speed improvement/multitasking ability
 
I'll just repost basically what I posted a week ago when a similar question was asked:

It works by allowing the processor to make better use of its computation resources, not adding additional resources. For example, if one thread is waiting on main memory for some data, this will not eject the thread from the core, so that core is stalled until the data arrives. In an HT system, another running thread may have instructions ready to execute immediately, and the processor will send them off to the execution unit, so work gets done during a period when the processor would normally be stalled. Also you might have a situation where you've got an FPU-heavy thread on one core and an integer heavy one on the other; it may be possible for the CPU to utilize both its FPU and ALU at the same time. The 'front end' (decoding, scheduling etc. hardware) of the processor keeps track of two or more threads simultaneously, but the 'back end' (execution hardware) stays largely the same, it's all about keeping the execution units as busy as possible and minimizing the amount of the time processor is not actually accomplishing work. Most processor developments in the past 10 years have been of this type - keeping the pipeline full.

You should read this article: http://arstechnica.com/old/content/2002/10/hyperthreading.ars and of course the various other related articles at Ars about superscalar architecture, out of order execution and etc.
 
If you have 2 people, and 2 dishwashers and for some reason the dishwashers run every single night without fail, sometimes a dinner has enough dishes to fill both dishwashers.
But sometimes you only have enough dishes from a dinner to fill one dishwasher.

Hyperthreading offers to wash your neighbors dishes too, instead of running one while its empty, as you have extra space available.
 
Oddly enough I ran a bunch of programs I use daily and I notice no difference in perception of HT on vs off. If I really care about benchmark numbers I would say that with HT off my programs where slightly faster in numbers.

I keep it off as it lowers my temps by 9c when running on load...
 
I was hoping Photoshop would have, used all 8 technically. But the results show that the 4 cores was faster then the HT enabled Core i7. So every program that I use that says, "multi-capable" never seemed to show any increase in performance. So I just turned it off. Considering when running games my temps drop by 9c during load.

I'm still scratching my head to figure out what consumer home based programs will use more then 4 cores?

It seems if you use any intensive FPU programs it'll show some difference. But none of my programs do.
 
Benchmarking generally shows HT giving a 0-20% improvement that depends heavily on the application in question. The reason your temperature is lower is because more of your CPU is idle (meaning you are working the CPU as hard). Thus, you have lower temps because your CPU is doing less work.

I don't know about Photoshop in particular, but I have yet to see a case where HT actually resulted in lowered performance. It is generally either no improvement or a very slight improvement.
 
I was hoping Photoshop would have, used all 8 technically. But the results show that the 4 cores was faster then the HT enabled Core i7. So every program that I use that says, "multi-capable" never seemed to show any increase in performance. So I just turned it off. Considering when running games my temps drop by 9c during load.

I'm still scratching my head to figure out what consumer home based programs will use more then 4 cores?

It seems if you use any intensive FPU programs it'll show some difference. But none of my programs do.
Individual programs won't take advantage of HT. What will benefit from it is heavy multitasking while running multiple CPU-intensive tasks at the same time. You'll see overall better performance in each program compared to having HT off, and the system will also be smoother and more responsive.
 
Individual programs won't take advantage of HT. What will benefit from it is heavy multitasking while running multiple CPU-intensive tasks at the same time. You'll see overall better performance in each program compared to having HT off, and the system will also be smoother and more responsive.

You never encoded much I can hear.
 
It's like waiting in line for the movies. Instead of there being 1 waiting line, there's 2 waiting lines. So more people get into the movies, faster.
 
I just thought that Photoshop would of benefited from HT, but it looks like it doesn't. Plus, I guess since I don't usually have like 4 to 8 programs running (minimized) I wouldn't really see any benefit of HT enabled. Does it hurt to have it on? Not really, but what's odd is that the benchmarks I've ran showed lower scores like 2% slower then when HT was enabled. Far Cry 2 and Resident Evil 5 showed a slight improvement when HT was disabled. When I say slight. I'm saying like 2 to 4 fps on average. lol...
 
It's like waiting in line for the movies. Instead of there being 1 waiting line, there's 2 waiting lines. So more people get into the movies, faster.

No that is multicore you are describing....


If you want to take that approach, it is more like there being two lines (threads), but only one ticket taker (core)

the speedup would come from the ticket taker being able to take money and give tickets to whoever had their movie picked out and their money ready first.
 
If you want to take that approach, it is more like there being two lines (threads), but only one ticket taker (core)

the speedup would come from the ticket taker being able to take money and give tickets to whoever had their movie picked out and their money ready first.

That is more accurate...but not 100% right either.
 
If you want to take that approach, it is more like there being two lines (threads), but only one ticket taker (core)

the speedup would come from the ticket taker being able to take money and give tickets to whoever had their movie picked out and their money ready first.

Still wrong. Its more like this:

There's one ticket person (core), but this person has two cash registers (hyper thread doubles some computational bits), but only one cash drawer.

The two threads on the single core complete over some resources (mostly the memory bits), but when it comes to computation they can do that in parallel.
 
That last one was a really poor analogy on what they were trying to make it lol. Made it more confusing.
 
Haha, bad analogy is bad.

Just read the Ars article, it's explained very well.

Code that's already been heavily optimized to make good utilization of the computation resources and minimize cache misses (ie. Photoshop filters) is going to benefit very little from HT because it already does a good job at keeping the execution units busy. However code that is poorly optimized or otherwise doesn't make good use of the execution resources can benefit a good deal.
 
Here's a link to an older article regarding Hyperthreading on the P4:

http://ixbtlabs.com/articles/pentium4xeonhyperthreading/

Things have improved since then, and L2/L3 caches have grown leaps and bounds since those times.

Look at the charts to have a better visual representation of what HT does. This talk of cash registers and dishwashers can be a little confusing for the casual user.
 
I'm still hoping for a good analogy. My understanding of this can only be solidified by some kind of story/fable/limerick. More analogies, please!
I dislike analogies, because they are just that - analogies. They don't represent fully the situation. Sometimes that can be a good thing, but I've learned through experience that more often than not it's a bad thing. People don't fully grasp the differences between the analogy and the actual thing, and come out with a distorted view of the reality. Therefore, I'm going to try explain the actual situation, but as simply as possible.

What happens is that for a core to work on a piece of data it has to get it from memory. Your memory transfers it into a series of holding pools called caches. The processor proper then finds, reads, and works on the data from the cache before sending it back.

This was fine originally, when clock speeds were a lot slower. However, as clock speeds increased, it became clear that the memory was bottlenecking the processor. The problem was that in the time it took to find and send a piece of data to the processor, the processor itself could have executed hundreds or thousands of cycles. Obviously a big waste of computational power. Therefore, people added more caches, which got smaller and smaller, and faster and faster, to adequately feed the processor. Level 1 cache is the fastest and slowest, L3 is the largest and slowest (though still very, very fast in comparison to almost every other data storage solution we have), and L2 is sort of in between.

This worked fine for a while, but after a while, it became apparent that this was still holding the system back. So we got hyperthreading. What happens is that your processor works on one bit of data (let's call it Thread 0). While the data from Thread 0 is being sent back and another piece of data is retrieved for it to start work again, the processor puts Thread 0 to one side and works on another bit of data (say, Thread 1). When Thread 1 has finished being worked on, it gets sent back and another bit of data retrieved. In the meantime, the data for Thread 0 has come back, so the core starts work on Thread 1 again. In this way, the core switches between the two threads to make sure that it works as efficiently as possible, i.e. the minimum time is spent twiddling its thumbs, waiting for the caches to supply it with data.
 
What happens is that for a core to work on a piece of data it has to get it from memory. Your memory transfers it into a series of holding pools called caches. The processor proper then finds, reads, and works on the data from the cache before sending it back.

This was fine originally, when clock speeds were a lot slower. However, as clock speeds increased, it became clear that the memory was bottlenecking the processor. The problem was that in the time it took to find and send a piece of data to the processor, the processor itself could have executed hundreds or thousands of cycles. Obviously a big waste of computational power. Therefore, people added more caches, which got smaller and smaller, and faster and faster, to adequately feed the processor. Level 1 cache is the fastest and slowest, L3 is the largest and slowest (though still very, very fast in comparison to almost every other data storage solution we have), and L2 is sort of in between.

This worked fine for a while, but after a while, it became apparent that this was still holding the system back. So we got hyperthreading. What happens is that your processor works on one bit of data (let's call it Thread 0). While the data from Thread 0 is being sent back and another piece of data is retrieved for it to start work again, the processor puts Thread 0 to one side and works on another bit of data (say, Thread 1). When Thread 1 has finished being worked on, it gets sent back and another bit of data retrieved. In the meantime, the data for Thread 0 has come back, so the core starts work on Thread 1 again. In this way, the core switches between the two threads to make sure that it works as efficiently as possible, i.e. the minimum time is spent twiddling its thumbs, waiting for the caches to supply it with data.
That's not how HT works. In fact, that's closer to how a typical single-threaded CPU functions. HT allows a CPU to process two threads simultaneously by taking advantage of the fact that most single threads don't actually use up a CPU's entire set of execution resources. There's no switching between threads.
 
That's not how HT works. In fact, that's closer to how a typical single-threaded CPU functions. HT allows a CPU to process two threads simultaneously by taking advantage of the fact that most single threads don't actually use up a CPU's entire set of execution resources. There's no switching between threads.
exactly.

Hyperthreading allows the processor to work on 2 threads at a time. Some resources are shared, some are dynamically allocated per thread, and some resources are doubled to allow it. The actual execution units are shared, each thread can access all or none of them, depending on need/balance.
 
Here's what (I hope) is the best analogy to date:

Think of your system to be a big hotel restaurant. The restaurant is so big that it has 4 kitchens. The duty chef in each kitchen is an iron-fisted tight ass and only allows one dish to be made at a time. This is your standard quad-core CPU (C2Q or i5).

You go to another restaurant. It's about as big and also has 4 kitchens. This time around, the duy chef in each kitchen is under pressure from the management to cook faster and increase turnover, so he allows 2 dishes to be made at a time. Even so, the 2 dishes being made at the same time can't be using the same parts of the kitchen at once -- you can have one dish in the oven and one dish in the deep fryer, but you can't have 2 dishes in the oven or 2 dishes in the deep fryer. This is your Hyperthreaded quad-core CPU (i7).
 
Here's what (I hope) is the best analogy to date:

Think of your system to be a big hotel restaurant. The restaurant is so big that it has 4 kitchens. The duty chef in each kitchen is an iron-fisted tight ass and only allows one dish to be made at a time. This is your standard quad-core CPU (C2Q or i5).

You go to another restaurant. It's about as big and also has 4 kitchens. This time around, the duy chef in each kitchen is under pressure from the management to cook faster and increase turnover, so he allows 2 dishes to be made at a time. Even so, the 2 dishes being made at the same time can't be using the same parts of the kitchen at once -- you can have one dish in the oven and one dish in the deep fryer, but you can't have 2 dishes in the oven or 2 dishes in the deep fryer. This is your Hyperthreaded quad-core CPU (i7).
Closer, but not exactly correct.

Trying to find an analogy is really useless, since there won't be any that fit perfectly. Just directly look up how the technology works and that'll save you from a lot of confusion.
 
Back
Top