• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

How much difference does a MHz really make?

Cerulean

[H]F Junkie
2FA
Joined
Jul 27, 2006
Messages
9,484
This is a science fair project I am doing. Last year I did it on a comparison of three different and common cooling methods (passive and active air cooling, and water cooling).

My question this year is "how much difference a MHz in a CPU really makes?" I intend to calculate Pi to 32 million places, since Pi is an infinite number with no pattern, while having the program give a very precise and accurate time of how long it took. I will be using SuperPi specifically to calculate Pi.

Right now I am only in the stages of doing research. I did Google my question, and quite fortunately (for me at least) there are no results for my particular question.

So what do I need help with? Well I need help and sources to explain why calculating Pi is a logical way to determine the MHz power of a CPU. Since GPUs and PPUs are better at math than CPUs are, this is why I raise this thought.

CPUs have a mathematical side and logical side, don't they? The mathematical side literally carries out mathematical equations and instructions, etc, kind of like a GPU, but not as efficiently or as well. The thing that CPUs are more specialized in is logic -- or like carrying out a program's instructions and stuff; if a GPU took the place of a CPU, it would struggle to do these things, whereas the CPU would struggle to render material.

EDIT :: Actually, it wouldn't matter...almost. If Pi were calculated on a CPU and GPU, you would get a result that would give you an idea of how many more operations per second one could do over the other.

So in that case, how is the GPU different? How can the GPU do math better than a CPU? Pi -- how would the GPU be able to calculate (for example) 32 million digits faster than a CPU? What IS this factor or thing?

Wow, now I'm really curious. I might modify my topic question to have something to do with CPUs vs GPUs and the MHz difference. o_o
 
For a 1GHz cpu, 1MHz is .1% of its overall performance.
I'm sure you know this but I am unsure why you need to ask.
 
The GPU, by design, is massively parallel, and excels in number crunching (I think it's got lots of ALUs?), whereas the CPU excels in logic... or something or other...
 
I think a problem you're going to run into (as far as accuracy goes) is that 1MHz on one CPU will get more work done than 1MHz on another CPU, even from the same company. For example, a Pentium D at 3GHz is actually slower than a Core 2 Duo at 2GHz.
 
The difference between a CPU and a GPU is in how they divide the work to be done.

CPUs, being general purpose, will perform any task in a more-or-less standard way - one step after another. They can do some re-ordering of instructions to take shortcuts, but they basically do everything one step at a time. For numeric calculations, there is one numeric processor (FPU or floating point unit) for each CPU.

GPUs are not general purpose. They will have many components that can do the same kinds of actions on data in parallel. Since number crunching is critical to graphics, they are optimized to do this number crunching. This is like having multiple FPUs per device.

Some systems, such as the Cell Broadband Engine in the Sony Playstation 3, are combined to do both. The Cell has a single general-purpose CPU (a Power architecture computer) along with 8 special purpose processors on a single chip. These special processors are called SPEs and are basically a kind of FPU optimized for number crunching. In the PS3, only 7 of the 8 SPEs are active and one is reserved for the CPU to use, leaving 6 available for software to use. In this computer, the CPU makes sure the SPEs are stuffed with numbers to calculate. This is similar to a CPU/GPU, where the CPU does mundane work and the GPU does the number-intensive graphics processing.

With the Cell, the system is so optimized that it can outperform all other current CPUs at the same MHz on numeric-intensive software (like scientific computing). This is largely due to the fact that while there is one FPU per CPU, the Cell has up to 8 FPUs per CPU.

What it comes down to is how parallel the design is. If only one instruction can be performed per Hz, then the processing speed is directly related to the clock speed. If multiple instructions can be performed per Hz, then processing speed is greater than the clock speed. The actual performance is dependent on how well the software and the hardware design can keep the multiple processing streams busy. Using the example of the Cell, if a program can keep all 8 SPEs busy all the time, it will be noticeably faster than a program that can only keep, on average, 4 SPEs busy.

There are many other factors that can affect the apparent speed - how fast the memory is, how much cache memory the CPU has, how well the computer keeps the cache full of data that the CPU needs, etc. Different companies, like Intel, AMD, Via, etc, will use slightly different ways of implementing hardware algorithms. This means they will reveal slightly different performance for the exact same computer program.


Using the example of calculating pi, if the program can get several calculations to be done at once, then it will be able to use a parallel-processing-capable processor like a GPU faster than a single CPU plus FPU that has to do calculations one at a time. If you can't do this and the program will always do one calculation at a time, then there will be little difference in the speed of a CPU+FPU, a GPU or a Cell and the result will pretty much align with the speed in Hz.

Note: I've described CPUs as if there's one on a chip. Some CPUs have two, three or four processors (e.g. Core 2 Duo, AMD's triple CPU chip or the Core 2 Quad). These still have one FPU per processor and can operate in parallel, but not quite as efficiently as a proper parallel processing somputer. This is because they need one processing thread per processor, each doing one instruction in sequence rather than one parallel processor doing a single thread with multiple numeric calculations at a time.
 
Using the example of calculating pi, if the program can get several calculations to be done at once, then it will be able to use a parallel-processing-capable processor like a GPU faster than a single CPU plus FPU that has to do calculations one at a time. If you can't do this and the program will always do one calculation at a time, then there will be little difference in the speed of a CPU+FPU, a GPU or a Cell and the result will pretty much align with the speed in Hz.
So normally, mathematically calculations can only be calculated by 1 thread and 1 core at once, from start to finish (linear). But in order to use multiple threads, the program that is sending the mathematical instructions to the CPU would have to split up the work load and then direct each workload piece to specific cores, and then recombine the results to get things done faster.

I have a Q9450 clocked at 3400MHz (3.4GHz); it would still only have 1 FPU, not 4, because the "4 cores" are really just threads, and not actual processors/CPUs?

I need the following terms cleared up:
- core(s)
- CPUs (plural)
- thread(s)
 
So normally, mathematically calculations can only be calculated by 1 thread and 1 core at once, from start to finish (linear). But in order to use multiple threads, the program that is sending the mathematical instructions to the CPU would have to split up the work load and then direct each workload piece to specific cores, and then recombine the results to get things done faster.

I have a Q9450 clocked at 3400MHz (3.4GHz); it would still only have 1 FPU, not 4, because the "4 cores" are really just threads, and not actual processors/CPUs?

I need the following terms cleared up:
- core(s)
- CPUs (plural)
- thread(s)

Wrong. X2, X3, X4, Exxxx, and Qxxxx chips have multiple complete CPU cores on the package. That includes FPU, ALU, logic unit, bus control unit, prefetch, branch prediction, L1 caches, registers, etc. Everything is replicated.

Core is a complete CPU.
"CPUs" tends to mean multi-socket system.
A thread is a stream of instructions in a program that need to be processed.

For all intents and purposes there is very little difference between multi-core and multi-CPU system (there are some major difference at the hardware and OS level, but they're beyond the scope of this thread).
 
Based on the conversations I'm having with the three buddies (programmers) I know, it seems that this topic question is massively flawed to the point where I cannot perform this without breaking or bending the scientific method. Maybe not, I'll see about this soon. D;

PS. I'm Qwerty and hlrsecom.

Conversation with Snrrrub
Code:
[17:26] Qwerty: http://hardforum.com/showthread.php?p=1033021929#post1033021929 
[17:27] Qwerty: You know how to program, so I was wondering if you have any input in regards to my thread. :) 
[17:37] Snrrrub: I think it's very difficult to give a definitive answer to the question you pose (in your science fair project)
[17:37] Qwerty: haha :P 
[17:37] Snrrrub: I think the issue is the form of the question itself
[17:38] Qwerty: I'm currently also talking to RJ (that other buddy of mine I told you about), and my cousin (also a programmer) 
[17:38] Qwerty: and you ;o 
[17:38] Snrrrub: "How much difference does a MHz make" - difference in terms of?
[17:38] Qwerty: I'm trying to get my confusions and stuff sorted out 
[17:38] Qwerty: because I really want to understand this 
[17:38] Qwerty: I was thinking like 
[17:38] Qwerty: If I clocked my CPU to 800MHz 
[17:38] Qwerty: and increased the clock by intervals of 8 MHz or something 
[17:39] Qwerty: and then benchmarked it based on a calculation of Pi to 32 million digits with an accurate/precise timer to milliseconds 
[17:39] Qwerty: something along those lines 
[17:39] Snrrrub: In general, you'll get a linear increase in performance
[17:40] Snrrrub: In reality, there are too many other factors that will add noise to your measurements
[17:41] Snrrrub: In particular, issues like CPU scheduling, page faults, branch prediction, etc. will all have an impact on your measurements
[17:41] Snrrrub: The question you're trying to answer (to me) is more about where the bottleneck is.
[17:43] Snrrrub: One potential bottleneck could be the FSB so even if you increase the MHz of your CPU, your FSB won't be able to get data to/from your CPU fast enough and you'll be stuck.
[17:44] Qwerty: :\ 
[17:44] Snrrrub: There's a lot of complexity in the system.
[17:45] Snrrrub: Hardware and software engineers have done a lot to extract performance out of machines. Caches and prefetching play a huge role. The *type* of CPU plays a huge role too.
[17:47] Snrrrub: RISC vs. CISC, pipeline depth... all I can say is that performance measurement is *extremely* difficult.
[17:49] Qwerty: o_o 
[17:49] Qwerty: simultaneously asking questions to 3 programmers is very interesting 
[17:50] Snrrrub: And that's just a single CPU, single core, single thread (i.e. not a hyperthreaded system)
[17:50] Qwerty: at the moment, with the other two, we're discussing calculating pi on multiple cores 
[17:50] Qwerty: dang, really? 
[17:50] Qwerty: So with a quadcore CPU (Q9450 for example), what else? 
[17:51] Snrrrub: Yeah, with multi-core you're going to have cache contention issues, even bigger CPU scheduling issues, data consistency, inter-processor synchronization...
[17:51] Qwerty: But this is for if a person wanted to get really technical about measuring performance per MHz, right? On regular usage, quadcore > single core? 
[17:52] Qwerty: Or does Quadcore actually slow you down? 
[17:52] Snrrrub: (I worked in the Virtual Machine Performance team at MS so these issues are near and dear to my heart)
[17:52] Snrrrub: No, generally speaking, quadcore is not faster than single core.
[17:52] Snrrrub: On *some* workloads, yes. Likewise, on others, no.
[17:53] Snrrrub: In fact, depending on how the algorithm is implemented, quadcore can be slower than single core.
[17:54] Qwerty: Application-wise, it just depends whether a program was designed to work with multiple cores or not, and if it was designed efficiently? 
[17:54] Snrrrub: Right. But frankly, it doesn't even matter because there are virtually no applications in the "average user" domain that use the CPU to its fullest (unless they're poorly written)
[17:55] Snrrrub: So whether you have 1 or 2 or 10, it doesn't matter - it's not like that 1 CPU is running at 100% :)
[17:55] Snrrrub: Right now, the major issues are at the bus and interconnect level

Conversation with Abe
Code:
[17:26] hlrsecom:  http://hardforum.com/showthread.php?p=1033021929#post1033021929
[17:27] hlrsecom: You know how to program, so I was wondering if you have any input in regards to my thread. :) 
[17:27] Abe♥: whats' the question?
[17:28] Abe♥: gpu = massively parallel, cpu = serial (as in; one thread at a time)
[17:29] hlrsecom: my CPU has 4 threads (aka 4 "cores") 
[17:29] Abe♥: um no
[17:29] Abe♥: it has 4 cores
[17:29] hlrsecom: explain D: 
[17:29] hlrsecom: the core / thread thing is confusing 
[17:29] Abe♥: which each can process one thread at a time
[17:29] Abe♥: a process or a thread is is something a processor executes
[17:29] hlrsecom: and at the same time, it shares the MHz bandwidth? 
[17:30] hlrsecom: ie. I have my CPU clocked at 3400MHz 
[17:30] hlrsecom: and that would be "shared bandwidth" 
[17:30] hlrsecom: amongst the 4 cores 
[17:30] Abe♥: no
[17:30] hlrsecom: O_o 
[17:30] Abe♥: it shares the clock
[17:30] Abe♥: 4 cores use the same clock
[17:30] Abe♥: but each one can execute a separate process or a thread
[17:31] Abe♥: if you have 2 cores, you can excecute 2 parallel processes/threads
[17:31] hlrsecom: yeah? 
[17:31] Abe♥: if you have one core and still the same two process/threads, the core just switches back and forth
[17:31] Abe♥: for example;
[17:31] hlrsecom: lol yes...just like how Windows gets away with a fake multithreading refresh 
[17:32] Abe♥: my 8600m GT has 32 stream processors
[17:32] Abe♥: so it can excecute 32 parallel "threads", although they aren't considered threads on a GPU
[17:32] hlrsecom: ah 
[17:32] hlrsecom: so that's like "32 cores" if it were CPU 
[17:32] Abe♥: kinda
[17:32] Abe♥: but they are much simpler
[17:33] Abe♥: they are not entire cores by themselves
[17:33] Abe♥: it's kinda like the cell processor, it has the main arithmetic logic unit (I think) and 8 "processors" or "cores" that share it
[17:33] Abe♥: the gpu is similar
[17:34] Abe♥: I think you have the definition of "thread" all wrong (judging by your post)
[17:34] hlrsecom: lol. 
[17:34] Abe♥: a thread is just a sequential piece of code
[17:34] Abe♥: say, a loop
[17:34] Abe♥: that does something
[17:35] Abe♥: and a thread is part of a process
[17:35] Abe♥: a process is more like a program
[17:35] Abe♥: the process can spawn threads
[17:35] hlrsecom: each thread is like a command? 
[17:35] Abe♥: well, so can the threads, but a thread is just a light-weight version of a process really
[17:35] Abe♥: for example, yes
[17:35] Abe♥: or part of a command
[17:36] Abe♥: that has to run at the same time as something else
[17:36] hlrsecom: now what about calculating Pi 
[17:36] hlrsecom: is it possible to take advantage of multiple cores? 
[17:37] Abe♥: as far as I know, calculating pi is a sequential process 
[17:37] Abe♥: possibly
[17:37] hlrsecom: hmm 
[17:38] Abe♥: pi = integral(1/sqrt(1-x^2), dx, -1, 1)
[17:39] hlrsecom: That isn't simple though, when calculating? 
[17:40] Abe♥: sorry
[17:40] Abe♥: back
[17:40] Abe♥: I guess you can split up the integral
[17:40] Abe♥: or calculate it using some sort of series
[17:40] hlrsecom: I read somewhere, I think from a research paper on Pi, that it took 300 calculations just to find one digit of Pi accurately 
[17:40] Abe♥: um, no
[17:41] Abe♥: thats rediculou
[17:41] Abe♥: s
[17:41] hlrsecom: well it would make sense to me, but anyway :o 
[17:42] hlrsecom: Because Pi is infinite 
[17:42] hlrsecom: well 
[17:42] hlrsecom: the more I look into this 
[17:42] Abe♥: haha
[17:42] Abe♥: try this:
[17:42] hlrsecom: and research Pi 
[17:42] hlrsecom: I find out 
[17:42] hlrsecom: that because it is infinite 
[17:42] hlrsecom: you cant come up with a definite, absolute formula that would prove it all 
[17:42] hlrsecom: but rather a concept 
[17:42] Abe♥: 1-1/3+1/5-1/7+1/9.... = pi/4
[17:42] hlrsecom: using variables, etc 
[17:42] Abe♥: no, it's not infinite
[17:42] Abe♥: it's irrational
[17:42] hlrsecom: yeah, that's what I said 
[17:42] hlrsecom: or meant 
[17:42] hlrsecom: lol 
[17:43] hlrsecom: I didn't mean infinite, as in from 0 to 999999999999999999999999999999999 and so on 
[17:43] hlrsecom: but rather 
[17:43] hlrsecom: irrational X-P 
[17:43] hlrsecom: the digits go on forever 
[17:43] Abe♥: yeah
[17:43] hlrsecom: pi cant be made a fraction 
[17:43] hlrsecom: because fractions are absolute 
[17:43] Abe♥: using summation you can approximate pi;
[17:43] hlrsecom: approximate. 
[17:44] Abe♥: sum((-1)^n/(2n+1), n, 0, infinity) * 4
[17:44] hlrsecom: I have still yet to learn integrals and summations 
[17:44] hlrsecom: maybe we'll learn that in Calc this year :D 
[17:44] Abe♥: sum(x, 1, 3) = 1 + 2 + 3 = 6
[17:45] hlrsecom: x = (add everything from 1 to 3) 
[17:45] hlrsecom: something like that 
[17:45] hlrsecom: how do you represent infinity? 
[17:45] Abe♥: means; "the sum of x, as it goes from 1 to 3"
[17:45] Abe♥: using integer steps naturally
[17:45] Abe♥: no
[17:45] Abe♥: x is the variable that goes from 1 to 3
[17:45] Abe♥: so...
[17:46] Abe♥: this is a better example
[17:46] hlrsecom: 1 + (1+2) + (1+2+3) 
[17:46] Abe♥: sum(2*n, 1, 3)
[17:46] Abe♥: =
[17:46] Abe♥: 2*(1) + 2*(2) + 2*(3)
[17:46] hlrsecom: ah 
[17:46] Abe♥: = 12
[17:46] hlrsecom: now I understand that :D 
[17:47] hlrsecom: Now calculating Pi accurately is a different story, isn't it? :P 
[17:47] Abe♥: yeah....
[17:47] hlrsecom: That's what I'm meaning 
[17:47] hlrsecom: if you were to calculate pi accurately 
[17:47] Abe♥: but you can split up the sumations
[17:47] hlrsecom: you wouldnt be able to use multiple cores, right? 
[17:47] Abe♥: you'd need to write a program for it
[17:47] hlrsecom: or would you? D: 
[17:47] Abe♥: you could...
[17:47] hlrsecom: RJ says it wouldnt help speed things up 
[17:47] Abe♥: write a multithreaded python program (for example) to calculate it
[17:48] hlrsecom: [17:40] RJ♥: let me see if I can explain this with an example... say you want to calculate 40 digits of Pi... to achieve parallelism.. you need to split up the workload.. so if you have 4 CPUs.. 40 / 4 = 10 digits... so you could send 10 tasks (4 batches) to each CPU... but the problem with the 2nd batch is that they are dependent on the 10th task digit of the 1st batch (and it continues onward.. 3rd batch depends on the value of the 10th task of the 2nd batch)
[17:48] hlrsecom: and that to me makes sense 
[17:48] Abe♥: yes, unless you use the integration method
[17:48] hlrsecom: but maybe I dont understand summation altogether correctly :P 
[17:48] Abe♥: in which case, they dont' depend on each other
[17:48] Abe♥: actually, even in the summation, they don't matter
[17:49] Abe♥: superposition still applies
[17:49] Abe♥: sum(x, 1, 4) + sum(x, 5, 8)
[17:49] Abe♥: is the same as sum(x, 1, 8)
[17:50] Abe♥: except the latter can only be executed (in most cases) by one process, however the former can be executed by two at the same time, when they are done, add them together
[17:50] Abe♥: brb
[17:52] Abe♥: back
[17:53] hlrsecom: im impressed with you and the other two conversaitons im having 
[17:53] hlrsecom: I'm simultaneously conversing with you, RJ, and Snrrrub (soon to be officially a Google Engineer) 
[17:53] hlrsecom: I'm enjoying this too 
[17:53] Abe♥: what did they say?
[17:54] hlrsecom: LOL. 
[17:54] Abe♥: in the mean time, I'll be working on coding for my UAV
[17:54] hlrsecom: hahahaa 
[17:54] hlrsecom: lots of info 
[17:54] hlrsecom: I'll e-mail you the logs after I'm done with these lengthy conversations 
[17:54] hlrsecom: You'll probably enjoy Snrrrub's 
[17:55] Abe♥: did snrrrub say it's possible?
[17:55] Abe♥: use the summation method
[17:55] hlrsecom: actually 
[17:55] hlrsecom: with Snrrrub 
[17:55] hlrsecom: He first questioned my topic question 
[17:55] Abe♥: why don't we just get in a chat room
[17:55] hlrsecom: [17:26] Qwerty:  http://hardforum.com/showthread.php?p=1033021929#post1033021929 
 [17:27] Qwerty: You know how to program, so I was wondering if you have any input in regards to my thread. :) 
 [17:37] Snrrrub: I think it's very difficult to give a definitive answer to the question you pose (in your science fair project)
 [17:37] Qwerty: haha :P 
 [17:37] Snrrrub: I think the issue is the form of the question itself
 [17:38] Qwerty: I'm currently also talking to RJ (that other buddy of mine I told you about), and my cousin (also a programmer) 
 [17:38] Qwerty: and you ;o 
 [17:38] Snrrrub: "How much difference does a MHz make" - difference in terms of?
[17:55] Abe♥: your topic is iffy
[17:55] hlrsecom: lol. 
[17:55] hlrsecom: Snrrrub doesn't have AIM :\ 
[17:55] hlrsecom: let me check though 
[17:56] Abe♥: you can measure performance in other ways too
[17:56] Abe♥: and the increase from parallelism
[17:56] hlrsecom: He only has MSN 
[17:56] hlrsecom: do you have MSN? 
[17:56] Abe♥: yeah

Conversation with RJ
Code:
[17:26] hlrsecom:  http://hardforum.com/showthread.php?p=1033021929#post1033021929
[17:27] hlrsecom: You know how to program, so I was wondering if you have any input in regards to my thread. :)\ 
[17:30] RJ♥: I'll say this much... they do a good job of explaining stuff
[17:30] RJ♥: better than I could write :P
[17:30] hlrsecom: i still have a little confusion 
[17:30] hlrsecom: I have a quadcore CPU 
[17:31] hlrsecom: isnt each core just a therad 
[17:31] hlrsecom: thread* 
[17:31] hlrsecom: and all inside my 3400MHz spectrum (that's what my CPU is clocked at) 
[17:32] RJ♥: well if they're really 4 separate CPUs.. each one will have its own registers and stuff...
[17:33] hlrsecom: but then why isnt my CPU 4x as powerful 
[17:33] hlrsecom: like 
[17:33] hlrsecom: 3400 MHz * 4 
[17:33] RJ♥: all having 4 CPUs mean is you can do 4 things at a time... and that's if the threads are truly independent
[17:34] RJ♥: if one thread has to wait on another thread... well.. you just lost hte advantage of parallelism
[17:35] hlrsecom: then Pi 
[17:35] hlrsecom: it would be possible to take advantage of multiple cores? 
[17:36] RJ♥: if the value of Pi is to calculated linearly... no
[17:36] RJ♥: at least that's what I think by simple logic
[17:37] hlrsecom: So it could be done...? 
[17:40] RJ♥: let me see if I can explain this with an example... say you want to calculate 40 digits of Pi... to achieve parallelism.. you need to split up the workload.. so if you have 4 CPUs.. 40 / 4 = 10 digits... so you could send 10 tasks (4 batches) to each CPU... but the problem with the 2nd batch is that they are dependent on the 10th task digit of the 1st batch (and it continues onward.. 3rd batch depends on the value of the 10th task of the 2nd batch)
[17:40] hlrsecom: ah 
[17:40] hlrsecom: That makes sense 
[17:41] hlrsecom: It's very linear 
[17:41] hlrsecom: so then 
[17:41] RJ♥: so unless someone knows something I don't... calculating Pi takes Linear time
[17:41] hlrsecom: the other cores would be waiting on the first 
[17:41] hlrsecom: and the ones before each other 
[17:41] hlrsecom: Would it make a performance difference if you calculated Pi using 1 core vs 4 cores? 
[17:41] RJ♥: Linear time which means.. no
[17:41] RJ♥: Speed is all that matters
[17:42] RJ♥: the reason why video encoding can be done in parallel.. is b/c frames are splited up into different batches.. and each one will start with a reference frame...
[17:43] RJ♥: that doesn't rely on any resulting data from any other batch
[17:44] hlrsecom: ah 
[17:48] RJ♥: btw.. threads are a pain in the ass for me as a programmer
[17:48] hlrsecom: ok 
[17:48] hlrsecom: for pi 
[17:49] hlrsecom: what if you used an integration method 
[17:49] hlrsecom: copied and pasted that explanation to my cousin (programmer too) 
[17:49] hlrsecom: [17:48] Abe♥:  yes, unless you use the integration method
 [17:48] hlrsecom: but maybe I dont understand summation altogether correctly :P 
 [17:48] Abe♥: in which case, they dont' depend on each other
 [17:48] Abe♥: actually, even in the summation, they don't matter
 [17:49] Abe♥: superposition still applies
[17:49] hlrsecom: [17:49] Abe♥:  sum(x, 1, 4) + sum(x, 5, 8)
[17:49] RJ♥: well then.. there you go
[17:49] RJ♥: I knew someone would pick up on something that I couldn't think of
[17:50] hlrsecom: [17:49] Abe♥:  is the same as sum(x, 1, 8)
 [17:50] Abe♥: except the latter can only be executed (in most cases) by one process, however the former can be executed by two at the same time, when they are done, add them together
 [17:50] Abe♥: brb
[17:52] RJ♥: there's a reason why I never assumed/said that splitting the calculation of Pi isn't possible.. I knew someone would know something about it :P
 
Sorry for double post. I went way over the 20,000 character limit, which is why I am double posting. :S

Group Conversation with Snrrrub, Abe, and RJ
Code:
[17:58] *** xxxxxx@xxxxxx.com (Abe) has joined the conversation.
[17:58] Qwerty:  This is Abe (cousin)
[17:58] Qwerty:  He's also a programmer
[17:58] Qwerty:  RJ didn't have as interesting thoughts, but I'm trying to get him onto MSN too
[17:58] Snrrrub: Also, just FYI: if you're looking for a number to calculate, 'e' (natural logarithm) might be a little easier.
[17:58] Qwerty:  Snrrrub, so far you've had the most interesting O_O
[17:59] Qwerty:  Abe and RJ -- we mostly talked about Pi, multicore, multithreading, etc
[17:59] *** xxxxxx@xxxxxx.com (RJ) has joined the conversation.
[17:59] Qwerty:  Here's RJ :P
[17:59] Snrrrub: Good evening
[17:59] Abe : why is e easier?
[17:59] Qwerty:  All three of you are aware of http://hardforum.com/showthread.php?p=1033021929#post1033021929
[17:59] Abe : *sorry, not so familiar with the approximations of e
[18:00] Snrrrub: The Taylor polynomial for e is extremely simple
[18:00] Abe : The leibniz formula for pi seems to be fairly simple as well, it seems like it could be split up nicely
[18:01] RJ : I'll be here when I get caught up with who's on my list and whatnot (I haven't signed into MSN in a year)
[18:01] Qwerty:  SuperPi is like the only Pi program out there
[18:01] Qwerty:  there was another one, but that one was unstable. It did calculations for square root of 2, e, and pi
[18:01] Qwerty:  Would 64-bit have any impact on calculations?
[18:01] Abe : ah, may as well, write your own
[18:01] Snrrrub: I think Fabrice Bellard has the best algorithm for PI at the moment.
[18:02] Snrrrub: http://bellard.org/pi/
[18:02] Qwerty:  The machine configuration I'm wanting to run this experiment on (and at the moment, based on indiivdual conversations with all three of you, may have to rework/modify the topic question or how this experiment will perform)
[18:02] Qwerty:  Has 4GB of RAM, although 3.5GB might be allocated for use...
[18:03] Qwerty:  Q9450 Yorkfield Quadcore w/ 12MB Cache
[18:03] Snrrrub: You shouldn't need much RAM
[18:03] Qwerty:  And a totally stripped version of XP
[18:03] Snrrrub: And if you do, you're hosed ;)
[18:03] Qwerty:  I'll probably take out 2GB
[18:03] Qwerty:  just so that it remains constant
[18:03] Abe : do you know anything about the computation of pi with fft multiplication?
[18:03] Qwerty:  32-bit version of XP doesn't fully allocate the 4GB
[18:03] Qwerty:  Although the system does infact use 4GB
[18:04] Abe : 32-bit xp has some trouble with that, it can only see about 3.5 gigs
[18:04] Snrrrub: You should enable PAE
[18:05] Qwerty:  From what I read, the system first reserves a certain amount of your RAM based on your hardware configuration, especially the kind of video card you're using
[18:05] Snrrrub: No, I don't know much about Pi using the FFT
[18:05] Qwerty:  And then it will leave the remaining to the user to use; 3.5GB in my case. When I had my ATI 4850, it was only leaving me with about 3.2GB
[18:05] Abe : Qwerty, do you just want to demonstrate the performance gain of multi-cores, or does the calculation matter to you?
[18:05] Qwerty:  At the moment I'm using a temporary replacement NVIDIA FX5200
[18:05] Qwerty:  the multicores do not matter
[18:06] Qwerty:  it would get overly complicated at that point :S
[18:06] Abe : oh, so the calculation does?
[18:06] Qwerty:  What I wanted to do
[18:06] Qwerty:  Was to see the performance gain or loss per MHz of CPU Clock
[18:06] Qwerty:  Snrrrub knows a lot about this...
[18:06] Qwerty:  and if you get really technical
[18:06] Qwerty:  it's extremely difficult to accurately ...
[18:06] Abe : I guess I don't understand
[18:06] Qwerty:  uh
[18:07] Qwerty:  to accurately show how much difference in performance does 1 MHz make
[18:07] Abe : oooh
[18:07] Qwerty:  In my experiment, the original intention was to begin at 800 MHz
[18:07] Qwerty:  and work my way up to 3000 MHz at 8 MHz intervals
[18:07] Abe : can you scale the frequency of your cores?
[18:07] Qwerty:  and calculate Pi using SuperPi, which would also tell me how long it took to calculate to 32 million decimal places in milliseconds
[18:08] Qwerty:  the MHz of each core?
[18:08] Abe : well, you can't change the clock for one at a time, they use the same source
[18:08] Qwerty:  yeah
[18:09] Qwerty:  Do you mean like dynamically adjusting my clock using software?
[18:09] Qwerty:  My CPU clock, multiplier, FSB, and everything can be adjusted through a program Gigabyte sent with my motherboard
[18:09] Abe : yeah, but I don't know how to do that
[18:09] Qwerty:  Could do it through Windows
[18:09] Abe : oh, convenient 
[18:09] Qwerty:  The only concern I have
[18:09] Qwerty:  is stability
[18:09] Qwerty:  Last time I tried it
[18:10] Qwerty:  I don't know if it was the board's overclocking enhancer being enabled or not
[18:10] Qwerty:  but it made everything unstable quickly
[18:10] Qwerty:  It might have been the enhancer. I'm supposed to turn off all enhancers before manually overclocking myself
[18:10] Qwerty:  Might be fine now.
[18:10] Snrrrub: Keep in mind that SuperPI is single threaded so the other cores shouldn't matter much (unless the OS migrates your thread to another core)
[18:11] Qwerty:  That would slow the calculation down if the OS did that?
[18:11] Snrrrub: Yup
[18:11] Qwerty:  oh wait
[18:11] Qwerty:  yeah
[18:11] Snrrrub: Threads are usually given a processor affinity but that doesn't guarantee that the thread will run on that specific core until the end of the computation
[18:11] Qwerty:  Is it possible to accurately perform this experiment, Snrrrub? How might you setup the experiment?
[18:12] Qwerty:  I know I would have to make sure my RAM's clocks and timings would be enforced, as well as many other things on my motherboard, such as the CPU and RAM FSB
[18:13] Abe : that may be unavoidable, e.g. my cores seem to juggle a single thread apps back and forth, possibly to maintain even heat dissipation?
[18:13] Snrrrub: Yup
[18:13] Qwerty:  oh yeah
[18:13] Qwerty:  gotta be
[18:13] Qwerty:  I've noticed that sometimes two or three of my cores are significantly hotter than the other
[18:14] Qwerty:  not sure why it usually leaves 1 of the cores just out hanging (and it isnt the first core, or Core 0)
[18:15] Snrrrub: To be honest, to get an accurate idea of what's going on, I think it would be best to run a CPU-only computation on a "bare bones OS" - one that doesn't support multiprocessing.
[18:15] Snrrrub: DOS would be a good example. :)
[18:15] Abe : hehe, freedos may still be available
[18:16] Qwerty:  :D
[18:16] Qwerty:  I don't quite know of any program to calculate Pi or e for DOS though, so I would have to find that somehow
[18:17] Abe : hehe, maybe you should reconsider the goal of the project?
[18:17] Qwerty:  >.<
[18:18] Snrrrub: I think you'll find the results to be "obvious" in many ways.
[18:18] Qwerty:  I'm out of ideas of what to experiment or test, that I could easily do but something nobody has really done or isnt of public/common knowledge
[18:18] Qwerty:  ryan_975 replied to my thread on HF
[18:18] Qwerty:  "Wrong. X2, X3, X4, Exxxx, and Qxxxx chips have multiple complete CPU cores on the package. That includes FPU, ALU, logic unit, bus control unit, prefetch, branch prediction, L1 caches, registers, etc. Everything is replicated. 
  
 Core is a complete CPU. 
 "CPUs" tends to mean multi-socket system. 
 A thread is a stream of instructions in a program that need to be processed. 
  
[18:18] Qwerty:  For all intents and purposes there is very little difference between multi-core and multi-CPU system (there are some major difference at the hardware and OS level, but they're beyond the scope of this thread)."
[18:18] Qwerty:  That's something new.
[18:18] Qwerty:  to me atleast
[18:18] Qwerty:  Either he has something confused too
[18:18] Qwerty:  Or something I'm still confused about and didn't know
[18:19] Qwerty:  Remember, I have a Q9450
[18:19] Snrrrub: No, that's accurate. Hyperthreaded processors are a totally different animal though.
[18:19] Qwerty:  wow, now I'm even more confused :D
[18:19] Qwerty:  CPUs aren't really as I thought them to be :\
[18:20] Snrrrub: Yes, they're getting pretty interesting these days. ;)
[18:20] Qwerty:  Seeing as how this is topic question (MHz making a difference) is ultra technical
[18:21] Qwerty:  I could do two or three things; I could continue the experiment, but would have its fallacies (which I would have to mention in a journal), manipulate the topic question to something that is more doable and legitimate, or get ideas for a different topic question
[18:22] Qwerty:  Last year I did my science fair experiment on comparing three different cooling methods: active and passive air cooling, and water cooling
[18:22] Qwerty:  I was tied with 1st place, so they put me as 2nd
[18:22] Qwerty:  Of course, if you got really technical about that experiment
[18:22] Qwerty:  :P
[18:22] Abe : make a pair of plasma speakers, that'll get you 1st
[18:23] Snrrrub: I think it would be really cool if you looked into the impact of teaching functional languages vs. imperative languages in undergraduate cirricula ;)
[18:23] Snrrrub: But that's just me.
[18:23] Abe : lol
[18:23] Qwerty:  Haha, well it has to be something I can do and learn within 4 to 6 minutes of time
[18:23] Qwerty:  months*
[18:24] Qwerty:  One other idea I had
[18:24] Qwerty:  was comparing CPUs and GPUs, despite the obvious
[18:24] Qwerty:  again, in math calculation. Either that or rendering (which might be more difficult or controversial in the setup)
[18:25] Qwerty:  If I did it in comparison of CPU to GPU, I'd have to figure out how to offload mathematical equations to my ATI card
[18:25] Snrrrub: Or how they started off in different areas entirely and are now meeting in the middle
[18:25] Qwerty:  Explain D:
[18:26] Snrrrub: http://libsh.org/
[18:26] Qwerty:  In the experiment, I have to apply the scientific method and get results, and report the results in the form of some sort of illustration or chart
[18:28] Qwerty:  hm
[18:28] Qwerty:  my mind is boggled...as in not sure what to do
[18:29] Qwerty:  I was pretty set with the MHz experiment, but I can see the main fallacies in that
[18:29] Qwerty:  what exactly is an ALU (arithmetic logic unit)
[18:29] Qwerty:  or does it do
[18:30] Snrrrub: Performs operations like addition, subtraction, multiplication, division, and some bit operations. Everything it does is in the domain of integers exclusively
[18:31] Qwerty:  And what do FPUs do?
[18:31] Snrrrub: Same deal except with real numbers
[18:32] Snrrrub: (minus the bit operations)
[18:32] Qwerty:  I might just continue with the original intention of finding the difference per MHz
[18:32] Qwerty:  But would have to clarify the fallacies, as I said earlier
[18:33] Qwerty:  The cooling experiment I did -- could the same fallacies similarly apply?
[18:33] Qwerty:  The CPU was overclocked to 2GHz throughout the whole experiment I think
[18:33] Qwerty:  (i think!)
[18:34] Snrrrub: I think there are fewer factors that affect a cooling system
Code:
[19:05] Qwerty:  Hmm
[19:05] Qwerty:  I have a question regarding the MHz thing
[19:05] Qwerty:  3DMark
[19:05] Qwerty:  The CPU scores...is that even accurate ?
[19:05] Qwerty:  Based on what you told me in the individual conversation, I'm starting to think that the CPU scores in 3DMark are balony
[19:05] Qwerty:  not totally though
[19:06] *** Abe has left the conversation.
[19:06] Snrrrub: Some of those benchmark suites are pretty decent but they're specific to the workloads so it doesn't mean ALL that much.
[19:07] Snrrrub: They're good for a relative measure if you know what exactly those numbers are testing. :)
[19:07] Qwerty:  I might be able to use that in my experiment next to Pi then
[19:08] Snrrrub: Yeah, choose a CPU-intensive benchmark suite.

The purpose of publishing these logs is for public knowledge, and anyone that is interested in an experiment like this -- like another high school kid like me. ;P
 
I have a Q9450 clocked at 3400MHz (3.4GHz); it would still only have 1 FPU, not 4, because the "4 cores" are really just threads, and not actual processors/CPUs?

I need the following terms cleared up:
- core(s)
- CPUs (plural)
- thread(s)

Core - traditionally meant memory (based on magnetic cores - donut shaped magnetic rings that were threaded onto a mesh of wires). Now it is being used to refer to a single processor on a multi-processor chip as in Core 2 Duo (two cores or two CPUs on one chip).

CPUs - more than one CPU. There's a bit of ambiguity in the term's use in that some make CPU synonymous with chip (correct only when there's one CPU/core per chip), others use it to mean a processor and others (newbies) the entire case (box, MB, CPU, disks etc). There's a bit more confusion since when we say x CPUs per "chip" , it may really mean per package - some quads consist of two chips in one package with two processors per chip. Others are 4 processors on one chip in one package.

CPU = processor = core for most intents and purposes.

Thread - a set of instructions run on a CPU as a single entity. It can be one program or a part of a program. If a program runs on a multi-CPU chip, then the program can split itself into multiple threads and run each one on a separate processor. So on a quad core computer, a program can run four parallel tasks. However, this must be done by the programmer - there are few tools available so far for automatically splitting a program into threads. Since threads can take different lengths of time to execute, a program may split itself into more threads than processors. For example, a program can split into, say, ten threads. The OS will schedule these on the processors according to various priority schemes - some must be finished before another can start, for example. That can result in an uneven distribution of work on the processors on the chip. A programmer will try to write the program so that the work is balanced as much as possible.

So - your Q9450 has four CPUs = four cores = four processors. Each one has one FPU so there are four in total. An E6600 (dual core) would have two CPUs and two FPUs. However, only one CPU can use its FPU - the FPU cannot be shared with the other CPUs on the chip/package. So, on a quad core, you can't have three CPUs running non-numeric threads with one CPU running four FPU calculations at once - it's strictly one-to-one. Only the Cell computer can do this sort of thing, since it has one general CPU and up to 8 SPEs.

I don't know pi algorithms offhand. I don't know if they are easy to split into multiple threads. If it's possible to calculate pi by doing many partial calculations in parallel and then combining the result, then you could run up to four threads simultaneously doing calcs on a quad core.

If pi cannot be calculated with multiple threads or if the programmer has written it a single-threaded program, then there will be little performance difference whether it's run on a single core, dual core or quad core computer.
 
What exactly is the difference between an FPU and ALU?

And what does an SPE do?
 
What exactly is the difference between an FPU and ALU?

And what does an SPE do?

ALU (Arithmetic and Logic Unit) deals only with integers and bit functions

FPU (Floating Point Unit) deal only with floating points (0.2 0.33234 0.0000000003242, etc)

SPE is what Sony/IBM is calling the slave processors in their Cell CPU.
 
Back
Top