• A Great friend to the HardForum with a great kid that he is trying to get a scholorship to continue his schooling. Please give hime a vote! Only 24 hours left! Thanks.
    If you have an VOTE FOR KEENAN!

From SuperComputer to GPU in 10 years.

PrincessFrosty

Supreme [H]ardness
Joined
May 6, 2009
Messages
5,905
Just looking through the super computer timeline on wikipedia and noticed that in the year 2000 the fastest super computer was the IBM ASCI White at 7.226 TFlops.

After watching the AMD video presentation on the 5970 they're saying its just a hair off 5 TFlops, overclocked from 725Mhz core to 1000Mhz ought to net it close to 7 TFlops at an estimate.

Thats from SuperComputer to GPU in 10 years :D

Just think the Deep Blue supercomputer to beat the worlds chess champion was 11.38 GFlops, thats 0.01136 TFlops, do the maths and a stock speed 5970 could basically play 440 world class chess champions at the same time.
 
Not quite the same thing. GPUs are high performance in a very narrow set of problems. Now those problems happen to be the sort of thing that graphics processing is, hence they are good at it. However they are not the same as real supercomputers, or even CPUs. I mean if a GPU was just a CPU but faster, well we'd just make GPUs in to CPUs. They aren't though, there is a subset of things they are good at.
 
Well yes of course, I don't think that devalues the comparison though, thats what supercomputers are built for. GPU's compared CPUs there's simply a lot more hardware inside modern GPUs than modern CPUs probably something along the lines of 3-4x the transistor count from what I remember reading.
 
Well yes of course, I don't think that devalues the comparison though, thats what supercomputers are built for. GPU's compared CPUs there's simply a lot more hardware inside modern GPUs than modern CPUs probably something along the lines of 3-4x the transistor count from what I remember reading.

That is what Nvidia is trying to do. GPUs excel at parallel processing - sometimes over 20 times faster than the fastest CPU.

Fermi has over 3 billion transistors and it will be featured in supercomputers alongside CPUs.
 
Deep Blue supercomputer to beat the worlds chess champion was 11.38 GFlops, thats 0.01136 TFlops


I saw a documentary that showed that the worlds chess champion was not only playing against the computer. He was also playing against the 2 people moving and observing. Also a chess master and a at the computer operator through a mic/camera communication system. So technically he was up against 1 cpu + 4 humans.

The chess master still manage to win in the rematch. lol.
 
What's funny to me is how a machined hunk of metal becomes more valuable than the state-of-the-art CPU it's sitting on in a shorter period of time.

At this accelerating pace of available processing power, which trickles down to other areas, I fear the day my coffee mug can do the job without me.
 
What's funny to me is how a machined hunk of metal becomes more valuable than the state-of-the-art CPU it's sitting on in a shorter period of time.

At this accelerating pace of available processing power, which trickles down to other areas, I fear the day my coffee mug can do the job without me.
No worries, humanity is defined by its creation and utilization of tools. Rather than replace you, its more likely it'll become part of you. I'd personally welcome a chip implanted in my brain that could give me a photographic memory and enhance my math and physics skills heheh!
 
Great, then I fear the day I have to worry about erectile software corruption.
 
I saw a documentary that showed that the worlds chess champion was not only playing against the computer. He was also playing against the 2 people moving and observing. Also a chess master and a at the computer operator through a mic/camera communication system. So technically he was up against 1 cpu + 4 humans.

The chess master still manage to win in the rematch. lol.

According to Wikipedia Kasparov won the first game 4-2 and lost the rematch 3.5 to 2.5

The team weren't allowed to interfere or intercept moves.

The rules provided for the developers to modify the program between games, an opportunity they said they used to shore up weaknesses in the computer's play that were revealed during the course of the match. This allowed the computer to avoid a trap in the final game that it had fallen for twice before.
 
The team weren't allowed to interfere or intercept moves.

The rules provided for the developers to modify the program between games, an opportunity they said they used to shore up weaknesses in the computer's play that were revealed during the course of the match. This allowed the computer to avoid a trap in the final game that it had fallen for twice before.

You just defeated your own statement...

Modifying the program to help it learn is the ULTIMATE interference. If you don't believe so, you have a VERY bad definition of 'interference.' It was allowed by the rules, sure, but it is still interference. If a human does not learn how to not fall for that trap, but a computer can be reprogrammed to notice it, it can't really be fair.

The point of the thread, though... yeah, GPUs are that fast and growing at a very high rate, but its mainly the parallelism, not so much the general computing like regular CPUs. However comparing Power7 to those older supercomputing clusters IS a more valid comparison :)
 
No I didn't. I said they weren't allowed to interfere or intercept MOVES

The only tinkering they could do to the AI was inbetween full games where the change had no effect mid-game. This was allowed in the rules specifically so they could find flaws and correct it.

If Kasparov is allowed to "learn" during the match and alter his habits to beat a specific opponent, why isn't the computer?
 
No I didn't. I said they weren't allowed to interfere or intercept MOVES

The only tinkering they could do to the AI was inbetween full games where the change had no effect mid-game. This was allowed in the rules specifically so they could find flaws and correct it.

If Kasparov is allowed to "learn" during the match and alter his habits to beat a specific opponent, why isn't the computer?

Because there are multiple brains analyzing the movements and adjusting the machines movements in reaction to whatever kasparov does. That is directly influencing the machines moves, and done on a people vs person basis. Whether it is done per-match or per-movement really doesn't matter. That's just like me saying Kasparov having a team of people analyzing the matches and telling him "this is what the machine is most likely to do" in between matches is not influencing Kasparovs moves.
 
Because there are multiple brains analyzing the movements and adjusting the machines movements in reaction to whatever kasparov does. That is directly influencing the machines moves, and done on a people vs person basis. Whether it is done per-match or per-movement really doesn't matter. That's just like me saying Kasparov having a team of people analyzing the matches and telling him "this is what the machine is most likely to do" in between matches is not influencing Kasparovs moves.

Yeah...
 
Because there are multiple brains analyzing the movements and adjusting the machines movements in reaction to whatever kasparov does. That is directly influencing the machines moves, and done on a people vs person basis. Whether it is done per-match or per-movement really doesn't matter. That's just like me saying Kasparov having a team of people analyzing the matches and telling him "this is what the machine is most likely to do" in between matches is not influencing Kasparovs moves.

QFT. Gotta take in the whole picture here.
 
Its like the gpu is good at graphing while the cpu is good at algebra type thing. Different worlds :)
 
Fixed. Oh, and most super computers are multithreaded.

Not really. GPUs are good at highly parallel *independent* tasks. A super computer will have far more available bandwidth and cache space and can process far larger data sets much faster than a GPU. Same with CPUs. Look at modern multi-core CPUs. They have large caches that are shared between them. Now look at a GPU - it doesn't. There are caches that allow sharing in smaller units (the SM in Nvidia's architecture, for example), but outside of that requires hitting VRAM. Also, the cache space available to the GPU is very small. I would bet that a supercomputer from 10 years ago will walk all over a GPU at doing the work supercomputers are designed for. FP performance is a very, very small part of the overall performance of the system.

Also, GPUs are really only good at *SINGLE PRECISION* floating point. Performance tanks if you use double precision (hence why that is a large focus for Fermi).
 
Because there are multiple brains analyzing the movements and adjusting the machines movements in reaction to whatever kasparov does. That is directly influencing the machines moves, and done on a people vs person basis. Whether it is done per-match or per-movement really doesn't matter. That's just like me saying Kasparov having a team of people analyzing the matches and telling him "this is what the machine is most likely to do" in between matches is not influencing Kasparovs moves.

I'm not sure your argument makes any sense. I understand what you're saying, but it implies that there's some kind of artificially intelligent decision-making going on. Of course it's people vs person. The programmers were "influencing" the machine's moves the moment they started writing the chess program. The point of the exercise was to see if the program (running on incredibly powerful hardware) could beat Kasparov and makes this

"The rules provided for the developers to modify the program between games, an opportunity they said they used to shore up weaknesses in the computer's play that were revealed during the course of the match."

completely understandable.
 
Not really. GPUs are good at highly parallel *independent* tasks. A super computer will have far more available bandwidth and cache space and can process far larger data sets much faster than a GPU. Same with CPUs. Look at modern multi-core CPUs. They have large caches that are shared between them. Now look at a GPU - it doesn't. There are caches that allow sharing in smaller units (the SM in Nvidia's architecture, for example), but outside of that requires hitting VRAM. Also, the cache space available to the GPU is very small. I would bet that a supercomputer from 10 years ago will walk all over a GPU at doing the work supercomputers are designed for. FP performance is a very, very small part of the overall performance of the system.

Also, GPUs are really only good at *SINGLE PRECISION* floating point. Performance tanks if you use double precision (hence why that is a large focus for Fermi).
I'm speaking in terms of Femri. The cache sizes, DP, it's all being moved to destroy the supercomputer market.
 
If my argument doesn't make sense, you need to re-think that English or mathematics degree :) There really isn't a flaw in my argument, it's bare-metal logic based on the most simplified principles involved. It's not that the rules didn't allow for it, and there's no debate that it was within or against the rules, it's the principle that each side's moves weren't influenced when one side clearly was influenced in between matches. If the wording was changed around to be more precise, the whole 'debate' there wouldn't have taken place. LOL, I'm a nitpicker, and my job makes me do it out of habit now. No offense meant, unlike whoever it was that tried to tell me to QFT... how rude.

Fermi is an interesting subject here. It's going to be able to do most of the tasks a CPU can perform, so I'm pretty excited to see how well it will do these. So far the vibe I'm getting is that it's going to be between your standard x86 CPU and a GPU in terms of its capabilities. It won't be an all-out CPU replacement, but Fermi should be in a good position to where an OS can be written and compiled to run on it, most likely run it fast, and with a good amount of eye candy. While I understand what Nvidia is trying to do with Fermi, if it doesn't really impress and the new features aren't utilized enough, this whole bit about GPUs becoming more like CPUs in their functionality and purpose will become fodder. Something like this needs adoption. Its not as cheap to implement for a niche user base, whereas physx is really just software so even though Ageia cost Nvidia some cash, physx is practically free to implement in new cards.
 
Computers better than humans at chess? Of course

Computers better than humans at Go.... not even close.


Interesting. It makes me believe that there will be a huge jump in technology the day we understand how our own brain actually works. And of course once we understand how our body creates body parts and even life forms will also cause a great jump in technology.
 
Fermi is an interesting subject here. It's going to be able to do most of the tasks a CPU can perform, so I'm pretty excited to see how well it will do these. So far the vibe I'm getting is that it's going to be between your standard x86 CPU and a GPU in terms of its capabilities. It won't be an all-out CPU replacement, but Fermi should be in a good position to where an OS can be written and compiled to run on it, most likely run it fast, and with a good amount of eye candy. While I understand what Nvidia is trying to do with Fermi, if it doesn't really impress and the new features aren't utilized enough, this whole bit about GPUs becoming more like CPUs in their functionality and purpose will become fodder. Something like this needs adoption. Its not as cheap to implement for a niche user base, whereas physx is really just software so even though Ageia cost Nvidia some cash, physx is practically free to implement in new cards.

Femri is not made to replace a CPU for the OS, nor should it be. A 100$ CPU can run the OS and most programs just fine. The difference is when you have multi-thread friendly problems like folding at home, decryption, FEA/FEM problems, etc.

GPUs aren't becoming more like CPUs in any sense of the word. They are going down the path of more cores, not higher clock speeds. The big change with Femri is going from a high single precision capabilities with small amounts of double precision to a mix of single precision and double precision. For computing needs, double precision is required.
 
Interesting. It makes me believe that there will be a huge jump in technology the day we understand how our own brain actually works. And of course once we understand how our body creates body parts and even life forms will also cause a great jump in technology.

Yeah, chess and go are very different problems. The simplest way to explain it is like this.

The first move on go has 361 moves (on a standard 19x19 board). The second move has 360 possible moves. The third move has 359 possible moves. That gives the first 3 moves 46655640 moves (removing the "symetric" moves, it is "only" 5831955, the 4th move will still yield 358 possible unique locations).

While chess has 64 squares, there are only 12 opening moves. There are 12 responses. There are 14 second moves for white. This gives a total moves of 2016 unique openings through three moves.

The really sick part about Go is the top level players can be looking at well over 50, 60, or even 100 moves in some cases ahead.

It's a great game and a game I am NOT very good at. However I still have a great appreciation for it. Especially as it is such a simple game.
 
The big change with Femri is going from a high single precision capabilities with small amounts of double precision to a mix of single precision and double precision. For computing needs, double precision is required.

That's hardly the big change. The caching system is, along with how the computing units work, compared to previous architectures. It's still a scalar architecture, much like G80 and GT200, but there's quite a bit more under the hood.
 
Femri is not made to replace a CPU for the OS, nor should it be. A 100$ CPU can run the OS and most programs just fine. The difference is when you have multi-thread friendly problems like folding at home, decryption, FEA/FEM problems, etc.

Are you saying that it would not be freaking awesome to see it run linux?
Why? Just because it can.
 
That's hardly the big change. The caching system is, along with how the computing units work, compared to previous architectures. It's still a scalar architecture, much like G80 and GT200, but there's quite a bit more under the hood.
Yes, the cache is a huge change, one that was necessary to feed the DP capabilities. One that is worthless without the 2 order of magnitude improvemnt in DP. Forgive me if when trying to explain to someone who does not appear to be "up to speed" or architectually proficent and I use simple terms. :rolleyes:

Are you saying that it would not be freaking awesome to see it run linux?
Why? Just because it can.
Oh, it would be cool, no question there. It's cool to see linux run on a PS3 or a gameboy. Not necessarily useful, but neat! :D
 
Yeah, I'm basically saying that I could see Fermi being the definitive first move into GPU-based OSes, but it is definitely not yet meant to do it alone. My expectation is that if Fermi's extra capabilities do take off with developers, we would most likely see an Nvidia-powered netbook running strictly off a modified or next generaton Fermi-type GPU. It Isn't going to happen in the next year, but it is definitely not out of the realm of possibility that an attempt will be made within the next 3 years.
 
Yeah, I'm basically saying that I could see Fermi being the definitive first move into GPU-based OSes, but it is definitely not yet meant to do it alone. My expectation is that if Fermi's extra capabilities do take off with developers, we would most likely see an Nvidia-powered netbook running strictly off a modified or next generaton Fermi-type GPU. It Isn't going to happen in the next year, but it is definitely not out of the realm of possibility that an attempt will be made within the next 3 years.

No, it isn't. Its nowhere near possible. x86 CPUs do *tons* of work and housekeeping for the OS. Virtual memory, pre-emptive multithreading, data/code seperation, paging, etc... These are all features made possible by the *CPU*. GPUs (even Fermi) simply lack the most important parts of a CPU, because they are designed to be managed by a driver. Also, they suck at integer work - and guess what CPUs primarily do? Integer work. OS's don't give two shits about floating point performance, it doesn't help. I don't even know if GPUs have any real form of internal exception handling (CPUs generate an interrupt if you divide by zero or access unmapped memory so that the OS can gracefully handle it).

Now, that said, *COULD* an OS be written for a GPU? Sure, it just wouldn't be usable or do anything. It would be about as advanced and useful as a corrupt install of MS-DOS 1.0.
 
No, it isn't. Its nowhere near possible. x86 CPUs do *tons* of work and housekeeping for the OS. Virtual memory, pre-emptive multithreading, data/code seperation, paging, etc... These are all features made possible by the *CPU*. GPUs (even Fermi) simply lack the most important parts of a CPU, because they are designed to be managed by a driver. Also, they suck at integer work - and guess what CPUs primarily do? Integer work. OS's don't give two shits about floating point performance, it doesn't help. I don't even know if GPUs have any real form of internal exception handling (CPUs generate an interrupt if you divide by zero or access unmapped memory so that the OS can gracefully handle it).

Now, that said, *COULD* an OS be written for a GPU? Sure, it just wouldn't be usable or do anything. It would be about as advanced and useful as a corrupt install of MS-DOS 1.0.

So you're saying that even though Fermi will be capable of running things written in multiple languages (CUDA, C++, Python, Java, .NET, Fortran), it would be absolutely useless as the single chip running a system. I'm not disputing that (too much), but what I am saying is Fermi is the first major step toward Nvidia putting a more solid feature set in place that would make it perfectly suitable for mobile devices that are more general-purpose like a netbook. If you are really disputing that when we can already render webpages, accelerate video encoding/decoding, and do virus scanning with a GPU, then I think you might be a bit out of whack. Fermi is quite literally one engineering step away from being given the ability to do basic CPU tasks and offload the rest of the operation to regular stream processors. I've given it 3 years before we start hearing about some tard giving this a serious go, that's definitely more than enough development time for Nvidia and recently that seems to be where they want to go is running systems off a monolithic Nvidia core of some kind.
 
So you're saying that even though Fermi will be capable of running things written in multiple languages (CUDA, C++, Python, Java, .NET, Fortran), it would be absolutely useless as the single chip running a system. I'm not disputing that (too much), but what I am saying is Fermi is the first major step toward Nvidia putting a more solid feature set in place that would make it perfectly suitable for mobile devices that are more general-purpose like a netbook. If you are really disputing that when we can already render webpages, accelerate video encoding/decoding, and do virus scanning with a GPU, then I think you might be a bit out of whack. Fermi is quite literally one engineering step away from being given the ability to do basic CPU tasks and offload the rest of the operation to regular stream processors. I've given it 3 years before we start hearing about some tard giving this a serious go, that's definitely more than enough development time for Nvidia and recently that seems to be where they want to go is running systems off a monolithic Nvidia core of some kind.

There is such a thing as a GPU render farm. They still use CPUs .. 1 CPU core to 4 GPUs, if i remember what i learned at Nvidia's GTC. If you like, i can look up my notes
- the idea is to use the GPU as a *co-processor* with the CPU.
 
Back
Top