• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Vega countdown has appeared

So all we got were some new videos? This is as bad as the Nvidia presentation from last night. Well the poor CEO of Nvidia put my gaming buddy that bought a GTX 1080 to sleep. I don't know which is worse. He did chirp up when I showed him a $1,300 GSYNC monitor. Then he thought about spending $1,300 on a monitor and called it a night. :)


https://www.youtube.com/user/amd/videos
 
Courtesy of guru3d.com:

index.php


The photo shows a working VEGA setup in a demo room running Doom at ultra quality settings in an Ultra HD resolution, this was a demo room at the event. The photo comes from Golem.de and as explained is Doom running at 4K Ultra. The card is outing an FPS of 68 fps, with the "Vulkan" API enabled and what probably is a very early unoptimized engineering sample. That roughly is indicative for the performance of a GeForce GTX 1080.

I really hope that's not the performance it ends up delivering. IMHO the big chip NEEDS to match or rather slightly exceed the TXP, otherwise it will be a disappointment, given the fact that they'll be somewhere between 10 and 12 months late.
 
It was a Vega ES faster than a stock 1080 back then.. it's nearly 3/4 of a year old.. this is the point I'm trying to make. Thank you. It will depend on the title but it seems it should be ~1080 speed or better depending. Maybe not closer to Pitan but around that spread as drivers mature and developers come onboard.
You miss the point. The AotS benchmark in article you reference IS A MONTH OLD. And the article you talk about is month old (exactly month, actually) old as well. And obviously it has to be, because Vega's milestone was only hit in late June, not May.

Anyways, the good point i suppose is that early ES of Polaris 10 were easily 30% slower than release version due to clock problems on first stepping.
 
Courtesy of guru3d.com:

index.php


The photo shows a working VEGA setup in a demo room running Doom at ultra quality settings in an Ultra HD resolution, this was a demo room at the event. The photo comes from Golem.de and as explained is Doom running at 4K Ultra. The card is outing an FPS of 68 fps, with the "Vulkan" API enabled and what probably is a very early unoptimized engineering sample. That roughly is indicative for the performance of a GeForce GTX 1080.

I really hope that's not the performance it ends up delivering. IMHO the big chip NEEDS to match or rather slightly exceed the TXP, otherwise it will be a disappointment, given the fact that they'll be somewhere between 10 and 12 months late.


In one of their videos, they say Radeon is back, I think we can safely assume right now its going to be similar to Pascal, not sure about gp102 but at least the 1080. Now that isn't that impressive, but still, a good leap up from what they had before, which was nothing lol. Lets see at what power consumption its at and then we can really say if they are back or not.
 
There's no way they're going to get TXP performance. That gap is just too big, but if they can deliver 1080 performance for a price near $450 or so, they could have a winner. It's really can't be more expensive than that since everyone expects Nvidia to cut 1080 prices to $499 if Vega is competitive. If both cards were the same price, people will just keep buying Nvidia.
 
What little was released seems interesting. We may in fact be looking at a GP100 competitor and the size may be misleading.
Yep.
And fits with exactly what I was saying a little while ago that Vega would be a large die and a cut core model would be for consumer and the full core for HPC, which IMO is big Vega and as I mentioned at the time same as what Nvidia did with GP102.
Also within the size some of us (including you) discussed.

Cheers
Difference being it may be positioned as a G/P100 competitor just based on the few specs we have so far. FP16 is in theory already ahead of P100. I'm wondering if effective FP16 and FP32 performance are identical with the fused instructions. That would take some serious compiler work to get working. They also mentioned Vega will be removing all resource management work from the developer. So it explains the high bandwidth cache. Could be a reference to added storage beyond HBM on the GPUs.
 
  • Like
Reactions: N4CR
like this
You miss the point. The AotS benchmark in article you reference IS A MONTH OLD. And the article you talk about is month old (exactly month, actually) old as well. And obviously it has to be, because Vega's milestone was only hit in late June, not May.

Anyways, the good point i suppose is that early ES of Polaris 10 were easily 30% slower than release version due to clock problems on first stepping.

12/05/2016 07:02 PM is the article date - the link I posted confirming Vega AOTS was leaked before mid last year. Lets leave it at that okay.

Your second point, yes this is what I'm really trying to point out. If we're seeing these results supposedly 3 months out from launch on current ES demo and the same 3/4 of a year ago, pretty consistent performance 'ranking' on two different titles... it's a good sign of what to expect at minimum.
 
Well in any case dont expect consumer vega to come any time soon, word on the street is that they are targetting researchers first as a high margin market most likely to acquire the full yield of vega chips if they do bring the performance hinted for partial precision and integer workloads= even if the first chips are ready in 3 months i would be looking at late 5 as a possibility.

Also let's wait until proper review samples are distributed around, vega does sound nice but a dash of salt is required with any and all technology announcements.
 
What little was released seems interesting. We may in fact be looking at a GP100 competitor and the size may be misleading.

Difference being it may be positioned as a G/P100 competitor just based on the few specs we have so far. FP16 is in theory already ahead of P100. I'm wondering if effective FP16 and FP32 performance are identical with the fused instructions. That would take some serious compiler work to get working. They also mentioned Vega will be removing all resource management work from the developer. So it explains the high bandwidth cache. Could be a reference to added storage beyond HBM on the GPUs.

The GP100 is a tricky situation because it has a lot of FP64 performance along with pretty high FP32\FP16.
Nvidia make the comparison a headache because the real comparison should be big Vega vs Tesla P40 (full GP102), but the P40 is only FP32 and accelerated int8, so Vega has an FP16 advantage there but its FP32 and Int8 will probably be the same in theory.
The reason for the lack of FP16 in the P40 could be either technical related to cuda/register-BW related design, or just to protect and differentiate the P100 more than just by FP64, if the latter a stupid move IMO by Nvidia but then Volta is on the horizon for Tesla and will probably change this product structure.
Anyway the highest performing FP32 GPU for Nvidia is the P40 rather than P100.

Now one big caveat, it is impossible to get a 100% accelerated dot product (FP16 or Int8) performance, theory definitely does not match reality (it is an acceleration of anywhere between 65% to 90% depending upon function-algorithm-etc).
So will be interesting to see how this ends when comparing Nvidia Cuda to GCN, problem is which scientific-modelling benchmarks should be used and it really needs to use optimised code for both to have any relevance.
Cheers
 
Last edited:
What little was released seems interesting. We may in fact be looking at a GP100 competitor and the size may be misleading.

Difference being it may be positioned as a G/P100 competitor just based on the few specs we have so far. FP16 is in theory already ahead of P100. I'm wondering if effective FP16 and FP32 performance are identical with the fused instructions. That would take some serious compiler work to get working. They also mentioned Vega will be removing all resource management work from the developer. So it explains the high bandwidth cache. Could be a reference to added storage beyond HBM on the GPUs.


Ok gotta stop believing in some of the things they are saying (we don't know anything about performance so expecting it to go against the gp100, I find it hard to believe at this point at least from what they have shown in games) Yeah Gaming work loads are different but those same bottlenecks will appear (shader, throughput) to some degree in compute work loads, not to mention the entire software ecosystem. Resource management, is a necessity on developers end with HPC applications.

http://www.ncsa.illinois.edu/People/kindr/papers/ppac09_paper.pdf

I'm going to point to this.

There is a need to do this in software vs. hardware, because

The Torque batch system used on AC considers a CPU coreas an allocatable resource, but it has no such awareness for GPUs. We can use the node property feature to allow users to acquire nodes with the desired resources, but this by itself does not prevent users from interfering with each other, e.g.,accessing the same GPU, when sharing the same node. In order to provide a truly shared multi-user environment, we wrote a library, called CUDA wrapper [8], that works in sync with the batch system and overrides some CUDA devicemanagement API calls to ensure that the users see and have access only to the GPUs allocated to them by the batch system. Since AC has a 1:1 ratio of CPUs to GPUs and since Torque does support CPU resource requests, we allow users to allocate as many GPUs as CPUs requested. Up to 4 users may share an AC node and never “see” each other’s GPUs.

Now this is just one example,

But when you have multiple users, multiple data sets, multiple loads, the need for this to be in software vs hardware is a bit more stringent as the hardware won't see the needs coming, I can see the "reduced" amount of work from the dev's side, but not a total drop.

Now this is why even on a single system with a single user, its always been hard to get rid of programmers intervention with OS and application system resource management, because sometimes (many times, the programmer needs access to that for what they are doing).
 
Ughhh this is disappointing.... was hoping Vega would be something special but it can't even outperform a 1080.

Did you see the Battlefront PC demo was locked at 60 fps and the system couldn't even maintain that lol
 
I expected the reveal a near launch time.
But Fury X was released in June 2015
Polaris released in June 2016
Expect Vega in June 2017

Until then... meh
 
530mm² and it isn't faster than a Titan? Surely you must be joking. Have we actually learned anything new about Vega? Was quad packed 8bit not already known?

This is rather anticlimactic

NVidia doesn't even have to try these days. They are sitting around on cards because nothing is competing with them.
 
530mm² and it isn't faster than a Titan? Surely you must be joking. Have we actually learned anything new about Vega? Was quad packed 8bit not already known?

This is rather anticlimactic

Doesn't that include the on die HBM memory?
 
Ryzen is coming out in a few months and they're still demoing on an Intel CPU?

This.

Let's do something interesting, we should try and list all the possible reasons they would be doing this instead of using Zen.

Once that's done we look at how many people took it is a bad sign, compile our data and mail it over to AMD under the title ;

A MARKETER'S DREAM
 
That stuttering, though...

Funny, since both id Tech 4 and 5 ran horribly on ATi/AMD hardware.

Ryzen is coming out in a few months and they're still demoing on an Intel CPU?
Glad you noticed.
Remember awhile back I mentioned it seemed there was a conflict between the CPU and GPU division since the team structure split and a new VP dedicated to Radeon (Raja), with the VP of the CPU division not only demo'ing with Titan Pascal but also having it on the official AMD news brief lol, this was after GPU team working with Intel in providing a deal of i5+480.
And now this lol.
Yeah I get the feeling they are not playing nice with each other these days, will be interesting to see what the next demo with RyZen will be using when done by the CPU team.
Cheers
 
Doesn't that include the on die HBM memory?

No its separate. The die may be as much as 553mm2. 2 stacks is also going to be a limiting factor. Same memory speed as Fiji if clocked to the max and only 8GB.
 
Well their tech support department, stated failure rates and other issues, the co founder stated, crap support from AMD, so I think its a combination of both. Or one leading to the other.

http://www.pcworld.com/article/2052...ion-to-so-publicly-dump-amd-video-cards-.html

This is what tech support stated



Independent tester




This isn't something that just came out of the blue.

Then add this to the AIB partner failure rates, you have more than three separate entities, all saying the same thing. What do they say two is a coincidence three is a pattern? And I'm not saying nV is off the hook either, some of their partners have higher then normal failure rates too.

Outside of them moving it forward to 16nm because 10nm is too far out, nothing else right now.
regardless of what happened over 3 years ago, builders are all currently using the rx400 series. so they must have changed their minds over the last couple years as drivers/support have massively improved.
 
12/05/2016 07:02 PM is the article date - the link I posted confirming Vega AOTS was leaked before mid last year. Lets leave it at that okay.

Your second point, yes this is what I'm really trying to point out. If we're seeing these results supposedly 3 months out from launch on current ES demo and the same 3/4 of a year ago, pretty consistent performance 'ranking' on two different titles... it's a good sign of what to expect at minimum.
12/01/2016 is the screenshot date. Do you want to claim it took Hilbert 5 months to pen an article about it and that Vega ES was working a full year ago? Or that it was simply MM/DD/YY format, and not DD/MM you think of?
Difference being it may be positioned as a G/P100 competitor just based on the few specs we have so far. FP16 is in theory already ahead of P100
Fairly positive their FP16 is in the same range.
 
The GP100 is a tricky situation because it has a lot of FP64 performance along with pretty high FP32.
Nvidia make the comparison a headache because the real comparison should be big Vega vs Tesla P40 (full GP102), but the P40 is only FP32 and accelerated int8, so Vega has an FP16 advantage there but its FP32 and Int8 will probably be the same in theory.
The reason could be either technical related to cuda/register-BW related design, or just to protect and differentiate the P100 more than just by FP64, if the latter a stupid move IMO by Nvidia but then Volta is on the horizon for Tesla and will probably change this product structure.
Anyway the highest performing FP32 GPU for Nvidia is the P40 rather than P100.
Just based on the die size it's comparable to GP100. FP16 performance is listed just above P100. FP32 is still a bit of an unknown, but it's a surprisingly ambiguous stat given other known details. If MI25 was all but equated to 25TFLOPs FP16, why not list Vega at half of that? Why is FP64 "configurable"? With all the packed math, actually throwing in some fused FP32 instructions would be reasonable. Even Zen was supposed to get FMA4 which covered 4 operand instructions prior to getting axed. GCN had a similar setup, but the 4th operand was feeding the scalar and moving data about as I recall. Sure it would be a subset of the entire instruction set, but that design I was theorizing would probably allow for it. That design is looking surprisingly accurate to what we've seen so far.

At the very least the specs seem to match GP102 setting aside FP64 performance as unknown. Similar die size, FP16, INT8, +50% memory bandwidth etc. Regular rate FP32 it's in theory on par with GP102. If a 480 and 1060 are comparable, Vega with superior stats in addition to architectural tweaks I would assume outperforms it's competitor.

Ok gotta stop believing in some of the things they are saying (we don't know anything about performance so expecting it to go against the gp100, I find it hard to believe at this point at least from what they have shown in games) Yeah Gaming work loads are different but those same bottlenecks will appear (shader, throughput) to some degree in compute work loads, not to mention the entire software ecosystem. Resource management, is a necessity on developers end with HPC applications.

http://www.ncsa.illinois.edu/People/kindr/papers/ppac09_paper.pdf

I'm going to point to this.

There is a need to do this in software vs. hardware, because
The memory management thing isn't that unreasonable. If VRAM was effectively unlimited and all resources loaded, why would a developer be required to manage it? Link

The entire reason for direct management is to intelligent work around bus limitations. Limitations that largely don't exist in this setup. In the case of gaming, so long as the resources don't completely swap every other frame the caching scheme should work well. It also leaves the entire PCIE3 bus open for communication with other adapters as it won't be streaming textures or other resources.
 
regardless of what happened over 3 years ago, builders are all currently using the rx400 series. so they must have changed their minds over the last couple years as drivers/support have massively improved.


Wasn't pointing to the rx 480 in that regard, was specific to why the FE was done the reasons why OEM's wanted the FE and nV tried to use the same reasons to push to consumers, which that didn't work.
 
Fairly positive their FP16 is in the same range.

Just looking at it from a theoretical standpoint, Vega would be faster than the P100 as its FP32 is similar to P40 at 12 TFLOPs FP32.
The P100 was 10.6 TFLOPS FP32 with NVLink and 9.3 with PCIe.
Real-world coding and use could change this though from a dot product acceleration perspective and no way to know until tested by some like Baidu/Google/some of the scientific modelling tools/etc.
Cheers
 
Last edited:
The memory management thing isn't that unreasonable. If VRAM was effectively unlimited and all resources loaded, why would a developer be required to manage it? Link

The entire reason for direct management is to intelligent work around bus limitations. Limitations that largely don't exist in this setup. In the case of gaming, so long as the resources don't completely swap every other frame the caching scheme should work well. It also leaves the entire PCIE3 bus open for communication with other adapters as it won't be streaming textures or other resources.

Well this also has been done through drivers in the past, but engines have incorporated memory management techniques (streaming textures) for years now. It might take away from future development, to some degree but right now, nope not that useful. Not only that, older games that pushed vram hard it was always handled by the driver teams to allocate the necessary resources to stream textures. End result, dev's never need to worry about it to begin with. The reason why current engines have it is because of large open worlds, and they will not be able to reduce everything they are talking about without the programmer involved, because they WILL need to tell the gpu what is going on from the camera position in game.
 
Just based on the die size it's comparable to GP100. FP16 performance is listed just above P100. FP32 is still a bit of an unknown, but it's a surprisingly ambiguous stat given other known details. If MI25 was all but equated to 25TFLOPs FP16, why not list Vega at half of that? Why is FP64 "configurable"? With all the packed math, actually throwing in some fused FP32 instructions would be reasonable. Even Zen was supposed to get FMA4 which covered 4 operand instructions prior to getting axed. GCN had a similar setup, but the 4th operand was feeding the scalar and moving data about as I recall. Sure it would be a subset of the entire instruction set, but that design I was theorizing would probably allow for it. That design is looking surprisingly accurate to what we've seen so far.

At the very least the specs seem to match GP102 setting aside FP64 performance as unknown. Similar die size, FP16, INT8, +50% memory bandwidth etc. Regular rate FP32 it's in theory on par with GP102. If a 480 and 1060 are comparable, Vega with superior stats in addition to architectural tweaks I would assume outperforms it's competitor.


The memory management thing isn't that unreasonable. If VRAM was effectively unlimited and all resources loaded, why would a developer be required to manage it? Link

The entire reason for direct management is to intelligent work around bus limitations. Limitations that largely don't exist in this setup. In the case of gaming, so long as the resources don't completely swap every other frame the caching scheme should work well. It also leaves the entire PCIE3 bus open for communication with other adapters as it won't be streaming textures or other resources.
I still think it fits with what I mentioned in our earlier discussions.
The size of Vega aligns with GP102 for reasons we discussed at length in the past, this was when we were coming from different perspectives and I was saying it would be a big die and basing it upon various factors which included a die size of being larger than the GP102, which is 471mm2.
Unfortunately there is no point bringing the P100 into this as it also has a serious amount of FP64 cores (in a ratio of 2:1), and Vega has minimal FP64 just like Fiji, it is closer to the P40 than P100.
Cheers
 
Back
Top