This strikes me as one of those features that if pulled off can be great, but Radeon doesn't exactly have a track record to pulling off this type of big level optimization.
Never seen it done by anyone in any silicon market
Follow along with the video below to see how to install our site as a web app on your home screen.
Note: This feature may not be available in some browsers.
This strikes me as one of those features that if pulled off can be great, but Radeon doesn't exactly have a track record to pulling off this type of big level optimization.
No way. This thread is a fantastic exercise in how much effort people are willing to put into petty arguments online.
Considering you can be banned for trolling, this should have been locked if [H] wants to follow their own logic.
Too piece meal for me to put together the ramblings. I would think most folks would just look at the bottom line - how does it perform for what I am going to use it for - gaming, mining, specific applications - look at price and then decide. Future performance or hopefulness may be very disappointing in the end.The casual reader could learn a lot from some of the posts in the this thread...I quite enjoy the technical level in some of the posts...even if they only are there to combat PR-FUD-Troll posts....they contain valuable technical information![]()
Primitive Shaders are very hard to develop and code for, one of AMD guys told Anandtech it's like writing an assembly code, you have to always outsmart the driver, and use some inefficient driver paths well.
This is a mischaracterization of the posts made by Ryan Smith about the subject on the Beyond 3D Forum. I'm going to quote all three relevant posts in full. The first is here by Ryan Smith:it also needs extensive developer involvement, judging by his comments. Here are the quotes:
Please note "The manual developer API is not ready, and the automatic feature to have the driver invoke them on its own is not enabled." This clearly states that there will be an automatic mode for primitive shaders in addition to creating an API to allow developers to override the default primitive shader implementation themselves.Quick note on primitive shaders from my end: I had a chat with AMD PR a bit ago to clear up the earlier confusion. Primitive shaders are definitely, absolutely, 100% not enabled in any current public drivers.
The manual developer API is not ready, and the automatic feature to have the driver invoke them on its own is not enabled.
Let me note again for emphasis that leonazzurro specifically says "there will be a "automatic mode" in the drivers that will specifically work for increasing the primitive discard rate." and that in addition "Then, there is the possibility to expose completely the primitive shaders to the developers, allowing other options (that means: new feature/possibilities in game engines)." It is crystal clear here that leonazzurro is talking about an automatic default implementation of primitive shaders on the one hand, and a manual developer override of that default implementation on the other.As far as I understood, there will be a "automatic mode" in the drivers that will specifically work for increasing the primitive discard rate and thus speeding up the geometry processing. I can only guess that is taking a lot of time to implement it because of possible compatibility issues with existing software. Then, there is the possibility to expose completely the primitive shaders to the developers, allowing other options (that means: new feature/possibilities in game engines). But Rys has confirmed that that is not planned yet in a tweet some time ago.
Ryan did not correct leonazzurro's description of there being a default automatic mode for primitive shaders that would be could eventually be overriden by developers once the API is implemented, but instead adds to leonazzurro's post by explaining that it will be very difficult for developers to outperform the automatic driver implementation of primitive shaders with manual control. That is what Ryan is referring to when he says "You have to be able to outsmart the driver (and the driver needs to be taking a less than highly efficient path) to gain anything from manual control."AMD is still trying to figure out how to expose the feature to developers in a sensible way. Even more so than DX12, I get the impression that it's very guru-y. One of AMD's engineers compared it to doing inline assembly. You have to be able to outsmart the driver (and the driver needs to be taking a less than highly efficient path) to gain anything from manual control.
Yes it is!!
No it isn't!!
Yes it is!!
No it isn't!!!
This thread has become a joke. I can't believe you guys behave like this.
This is a mischaracterization of the posts made by Ryan Smith about the subject on the Beyond 3D Forum. I'm going to quote all three relevant posts in full. The first is here by Ryan Smith:
Please note "The manual developer API is not ready, and the automatic feature to have the driver invoke them on its own is not enabled." This clearly states that there will be an automatic mode for primitive shaders in addition to creating an API to allow developers to override the default primitive shader implementation themselves.
Another B3D Forum user then asks Ryan to clarify whether he is referring to an automatic primitive shader implementation and a manual developer override of that automatic primitive shader implementation:
Let me note again for emphasis that leonazzurro specifically says "there will be a "automatic mode" in the drivers that will specifically work for increasing the primitive discard rate." and that in addition "Then, there is the possibility to expose completely the primitive shaders to the developers, allowing other options (that means: new feature/possibilities in game engines)." It is crystal clear here that leonazzurro is talking about an automatic default implementation of primitive shaders on the one hand, and a manual developer override of that default implementation on the other.
Finally, Ryan's second post is made directly in response to leonazzurro's post above and quotes it, where Ryan says:
Ryan did not correct leonazzurro's description of there being a default automatic mode for primitive shaders that would be could eventually be overriden by developers once the API is implemented, but instead adds to leonazzurro's post by explaining that it will be very difficult for developers to outperform the automatic driver implementation of primitive shaders with manual control. That is what Ryan is referring to when he says "You have to be able to outsmart the driver (and the driver needs to be taking a less than highly efficient path) to gain anything from manual control."
Ryan's posts explicitly support that there will be a default implementation of primitive shaders in drivers by AMD that will not require any developer input to use, and that AMD is working on implementing this default implementation first, and only later exposing the programmable geometry pipeline to manual developer control. It's also worth noting that Ryan's comments are also entirely consistent with Rys Sommefeldt's comments on twitter that the default version of primitive shaders will work automatically and that there are no plans yet to enable manual developer override of the default primitive shader implementation for the reasons Ryan Smith noted:
![]()
So you agree that an automatic mode for primitive shaders can be implemented in a subsequent driver update to RX Vega by AMD?now AMD needs to do the work in drivers
Ryan Smith makes it crystal clear that the automatic mode for primitive drivers is not yet enabled in any public driver:That reads like the drivers are already doing it. Unless I am reading that wrong?
Primitive shaders are definitely, absolutely, 100% not enabled in any current public drivers.
So you agree that an automatic mode for primitive shaders can be implemented in a subsequent driver update to RX Vega by AMD?
Ryan Smith makes it crystal clear that the automatic mode for primitive drivers is not yet enabled in any public driver:
Now do you understand why they didn't tell you everything? This is not as simple as a slide or two or a white paper that doesn't go in to details how they are doing it via drivers *shit if they said drivers take care of world hunger would you believe them? Software needs to be driving something right? I have worked with tesselators (software since the late 90's) and did a lot of work with truform and tried to emulate it on nV hardware but that didn't work too well, since DX or Ogl didn't expose or rather hardware didn't have it most likely as ATi's did. Did it through CPU though and worked well, since at the time skinning and animation were all CPU side anyways, it increased the bandwidth needs across the AGP port though which in some cases (outliers) caused slow downs.
Worse yet you will not get predictable performance increases while emulating, and to the contrary of that there are always performance drawbacks when emulating and chances are high of performance pitfalls when emulation is stressed enough. This is what happened with the emulation of the tesselator with displacement for the 9700 pro. The question is is Vega's shader array enough to manage current workloads and emulate the tessealtor and other shader functionality while delivering the same performance as before, lets not even talk about higher performance?
Maybe its possible in older games, in current games or games just about to come out that is questionable. Current games the chances are ok that you might get some benefits, not much because we can see its shader array is not scaling well in many tests. Future games or games coming out shortly forget about it cause yeah they will push Pascal's array more, and if that happens Vega can't even keep up with Pascal not unit for unit. And then you have the other bottleneck to worry about bandwidth. Virtually the same shader array in Fiji is getting hampered down but Vega is not going to have problems with doing things like this? Yeah gotta say not going to happen.
This is all if there is no draw backs for emulation of the tesselator and tessellation is used. If Tessellation isn't used then yeah they will get some benefits if the application is suitable to changed via driver to use primitive shaders. I can at this point say, I don't know any engine (AAA) that will be suitable for this type of replacement via drivers.
This all goes back to what I stated ALONG time ago one year ago about Vega's new features, they look like add ons after seeing what Maxwell had they were too far into the process of Vega that they couldn't do anything then "create something new from something old".
razor1 have you seen this yet?
https://hardforum.com/threads/amd-r...-in-ethereum-mining-efficiency-by-2x.1943165/
43MH for 130W! I believe this now contends with the 1070 for efficiency.
Yes I did see that, hmm its close to the efficiency of the rx 580, the 1070 is still a bit higher. But yeah its damn close, I would like to see the Vega 56 on this! At the moment this is only on ETH, other alt coins its much lower, I think this has to do with software more than anything else, but until that software is optimized for Vega, its not a good option right now, why would anyone keep mining ETH at this point while buying rigs for it? There is only about 4 to 5 months left to mine it.
Eth is going to POS in Feb of next year, no more mining, at least that is what the Eth founders are saying, I'm hoping it will be longer than that, but don't have my hopes up cause it was supposed to go to POS this year.
POS means?
The only POS abbreviation I know of is "Piece of Shite"...
Proof of Stake
but why does that mean no more mining?
Well the way it sounds like its going to be only proof of stake, no capability to mine anymore.
razor1 have you seen this yet?
https://hardforum.com/threads/amd-r...-in-ethereum-mining-efficiency-by-2x.1943165/
43MH for 130W. I believe this now contends with the 1070 for efficiency.
Yes I did see that, hmm its close to the efficiency of the rx 580, the 1070 is still a bit higher. But yeah its damn close, I would like to see the Vega 56 on this! At the moment this is only on ETH, other alt coins its much lower, I think this has to do with software more than anything else, but until that software is optimized for Vega, its not a good option right now, why would anyone keep mining ETH at this point while buying rigs for it? There is only about 4 to 5 months left to mine it.
LOL Rasterizer reads into marketing spiel like its gold.
Hold on a second. You've already said in three separate posts that you do believe that primitive shaders (NOT, I repeat NOT RPM) can achieve "x2 the triangle through put":Never stated it won't be, I stated it won't show anything more then what we see already, and its going to be hard do with current games.
While using RPM and all these things too, that is the only way it can achieve x2 the triangle through put lol. otherwise its around Fiji per clock......
Doesn't matter if it uses FP 16 or FP 32, if using the fixed function pipeline, the triangle amounts will stall GCN's pipeline. The only way around that is to use its primitive shaders. FP 16 calculations are done in the shader array. But the GU's have to do all the work after *this is where the problem is* This is where AMD's primitive shaders come in, they communicate with the programmable GU's of Vega.
So if using FP 16 or FP 32, if the bottleneck is the fixed function GU's, would it matter if that demo used FP 16 or FP 32? Cause the GU's can handle only so many triangles before they stall the pipeline.
Again, forget about RPM. I am not talking about RPM and I don't care about RPM. Aren't you saying in these posts that primitive shaders can actually achieve 2x the polygon throughput? As I understand it, RPM should contribute zero to polygon throughput even if it was working, so it has nothing to do with the polygon throughput. The polygon throughput gains you are talking about here should ALL be from primitive shaders, right?RPM for vertex processing gets not benefit unless you use primitive shaders for GCN! As it is right RPM isn't even available in any API and RPM alone will not give any benefit in geometry processing I should have been more clear about this, that is my fault but that doesn't change the fact that demo needed all of Vega's new pipeline to show its max through put which is coincidentally x2 the geometry through put that is in AMD's marketing material!.
I'll only address this point once. Please don't confuse me with someone that made outlandish, rambling claims with zero citations or evidence. I don't care if you agree with me or don't, but at least acknowledge that I make an effort actually use sources, to cite those external sources I am relying on, and to articulate my arguments coherently.the way he speaks and make claims sometimes remember to some anarchist that joined the heavy forces of the AMD rebellion on the forum. lol..
I actually probably understand about 10% of what you guys are arguing about, but I think I agree with you.I'll only address this point once. Please don't confuse me with someone that made outlandish, rambling claims with zero citations or evidence. I don't care if you agree with me or don't, but at least acknowledge that I make an effort actually use sources, to cite those external sources I am relying on, and to articulate my arguments coherently.
I'll only address this point once. Please don't confuse me with someone that made outlandish, rambling claims with zero citations or evidence. I don't care if you agree with me or don't, but at least acknowledge that I make an effort actually use sources, to cite those external sources I am relying on, and to articulate my arguments coherently.
Hold on a second. You've already said in three separate posts that you do believe that primitive shaders (NOT, I repeat NOT RPM) can achieve "x2 the triangle through put":
Again, forget about RPM. I am not talking about RPM and I don't care about RPM. Aren't you saying in these posts that primitive shaders can actually achieve 2x the polygon throughput? As I understand it, RPM should contribute zero to polygon throughput even if it was working, so it has nothing to do with the polygon throughput. The polygon throughput gains you are talking about here should ALL be from primitive shaders, right?
I think I finally understand the root of our disagreement. If I understand you correctly, you believe that primitive shaders will bypass the existing geometry engine entirely and do the geometry processing directly on the compute engine, thereby taking resources away from the rest of the graphics pipeline, yes?primitive shaders will need to use the shader array
I think I finally understand the root of our disagreement. If I understand you correctly, you believe that primitive shaders will bypass the existing geometry engine entirely and do the geometry processing directly on the compute engine, thereby taking resources away from the rest of the graphics pipeline, yes?
That's not my understanding at all. As far as I understand it, it is Vega's geometry engine itself with has replaced the formerly fixed function shaders of the geometry engine with programmable generalized non-compute shaders, and that right now those programmable non-compute shaders are operating in a legacy mode where they mimic the traditional GCN geometry engine. It's my understanding that it is those generalized non-compute shaders located inside the geometry engine that can be reprogrammed into primitive shaders, meaning there should be zero impact on shader array usage whatsoever from whether or not primitive shaders are enabled.
From my understanding, the new geometry engines are not emulating the old pipeline. I think it would be more accurate to say that Vega's geometry engines are presently configured like Fiji's were, or programmed like Fiji's were. Basically, the shaders inside of Vega's four geometry engines are no longer fixed function, instead they can be programmatically told what kind of shader behaviour to have (e.g. domain, hull, vertex, geometry), and once they have been programmatically configured to act as a certain type of shader, there should be minimal overhead for them continuing to do so (on anoher forum I saw a few people compare it to a shader based version of Larabee). The automatic form of primitive shaders being developed by RTG is going to be a driver defined geometry engine configuration that reconfigures the shaders within Vega's four next generation geometry engines to function as primitive shaders as RTG defines them in the Vega whitepaper.Ask yourself if emulated the old pipeline how would it be getting the same performance (it really is getting the same performance per clock), emulation always has overhead one way or another. Why would this new pipeline (emulated) have the same restrictions too? It shouldn't right?
Next-generation geometry engine
To meet the needs of both professional graphics and gaming applications, the geometry engines in “Vega” have been tuned for higher polygon throughput by adding new fast paths through the hardware and by avoiding unnecessary processing. This next-generation geometry (NGG) path is much more flexible and programmable than before.
I think it's pretty clear that they are talking about the four geometry engines functioning as primitive shaders, rather than bypassing the geometry engines and doing the geometry processing on the shader array.The “Vega” 10 GPU includes four geometry engines which would normally be limited to a maximum throughput of four primitives per clock, but this limit increases to more than 17 primitives per clock when primitive shaders are employed.⁷
The bottleneck is still there right now because Vega's four geometry engines are currently programmed by the drivers to act the same as the previous fixed function geometry engines of Fiji. I would assume this was done to not hold up development of the card while RTG worked on developing the new automatic primitive shader configuration for the four next generation geometry engines.cause the bottleneck for the polygon throughput is amount of geometry engines, nothing else. if those GE's are emulated already, then the bottleneck is removed they can just say use more shader units for GE operations and be done with. That is not what they are doing right now. That is what they will be done with primitive shaders though.
There is a reason why the API is there its to govern the way programming is done, the underlying hardware can do anything it wants to but if it can't do those steps without having problems its not going to get certified.
Actually if they are able to emulate it that well, then there would be no need for primitive shaders![]()
From my understanding, the new geometry engines are not emulating the old pipeline. I think it would be more accurate to say that Vega's geometry engines are presently configured like Fiji's were, or programmed like Fiji's were. Basically, the shaders inside of Vega's four geometry engines are no longer fixed function, instead they can be programmatically told what kind of shader behaviour to have (e.g. domain, hull, vertex, geometry), and once they have been programmatically configured to act as a certain type of shader, there should be minimal overhead for them continuing to do so (on anoher forum I saw a few people compare it to a shader based version of Larabee). The automatic form of primitive shaders being developed by RTG is going to be a driver defined geometry engine configuration that reconfigures the shaders within Vega's four next generation geometry engines to function as primitive shaders as RTG defines them in the Vega whitepaper.
The wording used by the Vega whitepaper itself makes it pretty clear that the next generation geometry engines are the new geometry engines in Vega, not uses of the shader array for geometry processing (apologies for quoting this again, but I think it's really relevant here:
I think it's pretty clear that they are talking about the four geometry engines functioning as primitive shaders, rather than bypassing the geometry engines and doing the geometry processing on the shader array.
The bottleneck is still there right now because Vega's four geometry engines are currently programmed by the drivers to act the same as the previous fixed function geometry engines of Fiji. I would assume this was done to not hold up development of the card while RTG worked on developing the new automatic primitive shader configuration for the four next generation geometry engines.
What exactly are you both arguing about? Who gives a shit if the white paper is full of promises. The feature either doesn't work as outlined or isn't worth implementing. The proof is in the performance, not documentation.
he is trying to say Vega's polygon throughput is going to fixed via drivers and it will return to glorious performance levels and his reasoning is based on things AMD has said but, never going to happen cause its just not, pretty much what you said there is nothing that AMD shown that would give any confidence that it will. Nor do I see any technical difference between CU's doing the fixed function part of the pipeline vs. primitive shaders.
You are both arguing beliefs and trying to back them up with facts and interpretations.
In a typical scene, around half of the geometry will be
discarded through various techniques such as frustum
culling, back-face culling, and small-primitive culling. The
faster these primitives are discarded, the faster the GPU
can start rendering the visible geometry. Furthermore,
traditional geometry pipelines discard primitives after
vertex processing is completed, which can waste computing
resources and create bottlenecks when storing a large batch
of unnecessary attributes. Primitive shaders enable early
culling to save those resources.
One factor that can cause geometry engines to idle is
context switching. Context switches occur whenever the
engine changes from one render state to another, such as
when changing from a draw call for one object to that of a
dierent object with dierent material properties. The
amount of data associated with render states can be quite
large, and GPU processing can stall if it runs out of
available context storage. The IWD seeks to avoid this
performance overhead by avoiding context switches
whenever possible.
Primitive shaders will coexist with the standard hardware
geometry pipeline rather than replacing it.
By definition, if you replace the fixed function shaders in the geometry engines with generarilzed non-compute shaders, they become programmable and thus of course they can be arbitrarily switched between behaving like Fiji's geometry engine and behaving like primitive shaders instead as needed.Come on Rasterizer it was right there. Page 7
This makes it very clear that Polaris' Primitve Discard Accelerators are inside the geometry engines prior to the rasterizers, yes?The Polaris geometry engines use a new filtering algorithm to more efficiently discard primitives. As figure 5 illustrates, it is common that small or very thin triangles do not intersect any pixels on the screen and therefore cannot influence the rendered scene. The new geometry engines will detect such triangles and automatically discard them prior to rasterization
It's hard to blame them. Radeon RX Vega 64 is a catastrophe in every single possible way.
That is no way anyone can sugarcoat this.
Compare to GeForce GTX 1080,
Radeon RX Vega 64 is...
1. Cheaper? ✗
2. Faster? ✗
3. More power efficient? ✗
______________________________________________________________________
It would be pretty hard to justify buying a Radeon RX Vega 64 over a Geforce GTX 1080.