Anarchist4000
[H]ard|Gawd
- Joined
- Jun 10, 2001
- Messages
- 1,659
Some of which may be provided by caching mechanisms. The acyclic graphs used by the DAG are trees. So it stands to reason the trunk could be cached or localized access patterns established over time. So even if random, there could exist a temporal access pattern a victim cache naturally discovers if one exists. Haven't seen any details on where all that SRAM went yet. Vega still has more unaccounted cache than P100 and a big L3 that works transparently would make sense.What does that mean, well pretty much ya need a butt load of bandwidth and memory more than processing power. That is why frequency more cores etc doen't do much for Eth mining if the bandwidth isn't there.
That was with the same memory speeds. No reason significantly more capable HBM2 won't exist at that time. That's also 6 months of driver improvements with a lot of new capabilities and even games that have already announced support for packed math. No guarantee gaming Volta has that because if segmentation.Vega won't be refreshed that quickly, why would it be, Polaris refresh was one year and did we see anything extra from that? 5% more performance at a cost of 30% more power?
There isn't a fixed amount of geometry units, but pipeline elements as the ALUs are doing the lifting. The 4SE arrangement seen more about binning triangles into specific pipelines. A triangle is a single thread in a wave, so AMD could push 4096/clock by the time the vertex shaders start up. They could go larger with some added hardware, but the 4SE part doesn't seem the concern. AMD was hiring new front-end engineers though, so maybe after Navi, unless it's a software issue. FPGA makes more sense there.But nV's architectue doesn't hit those limitations, only AMD's do...... That is because they only have so many geometry units.
Still think it'll take town Titan Xp once all the features are used. That part seems rather likely given possible performance gains from some abilities. Mantor explained how primitive shaders would make a difference, although it was limited to saving bandwidth. The culling mechanisms are well established, but with dynamic allocation they could speed it along with FP16. Not expecting huge gains until a dev really goes to town with it, but as I mentioned above, geometry isn't the biggest issue. The pixel shading is the bulk of the work where making z-culling more efficient has big gains. No idea if DSBR had that part enabled yet, but the primitive shaders likely assist. Converting g positions into 8/16bit to hopefully sort more efficiently. Lot of moving pieces that all need to be working and would prefer some tuning once they were.Look first off you throught Vega was going to be a 1080ti killer because of all the specs AMD was shouting out for close to a year now, Now that is not going to happen and its obvious. So you are going to tell me, what they just stated about primitive shaders are going to make any difference. If developers don't have that control over them, forget it, its polygon through put will be no better than Polaris for almost all titles. Past and near future, untill developers have access to it. AMD will not be able to do it through drivers as I stated unless initiation and propagation of vertices are done with FP16 there will be no way for primitive shaders to automatically done via drivers.