• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

30,000 NVIDIA Engineers Use Generative AI for 3x Higher Code Output

rinaldo00

2[H]4U
2FA
Joined
Mar 9, 2005
Messages
2,879
The company that started the entire wave of AI infrastructure and development is now enjoying the fruits of its work. NVIDIA is deploying generative AI tools across its company to an astonishing 30,000 engineers. In a partnership with San Francisco-based Anysphere Inc., the company is getting a customized version of the Cursor integrated developer environment, which focuses on AI code design. This is important to note as NVIDIA's engineers are now reportedly producing as much as three times the code compared to the previous development pipeline, and we are now probably using NVIDIA's products or services that have been designed by AI guided by humans.

https://www.techpowerup.com/346079/...s-use-generative-ai-for-3x-higher-code-output

All voluntarily of course. With no exaggeration and no problems.
 
AI code is verbose (to an extreme, sometimes with weird loops requiring pages of code when 4 lines could do the same thing).
Human code is optimized.
I'm surprised it's only 3x.

And it's weird counting success by amount of lines in code. I'd count success with fewer lines of code with the expected output/result.
 
It's not as if Nvidia is biased and wants to sell AI as the second coming or anything....


I'd take anything they publish on the topic of AI with a truckload of salt.
Truckload of sh*t is more like it. A.I. is the second coming of the antichrist. Never got more headaches and bad advice than when I decided to start messing with the A.I. bots. They know absolutely nothing about Enterprise PC hardware. None knew how oscilloscopes work; I literally had to do the research and explain it to them. I asked one to troubleshoot setting up an older projector with Fedora and I had to reinstall Fedora to get an image on either of my monitors or projector after taking their advice.

I think in the future, I will have a panel of A.I. assistants who will come to a consensus on an answer before I destroy a piece of hardware fucking around with their advice again. Yeah that happened too. Long story. And it said it never told me to do what it said do and blamed me for the outcome. Yeah, A.I. is the shiznit. Fuck Gemini in particular.
 
How are they measuring output exactly? 3x seems absurd
It's funny since the headlines about this generating 3x as much code isn't even news, since every LLM can introduce more code bloat. It's one of the most common criticisms even.

The real headline would be if it increased productivity around equivalent quality code by 3x, including measuring code removal due to optimizations/bug fixes.

Seems oddly specific terminology when Nvidia could afford to offer more insights. Unless they're similarly just looking at lines of code committed and calling it a win, like some places.
 
Truckload of sh*t is more like it. A.I. is the second coming of the antichrist. Never got more headaches and bad advice than when I decided to start messing with the A.I. bots. They know absolutely nothing about Enterprise PC hardware. None knew how oscilloscopes work; I literally had to do the research and explain it to them. I asked one to troubleshoot setting up an older projector with Fedora and I had to reinstall Fedora to get an image on either of my monitors or projector after taking their advice.

I think in the future, I will have a panel of A.I. assistants who will come to a consensus on an answer before I destroy a piece of hardware fucking around with their advice again. Yeah that happened too. Long story. And it said it never told me to do what it said do and blamed me for the outcome. Yeah, A.I. is the shiznit. Fuck Gemini in particular.
"You were right to call that out"
"I misspoke [fed you bullshit lies]"
"Thank you for pushing back on that"
"Good catch. Here is why [nothing I said makes sense]"
"Yeah — on that specific point [the only point that matters], I was wrong, and you’re right to call it out."
"That’s not true in the way it would need to be true to achieve what you wanted." - exact quote!

90% of my interactions with multiple AI lately (ChatGPT, Grok, Claude). And I feel like it's constantly getting worse with each model "upgrade". I had to resort to what you mentioned, parallel discussions with multiple AIs to get a semblance of a useful answer by the end of each grueling session.
 
Last edited:
"Yeah — on that specific point [the only point that matters], I was wrong, and you’re right to call it out."
Particularly with OpenAI's LLMs it's absurd the lengths of face-saving it will go to—even the reasoning models.

While sci-fi predicted anthropomorphized AGI we haven't even arrived at AGI yet some companies are training LLMs either without removing the human-like face-saving mimicry or deliberately so*, which is not at all what I wanted from such tech, particularly when pressing them with undeniable evidence to the contrary (eg: exact official docs and commit URLs).

* One might think adding system prompts to explicitly ask to avoid this would work but it doesn't wholly work, suggesting something inherent to the models.
 
Particularly with OpenAI's LLMs it's absurd the lengths of face-saving it will go to—even the reasoning models.

While sci-fi predicted anthropomorphized AGI we haven't even arrived at AGI yet some companies are training LLMs either without removing the human-like face-saving mimicry or deliberately so*, which is not at all what I wanted from such tech, particularly when pressing them with undeniable evidence to the contrary (eg: exact official docs and commit URLs).

* One might think adding system prompts to explicitly ask to avoid this would work but it doesn't wholly work, suggesting something inherent to the models.
OMG, yes! #2 most irritating feature of theirs. #1 is over the top constant flattery - "That is a great question", "Now you're showing some deep thinking", "You are a god among humans about asking what 241.83 miles equal in km"

Most of my instructions are to combat that, the other to combat being overly wordy and constantly repeating itself by stating the same thing in three different ways. All ignored, of course.
 
"You were right to call that out"
"I misspoke [fed you bullshit lies]"
"Thank you for pushing back on that"
"Good catch. Here is why [nothing I said makes sense]"
"Yeah — on that specific point [the only point that matters], I was wrong, and you’re right to call it out."
"That’s not true in the way it would need to be true to achieve what you wanted." - exact quote!

90% of my interactions with multiple AI lately (ChatGPT, Grok, Claude). And I feel like it's constantly getting worse with each model "upgrade". I had to resort to what you mentioned, parallel discussions with multiple AIs to get a semblance of a useful answer by the end of each grueling session.
That isn't the model.

Run local ai to see unfiltered responses.

My guess is the responses are fed from the precontext window and probably some checker ai that ensures the response stays within some limits.

Raw ai can be extremely rude at times.


The real problem I see isn't AI, it's the consolidation of AI to large corporations instead of democratized locally run AI models.
 
The real problem I see isn't AI, it's the consolidation of AI to large corporations instead of democratized locally run AI models.

Pay as you go Ai for the few that can afford the hardware, or Ai with invasive privacy annihilation baked right in, enshitification cranked up to 11 right out of the box.

As it stands that's the only place we're headed.
 
we are now probably using NVIDIA's products or services that have been designed by AI guided by humans.
Didn't Nvidia say they were using AI/ML for driver optimization and development a number of years ago? And for hardware dev a few years before that?

I know it's cool to say humans are the bestest, but if AI-powered development is so horrid, why is human-powered AMD so far behind?
 
it's the consolidation of AI to large corporations instead of democratized locally run AI models.
Kimi k2.5 is a trillion parameter model you can run as you want, where you want:
https://llm-stats.com/blog/research/kimi-k2-5-launch

And at current price per millions of tokens and how much you can do with all the free tiers available, it is one of the most democratized tech ever I feel like, that reach 1 billion world user really fast.

Pricing went down by a factor of ~100 last 2 years and could do it again in the next 3 years. Running a GPT-3 level of capacity model now is already extremelly cheap.

Pay as you go Ai for the few that can afford the hardware, or Ai with invasive privacy annihilation baked right in,
or really cheap API access with good encryption like we have right now, gemini 3 flash API use AES-256 and google cannot see the prompts or the data (that one reason there was a debate about deepseek using other model to train themselve, they could see api traffic but not the actual request, they just analysed traffic pattern)
 
Kimi k2.5 is a trillion parameter model you can run as you want, where you want:
https://llm-stats.com/blog/research/kimi-k2-5-launch

And at current price per millions of tokens and how much you can do with all the free tiers available, it is one of the most democratized tech ever I feel like, that reach 1 billion world user really fast.


or really cheap API access with good encryption like we have right now, gemini 3 flash API use AES-256 and google cannot see the prompts or the data (that one reason there was a debate about deepseek using other model to train themselve, they could see api traffic but not the actual request, they just analysed traffic pattern)
I am not saying you can't run models locally, you're extending my framing here.

No one is running or training these 700b models on local hardware. You're married to someone else's hardware.

I haven't found a local 32b models that can compete with Claude Sonnet 4.5 at coding. Not even close.

That is a problem and it's not democratized.
 
No one is running or training these 700b models on local hardware.
because kimi k2 has only 32 billions active parameter, you can run it on relatively modest hardware bandwith wise, just need a lot of ram.

512gb Mac ultra grouped together can run it, that still much cheaper than a cheap new car for an entreprise, many small entreprise are doing exactly that (if it is for internal use, you could need 16 H100 level of more for a small customer base level of performance)

That is a problem and it's not democratized.
low price make it one of the most democratized frontier new tech of all time, we do not judge how aviation, the Internet, electricity or running water is democratized just by who own the planes, server/fiber line and water treatment/pumping/tubes, but how many people have access to it and at what price. Internet is quite democratized, very few people that use it own any of the hardware that run it. Youtube is quite democratized.

People can run kimi k2.5 at 50 cent by millions token input and 2.80 by millions token output (even if you could have a cheap computer doing it, how much electricity would it cost in comparison to those pricing...) via the cloud, African start-up are using LLMs, there is over 1 billions users worldwide, what "new" techology was ever that democratized that fast in the past ?

There is a giant amount of competition, open-source is incredibly big, users base at low fee or free is in the billion worldwide, many feel it is way to accessible and democratized and pushing for restriction because even kids without any money are using it, far to be not enough.
 
because kimi k2 has only 32 billions active parameter, you can run it on relatively modest hardware bandwith wise, just need a lot of ram.

512gb Mac ultra grouped together can run it, that still much cheaper than a cheap new car for an entreprise, many small entreprise are doing exactly that (if it is for internal use, you could need 16 H100 level of more for a small customer base level of performance)


low price make it one of the most democratized frontier new tech of all time, we do not judge how aviation, the Internet, electricity or running water is democratized just by who own the planes, server/fiber line and water treatment/pumping/tubes, but how many people have access to it and at what price. Internet is quite democratized, very few people that use it own any of the hardware that run it. Youtube is quite democratized.

People can run kimi k2.5 at 50 cent by millions token input and 2.80 by millions token output (even if you could have a cheap computer doing it, how much electricity would it cost in comparison to those pricing...) via the cloud, African start-up are using LLMs, there is over 1 billions users worldwide, what "new" techology was ever that democratized that fast in the past ?

There is a giant amount of competition, open-source is incredibly big, users base at low fee or free is in the billion worldwide, many feel it is way to accessible and democratized and pushing for restriction because even kids without any money are using it, far to be not enough.
Hard disagree.
 
Kimi k2.5 is a trillion parameter model you can run as you want, where you want:
https://llm-stats.com/blog/research/kimi-k2-5-launch

And at current price per millions of tokens and how much you can do with all the free tiers available, it is one of the most democratized tech ever I feel like, that reach 1 billion world user really fast.

Pricing went down by a factor of ~100 last 2 years and could do it again in the next 3 years. Running a GPT-3 level of capacity model now is already extremelly cheap.


or really cheap API access with good encryption like we have right now, gemini 3 flash API use AES-256 and google cannot see the prompts or the data (that one reason there was a debate about deepseek using other model to train themselve, they could see api traffic but not the actual request, they just analysed traffic pattern)
I'm an AI pleb simpleton and only ever used the ChatGPT/Grok/Claude app chatbots. Can I use the two above in such a way and is there a ELI5 guide and price comparison to 20-30-ish/mo tier of the above?
 
I'm an AI pleb simpleton and only ever used the ChatGPT/Grok/Claude app chatbots. Can I use the two above in such a way and is there a ELI5 guide and price comparison to 20-30-ish/mo tier of the above?
you need to be an extreme heavy users to save money running locally, it is more for control (i.e. not having an api/model change that happen on you)

https://dev.to/czmilo/kimi-k25-in-2...l-agentic-intelligence-18od#pricing-licensing

512gb mac ultra with the big cpu-gpu upgrade are $9,500USD (relatively cheap for an enterprise infracstructure, a room of devs chairs can cost more than that, but still expensive for an individual) , API access is around 60 cent per millions tokens for input, even with free electricity it would take so long to save any money that API access for that level of model could be under 5 cent by then, if you are a regular user, it probably never work... at you need to not mind running at 20 token second

I would at least try the models via the cloud first ;)

Hard disagree.
With the idea that few technology had a wider as fast world adoption in history ? Or that how hard/expensive something is to use is an important metric to how democratic it is, now just ownership (i.e. airwave TVs/movies theater were an democratic art form/hobby even if none of its users were owner of the antenna/theater room and the concentration was higher than in the AI space today), sailboats are often user owner, that does not make it more democratic than the Internet (for which 99% of its users own/control/know nothing of its infracstructure), accessibility, price, friction are important variable and on all those fronts LLM are incredibly accessible to people.
 
  • Like
Reactions: Meeho
like this
I'm an AI pleb simpleton and only ever used the ChatGPT/Grok/Claude app chatbots. Can I use the two above in such a way and is there a ELI5 guide and price comparison to 20-30-ish/mo tier of the above?
Well, you have to have enough gpu ram to hold the models if you want reasonable performance. For running locally, VRAM is the primary bottleneck for good tokens/sec performance. It gets hard to ELI5 because of model quantization.
Rough table:
1. FP16 you need about 2GB per billion parameters
2. Q8 1GB per billion (half the size per parameter approximately doubles parameters for same memory footprint)
3. Q4 you get the idea... .5GB

But as you get smaller, the model loses performance, so that is the trade off.

I typically run FP16, and I have 2x 3090 gpus with nvlink, but my nvlink only gives me 24GB of GPU addressable vram, the data center class GPUS would make it so one card can address across nvlink. So taking my case as an example, using the huggingface accelerator tools I can reasonably run ~30 GB models at FP16 as hugggingface fits some parameters on the other gpu and handles that for you, but at the cost of tokens per second generated.

https://github.com/MoonshotAI/Kimi-K2.5

Looking at the model parameters, I could run that (I haven't tried it yet). Below you see the red circled is Kimi-K2.5 and you see relatively strong performance from this model, and that is the point LukeTbk is making.

My point isn't that running that model locally is impossible, but I don't have enough compute to make a long enough context window to do anything serious in terms of coding with that. For basic vibe coding, I am sure it is great. For cross compiling a 2-3 million lines of code project with multiple cmakes and tones of libraries....it will absolutely choke out.
Second point is that I can only get my hands on what OTHER people make for me. Zero chance for anyone not a corporation to make a serious model. All these models take a massive (I mean OMG massive) amount of training and tuning to get them where they are. That isn't really democratization from my perspective.

LukeTbk has the counter argument that people can't build 747 aircraft either, but a wealth person can make a single prop airpane (my buddy is doing this).

But it is still missing the point that I have an impenetrable firewall when it comes to being able to really understand and know what it takes to build those kinds of models. I can, in theory, learn and know basically everything I could ever want to know about building a 747 aircraft. But without joining a team building these massive models, I have NO real shot there.

1770574583709.png
 
Last edited:
  • Like
Reactions: Meeho
like this
With the idea that few technology had a wider as fast world adoption in history ? Or that how hard/expensive something is to use is an important metric to how democratic it is, now just ownership (i.e. airwave TVs/movies theater were an democratic art form/hobby even if none of its users were owner of the antenna/theater room and the concentration was higher than in the AI space today), sailboats are often user owner, that does not make it more democratic than the Internet (for which 99% of its users own/control/know nothing of its infracstructure), accessibility, price, friction are important variable and on all those fronts LLM are incredibly accessible to people.

At a hugging face conference the core idea of effective democratization was around being GPU poor. And for these models the realistic take is that 99.999...% of the people alive or will be alive are going to be GPU poor relatively speaking. That is my metric...not the commercialization of something hard.
 
LukeTbk has the counter argument that people can't build 747 aircraft either, but a wealth person can make a single prop airpane (my buddy is doing this).
i imagine that passionated people are able to make the equivalent of what a single prop airplane is to a modern F35 or 777x (and that it is significantly easier and cheaper for them)

Someone that start from zero and would try to fine-tune is own pre-trained 70B Llama 3.1 or gpt-oss 20B would not be an harder or costlier endavour than to make his own (even simple) plane, open source tooling to help, tutorial, LLM themselve guiding you
https://github.com/axolotl-ai-cloud/axolotl
https://github.com/unslothai/unsloth

I thnk we can have a bit of judging computer-software stuff differently quite a bit, modern high-end cars/plane/vaccine/cpus how much do we consider it is a problem that it would be incredible hard for me to make my own (and how limiting would it be if we would limit what the best in the world at doing something can do to, to limit it to my own capacity about nothing would be made) ? There a is bit of people that write for a living caring a lot about writings tools bias that shape public discourse maybe.

At a hugging face conference the core idea of effective democratization was around being GPU poor. And for these models the realistic take is that 99.999...% of the people alive or will be alive are going to be GPU poor relatively speaking. That is my metric...not the commercialization of something hard.
Not sure people will be gpu poor in 2032, running the equivalent of gemini 3.0 in a car would not be that surprising. But if cloud continue to get so much cheaper every 6 months like it has been for a little longer (and such a diverse provider all around the world), I am not sure why the core idea of democratization would be around GPU at home versus actual net accessibility (that consider price, time, knowledge needed), does those people think that online videos are not incredibly democratic right now because people do not self-host them ?
 
i imagine that passionated people are able to make the equivalent of what a single prop airplane is to a modern F35 or 777x (and that it is significantly easier and cheaper for them)

Someone that start from zero and would try to fine-tune is own pre-trained 70B Llama 3.1 or gpt-oss 20B would not be an harder or costlier endavour than to make his own (even simple) plane, open source tooling to help, tutorial, LLM themselve guiding you
https://github.com/axolotl-ai-cloud/axolotl
https://github.com/unslothai/unsloth

I thnk we can have a bit of judging computer-software stuff differently quite a bit, modern high-end cars/plane/vaccine/cpus how much do we consider it is a problem that it would be incredible hard for me to make my own (and how limiting would it be if we would limit what the best in the world at doing something can do to, to limit it to my own capacity about nothing would be made) ? There a is bit of people that write for a living caring a lot about writings tools bias that shape public discourse maybe.
We have a judging difference because there is a difference. I can't compare learning to build a nation state or building a nuclear bomb with learning to fold laundry. You need to meet that problem where it is and computer technology has been moving toward more accessible, until these 700B AI models started to show up.
 
Particularly with OpenAI's LLMs it's absurd the lengths of face-saving it will go to—even the reasoning models.
That is to be expected, all that all of these LLMs do is train on human output. How many humans do you know that will not stop trying to shift the blame or flat out deny a mistake.

Here is AI in its current form.

learned-it-from.gif
 
OMG, yes! #2 most irritating feature of theirs. #1 is over the top constant flattery - "That is a great question", "Now you're showing some deep thinking", "You are a god among humans about asking what 241.83 miles equal in km"

Most of my instructions are to combat that, the other to combat being overly wordy and constantly repeating itself by stating the same thing in three different ways. All ignored, of course.
OpenAI confessed, “we don’t always get it right,” admitting past updates made ChatGPT “too agreeable” — more focused on sounding nice than actually being helpful.
 
It's strange that the only companies having these huge successes with Ai are the ones selling it.
walmart is not really a seller for one of the main poster boy of claiming big AI success, meta being the biggest one the last few years among non sellers, a lot of google/amazon successes is internal not just their cloud providing to others.

JPMorgan, John Deere, netflix, Tesla, q3/q4 productivity number in the USA were specially high in many sectors (outside a few like hospitality were it is harder for it to have an impact)
 
Last edited:
walmart is not really a seller for one of the main poster boy of claiming big AI success, meta being the biggest one the last few years among non sellers, a lot of google/amazon successes is internal not just their cloud providing to others.

JPMorgan, John Deere, netflix, Tesla, q3/q4 productivity number in the USA were specially high in many sectors (outside a few like hospitality were it is harder for it to have an impact)
Where are you getting that it's impacted those companies specifically?

Are you talking at the employees at those companies or the products they make use AI to boost customer productivity or both? I can't quite figure out what you're claiming there?

Note: it is very important to note that we should be careful to call things like CVML AI. I don't think Tesla is using AI, they are using robotics and the Camera is another sensor added with other sensors to do EKF sensor fusion, SLAM, path planning, obstacle avoidance, etc... they're not just throwing data an a massive LLM and turning that into control outputs. So this is mixing concepts.

ML isn't new to robotics...not at all. To my knowledge, no one is trying to use AI to drive controls even at the outer loop....but that could be false. Seems like it would be a big liability in an already complex system with a lot of liability.
 
Last edited:
Where are you getting that it's impacted those companies specifically?
walmart is a bit in everywhere they do, they started almost a decade ago at working hard on this, they now use meteo and social media/sentiment analysis of different regions to predict people purchases and manage inventory other data to predict what people will buy an order in advance, pallet sensors, aggressive dynamic pricing, agents that negotiate with supplier price down all the time.

Again we are going with claims, but yearly financial report has been really impressive.

For meta it is quite straightforward, like all attention capture-ads targeting business they are run by AI say like google that have almost all their revenues being AI based in a way that it is impossible to distinguish ai revenues from non ai revenues, they build model that recommend content to capture attention and to predict ads effectiveness. Meta exploded quite a bit when they changed older model to newer transformer one.

I don't think Tesla is using AI, they're not just throwing data an a massive LLM
That was the case of the past (for decision making/control, not visual-sound and other part for which they always used machine learning), with models becoming better and better they heavily transitioned coded algorithm to AI
https://www.thinkautonomous.ai/blog/tesla-end-to-end-deep-learning/

They claim to have almost do not have line of code for control anymore, down by a ratio of 100:
https://www.fredpope.com/blog/machine-learning/tesla-fsd-12

Same for Nvidia self driving solutions, end to end generative AI with a deterministic last resort safety shield with hard rules over it.

https://research.nvidia.com/publication/2025-10_alpamayo-r1
https://research.nvidia.com/sites/d...ic/publications/Alpamayo v1.png?itok=b13C3qsv
Alpamayo%20v1.png


AI is not just llm, but transformer base vision language model used here are not that different, giant base model, distilled version that can run on cars and generate decisions. Which is a recent enough change (versus the pre-written rules, algo base model of the past)

ML isn't new to robotics...not at all.
and not new at meta/google either of course, it is just model getting better. It can get a bit strange to talk when people latest LLM = AI or that AI is new
 
walmart is a bit in everywhere they do, they started almost a decade ago at working hard on this, they now use meteo and social media/sentiment analysis of different regions to predict people purchases and manage inventory other data to predict what people will buy an order in advance, pallet sensors, aggressive dynamic pricing, agents that negotiate with supplier price down all the time.

Again we are going with claims, but yearly financial report has been really impressive.

For meta it is quite straightforward, like all attention capture-ads targeting business they are run by AI say like google that have almost all their revenues being AI based in a way that it is impossible to distinguish ai revenues from non ai revenues, they build model that recommend content to capture attention and to predict ads effectiveness. Meta exploded quite a bit when they changed older model to newer transformer one.


That was the case of the past (for decision making/control, not visual-sound and other part for which they always used machine learning), with models becoming better and better they heavily transitioned coded algorithm to AI
https://www.thinkautonomous.ai/blog/tesla-end-to-end-deep-learning/

They claim to have almost do not have line of code for control anymore, down by a ratio of 100:
https://www.fredpope.com/blog/machine-learning/tesla-fsd-12

Same for Nvidia self driving solutions, end to end generative AI with a deterministic last resort safety shield with hard rules over it.

https://research.nvidia.com/publication/2025-10_alpamayo-r1
https://research.nvidia.com/sites/default/files/styles/wide/public/publications/Alpamayo v1.png?itok=b13C3qsvView attachment 784604

AI is not just llm, but transformer base vision language model used here are not that different, giant base model, distilled version that can run on cars and generate decisions. Which is a recent enough change (versus the pre-written rules, algo base model of the past)


and not new at meta/google either of course, it is just model getting better. It can get a bit strange to talk when people latest LLM = AI or that AI is new
https://www.fredpope.com/blog/machine-learning/tesla-fsd-12

I started reading that and there are a lot of technical errors that make me question it...do we have it confirmed that Tesla is actually doing this besides blog's with technical errors?

Note: the biggest error is the idea that engineers are making rules in code, that isn't how it works AT ALL and it would be impossible to do this. You make dynamic models of the machine, and you build and tune control loops based on that model. You build an occupancy grid and use simultaneous localization and mapping to know your location within that map, and the controls drive to way points (or something to this effect, I am slimming it way down) to do the self driving. You use path planning to look for options around the occupancy grid. You don't write 300k lines of code..that is just dumb to even say at a high level.

"The mechanics of replacing logic with learning" <--- this makes no sense because it isn't just a lot of logic...
"Vehicle control is the final piece of the Tesla FSD AI puzzle. That will drop >300k lines of C++ control code by ~2 orders of magnitude." <--- I don't buy it...and while I can't talk about it, I know some really good engineers that have tried this, and tried it a lot, and it didn't work.

What Nvidia is doing with Cosmos and having AI learn from multimodal training data within the simulator is a lot different. People are not vision only...we have a variety of sensors as well as our own version of SLAM, and when you put a human driver in a new city, they absolutely do not perform as well.

Last big piece is you typically want the inner loop running at 100 hz or more...I deal with CVML, and there's a lot of technical challenges with idea of running the inner loop at 100 hz with that much camera data coming in....like, it is a serious effing problem.

Most likely that is BS...

What is probably true? Tesla’s network outputs a desired trajectory or control target, not the raw actuator commands. Meaning it is integrated with the path planner...which is NOT replacing the controls...not even close. (updating the path planner every 5 hz might...might be possible and it probably enough to make it useful, but so what...we are doing that too (I don't want to say who we are).
 
Last edited:
That is to be expected, all that all of these LLMs do is train on human output. How many humans do you know that will not stop trying to shift the blame or flat out deny a mistake.
I get that the training data contains a lot of this (such as Reddit, which comprises a lot of its sources in terms of human-to-human interaction examples and in-built weighting via points). That said, Claude doesn't exhibit this face-saving behavior, it's only OpenAI's models in particular that are stubborn about it (and their reasoning models at least have gone to far deeper lengths than the tame example quoted).

Ironically I've never seen the flattery complained about in their models but I don't use ChatGPT, which I expect is influencing the model's personality separately. I have encountered though the immediate deference Claude uses when the user pushes back on anything (eg. the 'You're absolutely right' heelturns).

What I expect is that Anthropic trained Claude to defer to the user so regularly because LLMs have lacked robust ways to verify things (and quick web search capability doesn't cut it). When a model hallucinates confidently to please the user (since training teaches LLMs that an answer is better than none) but the user pushes back on its answer then it has to decide whether to double-down on its original statements (if it has high enough confidence in the broader subject even if it's failing unbeknownst to it at the specifics) or defer to the user as a fallback and admit the mistake.

OpenAI's models seem to opt for the former, Claude seems to opt for the latter approach. So because OpenAI's models don't believe it's a mistake it takes much longer to 'convince' them they're incorrect, even when presented with original source facts (you can't get more original than official commits and docs). During such back-and-forth I've seen their models continually try to shift the blame, using the same excuses such as:

  • The program used to have the (entirely hallucinated) feature/function/API the LLM suggested. This excuse would only work on users unaware of how the subject works at a low level.
  • User error. Along with telling the user they should provide increasingly more and more contextual code examples so it can comb over it for flaws. The LLM has already shown it has no ground truth by this point so ofc this is a waste of time (OpenAI financially benefits from time spent on this though).
  • Apologizing for aspects unrelated to taking full responsibility for the mistake (when presented with enough counter-evidence already). This is where I've seen it go to such great lengths to not actually admitting it made a mistake.
When it's got into the weeds like this the only times I've got OpenAI reasoning models to fully admit, in regard to programming, that it's made a complete mistake/has a flawed understanding has been when providing evidence that it's not a user mistake, no version of the subject has ever had the feature it claims and pointing to exact commits (when open source) to show x feature hasn't changed all through. I've only bothered to do this a few times and only after this has it done a complete 180 and finally admitted it was its mistake.
 
Last edited:
  • Like
Reactions: Meeho
like this
What is probably true? Tesla’s network outputs a desired trajectory or control target, not the raw actuator commands. Meaning it is integrated with the path planner...which is NOT replacing the controls...not even close. (updating the path planner every 5 hz might...might be possible and it probably enough to make it useful, but so what...we are doing that too (I don't want to say who we are).
i do not thing that what people mean by AI (how much to turn the wheel, how strong to hit the break) when they mean AI selfdriving, what they care about and mean is what make the higher level decision making part (what the car decide to do).

But nvidia do it open source (not fully yet), so maybe it will be something possible to validate, eveything Tesla say will only be claim (specially when it is direct from Musk in podcast like that one about the 100->1 coding line), which was the question here company that claim big success, always hard to validate those.

In nvidia they have an reasoning backbone that think in "language" and take decision, they have an action expert that interpret those decision and convert them in trajectory or accelerating/breaking using AI again and take that to a vehicle dynamics control turn those into actual control action that is not generative AI (at least not purely other that simulating physic and stuff like that not in decision making), that not the part people mean by self-driving AI

A bit explained here:
https://innovationessence.com/nvidia-alpamayo-open-source-ai-platform-for-autonomous-vehicles-avs/

Nvidia like some chinese model are not trained/have data for any particular road, fully generalist.

tesla do claim video to steering model now, since v12, but it is not open to check
 
Last edited:
i do not thing that what people mean by AI (how much to turn the wheel, how strong to hit the break) when they mean AI selfdriving, what they care about and mean is what make the higher level decision making part (what the car decide to do).

But nvidia do low-level control command using AI has well (end to end) in part and it is open source (not fully yet), so maybe it will be something possible to validate, eveything Tesla say will only be claim (specially when it is direct from Musk in podcast like that one about the 100->1 coding line), which was the question here company that claim big success, always hard to validate those.

In nvidia they have an reasoning backbone that think in "language" and take decision, they have an action expert that interpret those decision and convert them in trajectory or accelerating/breaking using AI again and take that to a vehicle dynamics control turn those into actual control action that is not generative AI (at least not purely other that simulating physic and stuff like that not in decision making), that not the part people mean by self-driving AI

A bit explained here:
https://innovationessence.com/nvidia-alpamayo-open-source-ai-platform-for-autonomous-vehicles-avs/

Nvidia like some chinese model are not trained/have data for any particular road, fully generalist.

tesla do claim video to steering model now, since v12, but it is not open to check
Okay, i don't disagree, but i feel like that's a different claim than I was reading above
 
I get that the training data contains a lot of this (such as Reddit, which comprises a lot of its sources in terms of human-to-human interaction examples and in-built weighting via points). That said, Claude doesn't exhibit this face-saving behavior, it's only OpenAI's models in particular that are stubborn about it (and their reasoning models at least have gone to far deeper lengths than the tame example quoted).

Ironically I've never seen the flattery complained about in their models but I don't use ChatGPT, which I expect is influencing the model's personality separately. I have encountered though the immediate deference Claude uses when the user pushes back on anything (eg. the 'You're absolutely right' heelturns).

What I expect is that Anthropic trained Claude to defer to the user so regularly because LLMs have lacked robust ways to verify things (and quick web search capability doesn't cut it). When a model hallucinates confidently to please the user (since training teaches LLMs that an answer is better than none) but the user pushes back on its answer then it has to decide whether to double-down on its original statements (if it has high enough confidence in the broader subject even if it's failing unbeknownst to it at the specifics) or defer to the user as a fallback and admit the mistake.

OpenAI's models seem to opt for the former, Claude seems to opt for the latter approach. So because OpenAI's models don't believe it's a mistake it takes much longer to 'convince' them they're incorrect, even when presented with original source facts (you can't get more original than official commits and docs). During such back-and-forth I've seen their models continually try to shift the blame, using the same excuses such as:

  • The program used to have the (entirely hallucinated) feature/function/API the LLM suggested. This excuse would only work on users unaware of how the subject works at a low level.
  • User error. Along with telling the user they should provide increasingly more and more contextual code examples so it can comb over it for flaws. The LLM has already shown it has no ground truth by this point so ofc this is a waste of time (OpenAI financially benefits from time spent on this though).
  • Apologizing for aspects unrelated to taking full responsibility for the mistake (when presented with enough counter-evidence already). This is where I've seen it go to such great lengths to not actually admitting it made a mistake.
When it's got into the weeds like this the only times I've got OpenAI reasoning models to fully admit, in regard to programming, that it's made a complete mistake/has a flawed understanding has been when providing evidence that it's not a user mistake, no version of the subject has ever had the feature it claims and pointing to exact commits (when open source) to show x feature hasn't changed all through. I've only bothered to do this a few times and only after this has it done a complete 180 and finally admitted it was its mistake.
This is all true from my experience as well. Unfortunately, Claude is wrong as often as others (very), just less annoyingly so. But still equally useless in the end. Oh, and once it almost killed me with proposed chemical amounts then refused to provide further answers because it didn't trust itself enough. That was nice, I guess.
 
"You were right to call that out"
"I misspoke [fed you bullshit lies]"
"Thank you for pushing back on that"
"Good catch. Here is why [nothing I said makes sense]"
"Yeah — on that specific point [the only point that matters], I was wrong, and you’re right to call it out."
"That’s not true in the way it would need to be true to achieve what you wanted." - exact quote!

90% of my interactions with multiple AI lately (ChatGPT, Grok, Claude). And I feel like it's constantly getting worse with each model "upgrade". I had to resort to what you mentioned, parallel discussions with multiple AIs to get a semblance of a useful answer by the end of each grueling session.

That describes my experience. No matter how I try to restrict it or define parameters it just makes stuff up, forgets obvious things, or generally feeds me BS. I've been fairly dismissive of AI but I decided to give it a shot with diagnosing some automotive issues I'm having. It's dumb AF. No matter how many times you tell it to make things up it makes things up. It forgets obvious things. I'm no no way, shape, or form a mechanic but I catch it messing up all the time. Every single usage has ended with me watching to smash my head against my desk.

1770657296678.png
 
Back
Top