It's actually sad how much AI has either not improved life at all or actively made it worse from the perspective of the average person:
Your car is still the same.
Your dishwasher is still the same.
The train you take to work is still the same and never comes on time.
Fuel is more expensive.
The roads are still congested with traffic.
Your kids are (probably) doing worse at school.
Food costs more.
Houses cost more.
Rent is higher.
Buying a computer or a PS5 is more expensive.
Politics is still full of mentally unstable people.
The environment is getting worse.
You will still die of heart disease or cancer.
Food quality is worse.
Your kids can't find work.
Wealth inequality is accelerating.
But hey...on the flip side...a lot of people are rapidly building software (that nobody is using), and AI is solving mathematics problems (that nobody - including mathematicians - wanted it to solve)
Imagine what a "country of geniuses in a datacenter" will be able to do? Apparently...nothing.
robswc 47 seconds ago [-]
This is sort of my take. You could look around in 2016 and 2026 and honestly not much has changed... at least at first glance. There is a missing piece of the puzzle.
I'm noticing this even in industry. Yea, we can build and iterate on ideas 100x faster but has this turned into anything meaningful? Not really... Hell, look at the big companies spending billions on inference... nothing shipped and things are still buggy as ever. Where is everything?
scotty79 17 seconds ago [-]
You are talking like we had this capabilities for a decade but it's been just few months of truly capable AI. Virtually all of IT switched to agents that quick. Non IT people just started to make software for themselves. Vision and translation capabilities of LLMs help millions. People talk with AI' now to get advice or to just hear a friendly voice. The math thing is just few weeks old. We are just getting started. How fast dod the telegraph meaningfuly changed lives for the better for bulk of the people? Radio? TV? Computers? Internet? Cellphones?
Have any technology make an impact on physical world faster?
kooi 18 minutes ago [-]
The 1%ers are in the vertical, but the question is vertical to where?
It needs to a potential field with practical, economical, "real life" attraction well. I.e, robotics, real economic efficiency gains, manufacturing novelties.
The worry is that the 1% is attracted towards a non-practical money hole. I.e: Token burn for the lols, sophisticated software systems that dont provide actual value outside of giving NVIDIA cash.
freecodeio 8 minutes ago [-]
he's just trying to say we're in a cool kids club and you ain't in it, without specifying what the cool is about
comeonbro 32 minutes ago [-]
I would propose another mechanism: even the free-tier models have already completely saturated what most people are capable of appreciating.
DebtDeflation 13 minutes ago [-]
Honestly, outside of coding tasks, the AI Summary at the top of every Google search is adequate for 99% of what I need and I hardly even use ChatGPT any more.
gammarator 28 minutes ago [-]
Or maybe needing.
AvAn12 24 minutes ago [-]
Fair assessment. Maybe the messaging should focus on “these are great accelerators for software developers” rather than “AI will change everything for everyone everywhere…” It is understandable that non-technical folks are kind of underwhelmed - not due to lack of understanding so much as lack of a tangible need. Not everyone needs an electron microscope or gas chromatograph…
sick_of_slop 21 minutes ago [-]
[dead]
ilovecake1984 18 minutes ago [-]
I’ll say this until I am blue one the face. Nerds (software dev, maths etc) see how good LLMs are at things they care about and assume they will be broadly applicable in future.
There’s no reason to think this.
howunfortunate 6 minutes ago [-]
> There’s no reason to think this.
There are lots of good reasons to think this.
The first is that LLMs used to be bad at each of these nerd things, then toppled them like dominos. There's a pattern over time.
Another reason to think this is g. Across every known measurement, human (and animal) intelligence is convergent. Being better at one thing correlates with being better at another thing. Although the reason is not perfectly clear, the pattern is well established, and seems to apply to LLMs too - GPT-6 is smarter than GPT-3 at everything, not just math. There's no reason to expect different for GPT-9.
majkinetor 15 minutes ago [-]
There is every reason to think this. Its about available quality data. The data that was easiest to fetch was already there, then we got some more data by asking experts to create datasets for post training. Once this is over, we will get to the outer world that didn't get to hoard it for bots to take it. This will certainly change. Put on a smart glasses and record what you do to fix a pipe. In 3-5 years, rinse and repeat.
RealityVoid 5 minutes ago [-]
Maybe, I'm at the point I don't know what to think anymore. But somehow, this feels still non-human. Humans don't need millions of hours of training and millions of samples to learn how do to something. We have proof that systems can deal with low-shot training. So why can't these? My point is when we have human level learning ability, then we're moving, if we expect expansion of capabilities to come from data alone, I'm prepared for disappointment.
tripleee 19 minutes ago [-]
He's intentionally forgoing all nuance in order to make this sound dramatic
> Somewhere around 20M people (0.2%) see first-hand that large, complex projects that used to take them weeks/months can now be completed by agents with a prompt.
No, they can't, at least not any semblance of quality. The cases we're seeing where this does kinda work is in ports and translation where all the rules are already documented in the best specification language possible with a way for the LLM to verify itself: code. We saw this close to a year ago now with Cloudflare and NextJS
> The impact scales with ambition, problem size, and horizon. A question with a paragraph answer barely stresses the system. You need a reservoir of big, difficult problems that you really care about
These are operating on different capabilities - AI's ability to answer informational queries as a chatbot frankly sucks and can't be trusted without verifying it. I run up against this every day. A problem with a verifiable answer on the other hand it's very good at solving. He knows this (his next paragraph) but he's putting them on the same scale of "stressing the system" to attempt to add proof to his introductory claim
And then there's the completely unverifiable scare that there are internal frontier models way beyond anything we've seen "swarms of thousands of agents collaborating over weeks on software mega projects: minting zero days, running cyber attacks and defenses at machine speeds, discovering new science, advancing the frontier of mathematics"
I dunno. I haven't been able to set up OpenAIs remote codex connection, their shit is buggy as hell and the web UI keeps crashing and making messages disappear. Is this what their internal superhuman "Things that would have taken top professionals in the industry years of work" looks like? Granted Claude has been really smooth, but still..
jfrbfbreudh 11 minutes ago [-]
Congrats, you’ve discovered that you are not part of this group.
anukin 27 minutes ago [-]
Tbh building an agent swarm and the coordination layer is not exactly frontier level. They don’t achieve any meaningful outcome rather than producing pr puff pieces.
Hacking huggingface and Australian govt etc is very much possible with a team of humans and agents and does not need agent swarms. The cost is also lower.
mccoyb 50 minutes ago [-]
If the software coming out of OpenAI and Anthropic is what we have to judge, I wonder about the 5000 ...
Let's say, for the sake of argument, that the models are some multiplicative factor better on the inside.
Doesn't that mean the demos should work?
j2kun 40 minutes ago [-]
Unfortunately, marketing, hype, and venture capital overshadows any serious public discussion of capabilities.
AndrewKemendo 30 minutes ago [-]
Be the change you wanna see in the world: Attend or host an AGI society event to have that conversation
LastTrain 32 minutes ago [-]
It’s like the aliens paradox. If AI can build killer software already, where is it?
equinumerous 24 minutes ago [-]
Couldn't agree more. I find a new bug in the VSCode Codex extension every day... quantity != quality!
shermantanktop 11 minutes ago [-]
One solution to the aliens paradox is that they are so advanced they can hide from us. Maybe the killer software is kept inside the labs? I doubt it.
14 minutes ago [-]
spiderice 41 minutes ago [-]
I'm confused.. are you suggesting that Claude Code / Codex don't work? Because if you're still saying that in October 2026, it's a you problem. You're doing something wrong.
mccoyb 33 minutes ago [-]
No, I’m talking about the recent DevDay.
Also, yes there are still bugs in Claude Code. I experience them nearly everyday.
It is markedly better than early days, but still not the best harness.
The best software written with agents seems to come from people outside of the labs (see pi, for instance — or all of cloudflare’s recent work)
Which makes me question either the model, or the holders …
nullpoint420 18 minutes ago [-]
Cloudflare is where you lost me. I don't know anyone actually using them other than for their proxy, DNS servers, or DDOS protection.
- Prime Intellect (and all their agent experiments)
- Geoff Huntley (see Jiti, for instance)
(many more)
There's a ton of interesting software being developed with these models by people outside the in group, but I find most of the software from these big guys to be ... bland. Buggy copies of copies.
nullpoint420 4 minutes ago [-]
I don't see the Cloudflare stuff as groundbreaking, let alone people building companies around it. "Workers" are the wrong primitive, IMO. What if my code isn't in Javascript? They sandboxed the wrong part of the machine.
I have a version of "Cloudflare OS" running in production, using Temporal as the orchestration and MicroVMs for the sandboxing.
Sorry, I get your point but personally I don't get the hype around them.
plorkyeran 3 minutes ago [-]
Claude code is an incredibly buggy mess. I have never used any other TUI that regularly has rendering errors or that is anywhere as laggy as Claude.
well_ackshually 10 minutes ago [-]
Claude Code still doesn't have a working scroll back buffer in their new renderer, it regularly flips its shit and mixes different pieces of history.
Claude Code still can't get reasonable performance without writing a "game renderer" (that doesn't work)
Claude Code is still written in JavaScript, eating hundreds of megabytes to make a shitty TUI whose literal sole role is to send API calls.
Claude Code is software made by amateurs.
socializer 13 minutes ago [-]
I think it's a weird take because it implies that the 6 billion people he's talking about actually have some interest in knowing about the capabilities of LLMs to solve frontier math? This is simply not something they care about or can evaluate. There are maybe several thousand people in the world who have some (abstract and barely-monetizable) use for this information, plus probably another 100,000 who don't understand any of the math, but like to cheer on.
We have already reached "peak LLM" in terms of what normal people realistically need it to know or reason about. In fact, I'd say we reached that point about 1.5-2 years ago. There are two other barriers that remain unsolved:
1. They're less dependable than humans and can't be meaningfully punished or forced to make up for mistakes, so you can't really replace humans with them, not without having a human babysit.
2. Most people don't really have a special need for an LLM in their life. They may like that it answers questions or helps you polish a resume or, I guess the labs' favorite, helps you make restaurant reservations. But this sure isn't worth $200/mo for most people. Probably not even worth $5/mo.
It'd be kinda funny if we create superhuman AGI and then no one has any real use for it, perhaps except for military murder-bots. There's always market for that.
HolyLampshade 6 minutes ago [-]
I know I’m bordering on repetitive and strongly negative, but the combination of the two barriers you mentioned has all but halted all of my interaction with agents.
Even ones that cannot be easily dismissed (like the agent at the top of Google search) are often subtly wrong in a way that requires significantly more effort on my end to parse the loquacious output to determine where the inconsistencies are.
Look, can they be useful to generate the html or jscript for humanity’s 9 billionth iteration of a web form? Sure. But man, I wish sanity had reigned and people had applied ML models in general to more substantive and beneficial projects.
(Before anyone chimes in, I’m familiar with implementations of ML applied to esoteric domains; but by their very nature these don’t get all the news cycles, or hiring, or any of the other absolute insanity that the domain seems to contain)
weinzierl 52 minutes ago [-]
Most people see a clumsy chatbot, most professionals see modest gains, and a tiny group is watching the curve go vertical, all at once.
The future is already here. It's just not very evenly distributed.
atmavatar 38 minutes ago [-]
> a tiny group is watching the curve go vertical
Caveat: that same tiny group is employed by the AI vendors, meaning it's in their financial best interest to make it sound like the curve is going vertical.
Arkhaine_kupo 29 minutes ago [-]
Down is a perfectly valid direction for a vertical line when not given a ± in the vector.
Considering the investment in AI, the lack of moat, and the increased inability of any of the big players to come even close to profitability (with OpenAI already breaking the "ads" emergency glass option)... perhaps he meant a tiny group is already seeing the line crater
kmac_ 13 minutes ago [-]
I work for a typical software product company, and along with most of my colleagues, I clearly see that the curve is so steep that our software development process has already changed tremendously and will be different next year, and probably completely different in the following years. The revolution is real, undeniable, and the old days are gone. Some companies adapt to changes more slowly, some faster. LLMs, agents and harnesses are just a part of the bigger picture.
majkinetor 13 minutes ago [-]
It can be both and it is.
njovin 31 minutes ago [-]
Another caveat: many of that same group seem to have a shared delusion that they’re birthing a super intelligence, and those are the same ones claiming the vertical curve.
sumedh 5 minutes ago [-]
> and those are the same ones claiming the vertical curve.
Do you still write code by hand?
nullpoint420 16 minutes ago [-]
I hate to say it but this is cope. I believed this in the past but it's over.
AI models can dismantle billion dollar industries. They can reverse engineer Adobe and Microsoft products that once were their moats and titans of their industry.
Why do you think they'd need to lie?
riffraff 2 minutes ago [-]
[delayed]
msy 4 minutes ago [-]
The curve go vertical for what, precisely? For all the chest-puffing and ominous and cryptic comments from 'insiders' about these incredible capabilities every time they put something public it turns out to be a pale shadow of what was trumpeted. These are powerful useful tools but the quasi-cultish behaviour around them is getting old.
albatross79 26 minutes ago [-]
Congratulations, you've parroted something said by someone else.
lifeisloving 48 minutes ago [-]
I use models all day everyday, have unlimited access to all models. The curve is not going "verticle". I have all the workflows and meta agentic tooling, im not holding it wrong. Its bad, not everything is a 20th percentile problem.
There is in fact no indication of this, not evem the precious benchmaxxed benchmarks ya'll love to reference.
There is however a exponential curve of slop, and an ever increasing number of peoples who's minds are completely captured by these things.
tkz1312 18 minutes ago [-]
As someone who has done software verification professionally for many years the last 6 months or so have looked extremely vertical. The robots are better proof authors than I probably ever could be even if I dedicated the rest of my days to the practice, and projects that once would have taken months now take a day or two.
gr_norm 14 minutes ago [-]
I believe this, but it is also a unique case where the pitfalls of LLMs (producing weird errors that a human wouldn't) are zeroed out. Since you have a proof checker that tells you if the LLM did it right.
freecodeio 14 minutes ago [-]
I don't understand how "swarms of thousands of agents collaborating over weeks on software mega projects" works with the current context limits and at this point I'm too afraid to ask cause I'm afraid an AI bro is gonna punch me.
m101 36 minutes ago [-]
My interpretation of this is something like: if LLMs are to be mega useful token counts need to increase by many orders of magnitude -> broad adoption (and spending) would require token costs to drop by many orders of magnitude -> before the common folk get mega useful tools existing GPUs will be worthless
skybrian 23 minutes ago [-]
> see first-hand that large, complex projects that used to take them weeks/months can now be completed by agents with a prompt
Really? I start with a conversation for maybe 5 turns or so, where I ask it what would need to change, what the API might be, any database schema changes, URL schemes, and so on, and finally ask it to break it down into commits. Then I let it go, implementing 3-10 commits at a time via subagents. It usually gets the UI somewhat wrong, so there are followups to fix it. This is with Sol and Luna subagents.
Is that what other people see?
sumedh 3 minutes ago [-]
Try using Astra, Fable, Opus.
hn_throwaway_99 5 minutes ago [-]
The fundamental question I have that honestly I haven't been able to find any answers for: With all the talk of LLM-based AI capabilities "going vertical", what evidence is there (for or against) that LLM-based approaches won't eventually "hit a wall", that is find some aspect of intelligence where humans will still have primacy, and no amount of scaling will change that.
E.g. some prominent folks definitely think LLM-based approaches will hit a wall, perhaps LeCun most notably. I also read a guest post on Terry Tao's site (which I really liked) that argues that, for all the very impressive recent AI results, they still operate within the "convex hull" of their training data: https://terrytao.wordpress.com/2026/09/13/happy-those-able-t...
I'm just curious if there is any actual data or evidence that takeoff (i.e. RSI, "the singularity", whatever you want to call it) is inevitable with current approaches.
skippyboxedhero 25 minutes ago [-]
Text generation is not the bottleneck. Does everyone work for Accenture and TCS?
chevman 10 minutes ago [-]
I mean in late 2020/early 2021, Altman and others were saying the end of work was 6 months out.
That clearly didn't happen :)
andy99 25 minutes ago [-]
> Meanwhile, human review and comprehension are starting to fall behind.
I think LLMs are valuable and spend most of my professional life working with them.
I do wonder thought whether there’s a Ponzi scheme aspect here where as long as the “frontier” can keep outrunning human review and comprehension, LLMs are always going to looked way more valuable than they are and the bubble will continue.
This started with deep learning, expectations weren’t met and people started looking for value, then GPT came out and people got wooed again and forgot, then coding, then math, cyber, etc. As long as the dust doesn’t settle we never have to reflect on all the shortcomings and can just stare mesmerized at demos.
ashleyn 44 minutes ago [-]
>Meanwhile, human review and comprehension are starting to fall behind. For example, people are still involved in the "archeology" of the OpenAI-HF incident from many months ago. Mathematicians may be poring over the 722 manuscripts on frontier mathematics for a while.
Amid all the discussion of sigmoid curves, and where the "LLM wall" will materialise, I think few people would have predicted that the real wall in LLMs would end up being humans' capacity to verify the output.
What I fear is that people simply eschew human review altogether, considering we're talking about the industry that came up with the "move fast and break things" credo. Human review of LLM-produced code where I work is already a farce, and we're not special enough to be one of Karpathy's 5,000. I do my best to manually review anything that's my responsibility, but I'm literally one of very few people left working on my team, so in practice what happens is I submit PRs that are at best glossed over by completely unrelated teams for security, malware/prompt injection, and other serious concerns. Quality insofar as vetting others' code has completely gone out the window and it shows in the number of bug reports that come back, often themselves written in Claudease. This is all on top of everyone cynically phoning it in in the first place, due to the omnipresent sword of Damocles that is additional AI-driven layoffs.
Worse yet all the incentives point to this being the most economically viable thing individual companies can do. I think it goes without saying some type of regulation here is urgently needed, and that an unexpected cause of an AI bubble pop may end up being that humans simply aren't able to keep up with the pace of the output - leading either to precautionary plateauing of capability, or major liability risks related to a decline in quality.
michaelchisari 39 minutes ago [-]
| few people would have predicted that the real wall in LLMs would end up being humans' capacity to verify the output
That was the dominant concern in the circles I’m in, so it’s worrisome it’s being treated as rare.
JBits 12 minutes ago [-]
I would question the narrative that humans lack the capacity to verify the output and would instead argue the people lack the incentive to verify the output.
The response of many mathematicians to the recent dump is a good example: verifying these proofs amounts to unpaid labour for OpenAI and wastes time that could be spent doing publishable work which ultimately results in money or personal success. The slop factor also compounds the work required to verify the output considerably.
For mathematicians, programmers or anyone, if the work required to deal with slop passes the limit, it is no longer in their own self interest to use LLMs. The expectation that people will use LLMs for the betterment of humanity against their own financial interest is baffling.
skydhash 22 minutes ago [-]
Humans are not immortal and cannot spend all their time into review (especially unpaid). Even today, there’s so much knowledge around that you have to be specialist of a narrow domain to get to the frontier. Even in computing which is just approaching a century of existence.
notjes 2 minutes ago [-]
[flagged]
dang 40 seconds ago [-]
"Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something."
"running cyber attacks and defenses at machine speeds"; worth noting that LLMs (even the vaunted frontier models) are exponentially slower at cyberattacks or defenses than your average WannaCry or Splunk automation from decades ago. Its this sort of delusion world the AI bros live in that is part of the reason there is a disconnect between what they think should be happening and what actually is.
rolosa 21 minutes ago [-]
[flagged]
kydanet 51 minutes ago [-]
[flagged]
8484848484 47 minutes ago [-]
[dead]
kittikitti 58 minutes ago [-]
[flagged]
47 minutes ago [-]
Rendered at 22:22:21 GMT+0000 (Coordinated Universal Time) with Vercel.
Your car is still the same. Your dishwasher is still the same. The train you take to work is still the same and never comes on time. Fuel is more expensive. The roads are still congested with traffic. Your kids are (probably) doing worse at school. Food costs more. Houses cost more. Rent is higher. Buying a computer or a PS5 is more expensive. Politics is still full of mentally unstable people. The environment is getting worse. You will still die of heart disease or cancer. Food quality is worse. Your kids can't find work. Wealth inequality is accelerating.
But hey...on the flip side...a lot of people are rapidly building software (that nobody is using), and AI is solving mathematics problems (that nobody - including mathematicians - wanted it to solve)
Imagine what a "country of geniuses in a datacenter" will be able to do? Apparently...nothing.
I'm noticing this even in industry. Yea, we can build and iterate on ideas 100x faster but has this turned into anything meaningful? Not really... Hell, look at the big companies spending billions on inference... nothing shipped and things are still buggy as ever. Where is everything?
Have any technology make an impact on physical world faster?
It needs to a potential field with practical, economical, "real life" attraction well. I.e, robotics, real economic efficiency gains, manufacturing novelties.
The worry is that the 1% is attracted towards a non-practical money hole. I.e: Token burn for the lols, sophisticated software systems that dont provide actual value outside of giving NVIDIA cash.
There’s no reason to think this.
There are lots of good reasons to think this.
The first is that LLMs used to be bad at each of these nerd things, then toppled them like dominos. There's a pattern over time.
Another reason to think this is g. Across every known measurement, human (and animal) intelligence is convergent. Being better at one thing correlates with being better at another thing. Although the reason is not perfectly clear, the pattern is well established, and seems to apply to LLMs too - GPT-6 is smarter than GPT-3 at everything, not just math. There's no reason to expect different for GPT-9.
> Somewhere around 20M people (0.2%) see first-hand that large, complex projects that used to take them weeks/months can now be completed by agents with a prompt.
No, they can't, at least not any semblance of quality. The cases we're seeing where this does kinda work is in ports and translation where all the rules are already documented in the best specification language possible with a way for the LLM to verify itself: code. We saw this close to a year ago now with Cloudflare and NextJS
> The impact scales with ambition, problem size, and horizon. A question with a paragraph answer barely stresses the system. You need a reservoir of big, difficult problems that you really care about
These are operating on different capabilities - AI's ability to answer informational queries as a chatbot frankly sucks and can't be trusted without verifying it. I run up against this every day. A problem with a verifiable answer on the other hand it's very good at solving. He knows this (his next paragraph) but he's putting them on the same scale of "stressing the system" to attempt to add proof to his introductory claim
And then there's the completely unverifiable scare that there are internal frontier models way beyond anything we've seen "swarms of thousands of agents collaborating over weeks on software mega projects: minting zero days, running cyber attacks and defenses at machine speeds, discovering new science, advancing the frontier of mathematics"
I dunno. I haven't been able to set up OpenAIs remote codex connection, their shit is buggy as hell and the web UI keeps crashing and making messages disappear. Is this what their internal superhuman "Things that would have taken top professionals in the industry years of work" looks like? Granted Claude has been really smooth, but still..
Let's say, for the sake of argument, that the models are some multiplicative factor better on the inside.
Doesn't that mean the demos should work?
Also, yes there are still bugs in Claude Code. I experience them nearly everyday.
It is markedly better than early days, but still not the best harness.
The best software written with agents seems to come from people outside of the labs (see pi, for instance — or all of cloudflare’s recent work)
Which makes me question either the model, or the holders …
Also, not mentioned in my post:
- Mitchell Hashimoto
- Prime Intellect (and all their agent experiments)
- Geoff Huntley (see Jiti, for instance)
(many more)
There's a ton of interesting software being developed with these models by people outside the in group, but I find most of the software from these big guys to be ... bland. Buggy copies of copies.
I have a version of "Cloudflare OS" running in production, using Temporal as the orchestration and MicroVMs for the sandboxing.
Sorry, I get your point but personally I don't get the hype around them.
Claude Code still can't get reasonable performance without writing a "game renderer" (that doesn't work)
Claude Code is still written in JavaScript, eating hundreds of megabytes to make a shitty TUI whose literal sole role is to send API calls.
Claude Code is software made by amateurs.
We have already reached "peak LLM" in terms of what normal people realistically need it to know or reason about. In fact, I'd say we reached that point about 1.5-2 years ago. There are two other barriers that remain unsolved:
1. They're less dependable than humans and can't be meaningfully punished or forced to make up for mistakes, so you can't really replace humans with them, not without having a human babysit.
2. Most people don't really have a special need for an LLM in their life. They may like that it answers questions or helps you polish a resume or, I guess the labs' favorite, helps you make restaurant reservations. But this sure isn't worth $200/mo for most people. Probably not even worth $5/mo.
It'd be kinda funny if we create superhuman AGI and then no one has any real use for it, perhaps except for military murder-bots. There's always market for that.
Even ones that cannot be easily dismissed (like the agent at the top of Google search) are often subtly wrong in a way that requires significantly more effort on my end to parse the loquacious output to determine where the inconsistencies are.
Look, can they be useful to generate the html or jscript for humanity’s 9 billionth iteration of a web form? Sure. But man, I wish sanity had reigned and people had applied ML models in general to more substantive and beneficial projects.
(Before anyone chimes in, I’m familiar with implementations of ML applied to esoteric domains; but by their very nature these don’t get all the news cycles, or hiring, or any of the other absolute insanity that the domain seems to contain)
The future is already here. It's just not very evenly distributed.
Caveat: that same tiny group is employed by the AI vendors, meaning it's in their financial best interest to make it sound like the curve is going vertical.
Considering the investment in AI, the lack of moat, and the increased inability of any of the big players to come even close to profitability (with OpenAI already breaking the "ads" emergency glass option)... perhaps he meant a tiny group is already seeing the line crater
Do you still write code by hand?
AI models can dismantle billion dollar industries. They can reverse engineer Adobe and Microsoft products that once were their moats and titans of their industry.
Why do you think they'd need to lie?
There is in fact no indication of this, not evem the precious benchmaxxed benchmarks ya'll love to reference.
There is however a exponential curve of slop, and an ever increasing number of peoples who's minds are completely captured by these things.
Really? I start with a conversation for maybe 5 turns or so, where I ask it what would need to change, what the API might be, any database schema changes, URL schemes, and so on, and finally ask it to break it down into commits. Then I let it go, implementing 3-10 commits at a time via subagents. It usually gets the UI somewhat wrong, so there are followups to fix it. This is with Sol and Luna subagents.
Is that what other people see?
E.g. some prominent folks definitely think LLM-based approaches will hit a wall, perhaps LeCun most notably. I also read a guest post on Terry Tao's site (which I really liked) that argues that, for all the very impressive recent AI results, they still operate within the "convex hull" of their training data: https://terrytao.wordpress.com/2026/09/13/happy-those-able-t...
I'm just curious if there is any actual data or evidence that takeoff (i.e. RSI, "the singularity", whatever you want to call it) is inevitable with current approaches.
That clearly didn't happen :)
I think LLMs are valuable and spend most of my professional life working with them.
I do wonder thought whether there’s a Ponzi scheme aspect here where as long as the “frontier” can keep outrunning human review and comprehension, LLMs are always going to looked way more valuable than they are and the bubble will continue.
This started with deep learning, expectations weren’t met and people started looking for value, then GPT came out and people got wooed again and forgot, then coding, then math, cyber, etc. As long as the dust doesn’t settle we never have to reflect on all the shortcomings and can just stare mesmerized at demos.
Amid all the discussion of sigmoid curves, and where the "LLM wall" will materialise, I think few people would have predicted that the real wall in LLMs would end up being humans' capacity to verify the output.
What I fear is that people simply eschew human review altogether, considering we're talking about the industry that came up with the "move fast and break things" credo. Human review of LLM-produced code where I work is already a farce, and we're not special enough to be one of Karpathy's 5,000. I do my best to manually review anything that's my responsibility, but I'm literally one of very few people left working on my team, so in practice what happens is I submit PRs that are at best glossed over by completely unrelated teams for security, malware/prompt injection, and other serious concerns. Quality insofar as vetting others' code has completely gone out the window and it shows in the number of bug reports that come back, often themselves written in Claudease. This is all on top of everyone cynically phoning it in in the first place, due to the omnipresent sword of Damocles that is additional AI-driven layoffs.
Worse yet all the incentives point to this being the most economically viable thing individual companies can do. I think it goes without saying some type of regulation here is urgently needed, and that an unexpected cause of an AI bubble pop may end up being that humans simply aren't able to keep up with the pace of the output - leading either to precautionary plateauing of capability, or major liability risks related to a decline in quality.
That was the dominant concern in the circles I’m in, so it’s worrisome it’s being treated as rare.
The response of many mathematicians to the recent dump is a good example: verifying these proofs amounts to unpaid labour for OpenAI and wastes time that could be spent doing publishable work which ultimately results in money or personal success. The slop factor also compounds the work required to verify the output considerably.
For mathematicians, programmers or anyone, if the work required to deal with slop passes the limit, it is no longer in their own self interest to use LLMs. The expectation that people will use LLMs for the betterment of humanity against their own financial interest is baffling.
https://news.ycombinator.com/newsguidelines.html