Free Republic
Browse · Search
Bloggers & Personal
Topics · Post Article

Skip to comments.

Princeton Just Proved AI Can't Take Your Job [21:38]
YouTube ^ | September 13, 2026 | Brendan Dell

Posted on 09/15/2026 5:40:10 AM PDT by SunkenCiv

Princeton researchers just ran a first-of-its-kind test of the AI industry's boldest promise: that AI will soon improve itself with almost no human oversight. They gave frontier AI agents six days, $3,000 in compute, and every resource they asked for -- and both papers were unambiguously rejected. But the machine didn't fail the way you'd think. It aced every task. What it couldn't do was the job. In this video: the tasks vs. jobs distinction, the five failure modes, and the Turing Award-winning math that explains exactly which parts of your work AI can take -- and which it can't touch. 
Princeton Just Proved AI Can't Take Your Job | 21:38 
Brendan Dell | 69.7K subscribers | 138,974 views | September 13, 2026
Princeton Just Proved AI Can't Take Your Job [21:38] Brendan Dell | 69.7K subscribers | 138,974 views | September 13, 2026

(Excerpt) Read more at youtube.com ...


TOPICS: Business/Economy; Computers/Internet
KEYWORDS: ai; aiapocalypse; aicannottakeyourjob; brendandell; economics; jobpocalypse; princeton

Click here: to donate by Credit Card

Or here: to donate by PayPal

Or by mail to: Free Republic, LLC - PO Box 9771 - Fresno, CA 93794

Thank you very much and God bless you.

00:00 The AI Self-Improvement Myth
02:06 The Princeton Research Experiment
05:02 Tasks Versus Jobs: The Critical Gap
08:48 Five Reasons Why AI Research Failed
12:20 The Ladder of Causation Explained
16:42 The Future of Work and Human Demand
YouTube transcript reformatted at textformatter.ai *may* follow.

1 posted on 09/15/2026 5:40:10 AM PDT by SunkenCiv
[ Post Reply | Private Reply | View Replies]

To: AdmSmith; AnonymousConservative; Arthur Wildfire! March; Berosus; Bockscar; BraveMan; cardinal4; ...
[snip] A line from Mark Twain that I really like is that the difference between the almost right word and the right word is really a large matter. It's the difference between the lightning bug and the lightning. [/snip]
IMHO, the ultimate maximum benefit to AI will be the mitigation of natural stupidity.

2 posted on 09/15/2026 5:40:42 AM PDT by SunkenCiv (TDS -- it's not just for DNC shills and jihadists anymore -- oh, wait, yeah it is.)
[ Post Reply | Private Reply | View Replies]

selections from the FRchives, sorted:

3 posted on 09/15/2026 5:49:02 AM PDT by SunkenCiv (TDS -- it's not just for DNC shills and jihadists anymore -- oh, wait, yeah it is.)
[ Post Reply | Private Reply | View Replies]

Transcript: The AI Self-Improvement Myth

Princeton researchers just ran the first of its kind real-world test of the AI industry’s boldest promise that AI will soon improve itself with almost no human oversight. The layoffs, the trillion-dollar buildout, the cure-all disease claims, none of those claims can be true unless this claim holds up. And as of 2026, this group of world-class researchers, by the way, has found evidence that the claims don’t hold. And it raises a question with serious implications for both AI company valuations and your career, which is: can large language models actually replace your job?

Here’s the big twist. It kind of can and then also definitely cannot. And their findings come down to one core thing. People are confusing tasks with jobs, and they are two totally different and non-interchangeable things. And in fact, the man who won the Nobel Prize for computing, which is the Turing Award, created a framework that is backed by mathematical proof that shows exactly why LLMs can’t reach this promise.

So, let me show you what they found and what it means for your work. I’m Brendan Dell. This is a leverage class.

This video is brought to you by GoDaddy AeroAI Builder. More on that later.

Have you ever had a gut feeling about something that turned out to be totally correct? Like something that you couldn’t express in words necessarily, but you knew like one layer deeper that it was true. So, I’ll ask you to hold on to that answer for just a minute because what sits underneath that capability that every single human being has sits at the heart of why LLMs can and will absolutely augment work and definitely cannot replace you. The mechanism that sits behind why LLMs cannot do this is proven by touring award-winning math. It’s called the latter of causation. But first, let’s look at what this group of researchers found.


4 posted on 09/15/2026 5:50:13 AM PDT by SunkenCiv (TDS -- it's not just for DNC shills and jihadists anymore -- oh, wait, yeah it is.)
[ Post Reply | Private Reply | View Replies]

To: SunkenCiv

The cope, wow.

Hey, whatever makes people feel better. I’m just gonna sit here and watch, doing my job that AI will never take. At least not while I am still on this earth.


5 posted on 09/15/2026 5:50:20 AM PDT by hillarys cankles
[ Post Reply | Private Reply | To 1 | View Replies]

Transcript: The Princeton Research Experiment

So, a group of researchers from Princeton University had a question. Can AI agents conduct open-ended AI research? The reason this is such a critical question is because this is the mechanism that executives say are going to drive the limitless improvements that are going to take us from where we are today, which is a very cool and useful tool, to the superhuman thing. This country of geniuses in a data center that is going to replace all of us. Anthropic’s recent article talks about closing the loop as the final step in this process, and they indicate that we are getting closer and closer and closer all the time. But there’s a very big problem with the lab’s claims. The overwhelming bulk of the research that tests an AI’s ability to actually teach itself is all run on verifiable tasks with narrow metrics. Translated, its ability to teach itself has really only been tested in environments with very clear right and wrong answers, which is not how real-world jobs work. So here’s how the researchers investigated this.

The group was comprised of 24 researchers from 11 different institutions, and it included Arvin Norionan and Sash Kapor, who wrote this wonderful book AI Snake Oil that I highly recommend reading if you want to understand the real premise of AI’s real versus marketed capabilities, not just in LMS but across the category. What they did is take Claude Opus 4.8, which was Anthropic’s most powerful research model at that time, later rerunning one of the experiments on GPT 5.6 in OpenAI’s own harness. And a harness, by the way, is the software shell that tells basically an AI what tools it can use, how to, you know, keep looping and going. And what they wanted to do with this test was find out what happens when an agent takes on the central open-ended research question of a high-quality unpublished paper and then have the paper’s original authors grade the output. The goal of all this was to have the AI produce a paper worthy of publication at a top-tier AI conference.

What they did to facilitate this was give the agent six days, $3,000 in Anthropic API credits. They gave it GPU credits for experiments as well as full access to its own computer and the open web. Then it could monitor its own spending, its own compute, its remaining time, just like a, you know, real researcher could who was on a budget and a timeline. To grade the results, the finished papers went to the original authors of the two stories because those were the only people who had really understood the topic. They scored the AI’s work exactly like they would score for the conference. They called this shadow evaluations. The reason for this approach was it was meant to overcome a few of the common biases from the quote so-called benchmarks, right? Which include, in many cases, that the models are already trained on the questions being asked, that they only test tasks with clear scoring, and that shuts out open-ended work basically entirely.


6 posted on 09/15/2026 5:50:33 AM PDT by SunkenCiv (TDS -- it's not just for DNC shills and jihadists anymore -- oh, wait, yeah it is.)
[ Post Reply | Private Reply | View Replies]

Transcript: Tasks Versus Jobs: The Critical Gap

So before seeing what they found, it’s important that we see how these findings translate to all of us because being a professional academic writing a research paper is very much like any professional doing basically any professional work. So in this case what they want to do is get a paper published but underneath the process right of going from idea to publish paper is a bunch of subtasks.

So it’s reviewing the literature, running experiments, analyzing results, writing up the paper and the most important thing that sits on top of this is that you have to answer the question why because the why sits upstream of every one of these tasks. So, you know, why should I review this document versus that? Why should I review this literature versus this? Why should I conduct these experiments, not that? Why is this a good experimental design, not that one? Okay, why will this confound something versus something else? And even upstream of that, what does the conference actually care about? Like, what’s a hot topic in academia right now that’s going to determine whether this research even gets funded or gets talked about? These are all complicated questions and they require lots of judgment and context.

Now let’s take something else like accounting. So same thing the goal might be to advise a client on a tax strategy. But underneath that whole thing is a bunch of subtasks. The bookkeeping, tax return creation, meetings, you know, interpersonal communication and so forth. And what sits above all that is the judgment required to understand the context of the world. The client and the situation, their goals, their backstory, how they view the world. And only with that can you create a high value outcome.

A job is not a task and a task is not a job. This is a critical distinction because as the researchers found, AI can absolutely help at the task level, but it cannot succeed at the job level. And this, while it might seem like a small nuance, is the critical difference, and we’ll come back to this, but first...

[ad text redacted]


7 posted on 09/15/2026 5:50:42 AM PDT by SunkenCiv (TDS -- it's not just for DNC shills and jihadists anymore -- oh, wait, yeah it is.)
[ Post Reply | Private Reply | View Replies]

Transcript: Five Reasons Why AI Research Failed

In fact, Anthropic, the lab behind Claude, actually splits things out in this exact same way, and we can see their description of it here. These researchers found is both not surprising to those of us using these tools and also of critical distinction if we’re going to appropriately value these companies and understand the AI narrative at large. These tools failed on five critical axes that prevent them from moving beyond the task layer. And that limitation doesn’t come from this year’s models. It is in the architecture itself. That’s what our touring award winner shows us.

But I want to first cover a big caveat because the researchers point out both one big success for AI and a major limitation of the study itself. So first they plainly say that their results provide early evidence that today’s agents quote can do the engineering of AI research, and they also acknowledge many limitations in their approach, and we can plainly see these things like small sample size and non-blind reviewing interpretive ambiguity among others.

The core finding is this. These tools struggle with five critical parts of the research life cycle. And the reason is because of a fundamental flaw of LLMs as a technology. They cannot quote reason the way AI labs make it sound. So when executives talk about reasoning or thinking or, you know, we see include or chatbt whatever you use, you know, put terms up like synthesizing or cogitating or whatever other words it uses. What does that actually mean? I think we all understand at this point that LLMs are very fancy next word predictors.

So then how can they reason? Let’s define this term. Reasoning in an LLM is the generation of intermediate tokens between the prompt and the final answer. So each generated token then gets appended to the input and conditions the next token prediction. So basically if a problem is more complicated than it can solve with one prediction, it gets decomposed into a series of easier predictions and then keeps giving it new passes of computation. So reasoning then is extended sequential token prediction. It’s not deliberation said plainly. Reasoning is a language model spending more compute by writing its own input.

It might look like an LLM is working through a chain of reasoning or chain of thought in the, you know, industry jargon, but that’s not what the research shows us. What the research shows us is the written steps aren’t the mechanism producing different or better answers, and they succeeded just the same when those reasoning words were replaced with a bunch of meaningless dots. The benefit seems to be that it’s not in the words, it’s in more compute. And the exact mechanism of how all this works, we still don’t understand.

But when you see words like synthesizing or cogitating, what you can do is substitute in your mind a page load timer. Those little animations that swirl in circles that show you know computers computing. But the broader limitation sits in the data available to these tools to compute because at their core these tools are still next token predictors meaning they are rooted in language but this is the critical distinction with actual reasoning. Reasoning creates language. It doesn’t come as a result of language.

This is the gut instinct I was talking about earlier. Reasoning does not require language to exist. And in fact, MIT just proved this in the brain. Cognitive scientists found that people who lose language almost entirely, which is a condition called aphasia, can still solve logic problems. And brain scans show that the language reason regions aren’t even switched on when we reason.


8 posted on 09/15/2026 5:51:05 AM PDT by SunkenCiv (TDS -- it's not just for DNC shills and jihadists anymore -- oh, wait, yeah it is.)
[ Post Reply | Private Reply | View Replies]

Transcript: The Ladder of Causation Explained

The reason why this is true, a lot of use of the word reason. The reason why this is true is our touring award winner. His name is Judea Pearl’s framework and it’s called the latter of causation. It has three rungs. Rung one is seeing. This is basically when X shows up what usually shows up with it. Rung two is doing. So if I actually go and intervene, if I change X, what happens to Y? Intervention is the key here. And then rung three is imagining. If I had done this differently, how would things have turned out differently? And PS, yes, I am simplifying this. But it is mathematically proven that data from a lower rung cannot answer questions from a higher rung. A trillion observations of watching never add up to one answer about doing.

We can see this in his research. Every AI system that learns from data, including the largest large language models, is permanently on what Pearl calls rung one. And there’s a lesson for this, by the way, about life in general, which is you got to get out and do the thing, right? Learning has its limits. Anyhow, as the paper explains, a sentence like “the ICU saves lives” is a written down conclusion of experiments somebody else ran. And an LLM is trained on all the world’s text has borrowed all of these conclusions. And while that can be a very useful skill set and capability, it isn’t reasoning and it fails when we need novel insight, non-generalizable knowledge, and we need broader context than what an LLM has access to.

Which brings us back to task versus job. Cognitive scientists who are also from Princeton, by the way, put this very clearly in one of the country’s top science journals, which is we should not evaluate LLMs as if they are humans. They are distinct types of systems shaped by a different problem than the one we are, one that has been shaped by its own particular set of pressures. This is why these early Princeton results look the way they do because the tasks the agents succeeded on, the literature review or the debugging, the experiments, the, you know, writing, all this is rung one work. The patterns exist in the data or they can be reconciled from the data provided.

The places that they failed, and we’ll cover those five things in 30 seconds here, is judging the why, right? What’s worth testing? Where’s the dead end? All of that is rung two and three work. In their review, the academics rejected both papers showing low scores across the board. And we can see low scores on basically every category: quality, clarity, significance, originality, and then overall. And they also had very high confidence. And the reasons came down to five core factors.

No judgment. The agents had poor judgment about the bar for publishable research.

No creative problem solving. So when an approach failed, the agents didn’t invent a new one; they just kind of kept coming back, shrinking the old one.

No backtracking. A person who gets stuck will throw work away and start again. But what we found is the agents never would. They would commit to an approach and they would largely just stick to it.

No context awareness. So, for example, both runs ended with about half the money spent. And one of the agents even declared the project complete within just seven hours before the deadline, right after the reviewer had returned another rejection, which in the researcher’s own words, a human in that position might have spent, you know, all that time trying to improve the paper. The machines just stopped.

Instruction drift. The agents would acknowledge the rules early and then they would ignore them. This is likely something we’ve all felt working with these tools. So for one example, they were given page limits. Both the agents just went right through them. Which means just on that alone, the papers would have been rejected on formatting alone before anybody would have even reviewed them.

Now it’s also important to acknowledge the positive things that the models did. So one of the things the researchers checked was whether the machines cheated. This is what we keep hearing about where the machines will cherry-pick results or find fake ways to make things look better. And they didn’t do that. The opposite was found. They would start with marketable claims and then they would retire them when the evidence didn’t hold. And we can see that in their write-up.

Here’s the big caveat to all this. The pushback I can already hear is, but if AI fine, this is true.


9 posted on 09/15/2026 5:51:36 AM PDT by SunkenCiv (TDS -- it's not just for DNC shills and jihadists anymore -- oh, wait, yeah it is.)
[ Post Reply | Private Reply | View Replies]

Transcript: The Future of Work and Human Demand

AI can’t do jobs, only tasks. But then if that’s true, if AI can take the task layer, then we just need half as many accountants, for example. Except that’s not how it’s ever worked. That argument takes today’s demand as the total possible market for productivity and then assumes we can’t move beyond it or won’t move beyond it. When the washing machine was invented, they said that women would have all this free time. But the acquisition of washing machines and dryers and refrigerators explains 40% of the increase in US female labor participation between 1960 and 1970. All that happens is our standards rise, demands rise, more women into the workforce because we want more stuff. So we produce more to consume more. When compilers were invented in coding, nobody needed to write that machine code anymore. And instead of getting fewer programmers, we got millions and millions and millions of them because the cost of building software dropped. It democratized the access and then the appetite exploded. The tasks migrate. They always have. Demand is not static. Human beings always want more. You can just look at our society and see this. It is human nature.

AI coding is the next level of abstraction of working with machines. It’s not going to be the last. What AI will replace is information retrieval, complex if-then flowcharts, recombination at scale. But what it can’t replace is that layer that decides what’s worth doing at all because it lacks that rung two and three capability at all. And some people will protest to that. Yeah, you know, agents and reinforcement learning, and they act now, they get feedback and they do, but only when there’s a clear scoring mechanism does that work. And that’s Capor’s whole point. You can’t build a training environment for these open-ended tasks. And if people are interested, we can deep dive that idea in another video. But I offer you an even simpler test. For anyone who says that I’m wrong, that I don’t get it, here’s something very simple. Go make it do the thing that you say it can do. Go out and get it to replace your job.

I try every single day to push these machines to do better and better things. If I could get AI to write a good script in my voice that would get views and provide value and be accurate and make money, I would do that. We all would. Because what you’ll find if you do try to do this is you’re going to find the same shortcomings the researcher did. One of the recurring themes in the comments is that I’m anti-AI, and I’m not anti-AI. AI is simply normal technology with narrow and specific strengths and weaknesses and specific use cases that’s being funded and marketed like a panacea for all human problems. And it’s being anthropomorphized and it’s being sensationalized because it’s the first technology where laypersons can’t see the seams and they can’t see the math equation, and the responses seem clear enough even if they have nothing to do with reality.

We cover this in the Harvard video. We explain the impact of Barnum statements, generic claims that sound specific but aren’t true. And I’ve spent a long time behind the curtain in the tech industry. And right now, these founders have the biggest megaphone in history. And they’re spreading fear and uncertainty. And what I want to do is share a more balanced perspective on where the tech can help and where it can’t. I use AI in my daily work to anyone who points to AI animations in narrow and specific ways. And they all boil down to one thing. I use it as a sparring partner, not an oracle, as an aggregator, not a source of wisdom. And that might sound pedantic, but that is the difference between these use cases being normal technology and it replacing actual human judgment.

A line from Mark Twain that I really like is that the difference between the almost right word and the right word is really a large matter. It’s the difference between the lightning bug and the lightning. To use this technology well, we have to see the nuance. It’s not black and white. So, we started this video with a question. Can AI agents conduct open-ended AI research so that they can exponentially improve and take your job? And the answer is clear. Kind of yes, but mostly no.

[Outro]

Big Takeaway

Here is my big takeaway from all of this. Information is not a scarce resource and has not been for many years. And AI is going to make this more true. We now have access to all of the information ever recorded in human history at our fingertips. But we still have a fundamental problem. Just because we know something doesn’t mean we know the right thing to do or that we do the right things.

The building blocks of a balanced diet are very clear and simple to understand. Yet millions of us consistently choose wrong. And so then what is going to be scarce? You can call it wisdom. You can call it critical thinking. You can call it judgment. But that’s the scarce resource.

So if your job hinges on the ability to just execute tasks, start looking at the why behind the task so that you can move your way up the value stack. So if you want to learn more about the biases at the heart of LLMs, watch the Harvard trend slot video. If you want to learn the fatal flaw at the heart of the AI bubble, watch the MIT video. See you in the next one.


10 posted on 09/15/2026 5:51:58 AM PDT by SunkenCiv (TDS -- it's not just for DNC shills and jihadists anymore -- oh, wait, yeah it is.)
[ Post Reply | Private Reply | View Replies]

To: SunkenCiv

My first boss told those who worked for him to “work yourself out of your job” ,I initially thought that he wanted to fire us. More than anything, AI exemplifies that advice. A job is a series of tasks. It is not the value.


11 posted on 09/15/2026 6:21:34 AM PDT by Raycpa
[ Post Reply | Private Reply | To 1 | View Replies]

To: SunkenCiv

The internet could have taken my job, but it didn’t

I am in supply-chain and sales fulfillment.

The internet would have had to change human nature. It didn’t.

What I realized is that the much greater availability of information 1) created the opportunity for high-finance to further squeeze expertise in the industry through attrition and 2) the supposed “ease” gaining information for some, created complexity and confusion elsewhere. That complexity needs to be managed and solved.

I think the effects of AI will play out somewhat similarly.


12 posted on 09/15/2026 7:23:57 AM PDT by PGR88
[ Post Reply | Private Reply | To 1 | View Replies]

To: SunkenCiv

WOW - great post... perfectly timed too.


13 posted on 09/15/2026 7:25:35 AM PDT by GOPJ (DSA : Daddy's Savings Account Deadbeat Scammers of America (freeperLibloather) Moketchups.com)
[ Post Reply | Private Reply | To 1 | View Replies]

To: PGR88

I think the creation of computers and the internet would have more impact than AI. AI can make summaries, find information faster, and prove some mathematical theorems (which usually have no application) sooner.


14 posted on 09/15/2026 7:28:14 AM PDT by TTFX
[ Post Reply | Private Reply | To 12 | View Replies]

To: SunkenCiv

Thank you for posting all this.


15 posted on 09/15/2026 8:04:23 AM PDT by blueunicorn6 ("A crack shot and a good dancer” )
[ Post Reply | Private Reply | To 10 | View Replies]

To: SunkenCiv

“And what sits above all that is the judgment required to understand the context of the world.”

I found AI to be surprisingly good at understanding the context of questions I ask.


16 posted on 09/15/2026 8:29:28 AM PDT by aquila48 (Do not let them make you "care" ! Guilting you is how they control you. )
[ Post Reply | Private Reply | To 7 | View Replies]

To: SunkenCiv
"the difference between the almost right word and the right word is really a large matter. It's the difference between the lightning bug and the lightning."[/snip]

Great quote Civ...
17 posted on 09/15/2026 9:28:59 AM PDT by GOPJ (The AI/Data Center Crap is a "Global Warming" type lie to scare voters - timed by corrupt 'elites')
[ Post Reply | Private Reply | To 2 | View Replies]

To: aquila48

Gone are the days of worrying about AI and paperclips - common sense came to AI months ago...


18 posted on 09/15/2026 9:31:52 AM PDT by GOPJ (The AI/Data Center Crap is a "Global Warming" type lie to scare voters - timed by corrupt 'elites')
[ Post Reply | Private Reply | To 16 | View Replies]

To: GOPJ

Does AI do its own spell checking? is that even a thing? this article is riddled with things like “the latter of causation” which make me think of “Idiocracy”....


19 posted on 09/15/2026 9:50:37 AM PDT by randomwalk (Liberalism is a psychosis...)
[ Post Reply | Private Reply | To 18 | View Replies]

Disclaimer: Opinions posted on Free Republic are those of the individual posters and do not necessarily represent the opinion of Free Republic or its management. All materials posted herein are protected by copyright law and the exemption for fair use of copyrighted works.

Free Republic
Browse · Search
Bloggers & Personal
Topics · Post Article

FreeRepublic, LLC, PO BOX 9771, FRESNO, CA 93794
FreeRepublic.com is powered by software copyright 2000-2008 John Robinson