Can AI take your job?
A popular video is making the rounds under the title “Princeton Just Proved AI Can’t Take Your Job”. It builds on a recent study by researchers at Princeton University, known for its high-ranking academic programs in mathematics and public policy, who tested whether AI agents (AI systems that work through a task on their own, using tools) can do original research. The video is well made, and part of it is right. But “proved” is doing a lot of work. We went to the primary sources, and the picture is more interesting, and less comfortable, than either the video or the AI hype.
What this means for you
The video’s most useful idea holds up: a job is more than a bundle of tasks. In the study, the researchers gave AI agents a research project to work through alone. The agents handled the routine work, but struggled with judgment: deciding what is good enough to publish, when to drop an approach that isn’t working, and why a step matters in the first place. Those are the areas where people add the most value.
But don’t treat that as a guarantee. The honest reading of the evidence is that today’s AI is strong at tasks and weak at open-ended judgment, that this gap is real, and that it may narrow. For entry-level workers, the pressure is already noticeable.
That makes the practical question personal: which parts of your work are tasks, and which are judgment? Where can you move toward the second?
What the Princeton study actually tested
The paper is Can AI agents conduct open-ended AI research? by a team led from Princeton, with co-authors from many institutions. The researchers gave frontier AI agents six days and thousands of dollars of compute to tackle the central research question of two unpublished conference submissions. The original authors then graded the results.
Both papers were rejected. The agents handled the engineering but made no real progress on the research questions. The team names five failure modes, among them poor judgment about what counts as publishable, weak backtracking from dead ends, and drifting away from instructions.
That is a genuine and useful finding. It is also a small one: two papers, five runs in total, reviewers who weren’t blind to the fact that an AI wrote the work. The authors call it “early evidence” and say the results may not generalize to other kinds of research. And the paper says nothing about employment. It is a study about AI research, not about your job.
The “mathematical proof” that isn’t
The video’s strongest-sounding argument is Judea Pearl’s ladder of causation: seeing, doing, imagining. Since language models learn from data, the video says, they are “mathematically proven” to be stuck on the lowest rung.
The underlying math is real. It says that purely observational data cannot, on its own, pin down answers about interventions or counterfactuals. But that is a statement about a specific kind of data, not about every system trained on text and then refined by acting in the world.
Pearl himself has said as much. In a 2023 interview, he said the ladder restrictions “do not hold anymore because the data is text, and text may contain information on levels two and three.” He still sees real limits. Reinforcement learning, he said, sits at “level one and three-fourths,” because machines can’t infer much about moves they haven’t tried. That is a considered position, not a proof that language models can’t do causal reasoning.
The empirical picture is mixed, which is different from “impossible.” In one benchmark study, GPT-3.5 and GPT-4 reached 97% on causal discovery and 92% on counterfactual reasoning tasks, with unpredictable failures. On a stricter formal test (CLadder), the best model scored about 83% on the lowest rung and 62% on counterfactuals. Performance drops as you climb, so the video isn’t wrong that higher rungs are harder. It just isn’t a wall.
“It’s just compute” isn’t a rebuttal
The video argues that AI “reasoning” is only extra computation, citing a paper showing meaningless filler dots can work as well as written-out steps. That paper is real, but its authors note the trick needed dense, specific training and hasn’t been shown to occur naturally.
More importantly, “extra compute” and “reasoning” aren’t opposites. A theoretical result shows that written intermediate steps give transformers serial computation they otherwise lack. Interpretability work on Claude found the model carrying out multi-hop steps internally before answering. And a Meta paper showed that reasoning in a continuous internal space beat written-out steps on two logic benchmarks (97% vs 77.5% on one), while doing worse on grade-school math. Reasoning without words is possible for machines too. Nothing in the neuroscience of aphasia says otherwise.
Jobs: reassuring on average, not for entry-level workers
On the job question, the video says AI automates tasks, not jobs. Today, on average, the data agrees. The Yale Budget Lab’s May 2026 analysis found no clear sign of AI effects on the labor market yet.
But averages hide the people at the edges. A Stanford analysis of payroll data finds employment among workers aged 22 to 25 in AI-exposed occupations is about 19% lower than it would be had it kept pace with young workers in less-exposed jobs, as of June 2026. Experienced workers show no comparable gap. The authors are careful: the pattern is descriptive, not proof of cause. They find that interest rates and remote work don’t explain it, but the gap shrinks once education is taken into account.
And “tasks, not jobs” describes where we are, not where the trend is heading. On the Remote Labor Index, which tests agents on real freelance projects, the best result went from 2.5% in late 2025 to 15.8% in mid-2026. Most work is still not automated, and the benchmark’s own authors warn that automated grading flatters newer models. Both facts are true at once.
The washing machine argument cuts both ways
The video’s optimistic case is that demand always expands: new technology creates more work than it removes. Economist James Bessen documented exactly this in textiles, steel and autos. The same paper also shows the second half of the story: after a period of growth, demand saturated and employment fell. Elastic demand is a phase, not a law. Whether writing, analysis and code are closer to 1850 textiles or 1990 steel is an open question that nobody has measured yet.
We also couldn’t find a reliable source for the popular claim that compilers created millions of programmers. And the appliance story is real, but it’s a story about time freed at home, not about displaced workers finding new jobs.
Working out which parts of your work are tasks and which are judgment is exactly what we do in coaching, with AI as a tool that takes over the tasks so you have room for the judgment. If you’d like to explore that for your role, let’s talk.
Sources are linked inline. Several are preprints or lab-authored studies that have not been peer reviewed, and the labor-market findings are early.