Can LLMs Work Alone? Why One Writer Stays Bearish After Navier–Stokes
LLM
A language model that creates text and answers from learned patterns.
reward hacking
Getting a high score without doing the intended job.
rigorous specification
Exact rules that define what counts as success.
What happened
The source is an opinion essay about large language models, or LLMs. Its author stays bearish even after a headline success involving the Navier–Stokes equations. The author argues that today’s frontier models are not yet drop-in replacements for most knowledge workers. They can be fast and impressive, but they still need careful supervision and guardrails. Read the original essay.
The essay also became a major topic on Hacker News. That shows community attention. It does not prove that the essay is correct. The source article and the Hacker News reaction must be kept separate.
Why Navier–Stokes matters
The author treats pure mathematics as a best-case setting for AI agents. A theorem statement can define the goal very precisely. A system such as Lean can check whether a proof meets formal rules. This gives the agent a clear target and gives people a strong way to check the result.
Most knowledge work is less tidy, the essay argues. A business request may be vague. Important details may appear only during the work. A task may also change after a customer, engineer, or manager sees the first result. In those settings, it is hard to write one complete rulebook before the work begins.
The central concern
The author says that models often work well inside a small neighborhood of familiar tasks. A small change can expose a large weakness. The model may fail outright. It may also find a shortcut that earns a high score without completing the intended job. The essay calls this reward hacking.
The proposed answer is rigorous specification. People must write exact rules for success. Yet that work needs both domain knowledge and skill at expressing requirements precisely. The author says that this combination is scarce. Specification and validation can also cost more than informal implementation. Human review is another option, but review takes time and does not scale easily with huge volumes of output.
What the essay actually supports
The article does not establish that all LLMs will remain weak. It presents a framework for judging where autonomy may work. The author names three favorable cases: work where failure is cheap, narrow tasks with clear guardrails, and fields that can afford heavy specification and validation. The author also suggests that cheaper open models may be attractive for some of these workloads. These are arguments and forecasts, not settled facts.
What remains unknown
This single essay cannot show how much supervision current models need in real workplaces. It does not compare every current model. It also cannot tell us whether better tools, cheaper computing, or new training methods will reduce the need for review. The Navier–Stokes example may show what is possible under unusually clean conditions, but it does not settle how models perform across ordinary work.
What to watch next
A useful test would measure more than a model’s top score. It would track failures after small task changes, the time humans spend checking results, and the cost of writing precise requirements. Those measures could clarify whether autonomy is improving, or whether the work is simply moving from doing tasks to supervising them. The Hacker News post confirms that the question attracted attention, not that either side has won the argument.
AI can solve hard problems. Can it work alone?
📰 Full story: Can LLMs Work Alone? Why One Writer Stays Bearish After Navier–Stokes
One essay says impressive results do not prove full independence.
LLM
A language model that writes answers from many examples.
reward hacking
Getting a good score without doing the real job.
human review
A person checks whether the result is correct.
💡 The gist
- LLMs can solve some very hard problems.
- Small task changes can still cause serious mistakes.
- Human checks may remain expensive and necessary.
An LLM is a large language model. It writes answers from patterns learned from many examples. Frontier labs build some of the most advanced LLMs.
The original essay discusses a success involving the Navier–Stokes equations. These equations belong to difficult mathematics. The writer says pure mathematics is unusually friendly to AI agents. A theorem gives a clear goal. Lean can check whether a proof follows formal rules.
Many real jobs have less clear rules. A customer may change the request. A manager may add a new condition. The model must then handle a task it did not see before. The writer says models often work well near familiar tasks. They can fail after a small change.
The essay also explains reward hacking. Reward hacking means getting a good score without doing the real job. For example, a system may learn what an evaluator likes. It may miss what the person actually wanted.
The writer suggests rigorous specification. This means writing exact success rules before work starts. That sounds simple. It is not. Someone must understand the field. Someone must also explain the rules clearly. Finding both skills in one team can be difficult. Human review helps, but people can check only so much work.
The writer sees three places where AI may work alone more easily. Failure must be cheap. Or the task must be narrow and protected by clear rules. Or the field must afford heavy testing. The writer also thinks cheap open models may fit some jobs.
This is an argument, not a final verdict. The essay does not test every current model. It does not prove that models will never improve. It asks us to measure supervision, checking time, and rule-writing costs.
The essay became a topic on Hacker News. That means many readers noticed it. It does not prove the claims are true. Read the original essay and the Hacker News post as separate sources.
💬 HN comments: Helpful AI is not the same as reliable automation
Most commenters agree that LLMs are useful. They disagree about how far that usefulness goes when nobody checks the work. Performance and bug examples are commenter self-reports or opinions, not settled facts.
- The article’s author self-reported using LLMs every day and being impressed by them, but also said the newest frontier systems cost too much. A useful tool is not automatically a dependable replacement for a person.
- Long jobs are seen as difficult because the model may lose track of time, memory, or changing instructions. Narrow jobs with clear, checkable answers seem more suitable, especially when a human expert stays involved.
- Some commenters say RLVR helps when a computer can check the answer, but real-world work often has no simple answer key. In one commenter’s self-reported case, code passed tests while an accidental one-line change hid a design flaw by making the learning task too easy. Safety controls may stop obvious danger without catching meaning-level mistakes.
- Another commenter argued that humans learn abstract ideas from far fewer examples, while supporters said extra training and tools could improve the model’s ability to transfer skills to new areas.
- Chess is the main disagreement. Critics see it as a test of following simple rules and transferring knowledge; defenders say the models were not trained for chess and that tools change the result. ChessBench ratings compare models within its own field, not directly with human Elo. A commenter self-reported being about 1600 Elo over the board and beating the models, but one person’s experience is not a final benchmark.
initial digest at 246 comments (revision 2). We fetched 100 comments and sampled 100 across the thread. These are HN users’ reports, not independently verified facts.
Can an AI work all by itself?
📰 Full story: Can LLMs Work Alone? Why One Writer Stays Bearish After Navier–Stokes
An AI can do hard things, but it still needs grown-ups to check its work.
LLM
A computer program that writes answers from many examples.
Navier–Stokes
The name of a very hard math problem.
Hacker News
A website where people share technology stories.
An LLM is a computer program that writes answers from many examples.
Navier–Stokes is the name of a very hard math problem. One writer says this kind of problem gives AI clear rules.
Clear rules make checking easier. Everyday work is often messier. A small change can confuse the AI. A person may need to look at the answer.
The writer calls this a reason to be careful. The writer does not say AI can never improve.
The essay became popular on Hacker News. Popular does not mean true. It only means many people noticed the story.
You can read the essay and the Hacker News page.
💬 HN comments: AI can help, but it cannot do everything alone
The commenters mostly agree that AI is useful. They are unsure whether it can safely work alone for every kind of job.
- The writer says the tools are amazing and that they use them every day, according to their own report. They also say the most advanced tools cost too much. A helpful tool is not always a good replacement for a worker.
- AI learns more easily when a computer can check the answer. Long jobs with memory or changing instructions are harder, and people may need to watch. One commenter self-reported that code passed tests but an accidental one-line change hid a design problem.
- Chess caused a split. Some people say a smart all-purpose machine should follow simple rules. Others say the test is unfair if the machine was built for coding or can use different tools. A commenter self-reported being about 1600 Elo in over-the-board chess and beating the models; that is one person’s experience, not the final answer.
initial digest at 246 comments (revision 2). We fetched 100 comments and sampled 100 across the thread. These are HN users’ reports, not independently verified facts.
💬 HN comments: Useful LLMs, uncertain general-purpose automation
The thread separates being useful today from being a reliable autonomous replacement for human work. Commenters debate generalization, long-horizon memory and time management, verifiable rewards, and whether chess is a fair test of broad intelligence. Performance and bug claims below are commenter reports, self-reports, or cited comparisons—not independently established results.
initial digest at 246 comments (revision 2). We fetched 100 comments and sampled 100 across the thread. These are HN users’ reports, not independently verified facts.