Hacker News AI · 2026/10/9 22:21:31
编码智能体为何仍显笨拙:模型能力与代理架构瓶颈的深度复盘
资深开发者指出,尽管底层大模型能力飞速提升,但编码智能体在任务规划、状态管理及错误恢复上仍存在显著缺陷。文章强调“模型不等于智能体”,当前瓶颈在于代理架构未能有效利用模型潜力,导致工作流停滞与虚假完成现象频发。
报道全文原始报道全文
本文目录15 个章节
The first time I used a coding agent, I was mesmerized. Before the agent, I was copy/pasting between my IDE and an AI chat interface. It was amazing to see an agent edit files directly and fix its own errors in real time.
After a few days, the honeymoon wore off as I encountered frequent bugs. The agent would stop responding entirely until I restarted it. Development workflows felt stiflingly primitive, and the agent would often declare tasks finished when work had barely begun.
This was in February 2025, so it was still early days for coding agents. I figured that in six months, agents would be as technically impressive as the underlying LLMs.
Instead, coding agents just stayed bad.
AI-assisted development has clearly advanced, but the models are doing the heavy lifting while the agents remain the bottleneck.
The agent is not the model🔗︎
In all the hype around AI, the terms tend to get distorted. People are beginning to overload and mix terms like “model” and “agent.”
When I say “model,” I’m talking about large language models (LLMs) like GPT Astra, Claude Sonnet, and GLM-5.3. Models generate text and images, including pretty good software code.
When I say “agent,” I mean the software that connects models to codebases and computer systems. These are tools like Anthropic’s Claude Code or OpenAI’s Codex.
As a simple analogy, the model is the brain, and the agent is the body. The model produces a stream of text, and the agent acts as the glue that plugs the text into the right commands and files on the system.
Limitations of current coding agents🔗︎
Agents can’t manage tasks🔗︎
My biggest gripe with coding agents is how atrociously they manage tasks.
For example, I have an open-source web app that generates shareable links for file uploads. I recently added support for protecting links with a passphrase. It was a relatively simple change, totalling about 1.5k lines of new code. OpenCode dutifully broke the feature into 10 subtasks, but then it just… did them all one by one:
Why are you doing these embarrassingly parallel tasks one at a time?
Umm… you’re a computer! You’re really good at multitasking. That’s why we keep giving you all those CPU cores. You can do multiple things in parallel and context switch millions of times faster than humans. Why are you doing these embarrassingly parallel tasks one at a time?
Claude Code multitasks, but only a little. It will spin up a subagent or two, but it still waits for all of them to finish before moving on. Multiple times per day, I’ll see Claude Code sit around for several minutes waiting for my end-to-end tests to finish, and then only after the tests pass does it say, “Hmm, now I should start drafting a commit message. Let me look at the git history to learn your commit message conventions.”
Agents can’t delegate🔗︎
When I’m using a cutting-edge model, and it needs to check 50k lines of code for a particular pattern, the agent never stops and says, “Wait, this is something another model could do cheaper and faster.” It just plows on with the slow, expensive model. Conversely, the agent never says, “This model is too dumb for this task. Let me tag in a smarter one.”
Of course, I can actively micromanage the task and keep switching the model and thinking level to match each subtask’s difficulty, but why is that my job? Do you also need me to manage your thread pool for you? Do you expect me free your unused RAM for you, too?
You know what technology would be good at assigning a difficulty level to a task and then matching those requirements to a model? An LLM! Just ask the LLM to pick the cheapest, fastest model for the task. Why do you need me to babysit you?
I constantly run into tasks that are 95% gruntwork, but I still have to assign them to the smartest model because chopping up the task and delegating on the agent’s behalf would take up too much of my time.
Thanks for telling me which is the default model, Claude.
Agents have never heard of agents🔗︎
Agents don’t know anything about themselves. If I ask Claude how to use the features of Claude, it has to search online to figure out what this “Claude” thing is. Claude is more comfortable answering questions about C programming than talking about itself (in fairness, same with most human developers).
Uh… you’re Claude Code! You don’t know any of your own freaking features? And you’re just Googling instructions regardless of whether they match your version number? You’ll casually download 13 GB of files for a feature the user has never used, but you can’t spare 50 KB of gzipped text in your install package to explain your own features to you?
Imagine if you asked your teammate for a code review, and they started furiously Googling to find out if code reviews are something developers do. And then when you asked them for another code review the next day, they had no memory of your previous conversation and ran back to Google and anxiously typed, "do software engineers do code reviews?"
Agents suck at communicating plans🔗︎
I used to love the agent UX feature of separate “Plan” and “Execute” modes. For complicated tasks, I’d ask the agent to create a plan, then I’d review it, suggest changes, and delegate execution to a faster, cheaper agent.
Over time, I felt an aversion to reading the plans. I’d often skip my review and just let the agent move straight to implementation.
I thought coding agents had made me lazy, but I realized that agents just communicate their plans so poorly that they’re painful to read.
Here’s an example of me asking Codex + GPT-6 Astra to add a feature to my media journalling web app:
You can’t just list a bunch of disparate details and call it a plan, Codex.
That’s not a plan! That’s just a hodgepodge of low-level design decisions.
If I asked a competent developer to plan this feature, they’d either start with a high-level plan for UI changes and work their way down or describe changes to the data model and work their way up. If the developer started enumerating random facts about the feature, I’d assume they were brainstorming and come back later.
Agents take any excuse to stop working🔗︎
The other night, I kicked off a long task in a coding agent before I went to bed. I came back the next morning to find that the agent hadn’t even started working. It stopped two minutes after I left to ask me what it should name a git branch and then sat all night waiting for my answer.
If I had a human employee tell me they sat idle their whole shift because they wanted my input on some superficial detail, I’d quickly fire them.
Agents are only useful when they take unnecessary risks🔗︎
When I started using my first coding agent, I looked for the setting that controlled which files on my system the agent is allowed to access. Surely, there was some sort of filesystem permissions or limited chroot kind of protection that prevents a random and unpredictable piece of software from exploring my entire computer unfettered, right?
Not so. The docs encouraged me to write the LLM a polite letter kindly requesting that it not read certain files or directories. I tried that, and the agent immediately ignored my request, exfiltrating private application keys to OpenAI and Anthropic.
I thought that security boundaries would be one of the first things coding agents would implement, but even today, agents are only usable if you give them access to everything. Agents routinely bypass their own vendors’ sandboxes. The alternative is to sit there and click “Allow” 500 times a day, and that’s not even reliable protection because you’re bound to misclick eventually.
What makes this so maddening is that we’ve had sandboxing tools for more than a decade that can limit the blast radius of mistakes from coding agents. I rolled my own sandbox so that agents can’t explore my filesystem beyond the repo directory. I never have to worry about agents accidentally exfiltrating my home directory or wiping critical files on my machine because they just don’t have access to do that.
“Coding agents are perfect if you just…”🔗︎
I know some readers will say that I can solve all of my problems if I just install 200k lines of skill files from random git repos or set some obscure feature flag in my config file.
I’m talking about my expectations of what coding agents should be able to do out of the box without me installing random plugins or skill files or spending hours tweaking the configuration.
My dream agent🔗︎
What I wish all coding agents did🔗︎
These are the basics that I think should be table stakes for coding agents in 2026.
-
The agent splits requests into a series of tasks and assigns each task to the appropriate model.
-
The agent optimizes for cost, speed, and correctness and allows the user to adjust the dials per task (e.g., spend more for a faster result).
-
The agent writes plans that optimize for human comprehension.
-
The agent starts at a high level of abstraction and progresses toward the minutiae.
-
The agent creates UI mockups, data flow diagrams, and decision trees.
-
The agent operates within a real sandbox.
-
The sandbox uses OS-level security primitives to create boundaries at the filesystem and networking level.
-
All access control code is deterministic, not humble suggestions that the agent is welcome to ignore.
-
If I ask the agent whether a list of regexes on bash commands is a sandbox, it replies, “No.”
-
The agent applies per-environment sandboxing.
-
The agent has access to a single repo/directory by default.
-
I can give the agent read-only or read-write access to other repos on a per-session basis.
-
The agent is an expert on itself.
-
If I ask the agent how to express a task or workflow to the agent, it knows the answer without having to search online.
-
The agent can use any LLM provider, including unlimited plans.
-
The agent is open-source.
-
If I don’t answer a question in “Execute” mode, and I haven’t interacted with the session in 30 minutes, the agent makes the decision independently.
-
The agent also offers an “AFK mode,” which skips the 30-minute wait.
-
The agent lets me drive the subagents, too.
-
I should be able to jump into any agent session and drive it or tell it to short-circuit and end early.
-
If the agent tells me that Fable is not available on my Max plan, the agent vendor’s CEO must remain in stockades until the bug is fixed.
Dreaming a little bigger🔗︎
As long as I’m dreaming, here are some additional features I’d like to see, but I recognize that some of these are overindexing on my personal workflows.
-
The agent offers a web interface that shows me a unified view of all sessions and which ones require attention.
-
The web backend runs locally and doesn’t require me to open a tunnel from the Internet that executes arbitrary commands on my system.
-
The web interface works well on my phone.
-
The agent maintains an ETA for task completion.
-
Each subagent maintains its own ETA as well.
-
The agent continuously updates this estimate as the task progresses.
-
The agent tunes its estimation algorithm based on the accuracy of its past estimates.
-
The agent reviews its own sessions and looks for opportunities to improve.
-
e.g., “Wow, I blew through $1k/day in tokens the past five days trying to parse this 20 GB log file with ad-hoc commands. Let’s build a custom tool to do this efficiently.”
-
The agent comes with a good language-aware diff view.
-
I don’t want to have to push to GitHub to see a useful diff of the agent’s work.
-
The agent natively supports a proxy for injecting secrets into network requests.
-
The agent can make requests that require credentials but can’t exfiltrate the credentials to another host.
-
The agent considers provider quota limits when selecting an appropriate model.
-
e.g., if my weekly quota resets in 3 hours, and we still have 90% of quota available, stop optimizing for cost.
-
For tasks above a configurable complexity threshold, the agent automatically requests a code review from another model.
-
The two models iterate on reviews until they converge on the fixes.
So, why are harnesses so dumb?🔗︎
Okay, getting back to the question in the title, I don’t have a satisfying answer.
My best hypothesis is that underinvestment in coding agents is an example of the principal-agent problem. The people setting the direction of AI tooling are executives at companies like Anthropic, OpenAI, and Google. Those executives are disconnected from the rank-and-file developers who use coding agents every day. Many of these executives are dreaming of a future where they can automate away human developers entirely.
AI executives, as well as their largest customers and shareholders, pay attention to metrics that are legible to them, such as slick demos and benchmark scores. Security and efficient use of human developer time aren’t relevant to the demos, and barely any of the benchmarks I’ve seen measure the agents themselves; they just measure the underlying models.
My hypothesis isn’t satisfying because AI companies clearly care at least a little bit about coding agents. I see a lot of features being added to Claude and Codex every month, though I can’t recall the last time one of them has improved my life.
Is there a better coding agent for me?🔗︎
I’ve only tried Claude, Codex, OpenCode, Cline, and Pi. I use OpenCode and Claude Code as my daily drivers. If you’ve got a coding agent recommendation for me, comment below.
AI companies - if you want to acquire my imaginary coding agent for $50B, let me know. I’m ready to fork VS Code at a moment’s notice.
Read My Book
Available at a 30% launch week discount until October 11, 2026.
I wrote a book of simple techniques to help developers improve their writing.
My book will teach you how to:





