不可名状之失败的常态化
The Normalization of Inexplicable Failures

原始链接: https://www.ihatethefuture.com/2026/09/the-normalization-of-inexplicable.html

作者批判了以“Jev”等 AI 工具为代表的现代大模型驱动开发趋势,即为了追求速度和集成便捷性而牺牲可靠性与问责制。其核心担忧在于“不可解释性的常态化”:由于开发者在未进行严谨测试的情况下急于上线“AI 赋能”的功能,系统故障正成为一种被预料到却未被深究的常态。 当这些黑箱系统出现问题时,开发者和用户往往会不约而同地以“这破玩意儿太烂了”来推卸责任,而不是去探究根本原因。作者认为,对“置信度分数”等功能的依赖往往只是“货物崇拜”,因为用户缺乏基于这些指标采取行动的框架。通过放弃本已更容易自动化的严谨评估,行业正在助长一种抛弃工程精度的文化。我们正在主动构建这样的系统:双方都不愿去检查门后的“障碍”,以一种被默认的、反复无常的混乱取代了问责机制。

这篇文章及其讨论串的核心议题是现代软件开发中“无法解释的故障常态化”现象,而人工智能辅助编程的兴起加剧了这一趋势。 批评者认为,AI 工具助长了“感觉编程”(vibe-coding)——即在缺乏深刻理解的情况下依赖概率性输出,这导致开发者交付的代码平庸、晦涩且不可靠。尽管许多开发者为 AI 辩护,称其为能帮助理清复杂遗留系统的生产力倍增器,但怀疑论者指出,它降低了偷懒式开发的门槛,导致“粗制滥造”的软件激增。 主要观点包括: * **标准流失:** 人们越来越倾向于将软件故障视为不可避免且“无法解释”的现象,而非需要解决的技术问题。 * **责任缺失:** 管理层日益将速度和成本效率置于稳定性之上,并常将不良结果归咎于 AI 以推卸责任。 * **“遗留”陷阱:** 若未经严格测试和人工专业知识验证,AI 生成的代码会制造出“即时遗留”系统,导致后续维护或调试几乎成为不可能。 * **呼吁严谨:** 质量拥护者强调,无论使用何种工具,解决方案始终在于严格的测试驱动开发、清晰的评估指标,以及拒绝为速度牺牲基础工程原则。
相关文章

原文
In a recent episode of President Curtis, the President struggles with opening a door on two separate occasions.
These doors don't work because there are obstructions in the way: a body initially, then roughly a billion dollars worth of gold.


In both instances, in response to the frustration, the character mutters "stupid thing sucks." This is not a reasonable model of doors! Doors should not "suck" inexplicably! I found these moments outrageously hilarious¹ but maybe my stupid brain just sucks.

Jev: Making more doors that suck

The Internet has been abuzz about Jev, an AI model developed by TypeSafe AI, which returns typed values with probability estimates. The important things about Jev are, as far as I can tell:

  • it is fast and cheap,
  • you can build on it quickly,
  • it is fast, and
  • it is cheap.

I'm not particularly good at understanding what technology will get adopted.

I still don't understand² Slack.

Wait.

Do you still have to do the hard part?

Maybe my problem is expecting products to work.

Nobody buying this is running evals. They're just handing opaque questions to Jev and getting opaque responses. Charitably, this allows them to check the "AI-powered" box and ship before Friday, and when this breaks downstream logic, they can always shrug and say "well, AI makes mistakes."

Error budgets? Failure modes? Test sets? All of those can be handled later. The user can discover the failure rate! You've already shipped!

False Confidence

"Oh," the strawman responding to my post responds, "you haven't considered the fact that Jev gives you confidence scores!"

What are you going to do with those?

For you to do something reasonable with confidence scores you need to have both an understanding of the calibration of those confidence scores and also a model for the costs of the uncertainty.

At best, people use confidence scores in a cargo cult manner. At worst, people use them as an excuse for why the API call failed. The model was only 73% confident! That means my error budget is 27%!

Accountability

When a button breaks on a website, I have a model about what should have happened. Somewhere a contract got broken. My DNS is broken. Somebody shipped some slop that has JavaScript syntax errors along only a certain path. A handler threw that wasn't expected to throw. I might not have access to debug just an HTTP status 500, but I expect there to be somebody whose job is to understand why the endpoint is 500ing. The ownership is well-defined albeit opaque³.

For many users, however, the actual experience is roughly just "stupid thing sucks." Software already feels capricious; more failures just change the rate of frustration. It seems like not much of a loss to remove the possibility of following a failure to a concrete cause. Sometimes things just suck.

This leads to a normalization of inexplicability.

My fear is not that more things will fail when things are accelerated by LLM-driven development. They will. They have. Such is part of the price of building things in a novel manner.

My fear is that "sometimes it just sucks" is going to be more and more the accepted endpoint of investigations. This is sad because LLM-accelerated development can indeed help us solve some of these issues. There are plenty of automated QA workflows that aren't written because of lack of engineering time. The very eval that would get you most of the way to replacing (or even justifying the use of) Jev can be a few prompts away.

The tragedy of software engineering today is that we are actively engineering systems where neither the user nor the builder seems to have any interest in checking whether or not there's a body behind the door.

We just shrug and conclude: stupid thing sucks.

This reminds me of a saying that I find similarly hilarious: "sometimes you get the elevator, sometimes you get the shaft." This is also not a reasonable model of elevators!!
The lock-in network effect makes sense to me but I'm still bewildered as to how people standardized on a product that does not even reliably deliver messages. I have seen messages dropped on free, paid, and enterprise instances that only show up weeks later.
Well, maybe "well-defined" is optimistic. After Bill Gates famously failed to download Movie Maker, everybody agreed that it was presumably somebody's problem, just not necessarily theirs. Ideally we can get even this level of accountability without the customer being Bill Gates.
联系我们 contact @ memedata.com