为什么纯文本仍然是我们拥有的最佳技术之一
Plain text is still one of the best technologies we have

原始链接: https://deadparrotbbs.com/why-plain-text-is-still-one-of-the-best-technologies-we-have/

纯文本之所以经久不衰,是因为它以编码形式存储字符,不依赖特定的应用、操作系统或呈现格式。文本文件便于携带、查看、搜索和通过脚本处理,也易于编辑、复制、备份或比较。源代码、日志、配置文件以及 HTML、XML、JSON 和 CSS 等格式都以纯文本为基础。 现代 Unicode 使纯文本能够表示全球范围内的各种书写系统,不过,旧式编码和不统一的换行符偶尔仍会带来问题。即便如此,与被锁定在晦涩专有格式中的文档相比,文本通常更容易恢复。 纯文本本身无法呈现复杂布局、电子表格、数据库或多媒体,因此,许多任务仍然需要专业软件。不过,增加复杂性应该用于解决实际问题,而不应造成不必要的依赖。 Markdown 提供了一种实用的折中方案:它在保持普通文本可读性的同时,增加了标题、列表、链接和强调等格式。即使渲染它的软件消失,其源代码仍然有用。 纯文本还能让用户掌控自己的文件:文件可以保存在本地、同步、进行版本管理或迁移,而无需依赖某个厂商的账户或订阅。它的功能也许很少,甚至显得乏味,但正是这种简洁性,使它能够持久存在、便于迁移,并且极为不易过时。

Hacker News 上的讨论聚焦于一篇文章,该文章认为,纯文本仍然是技术领域最持久、最有效的通信工具之一。它的可移植性、长期可用性、可读性和广泛兼容性,在很大程度上得益于 ASCII 的标准化;不过,Unicode 未来或许也能实现类似的普适性。 评论者指出了纯文本在实际使用中的一些限制,例如换行符不一致、编码问题,以及纯文本看似简单背后隐藏的复杂性。有些人特别提到 Ledger 和 hledger 等纯文本工具,尤其是其中的会计应用,以此证明纯文本至今仍然有用。另一些人则推荐了 Dylan Beattie 的演讲《Plain Text? Really?》,以便更深入地了解纯文本的优势与挑战。 另一条评论串则批评了日益增长的低质量 AI 生成内容,并质疑各个平台是否真正采取了有效措施来遏制这类内容。
相关文章

原文

There are not many computer file formats I would trust to still be readable fifty years from now, but plain text is one of them. That may sound like faint praise. A text file cannot embed a spreadsheet, preserve elaborate page layouts, run a presentation, or provide many of the conveniences we expect from modern applications. What it does instead is store text in one of the simplest and most widely understood forms in computing, and that simplicity is a large part of why it has lasted.

Almost Nothing to Go Wrong

The Unicode Standard defines plain text essentially as a sequence of character codes, without the additional formatting information associated with rich text. Fonts, colors, layout, and similar presentation details belong somewhere else.

From a user’s point of view, the important part is simpler: the file contains the text itself without requiring a particular application to make sense of it. I can create a text file on Linux, copy it to a Windows machine, put it on a web server, open it in a terminal, search it with command-line tools, edit it with any number of programs, or send it to somebody using an entirely different system. Very little about the file depends on how it was created.

That is different from many document formats, where the information is closely tied to the application or family of applications that understands the format. Sometimes that is harmless. Sometimes the format is well documented and broadly supported. In other cases, opening an old file means hoping that suitable software still exists.

Plain text has very little of that baggage.

We Have Been Using It Forever for a Reason

Plain text is not some forgotten technology that needs rescuing. It remains part of the basic plumbing of computing.

Source code is generally text. Configuration files are often text. Log files are commonly text. Unix and Linux systems are full of text files. Much of the machinery behind the web also involves text in one form or another. HTML, XML, JSON, CSS, shell scripts, programming languages, configuration formats, and plenty of other things are built around characters that humans can inspect.

These are structured formats rather than ordinary prose, of course. An HTML document is not the same thing as a note in a .txt file. The similarity is that the contents are still visible. Open the file in a basic text editor and you can see what is there, even if you do not understand every part of it.

That has practical value. When something goes wrong with a text configuration file, I can inspect it directly, copy it, compare versions, search for a particular setting, or make a backup without needing the program that created it. Data that remains intelligible outside its normal application is much easier to troubleshoot and preserve.

Plain Text Plays Well With Others

One of plain text’s greatest strengths is that an enormous collection of tools already knows how to work with it.

On a Unix-like system, a text file can be handed to grep, sed, awk, sort, diff, and many other utilities that have existed for decades. A script can process thousands of files without pretending to be a human clicking through menus.

The advantage is not limited to the command line. Graphical editors, programming environments, note-taking applications, file managers, search tools, browsers, and countless other programs can work with text too. The real advantage is that the information is not tightly coupled to one interface.

Modern software often takes the opposite approach. An application may determine where your information lives, how it is organized, how it synchronizes, how it is searched, and what other software is allowed to touch it. Those systems can provide useful features, but they also make the application increasingly central to the data.

Plain text leaves more of those choices to the user.

Markdown Found a Useful Middle Ground

The obvious weakness of plain text is formatting. Headings, emphasis, links, lists, quotations, and other structures make documents easier to read, and a basic .txt file does not provide a standard way to express most of them.

Markdown is a practical compromise. It is a plain-text format for structured documents, drawing on conventions that were already familiar from email and Usenet. It was introduced in 2004, and variations of it are now used for software documentation, websites, notes, books, and many other kinds of writing.

Its real strength is that the source remains useful even if the Markdown processor disappears. A heading still looks like a heading. A bulleted list still resembles a list. A link still contains both its description and destination. The formatting syntax is visible rather than buried inside a binary structure.

That is a good example of adding useful capability without giving up the main advantages of plain text. The rendered output may be convenient, but the source file remains ordinary text.

Text Ages Surprisingly Well

Longevity may be plain text’s strongest argument.

Computer history is full of abandoned file formats and applications. Sometimes recovering an old document means locating obsolete software, finding an import filter, running an emulator, or converting through several intermediate formats.

Text files are not completely immune to compatibility problems. Older files can have character-encoding issues, and differences in line endings have caused irritation for years. Still, these are usually manageable problems compared with trying to recover information from an obscure proprietary format.

If I find an old text file, there is a very good chance that some program on a current computer can open it. Even if the formatting is crude or the encoding needs attention, the words themselves are generally recoverable.

Modern Unicode has also made plain text far more capable than the old idea of ASCII text containing little more than English letters, numbers, and punctuation. Plain text today can represent writing systems and characters from around the world while retaining the same basic idea: the file contains encoded characters rather than a proprietary presentation format.

Simple does not have to mean primitive.

It Is Also Easy to Own

The advantages of plain text become even more noticeable as software moves toward accounts, cloud storage, synchronization services, and subscriptions.

A text file can simply exist in a directory on my computer. I can back it up with the rest of my files, synchronize it however I want, put it under version control, copy it to another machine, or store it on a server I control. None of that requires the company that made my editor to stay in business or continue supporting a service.

That does not make every note-taking service or cloud application a bad idea. Specialized applications provide features that plain text alone cannot provide easily. Collaboration, databases, embedded media, complex formatting, and relationships between different kinds of information all have legitimate uses.

There is no benefit in forcing every kind of data into a .txt file. The useful distinction is whether the additional complexity solves a real problem or simply becomes another dependency.

When plain text is sufficient, it removes a surprising number of things that can later get in the way.

The Limitations Are Real

Plain text is not the answer to everything.

I would not want to replace a photograph with a text description of its pixels, and I do not want a financial workbook turned into a pile of numbers that requires me to reconstruct all the formulas manually. Rich documents exist because presentation and structure sometimes matter.

Applications also provide useful abstractions. A database can enforce relationships that a directory full of text files cannot. A word processor can handle page layout that would be tedious to reproduce manually. Specialized software exists for good reasons.

The problem is not using richer tools. It is assuming that richer tools are automatically better even when the job does not require them. In a surprising number of cases, the simplest adequate representation remains the most durable one.

Boring Is a Feature

Plain text is one of those technologies that becomes almost invisible because it works so reliably. It has no company trying to increase engagement, no account requirement, no subscription tier, and no service that can be discontinued when the business model changes.

It is portable, inspectable, searchable, scriptable, easy to back up, and remarkably resistant to obsolescence. Those qualities are not exciting, but they are useful, especially over long periods of time.

After decades of watching software and file formats come and go, I have developed a fair amount of respect for technologies that do less and survive longer. Plain text does not solve every problem, and it should not. What it does provide is a stable way to store information without requiring very much from the software around it.

For something so basic, that is a considerable achievement.

联系我们 contact @ memedata.com