Pandoc 二十年
Twenty Years of Pandoc

原始链接: https://pandoc.org/twenty-years-of-pandoc.html

2006年8月3日,约翰·麦克法兰(John MacFarlane)发布了Pandoc的首个版本。这是一款基于Haskell语言的文档转换器,旨在利用抽象语法树在Markdown、HTML和LaTeX等格式之间进行转换。Pandoc最初是因其对Haskell的兴趣而产生的个人项目,现已发展成为学术和技术写作的全球标准,支持超过50种输入格式和76种输出格式。 在二十年的历史中,Pandoc历经了多次重大架构变革,包括集成Lua过滤器、建立强大的模板系统以及参与Commonmark规范的开发。该项目的成功得益于Haskell语言的人体工程学特性和类型安全性,使其能够在有限的资源下实现复杂且可靠的转换。如今,Pandoc拥有超过600名贡献者和数百万用户,已成为确定性、高效文档处理的重要工具。尽管大语言模型的兴起为文本转换带来了新的可能性,但麦克法兰强调,Pandoc的能源效率、可靠性和精确性确保了其持续的相关性。在庆祝20周年之际,麦克法兰回顾了这个成功消除无数学术和专业琐碎工作的项目,深感欣慰。

这篇 Hacker News 帖子庆祝了通用文档转换工具 **Pandoc** 诞生 20 周年。来自学术界和软件工程等领域的广大用户纷纷称赞该工具的可靠性、多功能性,以及它通过中间表示法(Intermediate Representation)处理“N 个读取器和 M 个写入器”这一巧妙的设计选择。 讨论的重点包括: * **工程质量:** 贡献者认为 Pandoc 的稳定性和简洁的输出归功于其基础设计和 Haskell 语言的使用,这吸引了一个高质量且深思熟虑的开发者社区。 * **实际效用:** 许多评论者强调了其关键的应用场景,例如通过 Git 对二进制文档(如 .docx)进行版本比对、自动化静态网站生成,以及将混乱的 HTML 规范化以便进行高质量转换。 * **持久价值:** 尽管大语言模型(LLM)兴起,用户仍强调 Pandoc 在确定性的批量文档处理工作流中依然具备不可替代的优势。 该帖是对创作者 John MacFarlane 的集体感谢。许多用户表示,二十年来,该工具已成为他们职业生涯、科研工作和技术流程中不可或缺的基础。
相关文章

原文

On August 3, 2006, I uploaded the first version of pandoc to my website, releasing it under the free GPL license. Pandoc 0.1 consisted of about 3000 lines of Haskell code, with no dependencies aside from GHC’s standard library. It could convert Markdown, reStructuredText, HTML, and LaTeX documents into any of these formats, plus RTF or S5. I had no idea at the time that this would just be the first of over two hundred releases over the next twenty years; that the project would become the most popular program written in Haskell; that I would spend countless hours on bug-fixes, improvement, and project management; that I would collaborate with programmers in many other countries; that pandoc would come to support over fifty document formats; that it would allow automatic generation of citations and bibliographies; that it would become integrated into academic writing tools like Quarto and Jupyter Notebook; that it would be installed on millions of computers around the world.

How did this happen? I want to take advantage of pandoc’s birthday to tell the story of the project, as best I can remember it.

John MacFarlane
August 2, 2026

Greg Restall. Of an introductory book on Haskell, he said: “I’m glad that this wasn’t the textbook in my introductory computer science course, long ago in 1986. If it were, I may have fallen in love with computing and never become a philosopher” (consequently.org).

Intrigued by this (and not heeding Restall’s warning about the potential effects on my future philosophical productivity), I read A Gentle Introduction to Haskell to get a basic understanding of the language. But the only way to really learn a programming language is to write something in it. I saw that Haskell was good for writing parsers and compilers, and it came with a really nice parser combinator library (parsec), so I decided to write a Markdown parser.

At that time, there were implementations of Markdown in Perl, Python, Ruby, and PHP; they all transformed Markdown directly to HTML through a sequence of regex transformations. Pandoc took a different approach. It parsed the Markdown using parser combinators and produced a real abstract syntax tree (AST), which it could then render to HTML or another format. This was a more reliable architecture (avoiding many quirks of the regex versions). It was also a more extensible one: by writing N parsers (“readers”) and M renderers (“writers”), one could support N × M conversions. Soon I added a reader for reStructuredText, because I kept a lot of my lecture notes and handouts in that format. And I added a writer for LaTeX, because I wanted to be able to produce PDFs. Then I added a writer for Markdown, so I could start to convert my reStructuredText notes to Markdown. And from there the project just snowballed.

Thus, a project that started out as nothing more than the product of procrastination was nurtured by the joy of writing in Haskell and by its increasing usefulness for my own academic work.

Hackage Haskell package repository, which was started in 2007. The Hackage archive and the new cabal-install tool, which automatically resolved and fetched dependencies, opened up the possibility of depending on external packages.

zip-archive), using the excellent binary package for binary parsing and serialization. Support for syntax highlighting required a syntax highlighting library, which also did not exist in Haskell. For this, I wrote highlighting-kate, which parsed the XML syntax definitions used by the Kate text editor and turned them into Haskell code highlighters. This allowed pandoc to support a large number of syntaxes right off the bat. This version also contained support for automatic generation of citations and a bibliography using CSL style, using Andrea Rossato’s citeproc-hs library.

Throughout this period, I was involved in discussions with other Markdown implementers on the (now defunct) markdown-discuss mailing list. The syntax for delimited code blocks, which pandoc supported long before GitHub popularized fenced code blocks, was worked out in collaboration with Michel Fortin, the maintainer of PHP Markdown Extra. I took care when adding extensions to pandoc’s Markdown to pay attention to prior art, for example copying PHP Markdown Extra’s definition list syntax. During this period, I also became aware of many ambiguities in Markdown’s syntax—a situation I would later try to improve in the commonmark project.

The next big change to pandoc came in version 1.4 (released in January 2010), which introduced a flexible template system, replacing hard-coded headers and making pandoc’s output much more customizable.

In 2010, we moved from Google Code to GitHub, which would do even more to increase the visibility of the project. Further releases in 2010 and 2011 added support for EPUB output, Org-mode output (due to Puneeth Chaganti), and Textile input (due to Paul Rivier). Pandoc also gained support for converting TeX math to MathML (for DocBook or HTML), via my texmath library.

Pandoc 1.9, published in 2012, finally made it possible to produce Word docx output. To handle the equations properly, I added support for Word’s OMML format to texmath. This release also added an AsciiDoc writer and support for Beamer and DZSlides, and in 1.9.3 we gained a DocBook reader (with contributions from Mauro Bieg, who became a long-time contributor).

In 2013, we focused on several features that made pandoc much more flexible and customizable. The first was a fine-grained system of Markdown “extensions,” allowing support for the many variants of Markdown that were then proliferating. The second was the ability to include YAML metadata blocks in Markdown, with arbitrary structured fields that populate template variables. The third was the ability to create custom writers in Lua, allowing ad hoc output formats to be supported by users. The fourth was the introduction of JSON filters—user-created programs that transform a JSON serialization of the pandoc AST, allowing the document to be customized between the parsing phase and the rendering phase. Citation processing was moved from the core of pandoc into an external filter, pandoc-citeproc.

This era saw the addition of reveal.js, EPUB v3, DokuWiki, and FictionBook2 output; OPML input and output; and Haddock and MediaWiki input. Notable contributors include David Lazar (Haddock) and Sergey Astanin (FictionBook2).

The year 2014 saw the arrival of three new contributors who would go on to make many contributions to the project. Albert Krewinkel added support for Org-mode input; Jesse Rosenthal added a Word docx reader (complete with track-changes awareness); and Matthew Pickering (at the time a student at Oxford whom I “advised” as a Google Summer of Code Student) added support for EPUB and Txt2Tags as input formats. Supporting EPUB input required being able to convert MathML equations, so Pickering also worked on texmath. We were in very different time zones, and I remember waking up every morning to find all the work Pickering had done during the night. (Pickering has gone on to become one of the core maintainers of the ghc compiler.) All of these contributions were released in pandoc 1.13, together with Clare Macrae’s DokuWiki writer.

Since 2012, I had been involved in a working group that aimed to produce an unambiguous specification of Markdown’s syntax, initiated by Jeff Atwood and including representatives from GitHub, Reddit, and Stack Overflow. The group held intensive discussions in 2012, which petered out in 2013. I still believed in the project and didn’t want to let the work we’d done go to waste, so I sat down in August 2014, before the academic year began, and wrote up a spec for Markdown, as well as parsers in JavaScript and C. I sent the draft spec to John Gruber for comment and did not get a response, so a few weeks later we posted the spec. At this point, Gruber strongly objected and demanded that we not call the project “Standard Markdown,” so we changed the name to “commonmark.” The project has been a success, in that with a few exceptions, most Markdown processors implement the commonmark spec for their core rules. (Commonmark does not concern itself with extensions.)

Pandoc 1.14 (2015) added support for commonmark and a number of extensions (at first via bindings to the C library libcmark, but later, in 2020, via my Haskell packages commonmark, commonmark-extensions, and commonmark-pandoc). I intend eventually to replace pandoc’s legacy Markdown parser with a commonmark core, but there are still a few key extensions that have not been implemented, so pandoc users must still choose between parsing their documents as markdown (Markdown with pandoc’s extensions) or as gfm or commonmark or commonmark_x (commonmark with a number of extensions). Ironically, although I was the author of the commonmark spec, pandoc still uses a pre-commonmark Markdown parser!

The next year brought some important changes in the pandoc AST, with the addition of image and link attributes, a SoftBreak element (enabling pandoc to preserve line breaks from the original source, or wrap, depending on a command line setting), and a LineBlock element. MarLinn added an ODT reader, Chris Forster added a TEI writer, and Ivo Clarysse added support for DocBook 5.

“pure” (that is, they had Haskell types that prevented them from having any side effects, including I/O operations). But some formats needed to be able to do I/O for a fully faithful conversion. (For example, reStructuredText has a syntax for including files, so the parser needs to be able to read files; in some other formats, images require explicit sizes, so a renderer has to be able to read image files, perhaps fetching them using HTTP, and determine their sizes.) We designed a system that allowed pandoc readers and writers to run in any instance of the PandocMonad typeclass, and we provided both a pure instance (which could be used for controlled testing, and in situations where we wanted to forbid I/O) and an instance that allowed I/O operations. The system also provided a way to handle images included as resources in formats like docx or EPUB.

The other big change was the introduction of Lua filters: filters running in an embedded Lua interpreter and operating directly on the pandoc AST, requiring no software other than pandoc itself and offering far better performance than JSON filters. This was made possible by the massive efforts of Albert Krewinkel, building on the hslua, a Haskell-Lua bridge library.

In addition, pandoc 2.0 introduced the raw attribute syntax in pandoc’s Markdown, and support for GitHub-flavored Markdown, Emacs Muse (Alexander Krotov), TikiWiki, Vimwiki (Yuchen Pei), Creole (Sascha Wilde), groff ms, and JATS. The old highlighting-kate was replaced by the new skylighting, which offered better performance and more accurate interpretation of KDE syntax definitions. A PowerPoint writer (due to Jesse Rosenthal) soon followed, as well as support for FictionBook2 (Krotov) and man (Yan Pashkovsky and me) as input formats.

In 2018, the project received a generous $100,000 donation from Handshake, which we used over the next five years to give small stipends to the most active maintainers.

In 2019, support for ipynb (Jupyter notebooks) was added, allowing pandoc to be used in data science workflows, and Jira wiki markup was supported as an output format. With pandoc 2.8, it became possible to specify collections of default options using defaults files.

Users had long complained that pandoc’s model of a table was too restrictive, not even supporting row and colspans. After extensive discussion of what was needed in a table format, Christian Despres designed the new types for tables and modified all of the readers and writers to use it (a big job).

At this point pandoc had supported citation resolution for many years, by means of the pandoc-citeproc filter that used Andrea Rossato’s citeproc-hs. This was slow and somewhat buggy, and Rossato had long since disappeared from the scene, so I wrote a Haskell citeproc library from scratch, using just the CSL spec and test cases. Pandoc 2.11 depended on this library and offered far better citation support: faster, more faithful to CSL, and with no need for an external filter. In order to get citations to sort properly, I had to write a another library (unicode-collation) implementing the Unicode Collation algorithm in pure Haskell.

During this era Pandoc came to support conversions between bibliography database formats: BibTeX, BibLaTeX, and CSL JSON, EndNote XML and RIS; conversion from CSV and TSV to pandoc table formats; conversion to Markua; and conversion from RTF. With pandoc 2.15 a --sandbox option was added, which guarantees that pandoc’s parsers and renderers have no I/O side effects. (This was possible because of the PandocMonad abstraction we added back in pandoc 2.0.) With pandoc 2.16.2 it became possible to write custom readers in Lua to complement the custom Lua writers that had been added in 2013. And with pandoc 2.19.1 it became possible to run pandoc as a web server exporting an API.

Typst, a modern LaTeX competitor with incremental compilation, were released in 2023. I wanted to help the project by providing an easy on- and off-ramp, making it easy for others to try Typst. It turned out that creating a Typst reader for pandoc required implementing an interpreter for a fairly full-featured programming language. The result was the typst package on Hackage. Typst support was added in pandoc 3.1.3.

In 2018 I had published an essay “Beyond Markdown” in which I described the six features of Markdown that I thought had created the most difficulties, both for writing a spec and for implementations, and I explained how I thought these flaws could be fixed in a future Markdown-like light markup syntax. In 2022, I published a syntax description for such a syntax, djot, together with code in Lua, JavaScript and (later) Haskell. Pandoc 3.1.12, published in 2024, added djot as both an input and output format.

Subsequent releases in 2024 and 2025 saw the addition of an ANSI writer for formatted terminal output and a reader for the mdoc and POD formats (all due to Evan Silberman), a reader and writer for an XML representation of the pandoc AST (massifrg), a vimdoc writer (reptee), a PowerPoint reader (Anton Antich), an Excel spreadsheet reader (Anton Antich), and a BBCode writer (reptee), and an AsciiDoc reader (supported by my asciidoc package).

Pandoc 3.9, released in February 2026, included support for compiling pandoc to WASM, which allowed a full-featured version of pandoc to run in the browser. Most of the key work was done by TerrorJack. The GUI interface “pandoc for the people” was designed with the help of Claude Opus.

I still work on pandoc almost every day. Most of this work doesn’t involve the kind of new features or architectural changes I have focused on in this narrative. Mostly it consists in fixing small bugs, making tiny improvements, reviewing issues and pull requests, repairing infrastructure (continuous integration, building releases, code signing, website), improving documentation, and engaging in discussions with maintainers and users.

rules for emphasis. What we found is that, no matter how complex we made the rules for nested emphasis, it was always possible to come up with cases where the algorithm diverges from the meaning a human would naturally find in the string. In such cases, I would often remark, “until our programs have AI, we are going to have edge cases like this; at some point we have to accept that and stop trying to develop more complex rules.” Interestingly, now we do have tools that can understand (or at least simulate understanding) of the meaning and intent of the text, and can potentially do better at recognizing the formatting intended by the author than any light markup syntax that could be designed.

Whatever the future may bring, I am proud of the 20-year history of this project, which has saved people all over the world countless hours of drudgery. Happy 20th birthday, pandoc!


In honor of this occasion, I have produced some pandoc mugs and stickers:

联系我们 contact @ memedata.com