AlphaGenome 绘制了 90 亿个 DNA 变体图谱
AlphaGenome maps 9B DNA variants

原始链接: https://spectrum.ieee.org/alphagenome-atlas

Google DeepMind 推出了 **AlphaGenome Atlas**,这是一个包含人类基因组中所有 90 亿种可能的单字母基因变异预计算预测结果的公共存储库。虽然大多数 DNA 是“非编码”的——它们负责调节基因活性,而非直接产生蛋白质——但理解这些片段的功能对于疾病研究至关重要。此前,研究人员必须自行运行计算量巨大的 AlphaGenome 模型;而现在,该图谱提供了即时、可访问的数据,并为每个变异提供了简化的“影响评分”。 通过利用先进的工程技术,DeepMind 处理了 PB 级的数据,为科学家优先开展实验室实验和识别潜在疾病驱动因素提供了捷径。尽管该模型存在局限性(例如难以预测远程调控相互作用,以及多变异疾病的复杂性),但基因组学家 Carl de Boer 等专家认为,它是一项领先的工具,将显著加速生物学研究。该图谱可在非商业用途下使用,它建立在 DeepMind 之前的 AlphaFold 等突破性成果之上,代表了破译“生命语言”的重要一步。虽然研究人员提醒简化的影响评分可能会被误读,但该资源为理解 DNA 变异如何影响健康和发育提供了前所未有的基础。

Hacker News 新内容 | 过往 | 评论 | 提问 | 展示 | 招聘 | 提交 登录 AlphaGenome 绘制了 90 亿个 DNA 变体图谱 (ieee.org) 9 分,由 ltononro 发布于 1 小时前 | 隐藏 | 过往 | 收藏 | 讨论 帮助 指南 | 常见问题 | 列表 | API | 安全 | 法律 | 申请 YC | 联系 搜索:
相关文章

原文

DNA is often explained as a codebook or set of instructions for producing proteins, and ultimately, life. Some stretches of DNA, called genes, code for proteins, but the vast majority of DNA is considered “noncoding.” Some of it has no known function, while other segments are critical to regulating gene activity.

These regulatory elements can interact in complicated ways, and their effects can vary across different cells and tissues. Some also influence genes located far away in the genome. Understanding how changes in DNA affect this regulation “is fundamental to understanding most disease,” says Carl de Boer, a genomicist at the University of British Columbia.

That’s why researchers are working to understand what every imaginable small variation in human DNA across the entire genome might mean for gene regulation. A recent AI tool built for that purpose from Google DeepMind, AlphaGenome, was originally announced in 2025. In January, a paper published in Nature provided more details, and the model was released for public noncommercial use. The AI model can compare an original DNA sequence with an altered one and predict how the change might affect gene expression and other regulatory activity. But researchers had to select the variants they wanted to test, write code, and run the computationally demanding model themselves.

Now DeepMind has done that work in advance for all 9 billion possible single-letter changes to a reference human genome. Today, on 8 September, DeepMind announced the creation and public release of the AlphaGenome Atlas, an online repository of precomputed predictions made using the AlphaGenome model. The Atlas offers a more approachable interface for scientists, without the need to write code or run the AlphaGenome model themselves. It also includes a much-requested new feature, a single-number impact score intended to show at a glance if a variant is likely to be meaningful.

“Understanding our DNA is a grand challenge,” says Pushmeet Kohli, VP of science at Google DeepMind. “Understanding this language of life can unlock so many things.”

The AlphaGenome predictions have some important limitations. For example, many diseases are associated with multiple genetic variants. And although AlphaGenome looks at a relatively large segment of DNA surrounding the variant in question—1 million base pairs—some DNA sequences, called enhancers, can regulate genes over very long distances, sometimes beyond the model’s field of view. Their effects are difficult to predict.

But the Atlas could still help scientists filter possibilities and prioritize lab experiments that would validate its predictions. In that way, it could greatly accelerate work in fundamental biology, disease research, and treatment development, says Žiga Avsec, the genomics lead at DeepMind.

“It seems like they made a useful resource for people,” says de Boer, who recently helped create a framework for better comparisons of computational models similar to AlphaGenome. He is not affiliated with DeepMind.

Although de Boer considers AlphaGenome the “field’s leading model,” he notes that it’s also “very slow and computationally intensive.” The Atlas could benefit people without access to newer hardware, or simply reduce the number of people repeating the same simulations.

The Atlas is freely available for noncommercial research, with the potential for commercial licensing.

The entire human genome contains roughly 3 billion base pairs. At each position there are three possible single-nucleotide substitutions, and therefore 9 billion variants in the Atlas. The complete dataset is around 1 petabyte.

“When we started thinking about this project, it seemed impossible to do that computationally,” says Avsec. Early estimates told the team they would need to improve their calculation speed by a factor of 80 in order to compile the Atlas in a reasonable amount of time.

To reach that target, the team gained advantages using a few different techniques, including model distillation, GPU kernel optimization, and the elimination of redundant calculations. “There was a lot of thought and engineering that we had to do in order to make this happen at this scale,” says Avsec.

AlphaGenome and the Atlas build on years of related work at DeepMind. In 2020, AlphaFold predicted the three-dimensional structure of proteins from amino-acid sequences. In 2023, AlphaMissense predicted whether 71 million possible variants that alter proteins were likely benign or pathogenic. Similar to the new Atlas, prediction results from those projects were made available in a public database.

The Atlas allows a scientist to look up a single variant and see more detailed information about the model’s prediction, including 11 different output types. But the top-line figure is a single-number impact score, which by its nature is a simplification of many aspects of those predictions.

“It has a clear use, but it also is probably going to be easily misinterpreted,” says de Boer. “We’re talking about a very complex system, and there’s a lot of moving parts.”

From Your Site Articles

Related Articles Around the Web

联系我们 contact @ memedata.com