使用交叉注意力条件机制学习爵士钢琴演奏风格
Learning Jazz Pianist Style with Cross-Attention Conditioning

原始链接: https://almostimplemented.github.io/jazz-pianist-style/

受 Dick Hyman 的《In the Styles of… The Great Jazz Pianists》启发,这项研究探讨生成模型能否复现 individual jazz pianists 个体化的爵士钢琴演奏风格,而不只是识别演奏者。作者使用十二位 PiJAMA 钢琴家的独奏录音,对钢琴 MIDI Transformer 模型 Aria 进行微调。他们在 Aria 上半部分加入门控交叉注意力层,使模型能够根据所选钢琴家的学习嵌入来控制生成内容。 研究使用钢琴家分类器而非困惑度来评估风格,因为困惑度几乎无法区分条件生成模型与无条件生成模型。在滑动长度为 300 个音符的窗口中,条件生成的续奏有 70% 被归因于目标钢琴家,而无条件模型仅为 37%。一项仅使用生成音乐训练的分类器,能够以 95% 的歌曲级准确率识别真实录音,这表明生成出的风格可以迁移到原始分类器之外。 不同钢琴家的结果差异显著:Hank Jones 和 Dick Hyman 的归因准确率达到 96%,而 Cedar Walton 仅为 29%。交互式示例包括基于共同提示生成的多个版本、按钢琴家分别生成的续奏,以及对典型风格片段的可视化。由于自动转录只能近似还原力度和踏板信息,并且示例是依据分类器筛选的,因此仍需开展人工听辨研究作为下一步。

Hacker News 最新 | 往期 | 评论 | 提问 | 展示 | 招聘 | 提交 登录 使用交叉注意力条件控制学习爵士钢琴家的演奏风格 (almostimplemented.github.io) 9 分 由 ishan0102 提交 2 小时前 | 隐藏 | 往期 | 收藏 | 1 条评论 帮助 shermantanktop 17 分钟前 [–] 了解那本启发了该项目的书——它的作者是一位深入研究其他人风格的艺术家——确实令人佩服。 阅读一个关于某种 AI 模型的 AI 生成网站:这个模型只要收到指令,就能胡编出假的比尔·埃文斯——这可不怎么样。只要有足够多的 token,任何蠢货都能做到。恭喜。 回复 可以考虑申请 YC 的 2027 年冬季批次! 申请截止至 11 月 2 日。 指南 | 常见问题 | 列表 | API | 安全 | 法律 | 申请加入 YC | 联系我们 搜索:
相关文章

原文

Drew Edwards · Akira Maezawa · Simon Dixon

ISMIR 2026, Abu Dhabi

Cover of Dick Hyman's In the Styles Of... The Great Jazz Pianists, listing the pianists from Oscar Peterson to Jimmy Yancey

In 1994 Dick Hyman published In the Styles of… The Great Jazz Pianists: fifteen original études, each written in the manner of one master, from Scott Joplin to Bill Evans. Rather than transcribing their solos, Hyman composed new music that carries their signatures — Tatum’s “rapid runs in both hands,” Garner’s “strumming, guitar-like left hand,” Peterson’s “tremolos and glissandi.” That book is the inspiration for this project. Can a model learn to do what Hyman did: not just recognize who is playing, but play in their manner? Tatum, Garner, and Peterson are among the twelve pianists we study — and so is Hyman himself.

We fine-tune Aria, a transformer pretrained on piano MIDI, on solo performances by twelve jazz pianists from the PiJAMA dataset, adding a gated cross-attention layer that reads a learned embedding for each pianist. To check whether the style comes through, we slide a pianist classifier along the generated music: conditioned continuations are attributed to the intended pianist 70% of the time, against 37% without conditioning. A second classifier trained only on generated music then identifies real recordings with 95% accuracy.

Listen first; how it works is further down.

One prompt, twelve pianists

The opening bars of “Ain’t Misbehavin’” are played by one of us (Drew). Everything after the dashed line is generated: twelve takes of the same opening, each conditioned on a different pianist. Pick a pianist to hear their take from the top.

Every take generates the same number of notes, so they end at different times: Erroll Garner packs them into 1:26, Cedar Walton spreads them across 2:45. That difference in density is itself part of a pianist’s signature.

Loading…

Twelve continuations, scored

A shared prompt pulls every pianist toward the same tune. Here each pianist instead continues a few bars of their own playing. The strip under each take shows what our classifier heard as it slid along the continuation, one cell per window of about 300 notes: gold where it named the intended pianist, mauve where it named someone else. The two takes per pianist are the best of eight we scored; the line under them says how the rest did. Or switch on the blindfold and guess for yourself.

classifier names the intended pianist names someone else windows that include what you hear now

How it works

The model

We start from Aria (Bradshaw et al., ISMIR 2025; code), a 16-layer transformer pretrained on a large corpus of piano MIDI. Into each of its last eight layers we insert a cross-attention block: the music attends to a small learned embedding for the chosen pianist, four vectors per pianist. A learned gate scales what the block adds, starting at 0.1, so fine-tuning begins from Aria’s own behaviour and learns how much to listen. Because the embedding is attended to at every step, the conditioning does not fade as generation goes on, the way a prompt prefix does.

One transformer layer with the added gated cross-attention block x Self-attention Add & norm Cross-attention ×g Add & norm Feed-forward Add & norm x′ keys, values pianist embedding 4 learned vectors
One of the eight adapted layers. The new pieces are in amber: the music supplies the queries, the pianist embedding supplies keys and values, and the gate g scales the result before it joins the residual stream.

Measuring style

How can we tell whether the model has learned a pianist’s style? The standard yardstick for a generative model, perplexity on held-out music, turns out to be nearly blind to it: given the real preceding notes, the next one is predictable whoever is playing, so conditioning barely moves the score. Style shows up when the model generates freely and has to stay in character on its own output. So instead we let it play, and ask a pianist classifier who it sounds like. Agreement is how often the classifier names the intended pianist, in windows slid along each continuation.

ModelPerplexityAgreement
Pretrained Aria11.4125%
Fine-tuned, no conditioning6.9637%
Fine-tuned with pianist conditioning6.8270%

Perplexity (lower is better) barely separates the two fine-tuned models; agreement nearly doubles. Continuations are 4096 tokens from 256-token prompts; chance agreement is 8%.

Three ways we use the classifier

  • Sliding-window agreement, above. The classifier identifies 98.8% of held-out songs, and the conditioned model’s lead holds from the start of a continuation to its end. The strips under the scored takes are this measurement.
  • Synthetic transfer. A fresh classifier trained only on generated music identifies real recordings: 87% of 1024-token chunks and 95% of songs, within nine points of one trained on real data. The scored takes are samples of that training data.
  • Characteristic regions. Turned on real performances, the classifier points to where a pianist’s style is most concentrated — where the style lives.

The paper has the details: per-pianist results, the mismatch experiment (prompting with one pianist and conditioning on another), memorization checks, and a from-scratch classifier that confirms the transfer result.

About these examples

Everything here is MIDI, rendered in your browser on a sampled piano. The model was trained on automatic transcriptions of commercial recordings, so dynamics and pedalling are approximate, and the rendering is plainer than the records.

The twelve pianists were chosen for separability: they are the twelve of PiJAMA’s thirty whose recordings a pretrained model already tells apart most easily. Within them the model imitates some far better than others — across the paper’s evaluation, continuations were attributed to the intended pianist 96% of the time for Hank Jones and Dick Hyman, but only 29% for Cedar Walton.

The scored takes are selected, not copied. They are samples from the corpus of generated music that the paper’s synthetic-only classifier learned from; for each pianist we show the two highest-scoring of eight candidates. Their prompts come from the training recordings, and the paper checks that the continuations do not copy them: they resemble their closest training performance less than real held-out performances do.

The scores come from a classifier, not from listeners. It is a strong one (98.8% of held-out songs), but it has habits: many of its mistakes on generated music land on Dick Hyman — fitting, perhaps, for a pianist who made a career of playing in everyone else’s style. A listening study is the natural next step.

Built on Aria and the PiJAMA dataset. Code, checkpoints, and evaluation scripts: on GitHub.

@inproceedings{edwards2026jazzstyle,
  title     = {Learning Jazz Pianist Style with Cross-Attention Conditioning},
  author    = {Edwards, Drew and Maezawa, Akira and Dixon, Simon},
  booktitle = {Proc. Int. Society for Music Information Retrieval Conf. (ISMIR)},
  year      = {2026},
}
联系我们 contact @ memedata.com