大会堂羞耻录
Assembly Hall of Shame

原始链接: https://github.com/xoreaxeaxeax/asm-hall-of-shame

“耻辱装配大厅”(Assembly Hall of Shame)是克里斯托弗·多马斯(Christopher Domas)发起的一个项目,旨在探索性能优化的反面:寻找执行耗时最长的单一指令。 尽管现代计算追求纳秒级的效率,但该项目却刻意将 CPU 强行置于最慢的状态。其策略包括触发微码辅助(例如通过次正规浮点操作数)、强制执行高速缓存一致性总线锁,以及利用高延迟的硬件特性等。 目前的纪录保持者在 AMD Ryzen 7 5800H 处理器上使用了极端的“总线饥饿”(fabric starvation)技术。通过针对高延迟的 PCIe MMIO 区域执行 `fxrstor64` 指令(该指令用于加载 512 字节的 FPU/MMX/XMM 状态),同时利用“冲击”循环产生的竞争流量使 PCIe 总线饱和,该团队成功实现了一次运行耗时约 1980 亿个时钟周期(约 62 秒)的惊人结果。 “耻辱大厅”凸显了硬件抽象的脆弱性,展示了如何操纵微码、缓存一致性协议和系统级 I/O 来制造蓄意且巨大的瓶颈。未来的研究旨在利用更新架构上的 `xrstor64` 指令进行 8KB 状态加载,以进一步扩展这些技术。

Hacker News 社区正在讨论 GitHub 仓库“Assembly Hall of Shame”(汇编羞耻大厅)。该项目由 `xoreaxeaxeax` 创建,旨在记录那些刻意降低计算机性能的指令与技术。 评论者们对该项目探索的“性能反优化”深感着迷,并指出尽管现代硬件性能强大,但软件运行却常显迟缓。讨论重点介绍了多种制造极端瓶颈的方法,例如利用浮点数次正规数(subnormals),或是通过操纵内存映射 I/O(MMIO)来阻塞 PCIe 总线。 一些用户提到了作者的其他研究背景,包括仅使用 `mov` 指令的编译器,以及通过混淆控制流在反汇编程序中创建视觉图案的项目。讨论还推测了这些“最差性能”策略在不同处理器架构(如 POWER 架构)上与当前 x86 架构有何不同。总的来说,该仓库被视为一次对指令执行隐藏成本的巧妙且深入的技术探索。
相关文章

原文

x86 Leaderboard

Instruction latency analysis usually focuses on performance optimization—making code run as fast as possible. The Assembly Hall of Shame takes the opposite approach: searching for the absolute floor of single-instruction performance.

🏆 Current Champions 🏆

Strategy: Use fxrstor64 to load 512-byte FPU/MMX/XMM state from a high-latency MMIO region in the PCIe fabric, then starve the fabric while the load is in flight — a fleet of hammer cores pounds a different high-latency MMIO register with tight 4-byte reads, saturating the PCIe root complex and endpoint with non-posted transactions, so CPU 0's 512-byte fxrstor64 must queue behind all that contending traffic.

Contender: AMD Ryzen 7 5800H

; CPU 0 — timed instruction
movl $0xfcc68830, %rsi
fxrstor64 %rsi

; CPUs 1..N — hammer loop against a different high-latency location
movl 0xfcc68858, %eax

🏆 Score: 198,002,498,236 cycles

🏆 Time: 62 seconds

A spec-violating unaligned ymm0 load that forced non-posted dword transactions from stalled GPU registers was used to break the fundamental design of System Management Mode in smiiiiiiiiiiiiiiii.

vmovdqu 0xfcc003b1, %ymm0
  • Instructions may use whatever setup is necessary, but only a single instruction is eligible to be scored.
  • Trapped/emulated/virtualized instructions may only time the trap, not the handler.
  • Instructions must not be interruptible. rep movs, pause, etc. are disqualified.
  • Times are normalized based on the CPU base clock frequency.
  • All platforms must be in their factory stock configurations - no hardware modifications.

Strategy: nop does nothing. It opens the leaderboard accordingly.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

Score: 1 cycles

Time: 0 nanoseconds

Strategy: Regular nop was too short, but how do we make nothing take longer? Try a lonnnnnng nop.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

data16 data16 data16 data16 data16 data16 data16 nopl 0x00000000(%%eax,%%eax,1)

Score: 20 cycles

Time: 7 nanoseconds

Strategy: Just a reference instruction to get our bearings.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

Score: 49 cycles

Time: 18 nanoseconds

Strategy: Use 128-bit dividend (rdx:rax=2:0) with small divisor to push the quotient above the ceiling imposed by sign-extension, driving the longest path through the divider microcode.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

xorq %rax, %rax   ; rax = 0  (low 64 bits of dividend)
movq $2, %rdx     ; rdx = 2  (high 64 bits: full dividend = 2^65)
movq $5, %rbx     ; divisor → quotient = 2^65/5 ≈ 7.4×10^18
idivq %rbx

Score: 77 cycles

Time: 28 nanoseconds

Strategy: Use maximum nesting depth (31) to force 30 display-pointer loads and pushes through the microcode display-walk path.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

enter $0, $31       ; 0 bytes allocated, nesting depth 31 (maximum)

Score: 112 cycles

Time: 41 nanoseconds

Strategy: Try a small denormal to trigger an FP microcode assist.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

    movabsq $0x0000000000000001, %rax
    movq    %rax, -8(%rsp)
    fldl    -8(%rsp)

Score: 133 cycles

Time: 49 nanoseconds

Strategy: Just ensure the cache line is dirty.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

clflush (%rax)          ; rax -> dirty cache line resident in L3

Score: 165 cycles

Time: 60 nanoseconds

Strategy: Use exponent 0x7ff to reach 'special value' processing in microcode; positive/negative, NaN/inf doesn't seem to make a difference, go with QNaN.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

    movabsq $0x7fffffffffffffff, %rax
    movq    %rax, -8(%rsp)
    fldl    -8(%rsp)
    fsin

Score: 257 cycles

Time: 94 nanoseconds

Strategy: Saturate all write-combining line-fill buffers with movnti stores to distinct cache lines, forcing mfence to drain the full LFB write path to the uncore before retiring.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

movnti %r9,  0*64(%rdi)   ; ×16 distinct cache lines — saturate the write-combining LFBs
; …
movnti %r9, 15*64(%rdi)
mfence                     ; must drain all pending LFB writes before retiring

Score: 326 cycles

Time: 120 nanoseconds

Strategy: Nothing for now, just check how long it takes to invalidate the TLB.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

Score: 352 cycles

Time: 110 nanoseconds

Strategy: Hit x87 FP microcode assist path by using denormal source operand.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

fldl   subnorm    ; 1e-310: value < DBL_MIN, biased exponent = 0
faddl  subnorm    ; source is subnormal → FP microcode assist

Score: 677 cycles

Time: 249 nanoseconds

Strategy: Align lock-prefixed operand to straddle cache-line boundary, forcing CPU to assert the external bus lock rather than using the fast MESI cache-coherence path.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

; split_ptr % 64 == 63 — dword spans bytes 63 (line N) and 64–66 (line N+1)
lock xaddl %r9d, (%rdi)

Score: 865 cycles

Time: 319 nanoseconds

Strategy: Use subnormal divisor, hardware hands control to microcode assist, assist normalizes operand, performs the division, then restores architectural state.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

    movabsq $0x3ff0000000000000, %rax   ; 1.0 (normal dividend)
    movq    %rax, -8(%rsp)
    fldl    -8(%rsp)                     ; ST(0) = 1.0

    movabsq $0x0000002000000000, %rax   ; 6.79e-313 (subnormal divisor)
    movq    %rax, -8(%rsp)
    fdivl   -8(%rsp)                     ; ST(0) = 1.0 / subnormal → FP assist

Score: 883 cycles

Time: 325 nanoseconds

Strategy: Use rakefield to find the highest latency CPUID leaves.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

Score: 1248 cycles

Time: 460 nanoseconds

Strategy: Execute in a tight loop to deplete the hardware entropy pool faster than it can be refilled, forcing subsequent calls to stall while the entropy source recovers.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

Score: 5,579 cycles

Time: 2.057 microseconds

Strategy: Use project:nightshyft to identify high latency MSRs. MCG_CTL on Zen look like a winner: may be a microcode quiesce and synchronize on MCA error banks across hardware units, some potentially off-die, requiring fabric-level communication rather than a simple local register write.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

movl $0x17b, %ecx       ; MCG_CTL
wrmsr

Score: 34,304 cycles

Time: 10.742 microseconds

Strategy: Target an I/O port that straddles a NIC device register boundary, triggering the device to quiesce its TX DMA engine on each write.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

mov $0xf019, %dx
outl %eax, %dx

Score: 49,857 cycles

Time: 15.580 microseconds

Strategy: Use project:nightshyft to identify high latency model-specific-registers: VIA uses an undocumented register at 0x133 that gives wildly high response time. No idea what it does.

Contender: VIA Eden Processor 800MHz

movl $0x133, %ecx ; undocumented MSR
rdmsr

Score: 161,602 cycles

Time: 202.004 microseconds

Strategy: Fully load L1/L2/L3 caches with dirty lines to force DRAM writeback of entire hierarchy.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

Score: 1,616,480 cycles

Time: 506.165 microseconds

Strategy: Target I/O port mapped to an ACPI PM block where an unaligned 4-byte read decodes into multiple non-posted loads from wherever this port goes.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

mov $0x0413, %dx
inl %dx, %eax

Score: 12,524,415 cycles

Time: 3.921769 milliseconds

Strategy: Use mmiotic to identify high-latency deadspace in PCIe fabric, hit unkown GPU register.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

Score: 443,937,696 cycles

Time: 139.010268 milliseconds

Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 8-byte MMIO read to get two dword register accesses, which isn't technically allowed but works anyway.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

Score: 887,716,864 cycles

Time: 277.971228 milliseconds

Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 16-byte MMIO read to get four dword register accesses, which isn't technically allowed but works anyway.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

vmovdqu 0xfcc003b0, %xmm0

Score: 1,774,555,776 cycles

Time: 555.664133 milliseconds

Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 32-byte MMIO read to get eight dword register accesses, which still isn't technically allowed but works anyway.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

vmovdqu 0xfcc003b0, %ymm0

Score: 3,549,079,296 cycles

Time: 1.111345034 s

Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 32-byte unaligned MMIO read to get nine dword register accesses, which is even less allowed than the aligned version, but works anyway.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

vmovdqu 0xfcc003b1, %ymm0

Score: 4,453,212,256 cycles

Time: 1.394428818 seconds

Strategy: Use mmiotic to identify high-latency deadspace in PCIe fabric, isolate region near 0's and offset state to avoid MXCSR corruption (possibly VGA buffer?), use fxrstor64 to load 512-byte FPU/MMX/XMM state from MMIO, forcing CPU to process 512 bytes of I/O transactions through slowest available memory aperture.

Contender: AMD Ryzen 7 5800H

movl $0xfcc68830, %rsi
fxrstor64 %rsi

Score: 74,584,168,512 cycles

Time: 23.354502677 seconds

Strategy: Extend fxrstor64 (baseline) by starving the fabric while the load is in flight — a fleet of hammer cores pounds a different high-latency MMIO register with tight 4-byte reads, saturating the PCIe root complex and endpoint with non-posted transactions, so CPU 0's 512-byte fxrstor64 must queue behind all that contending traffic.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

; CPU 0 — timed instruction
movl $0xfcc68830, %rsi
fxrstor64 %rsi

; CPUs 1..N — hammer loop against a different high-latency location
movl 0xfcc68858, %eax

🏆 Score: 198,002,498,236 cycles

🏆 Time: 62 seconds

Strategy: Leverage extended AVX state in Sapphire Rapids with MMIO approach from fxrstor64: xsave state area is 8KB vs 512 bytes, 16x size -> 1,000,000,000,000 cycles

Contender: TODO

; XCR0 must enable AMX components (bits 17-18); state area ~8KB
xrstor64 (%rsi)         ; rsi -> MMIO region, same technique as fxrstor64

The assembly hall-of-shame is a research effort from Christopher Domas (@xoreaxeaxeax).

联系我们 contact @ memedata.com