跳转至

hardware

探究一个 USB Type-C 拓展坞的硬件实现

背景

最近在高强度使用 HDMI 和 DP 做视频输出,但在使用的时候遇到了各种细节问题,所以就研究整个链路上到底发生了哪些事情,在这个过程中,研究了一下手上的 Biaze KZ11 这款 Type-C 拓展坞,看看它内部有哪些芯片,又是怎样实现拓展坞的功能的。

One CPU Atomic Instruction, One Packaging Infinite Loop: The Story of the Lost Update on LA664

中文版本

TL;DR

In February 2026, Wang Miao ran into something strange while packaging normaliz for Debian on a LoongArch server: the math software's built-in test kept timing out, stuck in an infinite loop that it could not escape. Following the code, the problem pointed to a very ordinary operation: OpenMP's #pragma omp atomic accumulating into a shared variable. The loop's exit condition required the accumulated value to equal a certain number, but the accumulated result was always less than that number, causing the infinite loop. Because the program was large and the code complex, we never managed to reduce it to a minimal example a human could understand, so the matter was shelved.

Half a year later, in August, Wang Miao came to me again, wanting to pick it back up. This time we took a different approach: instead of having a human locate the problem, we let AI find a minimal reproduction, with the human directing the AI's investigation. About two days later, we had a stable reproducer, and only then discovered the root cause: the CPU's atomic add instruction occasionally fails to be atomic. This meant we had found a new CPU erratum, and after Loongson learned of it, only two weeks passed before they found a fix with almost no performance loss and provided us with test firmware. We confirmed that the test firmware resolves the issue, and Loongson told us the firmware is expected to be released before National Day (October 1), at which point readers will be able to upgrade their firmware to fix the problem.

一颗 CPU 的原子指令,一个打包死循环:LA664 丢失更新事件始末

本文同步发布到本人的知乎。

English version

太长不看版本

2026 年 2 月,王邈在龙架构服务器上给 Debian 打包 normaliz 时遇到一件怪事:这个数学软件的自带测试总是超时,现象是卡在死循环里出不来。顺着代码调查,问题指向一个很常规的操作:OpenMP 的 #pragma omp atomic 对共享变量进行累加。循环的退出条件要求累加后的值等于某个数,而累加的结果总是少于这个数,就导致了死循环。由于程序太大、代码又很复杂,始终没能把问题缩减成一个人类能看懂的最小例子,这件事就被搁置了。

半年后的 8 月,王邈再次找到我,想把它重新捡起来。这回我们换了个做法:不再由人来定位问题,而是让 AI 去找最小复现,人在这个过程中负责指挥 AI 调查的方向。大概两天后,我们拿到一个稳定的复现程序,才发现事情的根源是:CPU 的原子加法指令,居然偶尔会不原子。这意味着我们找到了 CPU 的一个新 erratum,而龙芯得知这件事后,仅仅过了两周,就找到了几乎没有性能损失的修复方法,并给我们提供了测试固件。我们确认了测试固件可以解决问题,并且龙芯告诉我们,该测试固件预计在国庆(10 月 1 日)之前发布,届时读者将可以升级固件以修复该问题。

Ampere Skylark 微架构评测

背景

Ampere eMAG 采用的是 Ampere Skylark 微架构,虽然是 2018 年的处理器了,但也顺带评测一下。其前身是 AppliedMicro 的 X-Gene 3 微架构,用在 Ampere eMAG 芯片上,用的是 TSMC 16nm FinFET+ 工艺。

SDRAM 在不同访存模式下的带宽分析与实验

背景

最近在和 @CircuitCoder 交流 SDRAM(通常简写为 DRAM,或更进一步简写为 DDR)的各种性能指标,于是想到利用现有的 DRAMSim3 和 Ramulator2 做一些模拟测试,看看各种访存模式下可以实现峰值带宽的多少比例,再结合时序验证理论与模拟结果是否吻合。实验相关代码已开源至 jiegec/dram-bench。

IBM POWER9 微架构评测

背景

继 IBM POWER8 之后,也来评测一下后续的 IBM POWER9 微架构。IBM POWER9 有 SMT4 和 SMT8 两种版本,我只有 SMT4 版本的测试环境,下列所有评测都是针对 SMT4 版本进行测试。

条件分支预测器逆向工程(以 Apple M1 Firestorm 为例)

背景

去年我完成了针对 Apple 和 Qualcomm 条件分支预测器(Conditional Branch Predictor)的逆向工程研究,相关论文已发表在 arXiv 上,并公开了源代码。考虑到许多读者对处理器逆向工程感兴趣,但可能因其复杂性而望而却步,本文将以 Apple M1 Firestorm 为例,详细介绍条件分支预测器的逆向工程方法,作为对原论文的补充说明。

AMD Zen 1 的 BTB 结构分析

背景

AMD Zen 1 是 AMD 在 2017 年发布的 Zen 系列第一代微架构。在之前,我们分析了 ARM Neoverse N1 和 V1 的 BTB,那么现在也把视线转到 AMD 上,看看 AMD 的 Zen 系列的 BTB 是如何演进的。

终端模拟器的文字绘制

背景

最近在造鸿蒙电脑上的终端模拟器 Termony,一开始用 ArkTS 的 Text + Span 空间来绘制终端,后来发现这样性能和可定制性比较差,就选择了自己用 OpenGL 实现,顺带学习了一下终端模拟器的文字绘制是什么样的一个过程。

鸿蒙电脑 MateBook Pro 开箱体验

购买

2025.6.6 号正式开卖,当华为线上商城显示没货的时候,果断去线下门店买了一台回来。购买的是 32GB 内存,1TB SSD 存储,加柔光屏的版本,型号 HAD-W32,原价 9999,国补后 7999。

分析 Rocket Chip 中 Diplomacy 系统

背景

Rocket Chip 大量使用了 Diplomacy 系统来组织它的总线、中断和时钟网络。因此,如果想要对 Rocket Chip 进行定制,那么必须要对 Rocket Chip 中 Diplomacy 系统的使用有充分的了解,而这方面的文档比较欠缺。本文是对 Rocket Chip 中 Diplomacy 系统的使用的分析。阅读本文前,建议阅读先前的 分析 Diplomacy 系统 文章,对 Diplomacy 系统的设计和内部实现获得一定的了解。