把博客生成器从 Mkdocs 迁移到 Zensical
距离上一次 从 Mkdocs 迁移到 Zensical 已经过去了三年,这次 Zensical 终于是补齐了原来 Mkdocs 用到的大部分插件,所以就当小白鼠,把博客从 Mkdocs + Mkdocs-Material 迁移到了 Zensical。
距离上一次 从 Mkdocs 迁移到 Zensical 已经过去了三年,这次 Zensical 终于是补齐了原来 Mkdocs 用到的大部分插件,所以就当小白鼠,把博客从 Mkdocs + Mkdocs-Material 迁移到了 Zensical。
In February 2026, Wang Miao ran into something strange while packaging normaliz for Debian on a LoongArch server: the math software's built-in test kept timing out, stuck in an infinite loop that it could not escape. Following the code, the problem pointed to a very ordinary operation: OpenMP's #pragma omp atomic accumulating into a shared variable. The loop's exit condition required the accumulated value to equal a certain number, but the accumulated result was always less than that number, causing the infinite loop. Because the program was large and the code complex, we never managed to reduce it to a minimal example a human could understand, so the matter was shelved.
Half a year later, in August, Wang Miao came to me again, wanting to pick it back up. This time we took a different approach: instead of having a human locate the problem, we let AI find a minimal reproduction, with the human directing the AI's investigation. About two days later, we had a stable reproducer, and only then discovered the root cause: the CPU's atomic add instruction occasionally fails to be atomic. This meant we had found a new CPU erratum, and after Loongson learned of it, only two weeks passed before they found a fix with almost no performance loss and provided us with test firmware. We confirmed that the test firmware resolves the issue, and Loongson told us the firmware is expected to be released before National Day (October 1), at which point readers will be able to upgrade their firmware to fix the problem.
本文同步发布到本人的知乎。
2026 年 2 月,王邈在龙架构服务器上给 Debian 打包 normaliz 时遇到一件怪事:这个数学软件的自带测试总是超时,现象是卡在死循环里出不来。顺着代码调查,问题指向一个很常规的操作:OpenMP 的 #pragma omp atomic 对共享变量进行累加。循环的退出条件要求累加后的值等于某个数,而累加的结果总是少于这个数,就导致了死循环。由于程序太大、代码又很复杂,始终没能把问题缩减成一个人类能看懂的最小例子,这件事就被搁置了。
半年后的 8 月,王邈再次找到我,想把它重新捡起来。这回我们换了个做法:不再由人来定位问题,而是让 AI 去找最小复现,人在这个过程中负责指挥 AI 调查的方向。大概两天后,我们拿到一个稳定的复现程序,才发现事情的根源是:CPU 的原子加法指令,居然偶尔会不原子。这意味着我们找到了 CPU 的一个新 erratum,而龙芯得知这件事后,仅仅过了两周,就找到了几乎没有性能损失的修复方法,并给我们提供了测试固件。我们确认了测试固件可以解决问题,并且龙芯告诉我们,该测试固件预计在国庆(10 月 1 日)之前发布,届时读者将可以升级固件以修复该问题。
最近频繁地和各家智算卡(GPU、NPU,或者统称为 xPU)厂商交流,讨论如何培养软件生态、如何进入校园。同样的观点我已经跟不同的人讲过至少五遍了,索性写成一篇博客,一次讲清楚。
上文 提到,我打算用采集卡来录制鸿蒙电脑的输出,作为 OBS 的输入来做软件导播,用的采集卡型号是采用了 MS2130S 芯片的绿联 UG307-95348 采集卡。在使用过程中,遇到了清晰度和颜色的问题,下面介绍我是怎么研究和解决的。
最近在做 PPT,用了一个在 Windows 上制作的 PPT 模板,它用到了 微软雅黑 Light 字体,在 macOS 上显示不正常,因此做了一些细致的研究和排查,找到了原因和解决方案。
最近在准备课堂展示,需要用到 Windows,于是翻出鸿蒙电脑,跑起了 Windows on ARM 虚拟机。结果 Windows 更新老是报 0x800703F1 错误,我做了不少自己都说不清的尝试,最后稀里糊涂地解决了。
近日,龙芯龙架构 CPU LA464/LA664 微架构部分步进的漏洞 LoongLeak 正式披露,龙芯官方也发布了 公告。作为公告中提到的“其他国内独立研究者”之一,我在此分享一下我的视角。
Ampere eMAG 采用的是 Ampere Skylark 微架构,虽然是 2018 年的处理器了,但也顺带评测一下。其前身是 AppliedMicro 的 X-Gene 3 微架构,用在 Ampere eMAG 芯片上,用的是 TSMC 16nm FinFET+ 工艺。
之前分析过 M1 和 M4,趁着机会,也评测一下 M2 的微架构,给出一个从 M1 到 M2 再到 M4 的发展脉络。
使用 ARM Neoverse V3 核心的 AWS Graviton 5 最近上线了,相比之前的 Neoverse V2 应该有一些改进,所以测试一下这个微架构在各个方面的表现。
Following the INT Rate article, this article continues with the workload analysis of SPEC FP 2026 Rate.
本文同步发布到本人的知乎。
继 INT Rate 篇 后,本文继续分析 SPEC FP 2026 Rate 的负载特性。
I've been running some benchmarks with SPEC CPU 2026 recently, and plan to do in-depth workload analysis combined with the test results. This article focuses on SPEC INT 2026 Rate workload characteristics. For SPEC FP 2026 Rate analysis, see the FP Rate article.
SPEC CPU 2026 官方只附带了 aarch64/ppc64le/riscv64/x86_64 指令集的预编译 tools,如果要在其他指令集上使用,就需要首先编译 tools,过程如下:
cd /mnt && tar xvf install_archives/tools-src.tar
wget -O config.guess 'https://git.savannah.gnu.org/gitweb/?p=config.git;a=blob_plain;f=config.guess;hb=HEAD'
wget -O config.sub 'https://git.savannah.gnu.org/gitweb/?p=config.git;a=blob_plain;f=config.sub;hb=HEAD'
cp config.* /mnt/tools/src/make-4.2.1/config/
# build tools
mkdir -p /mnt/config
cd /mnt && echo 'y' | SKIPTOOLSINTRO=1 FORCE_UNSAFE_CONFIGURE=1 MAKEFLAGS=-j16 ./tools/src/buildtools
mkdir -p /mnt/config
cd /mnt && . ./shrc && packagetools linux-loong64
例如下面是在 LoongArch 上编译 SPEC CPU 2026 的 Dockerfile,假设 SPEC CPU 2026 已经解压到 /mnt:
RUN cd /mnt && tar xvf install_archives/tools-src.tar
RUN wget -O config.guess 'https://git.savannah.gnu.org/gitweb/?p=config.git;a=blob_plain;f=config.guess;hb=HEAD'
RUN wget -O config.sub 'https://git.savannah.gnu.org/gitweb/?p=config.git;a=blob_plain;f=config.sub;hb=HEAD'
RUN cp config.* /mnt/tools/src/make-4.2.1/config/
# build tools
RUN mkdir -p /mnt/config
RUN cd /mnt && echo 'y' | SKIPTOOLSINTRO=1 FORCE_UNSAFE_CONFIGURE=1 MAKEFLAGS=-j16 ./tools/src/buildtools
RUN mkdir -p /mnt/config
RUN cd /mnt && . ./shrc && packagetools linux-loong64
RUN /mnt/install.sh -f
每次没有 UPS 或 UPS 容量不够用的倒闸对于运维来说都是一次鸡飞狗跳。这次很不幸,鸡飞狗跳终于轮到了我,还好花了一个半小时还是解决了。在这里做个简单的复盘。
前几天参加了系里的关于 AI 时代的 CS 教育的研究生论坛,在论坛上我分享了一些小的思考,也在论坛上得到了许多不同的想法,于是把一些想法记录下来,过一段时间再回来看看,到底 CS 教育应该怎么办。
最近在和 @CircuitCoder 交流 SDRAM(通常简写为 DRAM,或更进一步简写为 DDR)的各种性能指标,于是想到利用现有的 DRAMSim3 和 Ramulator2 做一些模拟测试,看看各种访存模式下可以实现峰值带宽的多少比例,再结合时序验证理论与模拟结果是否吻合。实验相关代码已开源至 jiegec/dram-bench。
最近有同学遇到这么一个问题:在 Nginx 反代后面搭了一个使用 SSE(Server Sent Events)机制的服务端,但客户端观察到请求延迟比较高,数据批量到达,而不是一行一行地出现。经过排查,发现是 Nginx 的 buffering 机制导致的。本文通过实验复现该问题,并探索了几种解决方法。
最近遇到一个运维场景,两个 SATA 盘组了一个 RAID1,Linux 的根系统也在上面,启动时能进内核,但是内核一直在报错 link is too slow to respond, please be patient 以及 COMRESET failed (errno=-16)。下面记录一下故障排查以及恢复的过程。
继 IBM POWER8 之后,也来评测一下后续的 IBM POWER9 微架构。IBM POWER9 有 SMT4 和 SMT8 两种版本,我只有 SMT4 版本的测试环境,下列所有评测都是针对 SMT4 版本进行测试。