~$ whoami
YJMSTR — I build systems for large models: attention kernels, training frameworks, inference engines, and the occasional RTL simulator. I like making "unsupported" hardware work, and I'm fascinated by agentic workflows. Off the keyboard: astrophotography, galgame, gadgets & desk setups.
~$ ▊
~/about
My interests are hardware–software co-design, performance optimization, and automated (agentic) workflows.
I am currently an intern at Z.ai, and I have also stayed at Baidu, NYU Shanghai, AMD, and ISCAS.
~/blog
-
从电路复杂性理论的角度分析大模型的可并行性和表达能力--Olmo Hybrid 系列
虽然这两篇有先射箭再画靶之嫌,但还是给大伙提供了一个不同的视角来看Linear RNN,值得一提
-
db-SP:稀疏注意力头并行和序列并行的负载均衡问题
考虑下面两个场景:
-
Mirage Persistent Kernel
2025.09.18 更新: Mirage 团队在 GPU MODE 上给了一次直播,介绍 Mirage 和 MPK 的实现:
-
FlashDMoE:分布式 MoE 执行范式的变革
继续尝试用沈向洋、华刚:读科研论文的三个层次、四个阶段与十个问题 - 知乎的十个问题读论文。