开源贡献
6 个已合并进上游项目的 PR,从新到旧。每一条都标了原始标题和编号,点进去能看完整 diff。
- 已合并 PR
- 53全部仓库合计
- 进入上游项目
- 6vLLM / Mooncake / lmdeploy / fla / TIRx
- 自建工具
- 46kernel_opt_agent
- 上游 PR Review
- 9DeepSpeed / TensorRT-LLM / vLLM 等 6 个仓库
数据来自 GitHub 公开接口,每日自动同步 ·
查看原始 JSON
进入上游项目的 PR
Report CUDA inspection tool timeouts through DumpResult
CUDA 检查流程里的各个 stage helper 已经会把「非零退出码」翻译成一个失败的检查结果。
但底层的 _run 只处理了「可执行文件不存在」这一种失败——
nvcc 或 cuobjdump 卡住时抛出的 TimeoutExpired 会直接逃出
dump_module,调用方拿不到 DumpResult.errors,
失败走了另一条通道,整条错误处理链就断了。
改成把超时也转成 CompletedProcess,复用既有的失败路径;同时保住部分
stdout/stderr(TimeoutExpired 在 text=True 下也可能携带 bytes,
需要显式解码)。原有的 300 秒上限不动。
+ except subprocess.TimeoutExpired as e:
+ # TimeoutExpired can carry bytes even with text=True. Keep any partial
+ # diagnostics, but let the existing stage helpers report a failed run.
+ stdout = e.stdout or ""
+ stderr = e.stderr or ""
+ if isinstance(stdout, bytes):
+ stdout = stdout.decode(errors="replace")
2 个文件 · +68 −1 · 合并于 2026-10-05
验证:3 个回归用例用真实卡住的 Python 子进程顶替 CUDA 工具并调短超时,
覆盖 ptx / cubin / SASS 三条路径,另外单独覆盖 bytes、text 和无部分输出三种情况。
原 main 上 6 failed / 5 passed,打上补丁后 11 passed。
[Fix] Preserve RWKV7 orthogonal initialization for low-precision weights
原来的写法把权重复制成 float 做正交初始化,再 .to() 转回原精度。
这个 round-trip 会重新分配张量,低精度下丢掉已经算好的数值。改成在 float 副本上初始化完
copy_ 回原张量,不再走一次转换。
- oringinal_dtype = weight.dtype
- weight = weight.float()
- nn.init.orthogonal_(weight, gain=gain)
- weight = weight.to(oringinal_dtype)
+ float_weight = weight.float()
+ nn.init.orthogonal_(float_weight, gain=gain)
+ weight.copy_(float_weight)
1 个文件 · +3 −4 · 合并于 2026-10-03
Keep ptxas resource counts scoped to the first kernel report
parse_ptxas 原来在整份日志里搜数字。一个 cubin 里有多个 kernel 时,
第一个 kernel 可能压根不报静态 smem,于是它的寄存器数会配上后面某个 kernel 的 smem 或 spill,
拼出一个从来没存在过的组合。改成先把日志切到第一个 report 再解析。PR 里带了 CUDA 13.0 SM103a 的真实编译日志做回归。
+ # Optional fields may be absent for the first kernel. Searching the entire
+ # log would then borrow a resource count from a later, unrelated function.
+ entries = list(re.finditer(r"(?m)^ptxas info\s*:\s*Compiling entry function", log))
+ if not entries:
+ entries = list(re.finditer(r"(?m)^ptxas info\s*:\s*Function properties for", log))
+ if entries:
+ end = entries[1].start() if len(entries) > 1 else len(log)
+ log = log[entries[0].start():end]
2 个文件 · +37 −0 · 合并于 2026-09-30
[TE] Bound pooled TCP admission by queued bytes
经典 TCP 路径按 item 条数限制每个连接组能排多少活,但每个 item 在真正开始或离开队列之前
都攥着自己的源 buffer 不放。条数相同,一次几 MB 的请求和一个几 KB 的请求占的资源差两个量级。
加了按字节计的上界 MC_TCP_MAX_QUEUED_BYTES_PER_PEER,超了返回 QUEUE_FULL。
这个 PR 花了 360 行,其中 266 行是可见性测试:验证被拒绝的传输不会改到目标内存、
已准入的字节在传输完成后正确释放、以及批量准入中途失败不会让调用方误以为源 buffer 已可用。
// Bytes whose source buffers remain owned by the queue or pending
// admission. Active lanes are excluded and already bounded by
// lanes_per_peer.
uint64_t queued_byte_capacity;
uint64_t admitted_bytes = 0;
5 个文件 · +360 −7 · 合并于 2026-09-24
fix: restore getenv after environment parsing errors
set_envs() 是个 context manager,解析模块级环境配置时会临时替换 os.getenv。
原来 yield 后面直接跟还原语句,没包 try。任何一个环境变量取值非法导致解析抛异常,
os.getenv 就永远停在打补丁的版本上,后面所有读环境变量的代码都拿到错的东西。
- yield
- os.getenv = _origin_get_env
+ try:
+ yield
+ finally:
+ os.getenv = _origin_get_env
1 个文件 · +4 −2 · 合并于 2026-09-23
[Bugfix][Multimodal] Preserve DeepSeek V4 image block spacing
DeepSeek V4 的 renderer 让 parse_chat_messages 返回 flatten 成字符串的内容,
text 和 image block 之间的结构边界在 tokenizer 套用官方空行模板之前就没了。
修完之后,两张图的输入会正确产出下面这段 token 序列——图片前后各有独立分隔符,而不是被并成一片。
<|begin▁of▁sentence|>You are a helpful vision assistant.<|User|>请按"第一张、第二张"的顺序回答:第一张图
<|deepseek_image|>
和第二张图
<|deepseek_image|>
中分别是什么食材?<|Assistant|></think>
5 个文件 · +60 −12 · 合并于 2026-09-18