Prhub

#52003 [Mypy Fix] Mypy fix for "vllm/model_executor/models/[cC][dD]"

原始 PR 作者 yewentao256 合并时间 2026-08-13 11:33 文件变更 24 提交数 1 评论 3 代码增减 +153 / -66

执行摘要

修复 24 个 c/d 开头模型文件的 mypy 类型错误

PR 标题即目的:Mypy fix for "vllm/model_executor/models/[cC][dD]",作者在评论中说明这是长期工作(longtime work),旨在提升仓库代码质量并避免潜在 bug。当维护者 DarkLight1337 质疑在模型文件上强制类型检查的价值时,作者回应称这类修复借助 AI 工具成本很低,不会给模型厂商上游增加太多摩擦,最终获得批准。

值得快速阅读,作为 mypy 全量覆盖的样板 PR。重点看 dbrx.py 与 deepseek_v2.py 的加载器契约统一、cosmos3_edge.py 的命名对齐、cohere_asr.py 的潜在行为修正;若你维护下游模型或 fork,需要特别关注 DbrxExperts.weight_loader 签名变化。

讨论亮点

核心争论点是“是否值得在模型文件上强制 mypy”。维护者 DarkLight1337 明确表示不看好:模型厂商普遍不关心类型安全,强制检查会给上游贡献增加摩擦。作者 yewentao256 反驳称这是长期质量投入、能避免潜在 bug,且 AI 工具让这类修复成本极低。最终 DarkLight1337 以“Let's see how it goes then”的观望态度批准合并,说明该决策属于试运行性质,后续若模型厂商抱怨会重新评估。

实现拆解

  1. 方法签名与命名对齐:在 cosmos3_edge.py 中,Cosmos3EdgeVisionEncoder.forward 改名 encode,避免与 nn.Module.forward 的隐式协议冲突;Cosmos3EdgeAttention.forward 改名 forward_with_positions 并去掉 **kwargs,由 Cosmos3EdgeAttentionDecoderLayer 显式按名调用;cohere_asr.py 的 CohereASRAttention.forward 增加 encoder_hidden_states: torch.Tensor | None = None 默认参数,使其与子类 CohereASRCrossAttention 的签名一致,满足里氏替换检查。

  2. Optional/Union 类型收窄:在 deepseek_v2.py、deepseek_eagle.py、dots_ocr.py、cosmos3_edge.py、config.py 等文件中为 multimodal_config、cache_config、q_lora_rank、speculative_config、pooler_config 等可选值补 assert;为可能为 None 的成员变量声明显式类型(self.post_trunk_norm、self.xscale、self.indexer_rope_emb、self.indexer);对静态不可达分支补 raise AssertionError 以提供 never 类型。

  3. 权重加载契约统一:dbrx.py 的 DbrxExperts.weight_loader 从旧签名 (weight_name, param_name) 改为 (weight_name, shard_id, expert_id, return_success),与 FusedMoE 专家加载器约定对齐并返回加载成功标志;deepseek_v2.py、deepseek_eagle.py 将 expert_params_mapping 解包变量改名 expert_shard_id,并把 maybe_remap_kv_scale_name 的返回值先存入 remapped_name 再覆写 name,避免 Optional 污染后续流程。

  4. 类型标注修正与潜在 bugfix:dots_ocr.py 的 compute_attn_mask_seqlen 返回类型从 int | None 改为 torch.Tensor | None,grid_thw 局部变量改名 grid_thw_tensor 避免与参数遮蔽;deepseek_v2.py 的 get_attn_backend 返回类型收紧为 type[AttentionBackend];cohere_asr.py 用 is_list_of 校验 att_context_size 的嵌套/扁平结构,CausalConv1D 的 padding 参数类型拓宽为 str | int | list[int] | None,并在 get_encoder_outputs 的两个分支统一传入 seq_lens(input_features 分支此前漏传)。

  5. 配套:无新增测试文件,依赖 pre-commit --hook-stage manual mypy-3.13 校验通过;提交为单个 commit "fix mypy models cd",本 PR 是模型目录 mypy 清理的第一批样例。

文件 模块 状态 重要度
vllm/model_executor/models/cosmos3_edge.py 模型实现 modified 7.12
vllm/model_executor/models/cohere_asr.py 模型实现 modified 6.76
vllm/model_executor/models/deepseek_v2.py 模型实现 modified 6.33
vllm/model_executor/models/dots_ocr.py 模型实现 modified 6.36
vllm/model_executor/models/dbrx.py 模型实现 modified 6.0
vllm/model_executor/models/deepseek_eagle.py 模型实现 modified 6.16
vllm/model_executor/models/config.py 配置校验 modified 5.4

关键符号

encode forward_with_positions get_attn_backend compute_attn_mask_seqlen weight_loader _calc_context_sizes _create_custom_model patch_vit_for_tp

关键源码片段

vllm/model_executor/models/cosmos3_edge.py core-logic

最典型的 mypy 修复样本:视觉编码器 forward→encode、注意力层 forward→forward_with_positions,既消除与基类 nn.Module.forward 的签名冲突,也让 DecoderLayer 的调用点更明确。

class Cosmos3EdgeVisionEncoder(Siglip2VisionTransformer):
    """Adapts Cosmos (T, H, W) metadata to vLLM packed SigLIP2."""
​
    # 原方法名为 forward,会与 nn.Module.forward 的隐式协议冲突,且
    # 参数形状与基类预期不一致;改名 encode 后既通过 mypy 重载检查,
    # 也符合 vLLM 视觉塔“encode 提取特征”的命名习惯。
    def encode(
        self,
        pixel_values: torch.Tensor,
        grid_thw: torch.Tensor,
    ) -> torch.Tensor:
        # SigLIP2 对每一帧独立做注意力,因此把每个 THW 条目展开成
        # T 条独立的 HW 注意力序列,并据此构造 packed 输入的
        # spatial_shapes 与 cu_seqlens。
        grid_thw_cpu = grid_thw.to(device="cpu")
        spatial_shapes = torch.repeat_interleave(
            grid_thw_cpu[:, 1:],
            grid_thw_cpu[:, 0],
            dim=0,
        )
        lengths_cpu = spatial_shapes.prod(dim=-1).to(torch.int32)
        lengths = lengths_cpu.to(
            device=pixel_values.device,
            non_blocking=True,
        )
​
        cu_seqlens = torch.zeros(
            lengths.numel() + 1,
            dtype=torch.int32,
            device=pixel_values.device,
        )
        cu_seqlens[1:] = lengths.cumsum(dim=0)
        max_seqlen = lengths_cpu.max().reshape(1)
​
        # 底层复用 SigLIP2 的 packed 前向实现,传入打包后的像素与分段信息。
        return super().forward(
            pixel_values_packed=pixel_values,
            spatial_shapes=spatial_shapes,
            cu_seqlens=cu_seqlens,
            max_seqlen=max_seqlen,
        )
​
​
class Cosmos3EdgeAttentionDecoderLayer(nn.Module):
    """Pre-norm attention layer for the Cosmos3 Edge dense text model."""
​
    def forward(
        self,
        positions: torch.Tensor,
        hidden_states: torch.Tensor,
        residual: torch.Tensor | None,
        **kwargs,
    ) -> tuple[torch.Tensor, torch.Tensor]:
        if residual is None:
            residual = hidden_states
            hidden_states = self.norm(hidden_states)
        else:
            hidden_states, residual = self.norm(hidden_states, residual)
        # 直接调用 mixer.forward_with_positions,而不是 self.mixer(...),
        # 避免位置参数经 nn.Module.__call__ 分发时与基类签名产生歧义。
        hidden_states = self.mixer.forward_with_positions(positions, hidden_states)
        return hidden_states, residual
​
​
class Cosmos3EdgeAttention(NemotronHAttention):
    """Nemotron-H attention with interleaved multimodal RoPE."""
​
    # 原 forward 带 **kwargs,既掩盖真实参数,也让 mypy 无法校验
    # 与 AttentionLayerBase 的一致性;改为 forward_with_positions
    # 并显式列出 positions / hidden_states 两个参数。
    def forward_with_positions(
        self,
        positions: torch.Tensor,
        hidden_states: torch.Tensor,
    ) -> torch.Tensor:
        qkv, _ = self.qkv_proj(hidden_states)
        q, k, v = qkv.split([self.q_size, self.kv_size, self.kv_size], dim=-1)
        q, k = self.rotary_emb(positions, q, k)
        attn_output = self.attn(q, k, v)
        output, _ = self.o_proj(attn_output)
        return output
vllm/model_executor/models/cohere_asr.py data-contract

除类型标注外还包含几处潜在行为变化:get_encoder_outputs 两个分支统一传 seq_lens、_calc_context_sizes 改用 is_list_of 校验、CausalConv1D padding 类型拓宽,是本次改动中运行时影响最不确定的文件。

以下片段来自 cohere_asr.py 的 _calc_context_sizes 归一化分支:

    # att_context_size 允许扁平 [a, b] 或嵌套 [[a, b], ...] 两种写法。
    # 旧实现用 isinstance(att_context_size_all[0], int) 判断首元素,
    # 在空列表或混合结构下既不健壮也无法收窄类型;改用
    # is_list_of(..., check="all") 做全量检查后,mypy 能确定列表元素
    # 类型,混合 / 非法输入也会在进入后续换算前抛出 ValueError。
    if att_context_size:
        if is_list_of(att_context_size, int, check="all"):
            # 扁平写法:单个 [left, right] 上下文窗口,包成一层列表。
            att_context_size_all = [att_context_size]
        elif is_list_of(att_context_size, list, check="all"):
            # 嵌套写法:逐层指定上下文窗口,直接沿用。
            att_context_size_all = att_context_size
        else:
            raise ValueError("att_context_size cannot mix nested and flat values")
vllm/model_executor/models/deepseek_v2.py data-contract

DeepSeek 系列核心模型文件,包含 get_attn_backend 返回类型收紧为 type[AttentionBackend]、Indexer 分支的 cache_config/q_lora_rank 断言,以及 expert_params_mapping 解包变量改名等。

以下片段来自 deepseek_v2.py 的 DeepseekV32IndexerCache:

class DeepseekV32IndexerCache(torch.nn.Module, AttentionLayerBase):
    """DeepSeek V3.2 indexer 的 fp8 稀疏 KV 缓存层。"""
​
    def __init__(
        self, head_dim: int, dtype: torch.dtype, prefix: str, cache_config: CacheConfig
    ):
        super().__init__()
        self.kv_cache = torch.tensor([])
        self.head_dim = head_dim
        self.prefix = prefix
        self.cache_config = cache_config
        self.dtype = dtype
        compilation_config = get_current_vllm_config().compilation_config
        if prefix in compilation_config.static_forward_context:
            raise ValueError(f"Duplicate layer name: {prefix}")
        compilation_config.static_forward_context[prefix] = self
​
    def get_kv_cache_spec(self, vllm_config: VllmConfig) -> KVCacheSpec:
        # 稀疏索引缓存只存 K 方向的量化向量,因此 KV cache 规格
        # 是单向量 MLA spec,而非完整的 K + V 两份。
        return MLAAttentionSpec(
            block_size=self.cache_config.block_size,
            num_kv_heads=1,
            head_size=self.head_dim,
            dtype=self.dtype,
        )
​
    # forward 是空实现占位,仅满足 AttentionLayerBase 的抽象要求,
    # 真正的索引计算在 SparseAttnIndexer 算子内完成。
    def forward(self): ...
​
    # 返回类型从 AttentionBackend 收紧为 type[AttentionBackend]:
    # 基类契约返回的是 backend 类本身(用于后续实例化),
    # 旧标注会让 mypy 在实例化用法上报类型不匹配。
    def get_attn_backend(self) -> type[AttentionBackend]:
        return DeepseekV32IndexerBackend

评论区精华

在模型文件上强制 mypy 是否值得 设计

维护者 DarkLight1337 评论:"Tbh I don't see much value on getting mypy to pass on modeling files. Many model vendors don't care about type safety in their model implementations, enforcing this would just create more friction when they try to upstream their work." 作者 yewentao256 回应:"this is a longtime work, making higher code quality for the repo and avoid some potential bugs. I don't think taking care of this would add friction a lot as this is quite simple with AI now."

结论:作者说服合并者后获批,DarkLight1337 以 "You've made a point with AI being available to help with these issues. Let's see how it goes then" 批准,属于试运行性质的决策。 · 已解决(已合并)

风险与影响

  1. cohere_asr.py 存在两处可能影响运行时行为的变化:get_encoder_outputs 的 input_features 分支此前未传 seq_lens,补齐后编码路径对 padding 的处理可能改变(更接近修复而非回归,但无测试覆盖);_call_hf_processor 中 mm_data 从原地 pop 改为在副本上 pop,原始字典不再被修改,后续读取行为有细微差异。
  2. dbrx.py 的 DbrxExperts.weight_loader 签名变化是接口级破坏:仓库外直接调用该方法的下游代码或 fork 若未同步更新,会因缺少 shard_id 参数直接报错;仓库内调用点已同步。
  3. 大量使用 assert 做类型收窄(deepseek_v2.py、deepseek_eagle.py、config.py、dots_ocr.py 等),在 python -O 优化模式下断言被剥离,后续若依赖这些不变量会产生静默行为差异,24 个文件的规模放大了该风险面。
  4. cosmos3_edge.py 将视觉编码器方法从 forward 改名 encode,若仓库外代码以子类重写 forward 的方式扩展该编码器,改名后重写将失效并落到基类实现。
  5. 无新增测试文件,mypy 通过只证明类型层面自洽,不能证明行为等价;建议在 CI 模型回归中重点观察 cohere_asr 与 dbrx 两条路径。

对用户:模型推理行为理论上不变,但 cohere_asr 编码路径的 seq_lens 补齐可能小幅改变 ASR 输出质量。对开发者:模型文件首次纳入 mypy 门禁,模型厂商上游代码需满足类型检查,维护者对此持保留态度。对社区:建立了“分目录推进 mypy”的先例,后续 [eE]、[fF] 等目录很可能跟进,长期看会显著提升模型代码的可维护性并提前暴露潜在 bug。

无新增测试覆盖 cohere_asr.py 存在行为路径微调 dbrx.py 加载器契约变更 assert 收窄在 -O 模式下失效 24 文件跨模型大范围改动

关联 Issue

未识别关联 Issue

当前没有检测到明确关联的 Issue 链接,后续同步到相关引用后会出现在这里。

完整报告

参与讨论