Skip to content

fix(monkey_patch): patch Q/K norms on existing Qwen models - #1470

Open
luca-888 wants to merge 1 commit into
linkedin:mainfrom
luca-888:fix/qwen-qk-instance-patch
Open

luca-888 wants to merge 1 commit into
linkedin:mainfrom
luca-888:fix/qwen-qk-instance-patch

Conversation

@luca-888

Copy link
Copy Markdown
Contributor

Summary

Patch self_attn.q_norm and k_norm when applying Liger to existing Qwen3, Qwen3 MoE, Qwen3 Next, Qwen3.5 and Qwen3.5 MoE instances.

Reuse each model's RMSNorm helper and respect rms_norm. Extend the existing instance tests following the Qwen3 VL/VL MoE and Gemma3 pattern: construct native HF instances first, then check Q/K norm forward bindings before and after patching. Also verify preserved weights and epsilon, casting settings, and unchanged Linear Attention norms.

Testing Done

  • Hardware Type: H100 80GB
  • Instance tests: 21 passed. All 13 new Q/K regression assertions fail before the fix and pass after it.
  • RMSNorm numerical tests: 64 passed.
  • Convergence: 16 passed, 1 known baseline failure, 5 existing skips.

The BF16 Qwen3 MoE with-logits failure also occurs on the base commit.

These results are from prior validation; tests were not rerun after switching to explicit model classes.

PyTorch 2.9.1 / CUDA 12.8, Triton 3.5.1, Transformers 5.15.1. Hybrid convergence uses FLA 0.5.2 and TileLang 0.1.14.

Before uses the existing Liger instance patch with native Q/K norms; After adds this fix. Both use matched random weights and inputs, BF16, SDPA, identical other kernel settings, and no torch.compile. Each runs in a separate process with warmup and five measurement rounds. Speedup is Before / After.

For these configurations, the patch covers all attention layers in Qwen3-8B (36/36) and Qwen3-30B-A3B (48/48). In Qwen3.5-9B and Qwen3.5-35B-A3B, it covers only the Full Attention layers (8/32 and 10/40); Linear Attention norms are unchanged.

Q/K norms and a single Full Attention module on one H100 80GB, forward + backward median time in ms:

Model Batch × seq Q+K Before → After Speedup Attention Before → After Speedup
Qwen3-8B 4 × 2048 2.439 → 0.562 4.34× 8.883 → 6.705 1.32×
Qwen3-8B 1 × 8192 2.435 → 0.542 4.50× 13.552 → 11.269 1.20×
Qwen3.5-9B 4 × 2048 2.560 → 0.739 3.46× 10.763 → 8.802 1.22×
Qwen3.5-9B 1 × 8192 2.564 → 0.729 3.52× 16.063 → 14.108 1.14×
Qwen3-30B-A3B 4 × 2048 2.219 → 0.643 3.45× 7.071 → 5.280 1.34×
Qwen3-30B-A3B 1 × 8192 2.212 → 0.659 3.36× 11.597 → 9.858 1.18×
Qwen3.5-35B-A3B 4 × 2048 2.292 → 0.740 3.10× 8.248 → 6.324 1.30×
Qwen3.5-35B-A3B 1 × 8192 2.288 → 0.691 3.31× 13.475 → 11.539 1.17×
  • run make test (targeted tests above; full suite not run)
  • run make checkstyle (prior validation)
  • run make test-convergence (targeted tests above; full suite not run)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

1 participant