先给出可理解、可引用的结论
“Outlier 有影响”本身并不足够;必须说明影响发生在哪里、效应多大,以及它是否足以改变 production decision。本研究得到的是一个分层结论。
Outlier 会让更多小权重变成零
Tensorwise scale 由最大值决定。不同 output channels 的最大值越不均衡,小权重发生 nonzero-to-zero underflow 的比例越高;Spearman ρ=0.812,95% CI [0.773, 0.846],属于稳定的强单调关系。
但这些零值携带的能量很小
Weight amax/p99.99 与 E4M3 relative-L2 的 ρ=-0.046,95% CI [-0.129, 0.037],统计上不支持单调关系。全局 underflow 仅约 0.00960%,且主要位于低能量区。
E4M3 与 uniform INT8 的 outlier 机制不同
INT8 使用全 tensor 共享的固定步长,极值会线性放大主体分布的步长;其 tail-error ρ=0.817。E4M3 具有 exponent,各数量级保留近似相对精度,因此 outlier 代价更多转移到动态范围底部。
Activation outlier 是更有价值的后续方向
真实 activation tail 与 tensorwise local error 的 ρ=0.380,与 rowwise MSE gain 的 ρ=0.442。该效应为中等强度,且有显著 family structure,但必须经过更广 heldout 与端到端配对实验。
13.690B 权重上的 exact global relative-L2;逐层高度集中。
Output-channel weight scale 减少的 underflow count;MSE 仅改善 0.476%。
95% bootstrap CI [0.390, 0.493],比 weight-tail effect 更明确。
两个 calibration prompts 使用相同 seed,不能代表 population-level quality。
Interpretation boundary
本研究测量 weight reconstruction、sampled activation reconstruction 与 descriptive risk ranking。它不证明某一层导致最终视频或音频质量下降,也不批准 rowwise、clipping、sparse residual 或其他 production quantization modification。
衡量两个变量排名是否共同升降。ρ 接近 1 为强正关系,接近 0 表示没有稳定单调关系;它不是“解释了多少百分比”。
按 12 个 Linear families 分层重采样 4,000 次。区间跨 0 时,不将观察到的方向视为稳定关联。
Underflow 统计丢失了多少非零元素;relative-L2 按平方能量加权。大量极小值变零,仍可能只贡献很小总能量。
四张图概括主要发现
Scatter plot 回答关联强弱,heatmap 回答风险集中在哪里,temporal map 回答 streaming 过程中何时发生,joint-risk map 用于选择下一轮实验层。
从数据定义到工程决策的五步证据链
分析遵循 Data provenance → weight mechanism → counterfactual interventions → activation dynamics → joint-risk localization 的顺序,避免由单一相关系数直接跳到 production action。
冻结 baseline27FPS 的 576-layer production manifest
实现来源唯一锁定为 /workspace/week2-improved-fp8_scaled。后续 week 的代码候选不被采纳;week4 observer bundle 仅作为 measurement input,并与 week2 manifest 逐层交集验证。
48 Transformer blocks × 12 Linear families。
FP8 errors、underflow 与 scale counterfactuals 全元素计算。
2 cells × 576 modules × 55 forwards。
每个主相关按 12 families 分层重采样。
aten._scaled_mm 输出 BF16。研究分别观测 scale、rounding、underflow 与 local error。Exact measurements
- Weight amax/channel amaxall elements
- FP8 error & underflowall elements
- Scale grid errorsall elements
- Activation amaxfull call
Bounded deterministic probes
- Weight quantiles/kurtosis≤262,144/layer
- Activation p50/p99.9≤8,192/call
- Activation quant error≤64 rows/call
- INT8 control≤262,144/layer
Two 20 s requests: person-car and person-ceramics, seed 12345。
832×480、489 frames、11 chunks、55 forwards,与 576-layer manifest 一致。
模块数 n=576 不等于 576 个独立内容实验;activation population 只有两个 cells。
区分 weight tail、channel skew、underflow 与 energy-weighted error
Tensor tail T(w)=amax/p99.99 衡量单 tensor 极端值;channel skew S(w)=max channel amax / median channel amax 衡量通道间不均衡。两者对应不同的统计后果。
| Relationship | ρ | 95% CI | p | Interpretation |
|---|---|---|---|---|
| Weight channel skew → underflow | 0.812 | [0.773, 0.846] | 1.62e-136 | Strong |
| Weight amax/p99.99 → E4M3 RelL2 | −0.046 | [−0.129, 0.037] | 0.271 | Not supported |
| Channel skew → channelwise MSE gain | 0.200 | [0.122, 0.280] | 1.30e-6 | Small engineering effect |
| Weight tail → dense clipping gain | −0.065 | [−0.144, 0.013] | 0.121 | Not supported |
| Format | Median RelL2 | Tail ρ | 95% CI | Evidence scope |
|---|---|---|---|---|
| E4M3 | 2.648% | −0.046 | [−0.128, 0.040] | Exact all weights |
| Uniform INT8 | 4.849% | 0.817 | [0.784, 0.847] | Deterministic sampled control |
Global nonzero-to-zero rate 为 9.601×10⁻⁵,即 0.00960%。
E4M3 exponent 缓冲 max outlier 对 normal-range relative precision 的影响。
不依据 weight tail severity 单独批准 clipping 或更细 scale。
量化每一种 outlier treatment 的误差—存储代价
对十个 scale multipliers 计算 exact dense clipping 与 optimistic exact-FP32 sparse residual。Sparse 结果仅是离线上界,不包含真实索引布局、alignment、kernel overhead 或端到端质量。
| Strategy | RelL2 | MSE reduction | Bytes/W | BF16 compression |
|---|---|---|---|---|
| Production tensorwise FP8 | 2.6497% | 0% | 1.0000 | 2.000× |
| Output-channel FP8 | 2.6434% | 0.476% | 1.0008 | 1.998× |
| Per-layer best dense clip | 2.6478% | 0.148% | 1.0000 | 2.000× |
| Sparse residual α=.5 | 2.6479% | 0.135% | 1.0003 | 1.999× |
| Sparse residual α=.25 | 2.6097% | 2.995% | 1.0214 | 1.958× |
| Sparse residual α=.125 | 2.3477% | 21.495% | 1.2463 | 1.605× |
| Sparse residual α=.0625 | 1.7045% | 58.619% | 2.1408 | 0.934× |
Dense clipping result
即使使用“每层看完结果后选最佳 α”的 optimistic oracle,全局 MSE 也仅改善 0.1479%,且 57.47% layers 仍选择 α=1。α=.25/.125/.0625 分别使 dense MSE 变为 production 的 3.71×/34.88×/175.25×。
Underflow count −93.89%,但 global MSE 仅 −0.476%。
现有 E4M3 不形成有意义的全局收益。
仅保留 α=.25 作为未来 kernel prototype 的 reference point。
Activation tail 与局部误差关系更强,但存在 prompt dependence
每个模块先在每个 cell 内聚合 55 calls,再跨两个 cells 合并。该层级避免单次峰值支配模块排序,同时保留 NFE、chunk 与 family 的 streaming structure。
Complete 12-family statistical summary
| Family | W skew median | W underflow p95 | Channel MSE gain | Act tail p95 | Rowwise gain | Top-1% Jaccard |
|---|---|---|---|---|---|---|
| S1-Q | 4.585 | 0.0107% | 0.493% | 4.151 | 6.491% | 0.547 |
| S1-K | 4.739 | 0.0101% | 0.788% | 4.151 | 6.491% | 0.547 |
| S1-V | 3.121 | 0.0052% | 0.510% | 4.151 | 6.491% | 0.547 |
| S1-O | 3.999 | 0.0050% | 0.496% | 3.350 | 1.216% | 0.180 |
| S2-Q | 8.575 | 0.0151% | 0.619% | 11.947 | 49.129% | 0.783 |
| S2-K | 6.580 | 0.0073% | 0.417% | 4.320 | 6.965% | 0.673 |
| S2-V | 4.447 | 0.0080% | 0.518% | 4.320 | 6.965% | 0.673 |
| S2-O | 7.325 | 0.1325% | 0.780% | 22.821 | 9.011% | 0.180 |
| A2V-Q | 3.588 | 0.0119% | 0.550% | 7.960 | 31.978% | 0.708 |
| A2V-O | 2.495 | 0.0015% | 0.534% | 8.286 | 5.720% | 0.183 |
| FF-Up | 5.834 | 0.0167% | 0.620% | 4.103 | 2.299% | 0.426 |
| FF-Down | 4.656 | 0.0260% | 0.244% | 6.218 | 1.261% | 0.176 |
| Relationship | ρ | 95% CI | p | Interpretation |
|---|---|---|---|---|
| Activation tail → tensorwise local RelL2 | 0.380 | [0.317, 0.446] | 3.29e-21 | Moderate |
| Activation tail → rowwise MSE gain | 0.442 | [0.390, 0.493] | 6.67e-29 | Moderate |
| Weight tail ↔ activation tail | 0.085 | [0.014, 0.159] | 0.0415 | Weak; fails Holm |
Stable evidence
- Within-module channel rankmedian ρ=.913
- Top-1% channel overlapJaccard=.547
- Same argmax channel62.5%
Generalization warning
- Module-risk rank across cellsρ=.436
- Median error location shiftWilcoxon p=.741
- Independent seeds0
Why local rowwise gain is not a production decision
Week4 历史证据显示 full-rowwise activation 可在局部 MSE 上改善,但六格 trajectory/media gate 存在明显 tail regression;SmoothQuant 亦在 calibration 改善而 heldout 失败。因此本研究只支持缩小 candidate scope,不支持重新启用 global rowwise。
联合 weight 与 activation symptoms,选择可归因实验层
Joint score 是四个 percentile ranks 的平均:weight tail、weight underflow、activation tail 与 rowwise local gain。它仅用于 experimental selection,不能解释为质量失败概率。
完整 top-15 joint-risk result table:以下结果同时覆盖 late S2-O 与 S2-Q,用于构造下一轮 high-risk intervention set。
| Top joint-risk module | Risk pct. | W tail | Underflow | Act tail | Row gain |
|---|---|---|---|---|---|
| transformer_blocks.36.attn2.to_out.0 | 97.96 | 56.62 | 0.112% | 110.12 | 40.63% |
| transformer_blocks.41.attn2.to_out.0 | 96.01 | 52.09 | 0.079% | 47.69 | 34.04% |
| transformer_blocks.38.attn2.to_out.0 | 95.01 | 60.59 | 0.104% | 31.10 | 27.92% |
| transformer_blocks.37.attn2.to_out.0 | 94.92 | 58.82 | 0.143% | 45.96 | 21.52% |
| transformer_blocks.40.attn2.to_out.0 | 94.31 | 64.00 | 0.113% | 57.44 | 13.69% |
| transformer_blocks.35.attn2.to_out.0 | 93.75 | 45.40 | 0.087% | 72.91 | 13.30% |
| transformer_blocks.32.attn2.to_q | 93.58 | 7.14 | 0.012% | 16.67 | 67.10% |
| transformer_blocks.47.attn2.to_q | 92.88 | 4.50 | 0.013% | 49.72 | 78.52% |
| transformer_blocks.42.attn2.to_out.0 | 92.80 | 52.22 | 0.101% | 19.67 | 14.49% |
| transformer_blocks.45.attn2.to_out.0 | 92.45 | 39.17 | 0.079% | 13.76 | 28.92% |
| transformer_blocks.34.attn2.to_q | 92.06 | 4.77 | 0.016% | 17.32 | 69.64% |
| transformer_blocks.44.attn2.to_out.0 | 92.01 | 52.40 | 0.081% | 16.64 | 18.46% |
| transformer_blocks.35.attn2.to_q | 91.93 | 5.09 | 0.016% | 15.95 | 65.44% |
| transformer_blocks.39.attn2.to_q | 91.86 | 4.84 | 0.016% | 17.10 | 66.19% |
| transformer_blocks.38.attn2.to_q | 91.36 | 4.95 | 0.014% | 15.82 | 62.42% |
Late S2-O combines weight underflow and activation tail;S2-Q primarily exposes activation rowwise gain。
固定 block 35–45 的 high-risk set,并构造 family/block-matched low-risk controls。
Risk percentile 不等于 failure probability,也不代表 media importance。
当前允许做什么,不允许做什么
统计证据支持 candidate selection,而不支持直接修改 production quantization policy。
Evidence-supported actions
- 保留 production tensorwise E4M3 作为 paired baseline。
- 优先研究 late S2-O/S2-Q activation-side interventions。
- 使用 family/block-matched low-risk layers 作为 negative control。
- 若因果 lane 成立,再评估 α=.25 sparse residual kernel prototype。
- 以 heldout cells 和最差 cell 作为 production gate。
Not supported by current evidence
- Global dense clipping。
- 仅因 weight outlier 高就全局改 output-channel weight scales。
- 把 sparse-residual offline upper bound 当成可部署方案。
- 由两个 calibration prompts 重新启用 global rowwise activation。
- 将 local RelL2/MSE improvement 表述为 media quality improvement。
| 问题 | 定量结果 | 工程含义 |
|---|---|---|
| Weight underflow | ρ=.812;channelwise underflow −93.89% | Count-level effect 大且稳定。 |
| Weight energy error | ρ=−.046;global RelL2 2.6497% | Extreme tail 不是当前 E4M3 total error 的主导因素。 |
| Output-channel weight scale | Global MSE −0.476% | 修复大量低能量 underflow,但 energy gain 很小。 |
| Dense clipping | Optimistic oracle MSE −0.148% | 不构成有意义 production candidate。 |
| Sparse α=.25 | MSE −2.995%;compression 1.958× | 仅为最早可考虑的 counterfactual Pareto reference。 |
| Activation local error | ρ=.380 / .442;family ε²=.658 / .687 | 值得做定向 causal lanes,但不能全局推广。 |
下一轮:从 descriptive statistics 进入 causal validation
优先固定 block 35–45 的 S2-O/S2-Q high-risk layers,并构造 family/block-matched low-risk controls。每个 lane 只改变一个因素。
Baseline
576 层 production tensorwise FP8。
Weight rescue
仅 top-15 layers 使用 output-channel weight scale。
Activation rescue
仅 top-15 layers 使用 dynamic rowwise activation。
Oracle
仅 top-15 layers 恢复 BF16 Linear,估计可恢复上限。
Negative control
对 15 个 matched low-risk layers 做同样处理。
Minimum data matrix
至少 6 content prompts × 3 seeds = 18 paired cells。Calibration 仅用于固定 layer set 与 policy;最终结论只使用未参与选择的 heldout cells。Cell 是独立重采样单元,module calls 是嵌套观测;推荐 mixed-effects model 或 cluster bootstrap。
Required quality evidence
- 逐层 output RelL2/cosine 与首次分叉位置
- Video/audio latent 与 velocity paired error
- Decoded video/audio quality 与 worst cell
- AV consistency / perceptual review
Required systems evidence
- DiT generated FPS 与 formal wall
- Peak allocated/reserved memory
- Fallback/upcast/scale bytes
- 每 cell paired delta 与 determinism
证据边界与复现资产
Sampling boundary
Weight error、underflow 与 scale-grid 为全量精确值;weight quantiles 与 INT8 control 每层最多 262,144 samples。Activation amax 为 full call,percentiles 与 quant-error 为 bounded deterministic samples。
Population boundary
仅两个 prompts 且共享 seed,无法独立估计 prompt、seed 及交互。Family-stratified bootstrap 描述 layer-level association,不提供 system-quality approval。
Metric boundary
Weight-space isotropic L2 对 task importance 一视同仁;低能量 weight 仍可能被高幅 activation 放大。Local reconstruction 不可替代 end-to-end media metrics。
Counterfactual boundary
Sparse residual 假设 exact FP32 values + uint32 indices,不含 row pointers、alignment、kernel overhead 或 runtime;INT8 是 mechanism control,不是 production candidate。
Provenance summary
Experiment: catnip_fp8_outlier_statistics_v1
Checkpoint SHA256: d09adb8c85c1ae9681958e4c45b4f6bdcfc3e596d3782ab2f76bd9cd9824734c
Manifest SHA256: c085f53dbdac584a6f53b000ba5efb59892d512dcd9b42a4a8f8f53eb87acdf9
Raw JSONL SHA256: 9346a7013b56ba49724e1b7372756057ccdee4ff8f8ae7f225903a2be5a1add5
Random seed: 20260721
Embedded reproducibility artifacts
Technical PDF 与 primary analysis summary 可由页面顶部下载;以下附件同样直接内嵌于本 HTML。