Week 10 baseline-bound observational study

Catnip KV Cache
Attention Topology & Sparsity

对后续 steady-state chunks 的真实 video self-attention 边界进行可视化与统计分析:区分 structural sparsityeffective concentration 与 MXFP8 numerical zeros,并量化 sink / recent cache / current chunk 各自承担的 attention mass。

Capture PASS · 42/42 129,024 query-head rows 7 layers · 3 chunks · 2 NFE positions All 32 heads Generated 2026-07-28T04:47:24.652915+00:00

总体结论 / Primary conclusions

先给出可以由当前证据支持的结论,再展开实验方法。所有比例均来自同一个冻结 prompt/seed 的 observational capture,不冒充独立多次模型运行。

6.76%
Median effective support
≈ 422 / 6,240 keys
86.43%
Median mass carried by top 10% keys
40.01%
Mean sink + recent cached-history mass
1.52%
Median Q/K MXFP8 vs BF16 probability TV
结论 A — 存在显著的 effective sparsity,但不存在统一的 structural sparsity。 Softmax strict-zero rate 为 0.00%;然而 65.23% 的 rows 只需 ≤25% keys 即可解释其 entropy-equivalent support,53.85% 甚至只需 ≤10%。同时仍有 17.84% rows 的支持集 ≥75%,因此不能将整个 attention 声称为同一种稀疏模式。
结论 B — KV Cache 被实质使用,但相对 token count 被低配。 cached history 占 10/16 key frames,uniform expectation 是 62.5%;实际平均 mass 为 40.01%,current chunk 为 59.99%。history 不可忽略,却也不是均匀依赖:20.54% rows 对 history 的 mass <10%,43.23% rows 则超过 50%。
结论 C — 量化扰动通常小于 attention topology 自身的异质性。 Q/K MXFP8 重建相对 BF16 reference 的 median TV 为 1.52%,argmax agreement 为 94.33%;P-path numerical-zero median 为 0.00%,表明大多数 rows 的有效集中不是由 FP8 下溢凭空制造。
Executive metrics
Figure 1. 四个 headline metrics。注意 instrumented run 的 wall time 被明确排除在 baseline 性能证据之外。

Week 10 baseline binding

研究绑定到 Week 10 single_tile_compile 路径;observer 只在 compiled video-self-attention 边界读取数据,不改变 KV、Q/K/V、provider output 或 VAE 策略。

Accepted performance evidence

34.392 FPS
Week 10 paired evidence 的 candidate mean generated FPS

相对 Phase2 single-tile lane 的 paired median gain 为 1.837%。这是外部继承的 performance evidence;本研究不重新估计该数字。

Instrumented observational run

32.204 FPS
Observer-on run,仅用于确认完整 pipeline 与采集 gate

该值包含 Q/K 再量化、概率重建、统计与 D2H 落盘,不得作为 baseline regression。该 run 的 Week 10 correctness gate 为 PASS

Acceptance boundary. Week 10 的 performance lane 通过,但 quality panel 未执行;因此 quality_accepted=falseproduction_accepted=false。本报告同样不作“画质没有下降”的表述。
Sink · 4 LF1,560 tokens
Recent cache · 6 LF2,340 tokens
Current chunk · 6 LF2,340 tokens

steady later chunk 的 query 为 6 LF / 2,340 tokens;K/V 共 16 LF / 6,240 tokens,物理顺序为 [sink4, recent6, current6]。每个 latent frame 对应 15×26 = 390 video tokens。

实验设计 / From production boundary to statistics

选择层、chunk 与 query 的原则在运行前冻结。采集没有保存巨大的 raw hidden-state payload,而是在 GPU 侧即时约化为概率统计和小型 heatmap arrays。

1 · Real boundarypost-RoPE、KV concat 后的 production Q/K;selected compiled attn1
2 · Stratified sampleL0/7/15/23/31/39/47,chunks 4/7/10,NFE 0/3
3 · All heads32 heads;每个 query LF 取 4×4 spatial grid,共 96 queries/head
4 · Reconstruct Pproduction signed-H128 + physical E4M3/E8M0 K32 Q/K
5 · Reduceentropy、top-k mass、region mass、TV、temporal/spatial heatmaps

Structural

重建 FP32 softmax 中严格等于零的比例。它回答“矩阵是否真的有 zero entries”。

Effective

exp(H)/Sk 衡量 entropy-equivalent support;top-1/5/10% mass 衡量 concentration。

Numerical MXFP8

模拟 kernel 的 per-K32 E4M3/E8M0 probability path,统计 underflow zeros 与局部 TV。

为什么这不是 129,024 次独立实验?
129,024 是 query-head rows 的描述性样本量:42 records × 32 heads × 96 sampled queries。它们嵌套于一个 prompt、一个 seed、一个完整模型运行中,存在强相关;因此报告不基于 row count 构造虚假的独立置信区间,也不把它写成 129,024 次 inference。

KV Cache 的实际作用

最直接的答案不是“cache 有或没有用”,而是后续 chunk 把多少概率质量分配给缓存历史、这种依赖如何随 layer / head / NFE 变化。

13.98%
Mean sink mass · uniform 25%
26.03%
Mean recent mass · uniform 37.5%
59.99%
Mean current mass · uniform 37.5%
1.60×
Current per-token enrichment vs uniform
Temporal attention heatmap
Figure 2. 平均 frame-to-frame attention heatmap。每个 query-LF row 的 16 个 key-LF mass 加和为 100%。白线分隔 sink / recent / current。
Region mass by layer
Figure 3. 深度方向的 region mass。history reliance 从 L7 的 52.82% 下降到 L47 的 30.46%;这更支持 layer-aware policy,而不是统一裁剪。

Layer summary

LayerSinkRecentCurrentHistoryMedian support
L018.37%27.66%53.97%46.03%85.67%
L720.33%32.50%47.18%52.82%47.80%
L1517.14%29.66%53.20%46.80%38.29%
L2311.18%23.87%64.96%35.04%4.48%
L3112.38%24.57%63.05%36.95%1.31%
L3913.18%18.77%68.05%31.95%1.16%
L475.27%25.19%69.54%30.46%0.60%

Interpretation

  • history 平均 mass 为 40.01%,低于按 token count 的 62.5%,但仍是局部 attention output 的主要组成之一。
  • sink 平均 13.98%,recent 平均 26.03%;两者都被 under-attend,而 current 被显著 over-attend。
  • chunk 4/7/10 的 history mean 仅在 39.79%–40.28% 之间,说明该冻结内容的 later-chunk topology 相对稳定。
  • NFE 0 → NFE 3 时 history 从 42.48% 降至 37.54%,denoising 后段更依赖 current context。
Chunk and position stability
Figure 4. 跨后续 chunk 稳定,但随 denoising position 有系统变化。该结果提示 runtime policy 可能需要同时感知 layer 与 NFE。

“稀疏性很大吗?”——是,但必须说完整

Median row 高度集中;分布尾部却包含一批接近 uniform 的 heads/queries。有效稀疏性是潜在优化信号,不等于已有结构化 sparse pattern。

53.85%
Rows with effective support ≤10%
65.23%
Rows with effective support ≤25%
55.24%
Rows whose top 10% keys carry ≥80%
17.84%
Diffuse minority: support ≥75%
Sparsity distributions
Figure 5. 左:effective support 的 ECDF;右:top-ranked key mass。分布不是单峰,说明 average-only 结论会掩盖两类不同 attention rows。
Head heterogeneity
Figure 6. 每个 layer/head 的 support 与 cached-history mass。Head identity 带来明显差异;例如 H27/H8 高度集中,而 H0 更 diffuse。

Most concentrated heads

HeadMedian supportMean supportMean top10History
H270.37%8.97%88.44%21.97%
H81.07%4.22%94.59%25.13%
H291.11%15.32%82.30%27.75%
H261.21%14.43%83.56%27.24%
H71.33%12.27%87.03%27.34%
H191.45%18.77%80.38%37.83%
H231.68%17.74%81.10%44.17%
H211.70%10.80%85.48%30.12%

Most diffuse heads

HeadMedian supportMean supportMean top10History
H061.12%57.34%43.35%56.47%
H1345.77%49.57%51.13%45.75%
H240.64%53.06%46.46%51.65%
H128.18%42.80%56.77%66.63%
H425.94%46.26%55.69%56.29%
H1025.05%42.31%60.52%34.85%
H1125.01%34.73%61.67%39.71%
H1523.69%37.72%61.44%37.28%
不能直接推出 sparse-kernel speedup。 Top-k indices 目前不是预先结构化 block,选择与压缩本身有成本;17.84% diffuse rows 还需要 fallback。必须先证明可预测的 block pattern、局部 output error 与 end-to-end quality,再讨论 kernel promotion。

Q/K 与 probability quantization 的影响

这里比较的是同一真实 BF16 boundary:一边经过 production signed-H128 + E4M3/E8M0 K32 Q/K,另一边用 BF16 boundary tensors 做 FP32 reference matmul。TV 是局部 attention-row 指标,不是视频质量指标。

1.52%
Median Q/K probability TV
94.33%
Q/K argmax agreement
1.06%
Median P-path proxy TV
2.45%
Mean numerical-zero rate · median 0.00%
Quantization effect distributions
Figure 7. Q/K TV 与 P-path TV 通常较小;P zero-rate 强烈右偏,少量 rows 可出现很高 underflow,但 median 为零。

What the data supports

  • 1.20% rows 的 Q/K probability TV >5%;仅 0.03% >10%。
  • 36.63% rows 至少有一个 P-MXFP8 numerical zero;但只有 5.88% rows 的 zero-rate >10%。
  • P proxy row sum 的 median 是 0.999392,表明局部 mass conservation 误差较小。

What it does not support

  • 不能把 Q/K TV 等同于 DiT hidden-state error 或视频 perceptual quality。
  • observer 重建使用 production quantizer bytes,但不声称 bit-exact 复刻 fused kernel accumulator ordering。
  • 不能用平均 2.45% P-zero rate 宣称 kernel 已产生稳定可利用的结构化零块。

Temporal & spatial heatmaps

Temporal heatmap 回答“看哪一帧”,spatial heatmap 回答“在该 region 内看哪个 latent-grid location”。所有图与本页一同部署,不依赖任何外部站点。

Temporal heatmaps by layer
Figure 8. 七个 depth-stratified layers 的 6×16 frame topology。越深的层总体更偏向 current chunk,但不同 query LF 的模式并不完全相同。
Region spatial maps
Figure 9. 在 sink / recent / current 各自内部归一化后的 key spatial preference。该图用于观察空间集中,不应跨 region 比较绝对 mass;绝对 region mass 见 Figure 2–4。
Representative key heatmaps
Figure 10. L23 / chunk10 / NFE3 的代表性 query 对全部 16 个 key latent frames 的 15×26 heatmaps。颜色范围按该 record 的 99.5th percentile 截断,避免单点掩盖结构。

当前判断与下一步优化空间

当前数据支持“建立 selective hypothesis”,尚不支持“改 production KV policy”。最有价值的是将下一轮实验设计成可证伪的 causal ablation。

Recommended sequence

  1. Local output ablation. 对同一 sampled Q/K/V,分别 drop sink、drop recent、drop all history,并重归一化 P;测 pre-projection attention-output rel-L2 与 cosine。
  2. Block predictability. 按 LF×spatial blocks 统计 top-mass blocks 的 Jaccard stability,检查 indices 能否由 layer/head/NFE 预测。
  3. Selective policy. 先在晚层与高集中 heads 上试 cache compression;diffuse heads 保持 dense fallback。
  4. Causal full model. 至少多 prompt、多 seed 做 paired run;同时记录局部 output error、end-to-end quality panel 与真实 kernel timing。

Promotion gates

  • Observed: effective concentration, region mass, depth/head/NFE heterogeneity.
  • Observed: later-chunk temporal stability for this prompt/seed.
  • Not yet: causal proof that cached keys can be removed without output/quality loss.
  • Not yet: structured block pattern that amortizes sparse-index overhead.
  • Not yet: end-to-end FPS improvement on the Week 10 paired lane.
最合理的工程方向:不是把 4+6 KV Cache 整体砍掉,而是研究 layer/head/NFE-aware selective KV。当前层级差异很大:L7 history mean 52.82%,L47 只有 30.46%;这为分层策略提供了明确 hypothesis。

完整 record-level 结果

下表列出全部 42 records,每条包含 3,072 query-head rows。表格可滚动;所有 headline conclusions 均可回溯到这些 records 与内嵌 JSON summary。

RecordLChunkNFEMedian supportMedian top10SinkRecentCurrentHistoryMedian QK-TVMean P-zero
layer00_chunk04_position004077.66%27.06%16.46%25.04%58.50%41.50%0.36%0.01%
layer00_chunk04_position304387.81%22.18%19.33%29.96%50.71%49.29%0.26%1.48%
layer00_chunk07_position007077.91%27.02%17.66%25.08%57.26%42.74%0.38%0.01%
layer00_chunk07_position307388.03%22.24%20.64%29.94%49.41%50.59%0.26%1.49%
layer00_chunk10_position0010077.64%26.26%16.38%25.53%58.09%41.91%0.37%0.01%
layer00_chunk10_position3010388.73%21.86%19.78%30.39%49.82%50.18%0.26%1.46%
layer07_chunk04_position074053.04%44.47%21.71%35.91%42.38%57.62%0.69%0.02%
layer07_chunk04_position374339.12%54.19%17.47%29.77%52.77%47.23%0.72%0.08%
layer07_chunk07_position077052.23%45.32%22.35%34.36%43.29%56.71%0.71%0.01%
layer07_chunk07_position377342.88%51.49%18.59%30.04%51.38%48.62%0.68%0.08%
layer07_chunk10_position0710053.22%44.11%23.05%35.29%41.65%58.35%0.72%0.01%
layer07_chunk10_position3710343.02%51.79%18.78%29.61%51.60%48.40%0.70%0.08%
layer15_chunk04_position0154040.79%52.15%18.33%34.21%47.46%52.54%0.87%0.05%
layer15_chunk04_position3154331.45%59.18%14.98%26.13%58.90%41.10%0.79%0.24%
layer15_chunk07_position0157041.80%51.95%19.04%32.81%48.14%51.86%0.89%0.05%
layer15_chunk07_position3157331.67%58.31%15.55%25.27%59.17%40.83%0.75%0.24%
layer15_chunk10_position01510044.46%49.80%19.57%33.26%47.17%52.83%0.85%0.05%
layer15_chunk10_position31510334.58%57.09%15.37%26.26%58.37%41.63%0.75%0.25%
layer23_chunk04_position023405.34%87.71%11.75%27.82%60.43%39.57%1.68%0.15%
layer23_chunk04_position323433.03%91.30%8.46%24.13%67.40%32.60%1.66%0.16%
layer23_chunk07_position023706.19%86.51%14.72%24.37%60.91%39.09%1.67%0.13%
layer23_chunk07_position323733.36%89.99%9.97%21.09%68.94%31.06%1.65%0.15%
layer23_chunk10_position0231005.80%86.52%13.32%24.73%61.94%38.06%1.66%0.13%
layer23_chunk10_position3231033.29%90.39%8.84%21.04%70.12%29.88%1.66%0.17%
layer31_chunk04_position031401.63%97.04%13.90%25.40%60.70%39.30%2.24%2.12%
layer31_chunk04_position331431.00%98.28%11.63%26.94%61.44%38.56%2.22%3.07%
layer31_chunk07_position031701.80%96.38%14.84%21.81%63.35%36.65%2.22%1.74%
layer31_chunk07_position331730.98%98.34%10.77%24.15%65.08%34.92%2.24%3.20%
layer31_chunk10_position0311001.69%96.64%12.57%22.58%64.85%35.15%2.23%1.67%
layer31_chunk10_position3311030.98%98.24%10.56%26.53%62.90%37.10%2.18%3.31%
layer39_chunk04_position039401.66%94.72%14.10%21.39%64.51%35.49%2.11%3.45%
layer39_chunk04_position339430.66%98.44%10.10%16.63%73.26%26.74%1.96%5.20%
layer39_chunk07_position039701.87%94.14%16.35%20.96%62.69%37.31%2.10%3.34%
layer39_chunk07_position339730.69%98.35%11.59%17.38%71.02%28.98%1.97%5.44%
layer39_chunk10_position0391002.05%93.94%15.85%19.72%64.44%35.56%2.11%3.12%
layer39_chunk10_position3391030.72%98.39%11.08%16.55%72.37%27.63%1.99%5.34%
layer47_chunk04_position047400.63%99.39%6.01%25.04%68.95%31.05%2.26%8.04%
layer47_chunk04_position347430.44%99.69%3.47%21.01%75.52%24.48%2.20%9.80%
layer47_chunk07_position047700.73%99.20%7.61%27.37%65.02%34.98%2.33%8.46%
layer47_chunk07_position347730.46%99.70%4.53%25.11%70.36%29.64%2.23%10.56%
layer47_chunk10_position0471000.89%99.05%6.01%27.72%66.27%33.73%2.36%8.28%
layer47_chunk10_position3471030.50%99.70%3.97%24.88%71.15%28.85%2.25%10.29%
Embedded machine-readable analysis summary

下方数据同样写入 data/analysis_summary.json;HTML 内嵌副本可用于部署后的审计。

Evidence, failure handling & claim boundary

Capture gate

  • formal observer calls: 385 / expected 385
  • warmup observer calls: 175 / expected 175
  • idle calls: 0
  • opaque sink value: 560 / expected 560
  • raw Q/K/V persisted: no; only reduced NPZ statistics

Excluded infrastructure runs

r01 · stage load_te_gpu0 · OutOfMemoryError. 0 attention records; excluded before model forward.
Bound artifacts and SHA-256
ArtifactSHA-256BytesResolved path
Frozen study config783ecdd7fe81c6c2…1,986/workspace/catnip-week10-kv-cache-attention-study/configs/study_v1.json
Methodology7de9496102fb5e38…3,554/workspace/catnip-week10-kv-cache-attention-study/METHODOLOGY.md
Capture manifestb1523df0362ec4f9…210,915/workspace/catnip-week10-kv-cache-attention-study/results/capture_r02/capture_manifest.json
Week 10 accepted pair summary01f286b3609a8456…85,945/workspace/week10-vae-decode-optimization/results/phase3_full_model/pairs_r01/FINAL_SUMMARY_V2.json
Week 10 accepted report560b3cc7dca8cb60…17,550/workspace/week10-vae-decode-optimization/reports/PHASE03_FULL_MODEL_RESULT.md
Instrumented run.jsonecdb7d0a9bdffcdb…43,392/workspace/catnip-week10-kv-cache-attention-study/results/week10_observed_r02/run.json
Generalization boundary. 本报告对一个 frozen prompt/seed 做细粒度 descriptive analysis。它不证明其他内容、其他 seed、landscape geometry 或未来模型版本具有相同 sparsity;也不证明任何 KV pruning 的质量安全性。