Catnip Report
MaineCoon Portrait Pipeline · Dual RTX 5090 · Mixed Precision

Catnip Portrait v2: 37.56 stable-DiT / 34.39 generated FPS

当前正式工程基线为 catnip-k32-vdirect-single-tile-compile-portrait-480x832-v2: GPU1 承载完整 Catnip Streaming DiT,GPU0 承载 INT8 text encoder 与 BF16 video/audio VAE, 在 480×832、20 s、489-frame 的 frozen request 上,以 576 个 K32 MXFP8 Linear 与 48 个 V-direct-BSH video-self attention 达到 34.392314 generated FPS37.557043 stable-DiT FPS。 三对正式 candidate runs 给出权威 warm KPI,最终无争用 r03 smoke 以 34.408774 generated FPS 通过;这是一项 warm engineering-performance baseline, 不包含 cold-start speedup 或正式 ground-truth perceptual-quality claim。

Read the current findings → Download PDF
Abstract & principal findings

总体结论 / Primary Conclusions

本研究完成了 Catnip Streaming 模型从 capacity-oriented BF16 model parallelism 到 production-oriented mixed-precision pipeline parallelism 的迁移,并在此基础上完成了 quantization encoding、custom attention kernel 与 acceptance methodology 三次演进。 下表给出可由当前证据支持的最终结论。

Active deploymentPortrait v2 · 480×832 on 2× RTX 5090catnip-k32-vdirect-single-tile-compile-portrait-480x832-v2;GPU1 complete DiT,GPU0 INT8 text encoder + BF16 video/audio VAE。11 chunks、55 forwards、4+6 KV、depth-3 video decode。
Mixed precision576 selected Linear use native K32 MXFP8E4M3 qdata + K32 E8M0 scales + BF16 output/fused bias;13.69B selected weight elements / 72.14% all-Linear coverage。该配置不是 full-model FP8。
Attention48 video-self attn1 use V-direct-BSHStrict fail-closed provider;V-direct consumption is an accepted historical engineering optimization, while the Portrait repeats validate the integrated stack.
Current performance34.392314 generated / 37.557043 stable-DiT FPSThree paired Phase 3 candidate means; population CV 0.085863% / 0.151246%。generated FPS = sampling + pipelined video decode boundary。
Promotion gateFinal uncontended r03 PASS34.408774 generated / 37.566968 stable-DiT FPS;489-frame H.264、AAC、4+6 KV、576 K32、48 V-direct 与 no-new-graph gates PASS。
Historical v1Throughput regime preservedv1 Portrait versus 832×480 V-direct reference changes stable/generated means by only +0.231955% / +0.105325%。It is not a contemporaneous paired causal speedup claim。
Historical causal evidenceWeek 9 V-direct-BSH +1.289408% stable-DiTThree-pair 832×480 full-model engineering A/B; runtime inputs were semantically equivalent, but this evidence is neither a Portrait comparison nor a quality acceptance。
Claim boundaryWarm-throughput onlyFresh-process wall paired median +10.802722%;no cold-start speedup。已有 VAE 报告不构成 ground-truth 或 blind-panel quality acceptance。
34.3923
Portrait generated FPS

489 frames / sampling + pipelined video decode;three-candidate CV 0.085863%。

37.5570
Portrait stable-DiT FPS

288 / chunks 5–10 DiT wall;three-candidate CV 0.151246%。

576 + 48
K32 Linear + V-direct attn1

0 fallback、0 upcast、0 retained converted BF16 weight;strict fail-closed attention。

未验收
Formal perceptual quality

工程 runtime/media gate PASS;正式 blind perceptual-quality acceptance 尚未完成。

图 10|Historical Portrait v1 baseline。 三次 fresh-cache repeats 给出历史 v1;landscape 对比仅说明 throughput regime preserved,不是 paired causal speedup。

Methodological correction / 判据证伪

本版最重要的结论不是一个性能数字,而是一次判据证伪。Week 4–6 的全部精度判决都以端到端 latent trajectory rel-L2 作为 accept/reject 依据。2026-07-22 的因果 lane 实验给出三重反证:同一干预(BF15)在 chef 上 video −3.5075%、在 drummer 上 +10.7793%,符号翻转安慰剂 R15 在 chef 上全面优于处理组(audio +19.8825% 对比 +8.9475%);全部效应落在 ±3~20%,与轨迹混沌同量级。该度量因此降级为“轨迹分岔幅度”的描述性指标,局部单-forward oracle-L2 仍然有效。凡以旧判据作出的结论一律标注为“判据失效、需重审”,而不是静默保留。

Throughput definition / 吞吐口径

generated FPS 是 489 帧除以 sampling 加流水 video decode 的秒数;stable-DiT FPS 是 288 除以 chunks 5–10 的 DiT wall。两者分子分母都不同,不可混用。25 FPS 是 MP4 playback rate。Portrait headline 使用未插桩的 fresh-cache runs。Nsys 对历史 landscape K32 subject 的 stable/generated 口径分别引入约 1.533% / 2.040% 扰动,因此 trace 不得用来更新 Portrait FPS。历史计时器之外的 text、audio decode 与 media write 也不属于 generated-FPS 分母。

GPU 0 · conditioning / decode

Conditioning and asynchronous decoding

  • Gemma INT8 text encoder 与 embedding processor
  • BF16 video / audio VAE 与 vocoder
  • 稳态 kernel sum 中 VAE 卷积占 73.257%
GPU 1 · complete Catnip DiT

Complete Catnip DiT execution

  • 576 native K32 MXFP8 Linear(72.138575% all-Linear weight coverage)
  • 48 V-direct-BSH video-self attn1,48 blocks independently compile
  • 480×832 Portrait;bounded first-4 + latest-6 KV,shared PE
Experimental protocol

Workload Contract and Measurement Boundary

为消除 workload drift 引入的混杂因素,正式实验在 chunk count、KV semantics、随机输入与 timer scope 上保持一致。 Workload identity、evidence grade 与 promotion gate 均采用 fail-closed protocol。

Current Portrait Streaming ABI

480 × 832Portrait spatial resolution
20 s / 489 framesencoded duration 19.56 s; playback 25 FPS
11 chunks62 latent frames; 390 spatial tokens/LF
55 forwards每 chunk 5 次 Transformer forward
first 4 + latest 6persistent KV 上限 10 LF
fresh cachethree independent cache roots; frozen seed/noise

Historical paired-candidate protocol

  1. 绑定内容寻址的 release 作为控制组
  2. control/candidate、candidate/control、control/candidate 三对
  3. 每条 lane 独立且必须不存在的 cache root
  4. 独立验证器从原始 chunk timing 重算 FPS
  5. 拒绝身份、顺序、环境、cache 或 artifact 漂移

Gate:median gain ≥ 10%、每对 ≥ 8%、三类 CV ≤ 2%、GPU1 free ≥ 512 MiB。

Acceptance authority / 验收权威

解码媒体的 reference-free 伪影面板以 BF16 与生产 FP8 之间的自然差异带归一化:宽度取四个 calibration cell 上二者绝对差的最大值。孤立有害偏差严格大于 2.0 报警,同一有害指标在两个及以上 cell 超过 1.0 亦报警。面板经伪影注入标定——flicker_global 25.68×、audio_click 8.51×、sharpness 1.78×、blockiness 1.13×、seam 1.46×(仅可相对比较)。av-sync 代理标定失败(400 ms 延迟未检出),已降级为 diagnostic。

Heldout leakage / 一次必须记录的错误

Week 8 的第一版面板在构造自然差异带时使用了原本作为 final-heldout 的 scientist-greenhouse/42424violinist-station/77777 两格,构成泄漏。V2 协议已记录该泄漏,把两格重分类为 calibration(现共四格),并另行保留两个不透明的 final-heldout-v2 占位符,需 hash-bound 书面解锁。其代价是永久损失两格独立性。

Methods & empirical evidence

Technical Methods and Evidence Chain

以下按真实实验顺序展开,每个节点均按 Action → Evidence → Decision 报告, 以保持工程变更、测量结果与最终取舍之间的可追溯性。

1
Week 1–2 · capacity & native backend

从 storage-only FP8 到 native W8A8,以及双卡功能拆分

历史 24/24 配置属于 block model-parallelism:两卡串行执行同一次 forward 并在中点搬运完整 state。Legacy FP8-cast 在 GEMM 前把权重转回 BF16,属 storage-only quantization。改为完整 DiT 置于 GPU1、conditioning 与 VAE 置于 GPU0,并把 576 个 Linear 改成 native scaled-MM。

240→480→576 coverage72.138575% Linear weight4+6 KVO0 19.080 FPS
图 1|Storage evolution。 FP8 coverage 降低 static storage;20 s request feasibility 由 bounded KV、shared PE 与 fused bias 对 lifetime peak 的联合约束实现。
Action

逐步扩大 FP8 coverage,并把 persistent KV 约束为 first-4 + latest-6;用 host worker 修复 same-thread false overlap。

Evidence

480-module 在 chunk 1 因 append-before-evict lifetime peak OOM;bounded KV + shared PE + fused bias 后 11 chunks 跑通。

Decision

确立 O0 基线 19.080070887 generated FPS,进入 O1–O7 优化瀑布。

2
O0–O7 · throughput waterfall

完整请求驱动的优化瀑布,而非 microbenchmark 驱动

48-block 独立 compile、消除重复 V2A、扩大 FP8 coverage 与图稳定化依次叠加。O3 的 selective cuDNN SDPA 在单 shape 上快约 19.7%,完整请求却慢 220.053 ms,因此回退——这是本报告第一次出现“局部更快、整体更慢”。

O2 −2,892.619 msO4 −2,380.908 msO7 27.104924 FPSCV 0.046536%
图 2|Performance waterfall。 O3 的局部 cuDNN SDPA microbenchmark 更快,但完整请求比 O2 慢 220.053 ms,因此恢复 PyTorch Flash。
Action

每一步都以三次独立 warmed 20 秒进程验证,变化小于 0.25% 或三倍 baseline CV 视为 neutral。

Evidence

O7 三轮 27.104924035 generated FPS;formal graph delta 为空;0 fallback/upcast。

Decision

交付 O7,并进入 Week 3 的 profiling 驱动候选筛选。

3
Week 3 · profiling-driven selection

六类候选中只有一个通过三轮门禁

Nsys 归因显示当时 GPU1 kernel-duration 中 true FP8 GEMM 占 45.4866%。据此筛选的六类候选里,external M padding、fast-accum、shared activation quant、text K/V cache 与 simple concat 全部被数据否定,只有 16 MiB cuBLASLt workspace 通过。

FP8 GEMM 45.4866%workspace −0.865511%27.370839 FPS5/6 rejected
图 3|Week 3 full-request A/B。 Provider-only median wall 为 17.865729 s / 27.370839 FPS。Shared+provider 的 incremental gain 仅 0.0544%,低于 promotion threshold。
Action

先做 exact production shape microbenchmark,通过后才允许三次独立 20 秒进程;显存增量不得超过 128 MiB。

Evidence

provider-only median wall 17,865.729011 ms / 27.370839427 generated FPS;shared+provider 的增量仅 9.716 ms,低于门槛且 seam 更差。

Decision

Promote workspace(现已固化进 Week 8 冻结启动环境),其余五类进入拒绝账本。

4
Week 4–6 · precision generalization

Scale granularity 的系统评估,以及一次机制性拆分

rowwise、channelwise、SmoothQuant、K32/K64 与 temporal routing 被系统评估,全部因 worst-cell 或跨 prompt 反转失败。Week 6 把“更细分组”与“E8M0 scale 编码”拆开,发现 exact-FP32 K32 改善 local rel-L2 22.910%,而 ceil-pow2 / native E8M0 反而比 tensorwise 差 1.056%。

exact K32 +22.910%E8M0 −1.056%propagation 4/8survivor 0
图 4|K32 factorization。 左:只有 exact FP32 scale 保留 K32 局部收益;右:四条 paired block-probe lanes 中 video p50/p95 改善,audio p95 全部回退。
Action

对 320 strata × 8 构成 2,560 paired local samples,分别比较 tensor FP32、K32 exact、K32 pow2 与 native K32 E8M0。

Evidence

四条 video block-probe lanes 通过、四条 audio p95 失败;总 gate 4/8,Phase 5 锁定。

Decision

当时判定 native K32 不获部署资格——该判决其后被推翻

5
2026-07-21 · outlier statistics

把“离群值伤害 FP8”从直觉改写成量化结论

对 576 层的 13,690,208,256 个权重元素做全量精确统计,并用 63,360 条真实 activation call records 做激活侧分析。核心发现是一个零结果加一个对照:weight tail 对 E4M3 误差的 Spearman ρ 为 −0.046(CI 跨零),而同一指标在 INT8 对照下 ρ 为 0.817。

ρ = −0.046 (n.s.)INT8 ρ = 0.817underflow ρ = 0.812activation ρ = 0.442
图 5|Outlier statistics。 左:五个 Spearman 相关及 family-stratified 95% CI,weight tail 对 E4M3 误差跨零不显著而 INT8 对照 ρ 0.817。右:权重侧四条反事实的 MSE 改善与静态压缩率 Pareto;sparse residual 是离线上界,不是可部署 kernel。
Action

全量精确计算误差、amax、underflow 与 clip fraction;分位数与 kurtosis 用每层最多 262,144 个确定性等距样本;每个相关系数 4,000 次 family-stratified bootstrap。

Evidence

output-channel scale 消除 93.89% 的 underflow,却只换来 0.476% 的 MSE 改善——underflow 是 count 型损失,平方能量太小。dense clip oracle 上界仅 0.1478865618259051%。

Decision

把 INT8 outlier 文献(clipping / smoothing)排除出 FP8 权重侧路线;研究预算转向激活侧。

6
Week 7 · custom SM120 kernel

自研 MXFP8 attention:CUTLASS 特化 + FA3 血统 + 手写 producer

48 个 video-self attn1 被替换为自研 persistent kernel:signed normalized H128 → E4M3/UE8M0 K32 → native SM120 block-scaled QK → one-pass online softmax → 在线 P K32 量化 → block-scaled PV → BF16。fused core grid 170、block 384、168 registers/thread、84,992 B dynamic shared。

r04 28.084267 FPS主形状 1.44236×core 1.96373×2640 calls
图 6|Attention stage matrix。 左:隔离 stage 中位数的 20 秒加权投影。非可加——组件和 2,628.061 ms 大于独立测得的 provider e2e 投影 2,461.068 ms,这个差异正是非可加性的直接证据。右:四组形状的贡献占比,主形状 2340×6240 占 82.536%。
Action

QK/PV 核心特化自 CUTLASS 4.5.2 SM120 blockscaled example;attention 主体为 FA3 血统 warp-specialized persistent 核;量化 producer 为手写 CUDA;集成为 fail-closed custom op。

Evidence

整机 r04 formal wall 17,411.884168 ms。但 fused cooperative H128 虽达成逐 bit 一致与单 launch,四个生产 shape median 全部回归 +1.512%–+7.440%,因此不作生产默认。

Decision

保留“独立 producer + attention core”为生产默认;分布级画质证据仍缺失,且 r04 没有相邻 provider-off 对照。

7
Week 8 上 · falsification

验收判据的证伪与重建

为验证 outlier 统计挑出的 15 个高危层是否真的重要,设计因果 lane:BF15 把这 15 层从 FP8 manifest 移除(权重与激活两个误差源同时消失,因此是联合上界),R15 对 15 个 family 匹配的最低风险层做同样处理作为安慰剂。

BF15 跨 cell 翻转安慰剂 > 处理组flicker 25.68×K32 判决反转
图 7|Metric falsification。 左:BF15 处理组与 R15 安慰剂在两个 cell 上的 oracle-L2 改善——处理组跨 cell 符号翻转,安慰剂在 chef 上全面优于处理组。右:面板伪影注入标定倍数(对数轴);av-sync 未通过标定,不在图中。
Action

先修方法学缺陷:Week 4 的 runner 每条 lane 重跑 text encoder 且不硬绑定,正是让 Week 5 第一个候选作废的机制。新 wrapper 装上 fail-closed 冻结 conditioning,哈希与参考值逐字节一致。

Evidence

BF15/chef video −3.5075% 而 drummer +10.7793%;R15/chef 在同一度量下 audio +19.8825%,全面优于处理组。所有效应落在 ±3~20%。

Decision

端到端 oracle-L2 降级为描述性指标;建立并标定 reference-free 面板;K32 在新面板下无确凿质量损失,判决反转

8
Week 8 中 · K32 baseline V2

第一次取得严格配对 A/B

576 个 Linear 全部替换为 native K32 MXFP8。三对独立 fresh 进程按 B/K、K/B、B/K 交替顺序执行,每条 lane 独立且必须不存在的 compile-cache root——这是本报告此前一直缺少的严格因果对照。

median +14.103774%min pair +13.995281%ratio CV 0.045519%cold −23.34%
图 8|K32 progression。 左:同一 489/formal-wall timer 下的 generated FPS 演进,但前四项跨越不同代码栈,只有右图是严格因果配对。右:三对交替顺序运行的 stable-DiT FPS 与逐对增益。
Action

绑定内容寻址 release manifest;独立验证器重放控制组 baseline reference、从原始 chunk timing 重算 stable-DiT FPS 并拒绝任何身份或环境漂移。

Evidence

stable-DiT 32.319420887966025 → 36.866330113659890;generated 28.344309608927105 → 32.572401806507680(+14.916900%)。门禁全部 PASS。

Decision

Promote 为工程基线 V2;但 formal_baseline_accepted=false,且冷启动 compile warmup 慢 23.342073%。

9
Week 8 下 · historical device-separated profiling

两个层级的 P0,以及一堵近在 5–9% 之外的墙

对历史 832×480 K32 baseline 做 20 秒全请求 Nsys trace 与 attention 分阶段 CUDA-event benchmark。45 个分析窗口、12 个 stage gate、3 个 attention count gate、5 个 sequence-pair gate 全部 PASS。

scaled-mm 38.489%三 P0 shape 87.665%VAE 卷积 73.257%两卡 busy 99.640%
图 9|Historical device-separated taxonomy。 左:历史 GPU1 taxonomy。右:历史 GPU0 taxonomy。两卡 kernel sum 不可相加,且这不是 Portrait absolute benchmark。
Action

Nsys 2025.3 的 NVTX trigger 在本机损坏,改用验证过的 --start-later 会话并离线裁剪到 PROFILE_REQUEST 范围;应用仍跑完整冻结请求。

Evidence

稳态 DiT 1,358–1,366 ms/chunk 对比 VAE 1,248–1,250 ms/chunk,两卡同时 busy 87.327%。这些是优先级证据,不是当前 Portrait FPS。

Decision

将 V layout、VAE 和三大 scaled-mm shape 排为候选;随后 Week 9 必须用新的 A/B 而非将 profiling share 直接换算成收益。

10
Week 9 · profiling-guided boundary optimization

从非可加 stage 投影,到一项 full-model engineering PASS

V-direct-BSH 只改变 attention V 的消费边界:control 显式 materialize BSH→BHSD(padded case 另做 fill/copy),candidate 让 producer direct-consume strided BSH;Q/K、fused core、weights、KV、VAE 与 scheduler 不变。

provider +7.5686%full model +1.289408%all 3 pairs positive6/6 gates PASS
图 11|Week 9 validated and rejected experiments。 左、中为历史 832×480 K32 stack 的 provider 与 full-model evidence;右为正确早停的 heuristic-pinning negative result。
图 12|Pipeline balance and counter-free NCU。 历史 profiling 显示 DiT/VAE 接近平衡;counter-free NCU 将 stream-K/persistent 调度保留为可检验方向,而不是把理论 occupancy 误读为免费 headroom。
Action

先用 blocked ABBA/BAAB 证明 four-shape harness neutrality(31/31 checks PASS),再做 four-shape provider e2e、Nsys structure 与三对 20 s full-model A/B。

Evidence

micro 四 shape 全正;Nsys 显示 aligned 10→9、padded 14→11 launches;full-model stable-DiT 36.923232→37.400230,generated 32.624456→33.035966。另一条 dominant scaled-mm heuristic-pinning screen 仅 +0.058769%,低于 +1% gate。

Decision

Promote V-direct-BSH 为历史 frozen K32 stack 的 engineering-performance improvement;Reject frozen 1 MiB heuristic pinning;两者都不主张 quality acceptance。

11
2026-07-26 · historical Portrait v1

将已验证 stack 首次固化为 480×832 engineering baseline

保持 576 K32 MXFP8 Linear、48 V-direct-BSH attn1、GPU1 DiT / GPU0 text+VAE、4+6 KV 与 depth-3 decode,仅改变 output orientation,并在三个独立 absent-cache roots 上完整执行。

37.486982 stable-DiT33.070761 generatedCV < 0.065%quality not accepted
Action

对每个 repetition 执行 request、K32、V-direct、KV、compile 与 media gates;从原始 timing summary 聚合均值和 population CV。

Evidence

Portrait 相对 landscape V-direct reference 的差异仅 +0.231955% / +0.105325%,量级不足以解释为 orientation speedup。三次 runtime 皆无 fallback/upcast/retained BF16 converted weights。

Decision

Adopt catnip-k32-vdirect-portrait-480x832-v1;该版本已于 2026-07-28 被 v2 取代,保留为历史证据。

12
2026-07-28 · current Portrait v2 promotion

接纳已独立报告的 VAE 运行时改动并完成最终晋升门禁

保持 20 秒合同、576 K32、48 V-direct、4+6 KV 与双卡 placement;当前默认入口仅新增 single-tile latent-window decode 和 compiled video VAE decoder。VAE 实验过程不在本报告重复。

34.392314 generated37.557043 stable-DiTr03 smoke PASSwarm only
Action

以 Week 10 Phase 3 的三对 candidate runs 作为权威 warm KPI,再在无并发作业环境执行独立 20 秒 promotion smoke。

Evidence

Generated-FPS CV 0.085863%;paired median compile gain +1.837371%;final r03 为 34.408774 generated FPS,489-frame H.264/AAC media 与 runtime gates 全 PASS。

Decision

Promote catnip-k32-vdirect-single-tile-compile-portrait-480x832-v2。Fresh-process wall paired median +10.802722%,所以不声明 cold-start speedup;formal quality acceptance 仍未建立。

Decision ledger

Experimental Decisions and Rejection Rationale

本节汇总主实验链路之外的 candidate decisions。Rejected 表示实验本身有效但 hypothesis 被数据否定; Blocked 表示已取得有效 gate evidence,但后续 phase 按预注册协议保持锁定;Invalid 不产生科学结论; Not run 表示 protocol-compliant early stopping 或行政阻塞。

阶段尝试结果 / 最终原因状态
Week 2Bounded KV + shared PE + fused bias11 chunks 完成,reserved 回到 memory gate 内。保留
Week 2Host-worker async VAE修复 same-thread false overlap,形成 O0 19.080 FPS。保留
O248-block independent compileO1→O2 observed wall −2,892.619 ms。保留
O3Selective cuDNN SDPA单 shape 快约 19.7%,完整请求却慢 220.053 ms。拒绝
O4Fold duplicate V2AWall −2,380.908 ms;保留 cache 副作用与残差顺序。保留
O5576-module coverageWall −824.946 ms;Linear weight coverage 72.138575%。保留
O6/O75-shape warmup + strict entryFormal graph delta 为空;三轮 27.104924 FPS,CV 0.046536%。交付
Week 3External M padding三大 shape 都慢;预计每请求 +52.517 ms。拒绝
Week 3Shape-specific fast-accumSquare micro 1.0237×,full median 反而慢 2.402 ms。拒绝
Week 3Shared activation quantPhysical quant calls −25%,但 full wall 仅 −0.166418%。拒绝
Week 3Text K/V cache48 层需 3,840 MiB;9 层 720 MiB 也超 memory gate。拒绝
Week 3Simple QKV/KV concat乐观 request ceiling 只有 0.173474%。拒绝
Week 3cuBLASLt workspace 16 MiBWall −0.865511%,FPS +0.873067%;已固化进 Week 8 冻结启动环境
Week 4Row activation / tensor weight首个单 cell aggregate 已比 tensorwise 更差。拒绝
Week 4Hybrid-KOverall worst −20.320%,audio median 退化。拒绝
Week 4Full rowwiseOverall worst −7.563%;audio worst −61.777% / −65.655%。拒绝
Week 4SmoothQuantCalibration median +22.782%,heldout median −2.991%。判据已失效,需重审。拒绝
Week 5Sensitive top-48 / 96 / 192跨 cell 反转。判据已失效,需重审。拒绝
Week 5Native K32Week 5 观测值保留(chef +0.827%、drummer −7.064%);拒绝依据于 2026-07-22 被证伪,同栈配对 stable-DiT median gain 14.103774%。
Week 5Dynamic torch.cond routingCUDA overhead 10.19%–20.98%,在 full video 之前已超门槛。Blocked
Week 6Exact-FP32 K32 simulatorLocal +22.910%,四条 video block-probe lanes 通过。研究保留
Week 6K32 pow2 / native E8M0Local rel-L2 差 1.056% 的数值保留;部署判决已推翻——它正是基线 V2 的 scale 编码。
Week 6Exact-K32 propagationVideo 4/4,audio p95 0/4;总 gate 4/8。该判据已被证伪。Blocked
Outlier 统计576 层全量权重 + 63,360 activation recordsINT8 对照(ρ 0.817)证明 outlier 直觉不能移植到 E4M3(ρ −0.046)。研究保留
Week 7A3 signed H128 旋转每个 pooled 类别都改善(overall 2.1063%),但门槛为 5%–10%;是幅度不够而非回归。拒绝
Week 7A7 概率 K32非零下溢 −95.2250%,但 defect 仅改善 0.0267%;25 项 gate 中 11 项失败。拒绝
Week 7P1d learned rotationodd-chunk mean 仅 +0.246%–+0.623%,门槛 ≥5%;明显局部过拟合。拒绝
Week 7Fused cooperative H1285 shape 逐 bit 一致,但四个生产 shape median 全部回归 +1.512%–+7.440%。拒绝
Week 7FA3 K/V warp-release五 shape bitwise 一致但全部变慢 +0.79%–+2.83%。拒绝
Week 7自研 SM120 MXFP8 attention48 个 attn1 替换;r04 整机 28.084266773 generated FPS;主形状隔离 1.44236×。生产默认
Week 8质量面板证伪旧判据处理组跨 cell 符号翻转,安慰剂优于处理组;判据降级为描述性指标。
Week 8Reference-free 解码媒体面板伪影注入标定:flicker 25.68×、click 8.51×、sharpness 1.78×、blockiness 1.13×。
Week 8Native K32 MXFP8 baseline V2三对交替配对 stable-DiT median gain 14.103774%、最小配对 13.995281%;两 lane AUTO_PASS。
Week 8全局 Q/K single-launch仅最小 shape 快 17.117%;20 秒加权慢 0.934%。拒绝
Week 8全家族 rowwise(ROW192)面板 worst dev 1.709,seam spike 6.5 对比 3.1,本轮最强伪影信号。拒绝
Week 7MXFP8 r01 / r02 / r03外部作业占用 GPU 或 provider dtype gate 过严;均在 formal timer 之前停止,无 FPS。Invalid
Week 8BF15 首次 drummer 运行编译预热阶段被中断,无 run.json、无 probe、无对比;已隔离标注。Invalid
Week 8Anatomy capture 首次尝试与前一运行 teardown 撞车导致 GPU0 OOM;清理后重跑成功。Invalid
Week 8第一版面板的自然带使用了原 final-heldout 两格,构成泄漏;V2 已重分类并另设 heldout-v2。Invalid
AttentionNVML 100% / Kineto 301%指标或计时边界错误;被 25-shape complete-op V2 取代。Invalid
Week 8Calibration freezemarker status/quality_calibration.freeze.json 不存在。Blocked
Week 8面板 av-sync gate注入 400 ms 音频延迟未被检出,代理指标标定失败。降级 diagnostic
Week 8两人独立盲评公开包与响应验证器已就绪,等待两名独立评审。PENDING
Week 8Validation 4 runs需 freeze marker 解锁;已完成 0/4。LOCKED
Week 8Heldout-v2 两格需 hash-bound 书面解锁。EMBARGOED
NCU完整 tensor/stall/achieved-occupancy countercounter-free ncu --set none 已给出 launch/wave/theoretical occupancy;完整 PerfWorks counters 仍受权限阻断。部分解锁
Week 9V-direct-BSH attention boundary历史 832×480 K32 stack 上三对 full-model A/B:stable-DiT median +1.289408%,generated +1.284480%,6/6 gates PASS;不声称 quality acceptance。
Week 9cuBLASLt heuristic top-k pinning主 shape bitwise correct 但 paired median 仅 +0.058769%,低于 +1% gate;按三 shape conjunction early stop。拒绝
Week 9Counter-free NCU --set noneattention 1.00 wave;P0 GEMM 3.58 waves / 89.4% efficiency。关闭 tile/stage 扩张假设,保留 stream-K/persistent。方法学
Portrait v1catnip-k32-vdirect-portrait-480x832-v1三次 fresh-cache mean 37.486982 stable-DiT / 33.070761 generated FPS;现由 v2 取代。历史基线
Portrait v2Single-tile + compiled video VAE decoder三对 candidate mean 37.557043 stable-DiT / 34.392314 generated FPS;final r03 smoke PASS;warm-throughput only。
Portrait v2Promotion smoke r01/r02并发 Week 10 KV study 占用 GPU,VAE decode 前 OOM;清场后 r03 完整通过。Invalid
两类使用规则:数据否决 vs 判据失效
第一类,数据否决、不再回锅:external M padding、simple QKV/KV concat、text K/V cache、原 dynamic torch.cond 路由、全局 Q/K single-launch、H128 producer 融进 fused core、全家族 rowwise。

第二类,判据失效、必须重审:所有以端到端 trajectory rel-L2 作为 accept/reject 依据的结论——包括 Week 4 与 Week 5 的全部精度 reject 与 Week 6 的 propagation 阻断——都建立在一个已被证伪的度量上。它们不自动变成“其实可行”,而是回到未判定状态。K32 是第一个走完这条路径并回到工程基线的候选。
Role of Invalid records in the evidence hierarchy
Invalid runs 不支持 candidate efficacy 的统计结论,但可防止 infrastructure failure 被误归因于 algorithmic effect。修复后必须使用新的 run ID,从相应 phase boundary 重新执行;failure 前产生的 partial artifacts 不进入正式 gate。
Production snapshot

Current Engineering Baseline and Acceptance State

本节区分 current engineering baseline、paired control 与 historical delivery,以避免将实验 promotion 误表述为 production acceptance。

项目冻结配置
Quantization576 selected Linear:native K32 MXFP8 E4M3 qdata、K32 E8M0 scale、BF16 output + fused bias。0 fallback / 0 upcast / 0 retained converted BF16 weight;不是 full-model FP8。
Attention48 video-self attn1:V-direct-BSH strict fail-closed provider。Week 9 验证 V boundary 的历史 engineering A/B;Portrait repeats 验证集成 stack。
Streaming480×832 Portrait、20 s、489 frames、11 chunks、55 forwards、first 4 + latest 6 LF KV、390 spatial tokens/LF。
SchedulingGPU1 DiT + GPU0 text/VAE;per-chunk video decode depth 3,video drain 后单次 audio decode;每个 latent window 使用 single tile。
Compilation48 DiT blocks independently compiled;当前 v2 另编译 video VAE decoder forward。性能声明限于 warm steady state。
Current baselinethree-candidate mean 34.392314 generated FPS / 37.557043 stable-DiT FPS;final smoke 34.408774 generated FPS。
Repeatabilitypopulation CV:generated 0.085863% / stable-DiT 0.151246%;Phase 3 paired order 为 control/candidate、candidate/control、control/candidate。
Historical causal evidence832×480 K32/tensorwise +14.103774% stable-DiT,Week 9 V-direct-BSH +1.289408%;二者均不替代 Portrait absolute KPI。
Historical deliveryO7 27.104924 generated FPS;保留为 Week 2 delivery reference。
Release identitycatnip-k32-vdirect-single-tile-compile-portrait-480x832-v2;ACTIVE engineering baseline;adopted 2026-07-28。
Claim boundaryWarm-throughput only;fresh-process wall paired median +10.802722%,no cold-start speedup。Runtime/media PASS does not establish ground-truth perceptual-quality acceptance。

Unresolved evidence requirements

Portrait v2 formal quality acceptance 仍未建立:已有 VAE 报告没有 ground truth 或正式 blind panel,且不能由 runtime/media gate 自动替代。v2 cold-start speedup 也未建立。完整 NCU 的 tensor-pipe、warp-stall、achieved-occupancy 与 duration counters 仍受权限限制;matched commit/workload/timer 的 BF16↔FP8 performance A/B 仍未建立,因此本报告不给出官方 BF16→FP8 speedup 倍数。

Limitations & future work

Active Problems and Evidence-Constrained Optimization Space

Portrait v2 已接纳 GPU0 路径的首轮优化,当前 warm KPI 为 34.392314 generated FPS。下一轮不能继续把历史 profiling 占比当作当前精确瓶颈;排序遵循单变量、paired 20 s verification 与质量边界。

P0 · 冷启

Compile cache 与预热摊销

v2 warm throughput 已提升,但 fresh-process wall paired median 增加 10.802722%。

拆分 graph build、compile 与 first execution;不得用 warm KPI 代替首轮请求等待。
P0 · 边界

Q/K cast 或 direct-layout fusion

V-direct 已验证“先测四 shape、再测 20 s”的路线。下一步每次只改一个 boundary,禁止把独立 stage 投影相加成收益。

four-shape CUDA-event A/B → Nsys structure → three-pair full-model A/B。
P0 · GEMM

Epilogue fusion / stream-K

heuristic pinning 已以 +0.058769% 被拒。P0 GEMM 仍有 10.6% wave tail;结构性收益更可能来自减少 quant/layout 读写。

新 hypothesis 独立预注册;优先 stream-K/persistent,tile/stage 扩张已因资源余量降级。
P0 · 质量

Blind review 到 heldout

Portrait runtime/media PASS 不等于视觉质量验收。blind review → freeze → validation → heldout-v2 仍是正式身份的必经链路。

仅人评与锁定 marker 能升级 quality claim。
P1 · counters

完整 NCU counters

ncu --set none 已关闭 launch/wave/tile 路由问题;tensor pipe、warp stalls、achieved occupancy 与 duration 仍不可观测。

宿主机解除 profiling restriction 后,按 counter-to-action table 决定是否重写 core。
P2 · 用户等待

NVENC / streaming mux

历史 media write 约 7.4 s,不在 generated-FPS 分母,却影响用户等待。

单独测 warm request-to-final-MP4 wall;不得以降低编码质量换取模型性能。