489 frames / sampling + pipelined video decode;three-candidate CV 0.085863%。
总体结论 / Primary Conclusions
本研究完成了 Catnip Streaming 模型从 capacity-oriented BF16 model parallelism 到 production-oriented mixed-precision pipeline parallelism 的迁移,并在此基础上完成了 quantization encoding、custom attention kernel 与 acceptance methodology 三次演进。 下表给出可由当前证据支持的最终结论。
| Active deployment | Portrait v2 · 480×832 on 2× RTX 5090 | catnip-k32-vdirect-single-tile-compile-portrait-480x832-v2;GPU1 complete DiT,GPU0 INT8 text encoder + BF16 video/audio VAE。11 chunks、55 forwards、4+6 KV、depth-3 video decode。 |
| Mixed precision | 576 selected Linear use native K32 MXFP8 | E4M3 qdata + K32 E8M0 scales + BF16 output/fused bias;13.69B selected weight elements / 72.14% all-Linear coverage。该配置不是 full-model FP8。 |
| Attention | 48 video-self attn1 use V-direct-BSH | Strict fail-closed provider;V-direct consumption is an accepted historical engineering optimization, while the Portrait repeats validate the integrated stack. |
| Current performance | 34.392314 generated / 37.557043 stable-DiT FPS | Three paired Phase 3 candidate means; population CV 0.085863% / 0.151246%。generated FPS = sampling + pipelined video decode boundary。 |
| Promotion gate | Final uncontended r03 PASS | 34.408774 generated / 37.566968 stable-DiT FPS;489-frame H.264、AAC、4+6 KV、576 K32、48 V-direct 与 no-new-graph gates PASS。 |
| Historical v1 | Throughput regime preserved | v1 Portrait versus 832×480 V-direct reference changes stable/generated means by only +0.231955% / +0.105325%。It is not a contemporaneous paired causal speedup claim。 |
| Historical causal evidence | Week 9 V-direct-BSH +1.289408% stable-DiT | Three-pair 832×480 full-model engineering A/B; runtime inputs were semantically equivalent, but this evidence is neither a Portrait comparison nor a quality acceptance。 |
| Claim boundary | Warm-throughput only | Fresh-process wall paired median +10.802722%;no cold-start speedup。已有 VAE 报告不构成 ground-truth 或 blind-panel quality acceptance。 |
288 / chunks 5–10 DiT wall;three-candidate CV 0.151246%。
0 fallback、0 upcast、0 retained converted BF16 weight;strict fail-closed attention。
工程 runtime/media gate PASS;正式 blind perceptual-quality acceptance 尚未完成。
Methodological correction / 判据证伪
本版最重要的结论不是一个性能数字,而是一次判据证伪。Week 4–6 的全部精度判决都以端到端 latent trajectory rel-L2 作为 accept/reject 依据。2026-07-22 的因果 lane 实验给出三重反证:同一干预(BF15)在 chef 上 video −3.5075%、在 drummer 上 +10.7793%,符号翻转;安慰剂 R15 在 chef 上全面优于处理组(audio +19.8825% 对比 +8.9475%);全部效应落在 ±3~20%,与轨迹混沌同量级。该度量因此降级为“轨迹分岔幅度”的描述性指标,局部单-forward oracle-L2 仍然有效。凡以旧判据作出的结论一律标注为“判据失效、需重审”,而不是静默保留。
Throughput definition / 吞吐口径
generated FPS 是 489 帧除以 sampling 加流水 video decode 的秒数;stable-DiT FPS 是 288 除以 chunks 5–10 的 DiT wall。两者分子分母都不同,不可混用。25 FPS 是 MP4 playback rate。Portrait headline 使用未插桩的 fresh-cache runs。Nsys 对历史 landscape K32 subject 的 stable/generated 口径分别引入约 1.533% / 2.040% 扰动,因此 trace 不得用来更新 Portrait FPS。历史计时器之外的 text、audio decode 与 media write 也不属于 generated-FPS 分母。
Conditioning and asynchronous decoding
- Gemma INT8 text encoder 与 embedding processor
- BF16 video / audio VAE 与 vocoder
- 稳态 kernel sum 中 VAE 卷积占 73.257%
Complete Catnip DiT execution
- 576 native K32 MXFP8 Linear(72.138575% all-Linear weight coverage)
- 48 V-direct-BSH video-self attn1,48 blocks independently compile
- 480×832 Portrait;bounded first-4 + latest-6 KV,shared PE
Workload Contract and Measurement Boundary
为消除 workload drift 引入的混杂因素,正式实验在 chunk count、KV semantics、随机输入与 timer scope 上保持一致。 Workload identity、evidence grade 与 promotion gate 均采用 fail-closed protocol。
Current Portrait Streaming ABI
Historical paired-candidate protocol
- 绑定内容寻址的 release 作为控制组
- control/candidate、candidate/control、control/candidate 三对
- 每条 lane 独立且必须不存在的 cache root
- 独立验证器从原始 chunk timing 重算 FPS
- 拒绝身份、顺序、环境、cache 或 artifact 漂移
Gate:median gain ≥ 10%、每对 ≥ 8%、三类 CV ≤ 2%、GPU1 free ≥ 512 MiB。
Acceptance authority / 验收权威
解码媒体的 reference-free 伪影面板以 BF16 与生产 FP8 之间的自然差异带归一化:宽度取四个 calibration cell 上二者绝对差的最大值。孤立有害偏差严格大于 2.0 报警,同一有害指标在两个及以上 cell 超过 1.0 亦报警。面板经伪影注入标定——flicker_global 25.68×、audio_click 8.51×、sharpness 1.78×、blockiness 1.13×、seam 1.46×(仅可相对比较)。av-sync 代理标定失败(400 ms 延迟未检出),已降级为 diagnostic。
Heldout leakage / 一次必须记录的错误
Week 8 的第一版面板在构造自然差异带时使用了原本作为 final-heldout 的 scientist-greenhouse/42424 与 violinist-station/77777 两格,构成泄漏。V2 协议已记录该泄漏,把两格重分类为 calibration(现共四格),并另行保留两个不透明的 final-heldout-v2 占位符,需 hash-bound 书面解锁。其代价是永久损失两格独立性。
Technical Methods and Evidence Chain
以下按真实实验顺序展开,每个节点均按 Action → Evidence → Decision 报告, 以保持工程变更、测量结果与最终取舍之间的可追溯性。
从 storage-only FP8 到 native W8A8,以及双卡功能拆分
历史 24/24 配置属于 block model-parallelism:两卡串行执行同一次 forward 并在中点搬运完整 state。Legacy FP8-cast 在 GEMM 前把权重转回 BF16,属 storage-only quantization。改为完整 DiT 置于 GPU1、conditioning 与 VAE 置于 GPU0,并把 576 个 Linear 改成 native scaled-MM。
逐步扩大 FP8 coverage,并把 persistent KV 约束为 first-4 + latest-6;用 host worker 修复 same-thread false overlap。
480-module 在 chunk 1 因 append-before-evict lifetime peak OOM;bounded KV + shared PE + fused bias 后 11 chunks 跑通。
确立 O0 基线 19.080070887 generated FPS,进入 O1–O7 优化瀑布。
完整请求驱动的优化瀑布,而非 microbenchmark 驱动
48-block 独立 compile、消除重复 V2A、扩大 FP8 coverage 与图稳定化依次叠加。O3 的 selective cuDNN SDPA 在单 shape 上快约 19.7%,完整请求却慢 220.053 ms,因此回退——这是本报告第一次出现“局部更快、整体更慢”。
每一步都以三次独立 warmed 20 秒进程验证,变化小于 0.25% 或三倍 baseline CV 视为 neutral。
O7 三轮 27.104924035 generated FPS;formal graph delta 为空;0 fallback/upcast。
交付 O7,并进入 Week 3 的 profiling 驱动候选筛选。
六类候选中只有一个通过三轮门禁
Nsys 归因显示当时 GPU1 kernel-duration 中 true FP8 GEMM 占 45.4866%。据此筛选的六类候选里,external M padding、fast-accum、shared activation quant、text K/V cache 与 simple concat 全部被数据否定,只有 16 MiB cuBLASLt workspace 通过。
先做 exact production shape microbenchmark,通过后才允许三次独立 20 秒进程;显存增量不得超过 128 MiB。
provider-only median wall 17,865.729011 ms / 27.370839427 generated FPS;shared+provider 的增量仅 9.716 ms,低于门槛且 seam 更差。
Promote workspace(现已固化进 Week 8 冻结启动环境),其余五类进入拒绝账本。
Scale granularity 的系统评估,以及一次机制性拆分
rowwise、channelwise、SmoothQuant、K32/K64 与 temporal routing 被系统评估,全部因 worst-cell 或跨 prompt 反转失败。Week 6 把“更细分组”与“E8M0 scale 编码”拆开,发现 exact-FP32 K32 改善 local rel-L2 22.910%,而 ceil-pow2 / native E8M0 反而比 tensorwise 差 1.056%。
对 320 strata × 8 构成 2,560 paired local samples,分别比较 tensor FP32、K32 exact、K32 pow2 与 native K32 E8M0。
四条 video block-probe lanes 通过、四条 audio p95 失败;总 gate 4/8,Phase 5 锁定。
当时判定 native K32 不获部署资格——该判决其后被推翻。
把“离群值伤害 FP8”从直觉改写成量化结论
对 576 层的 13,690,208,256 个权重元素做全量精确统计,并用 63,360 条真实 activation call records 做激活侧分析。核心发现是一个零结果加一个对照:weight tail 对 E4M3 误差的 Spearman ρ 为 −0.046(CI 跨零),而同一指标在 INT8 对照下 ρ 为 0.817。
全量精确计算误差、amax、underflow 与 clip fraction;分位数与 kurtosis 用每层最多 262,144 个确定性等距样本;每个相关系数 4,000 次 family-stratified bootstrap。
output-channel scale 消除 93.89% 的 underflow,却只换来 0.476% 的 MSE 改善——underflow 是 count 型损失,平方能量太小。dense clip oracle 上界仅 0.1478865618259051%。
把 INT8 outlier 文献(clipping / smoothing)排除出 FP8 权重侧路线;研究预算转向激活侧。
自研 MXFP8 attention:CUTLASS 特化 + FA3 血统 + 手写 producer
48 个 video-self attn1 被替换为自研 persistent kernel:signed normalized H128 → E4M3/UE8M0 K32 → native SM120 block-scaled QK → one-pass online softmax → 在线 P K32 量化 → block-scaled PV → BF16。fused core grid 170、block 384、168 registers/thread、84,992 B dynamic shared。
QK/PV 核心特化自 CUTLASS 4.5.2 SM120 blockscaled example;attention 主体为 FA3 血统 warp-specialized persistent 核;量化 producer 为手写 CUDA;集成为 fail-closed custom op。
整机 r04 formal wall 17,411.884168 ms。但 fused cooperative H128 虽达成逐 bit 一致与单 launch,四个生产 shape median 全部回归 +1.512%–+7.440%,因此不作生产默认。
保留“独立 producer + attention core”为生产默认;分布级画质证据仍缺失,且 r04 没有相邻 provider-off 对照。
验收判据的证伪与重建
为验证 outlier 统计挑出的 15 个高危层是否真的重要,设计因果 lane:BF15 把这 15 层从 FP8 manifest 移除(权重与激活两个误差源同时消失,因此是联合上界),R15 对 15 个 family 匹配的最低风险层做同样处理作为安慰剂。
先修方法学缺陷:Week 4 的 runner 每条 lane 重跑 text encoder 且不硬绑定,正是让 Week 5 第一个候选作废的机制。新 wrapper 装上 fail-closed 冻结 conditioning,哈希与参考值逐字节一致。
BF15/chef video −3.5075% 而 drummer +10.7793%;R15/chef 在同一度量下 audio +19.8825%,全面优于处理组。所有效应落在 ±3~20%。
端到端 oracle-L2 降级为描述性指标;建立并标定 reference-free 面板;K32 在新面板下无确凿质量损失,判决反转。
第一次取得严格配对 A/B
576 个 Linear 全部替换为 native K32 MXFP8。三对独立 fresh 进程按 B/K、K/B、B/K 交替顺序执行,每条 lane 独立且必须不存在的 compile-cache root——这是本报告此前一直缺少的严格因果对照。
绑定内容寻址 release manifest;独立验证器重放控制组 baseline reference、从原始 chunk timing 重算 stable-DiT FPS 并拒绝任何身份或环境漂移。
stable-DiT 32.319420887966025 → 36.866330113659890;generated 28.344309608927105 → 32.572401806507680(+14.916900%)。门禁全部 PASS。
Promote 为工程基线 V2;但 formal_baseline_accepted=false,且冷启动 compile warmup 慢 23.342073%。
两个层级的 P0,以及一堵近在 5–9% 之外的墙
对历史 832×480 K32 baseline 做 20 秒全请求 Nsys trace 与 attention 分阶段 CUDA-event benchmark。45 个分析窗口、12 个 stage gate、3 个 attention count gate、5 个 sequence-pair gate 全部 PASS。
Nsys 2025.3 的 NVTX trigger 在本机损坏,改用验证过的 --start-later 会话并离线裁剪到 PROFILE_REQUEST 范围;应用仍跑完整冻结请求。
稳态 DiT 1,358–1,366 ms/chunk 对比 VAE 1,248–1,250 ms/chunk,两卡同时 busy 87.327%。这些是优先级证据,不是当前 Portrait FPS。
将 V layout、VAE 和三大 scaled-mm shape 排为候选;随后 Week 9 必须用新的 A/B 而非将 profiling share 直接换算成收益。
从非可加 stage 投影,到一项 full-model engineering PASS
V-direct-BSH 只改变 attention V 的消费边界:control 显式 materialize BSH→BHSD(padded case 另做 fill/copy),candidate 让 producer direct-consume strided BSH;Q/K、fused core、weights、KV、VAE 与 scheduler 不变。
先用 blocked ABBA/BAAB 证明 four-shape harness neutrality(31/31 checks PASS),再做 four-shape provider e2e、Nsys structure 与三对 20 s full-model A/B。
micro 四 shape 全正;Nsys 显示 aligned 10→9、padded 14→11 launches;full-model stable-DiT 36.923232→37.400230,generated 32.624456→33.035966。另一条 dominant scaled-mm heuristic-pinning screen 仅 +0.058769%,低于 +1% gate。
Promote V-direct-BSH 为历史 frozen K32 stack 的 engineering-performance improvement;Reject frozen 1 MiB heuristic pinning;两者都不主张 quality acceptance。
将已验证 stack 首次固化为 480×832 engineering baseline
保持 576 K32 MXFP8 Linear、48 V-direct-BSH attn1、GPU1 DiT / GPU0 text+VAE、4+6 KV 与 depth-3 decode,仅改变 output orientation,并在三个独立 absent-cache roots 上完整执行。
对每个 repetition 执行 request、K32、V-direct、KV、compile 与 media gates;从原始 timing summary 聚合均值和 population CV。
Portrait 相对 landscape V-direct reference 的差异仅 +0.231955% / +0.105325%,量级不足以解释为 orientation speedup。三次 runtime 皆无 fallback/upcast/retained BF16 converted weights。
Adopt catnip-k32-vdirect-portrait-480x832-v1;该版本已于 2026-07-28 被 v2 取代,保留为历史证据。
接纳已独立报告的 VAE 运行时改动并完成最终晋升门禁
保持 20 秒合同、576 K32、48 V-direct、4+6 KV 与双卡 placement;当前默认入口仅新增 single-tile latent-window decode 和 compiled video VAE decoder。VAE 实验过程不在本报告重复。
以 Week 10 Phase 3 的三对 candidate runs 作为权威 warm KPI,再在无并发作业环境执行独立 20 秒 promotion smoke。
Generated-FPS CV 0.085863%;paired median compile gain +1.837371%;final r03 为 34.408774 generated FPS,489-frame H.264/AAC media 与 runtime gates 全 PASS。
Promote catnip-k32-vdirect-single-tile-compile-portrait-480x832-v2。Fresh-process wall paired median +10.802722%,所以不声明 cold-start speedup;formal quality acceptance 仍未建立。
Experimental Decisions and Rejection Rationale
本节汇总主实验链路之外的 candidate decisions。Rejected 表示实验本身有效但 hypothesis 被数据否定; Blocked 表示已取得有效 gate evidence,但后续 phase 按预注册协议保持锁定;Invalid 不产生科学结论; Not run 表示 protocol-compliant early stopping 或行政阻塞。
| 阶段 | 尝试 | 结果 / 最终原因 | 状态 |
|---|---|---|---|
| Week 2 | Bounded KV + shared PE + fused bias | 11 chunks 完成,reserved 回到 memory gate 内。 | 保留 |
| Week 2 | Host-worker async VAE | 修复 same-thread false overlap,形成 O0 19.080 FPS。 | 保留 |
| O2 | 48-block independent compile | O1→O2 observed wall −2,892.619 ms。 | 保留 |
| O3 | Selective cuDNN SDPA | 单 shape 快约 19.7%,完整请求却慢 220.053 ms。 | 拒绝 |
| O4 | Fold duplicate V2A | Wall −2,380.908 ms;保留 cache 副作用与残差顺序。 | 保留 |
| O5 | 576-module coverage | Wall −824.946 ms;Linear weight coverage 72.138575%。 | 保留 |
| O6/O7 | 5-shape warmup + strict entry | Formal graph delta 为空;三轮 27.104924 FPS,CV 0.046536%。 | 交付 |
| Week 3 | External M padding | 三大 shape 都慢;预计每请求 +52.517 ms。 | 拒绝 |
| Week 3 | Shape-specific fast-accum | Square micro 1.0237×,full median 反而慢 2.402 ms。 | 拒绝 |
| Week 3 | Shared activation quant | Physical quant calls −25%,但 full wall 仅 −0.166418%。 | 拒绝 |
| Week 3 | Text K/V cache | 48 层需 3,840 MiB;9 层 720 MiB 也超 memory gate。 | 拒绝 |
| Week 3 | Simple QKV/KV concat | 乐观 request ceiling 只有 0.173474%。 | 拒绝 |
| Week 3 | cuBLASLt workspace 16 MiB | Wall −0.865511%,FPS +0.873067%;已固化进 Week 8 冻结启动环境。 | 已固化 |
| Week 4 | Row activation / tensor weight | 首个单 cell aggregate 已比 tensorwise 更差。 | 拒绝 |
| Week 4 | Hybrid-K | Overall worst −20.320%,audio median 退化。 | 拒绝 |
| Week 4 | Full rowwise | Overall worst −7.563%;audio worst −61.777% / −65.655%。 | 拒绝 |
| Week 4 | SmoothQuant | Calibration median +22.782%,heldout median −2.991%。判据已失效,需重审。 | 拒绝 |
| Week 5 | Sensitive top-48 / 96 / 192 | 跨 cell 反转。判据已失效,需重审。 | 拒绝 |
| Week 5 | Native K32 | Week 5 观测值保留(chef +0.827%、drummer −7.064%);拒绝依据于 2026-07-22 被证伪,同栈配对 stable-DiT median gain 14.103774%。 | 重开并 Promote |
| Week 5 | Dynamic torch.cond routing | CUDA overhead 10.19%–20.98%,在 full video 之前已超门槛。 | Blocked |
| Week 6 | Exact-FP32 K32 simulator | Local +22.910%,四条 video block-probe lanes 通过。 | 研究保留 |
| Week 6 | K32 pow2 / native E8M0 | Local rel-L2 差 1.056% 的数值保留;部署判决已推翻——它正是基线 V2 的 scale 编码。 | 判决推翻 |
| Week 6 | Exact-K32 propagation | Video 4/4,audio p95 0/4;总 gate 4/8。该判据已被证伪。 | Blocked |
| Outlier 统计 | 576 层全量权重 + 63,360 activation records | INT8 对照(ρ 0.817)证明 outlier 直觉不能移植到 E4M3(ρ −0.046)。 | 研究保留 |
| Week 7 | A3 signed H128 旋转 | 每个 pooled 类别都改善(overall 2.1063%),但门槛为 5%–10%;是幅度不够而非回归。 | 拒绝 |
| Week 7 | A7 概率 K32 | 非零下溢 −95.2250%,但 defect 仅改善 0.0267%;25 项 gate 中 11 项失败。 | 拒绝 |
| Week 7 | P1d learned rotation | odd-chunk mean 仅 +0.246%–+0.623%,门槛 ≥5%;明显局部过拟合。 | 拒绝 |
| Week 7 | Fused cooperative H128 | 5 shape 逐 bit 一致,但四个生产 shape median 全部回归 +1.512%–+7.440%。 | 拒绝 |
| Week 7 | FA3 K/V warp-release | 五 shape bitwise 一致但全部变慢 +0.79%–+2.83%。 | 拒绝 |
| Week 7 | 自研 SM120 MXFP8 attention | 48 个 attn1 替换;r04 整机 28.084266773 generated FPS;主形状隔离 1.44236×。 | 生产默认 |
| Week 8 | 质量面板证伪旧判据 | 处理组跨 cell 符号翻转,安慰剂优于处理组;判据降级为描述性指标。 | 方法学 |
| Week 8 | Reference-free 解码媒体面板 | 伪影注入标定:flicker 25.68×、click 8.51×、sharpness 1.78×、blockiness 1.13×。 | 新验收权威 |
| Week 8 | Native K32 MXFP8 baseline V2 | 三对交替配对 stable-DiT median gain 14.103774%、最小配对 13.995281%;两 lane AUTO_PASS。 | Promoted |
| Week 8 | 全局 Q/K single-launch | 仅最小 shape 快 17.117%;20 秒加权慢 0.934%。 | 拒绝 |
| Week 8 | 全家族 rowwise(ROW192) | 面板 worst dev 1.709,seam spike 6.5 对比 3.1,本轮最强伪影信号。 | 拒绝 |
| Week 7 | MXFP8 r01 / r02 / r03 | 外部作业占用 GPU 或 provider dtype gate 过严;均在 formal timer 之前停止,无 FPS。 | Invalid |
| Week 8 | BF15 首次 drummer 运行 | 编译预热阶段被中断,无 run.json、无 probe、无对比;已隔离标注。 | Invalid |
| Week 8 | Anatomy capture 首次尝试 | 与前一运行 teardown 撞车导致 GPU0 OOM;清理后重跑成功。 | Invalid |
| Week 8 | 第一版面板的自然带 | 使用了原 final-heldout 两格,构成泄漏;V2 已重分类并另设 heldout-v2。 | Invalid |
| Attention | NVML 100% / Kineto 301% | 指标或计时边界错误;被 25-shape complete-op V2 取代。 | Invalid |
| Week 8 | Calibration freeze | marker status/quality_calibration.freeze.json 不存在。 | Blocked |
| Week 8 | 面板 av-sync gate | 注入 400 ms 音频延迟未被检出,代理指标标定失败。 | 降级 diagnostic |
| Week 8 | 两人独立盲评 | 公开包与响应验证器已就绪,等待两名独立评审。 | PENDING |
| Week 8 | Validation 4 runs | 需 freeze marker 解锁;已完成 0/4。 | LOCKED |
| Week 8 | Heldout-v2 两格 | 需 hash-bound 书面解锁。 | EMBARGOED |
| NCU | 完整 tensor/stall/achieved-occupancy counter | counter-free ncu --set none 已给出 launch/wave/theoretical occupancy;完整 PerfWorks counters 仍受权限阻断。 | 部分解锁 |
| Week 9 | V-direct-BSH attention boundary | 历史 832×480 K32 stack 上三对 full-model A/B:stable-DiT median +1.289408%,generated +1.284480%,6/6 gates PASS;不声称 quality acceptance。 | Engineering pass |
| Week 9 | cuBLASLt heuristic top-k pinning | 主 shape bitwise correct 但 paired median 仅 +0.058769%,低于 +1% gate;按三 shape conjunction early stop。 | 拒绝 |
| Week 9 | Counter-free NCU --set none | attention 1.00 wave;P0 GEMM 3.58 waves / 89.4% efficiency。关闭 tile/stage 扩张假设,保留 stream-K/persistent。 | 方法学 |
| Portrait v1 | catnip-k32-vdirect-portrait-480x832-v1 | 三次 fresh-cache mean 37.486982 stable-DiT / 33.070761 generated FPS;现由 v2 取代。 | 历史基线 |
| Portrait v2 | Single-tile + compiled video VAE decoder | 三对 candidate mean 37.557043 stable-DiT / 34.392314 generated FPS;final r03 smoke PASS;warm-throughput only。 | 当前工程基线 |
| Portrait v2 | Promotion smoke r01/r02 | 并发 Week 10 KV study 占用 GPU,VAE decode 前 OOM;清场后 r03 完整通过。 | Invalid |
两类使用规则:数据否决 vs 判据失效
第二类,判据失效、必须重审:所有以端到端 trajectory rel-L2 作为 accept/reject 依据的结论——包括 Week 4 与 Week 5 的全部精度 reject 与 Week 6 的 propagation 阻断——都建立在一个已被证伪的度量上。它们不自动变成“其实可行”,而是回到未判定状态。K32 是第一个走完这条路径并回到工程基线的候选。
Role of Invalid records in the evidence hierarchy
Current Engineering Baseline and Acceptance State
本节区分 current engineering baseline、paired control 与 historical delivery,以避免将实验 promotion 误表述为 production acceptance。
| 项目 | 冻结配置 |
|---|---|
| Quantization | 576 selected Linear:native K32 MXFP8 E4M3 qdata、K32 E8M0 scale、BF16 output + fused bias。0 fallback / 0 upcast / 0 retained converted BF16 weight;不是 full-model FP8。 |
| Attention | 48 video-self attn1:V-direct-BSH strict fail-closed provider。Week 9 验证 V boundary 的历史 engineering A/B;Portrait repeats 验证集成 stack。 |
| Streaming | 480×832 Portrait、20 s、489 frames、11 chunks、55 forwards、first 4 + latest 6 LF KV、390 spatial tokens/LF。 |
| Scheduling | GPU1 DiT + GPU0 text/VAE;per-chunk video decode depth 3,video drain 后单次 audio decode;每个 latent window 使用 single tile。 |
| Compilation | 48 DiT blocks independently compiled;当前 v2 另编译 video VAE decoder forward。性能声明限于 warm steady state。 |
| Current baseline | three-candidate mean 34.392314 generated FPS / 37.557043 stable-DiT FPS;final smoke 34.408774 generated FPS。 |
| Repeatability | population CV:generated 0.085863% / stable-DiT 0.151246%;Phase 3 paired order 为 control/candidate、candidate/control、control/candidate。 |
| Historical causal evidence | 832×480 K32/tensorwise +14.103774% stable-DiT,Week 9 V-direct-BSH +1.289408%;二者均不替代 Portrait absolute KPI。 |
| Historical delivery | O7 27.104924 generated FPS;保留为 Week 2 delivery reference。 |
| Release identity | catnip-k32-vdirect-single-tile-compile-portrait-480x832-v2;ACTIVE engineering baseline;adopted 2026-07-28。 |
| Claim boundary | Warm-throughput only;fresh-process wall paired median +10.802722%,no cold-start speedup。Runtime/media PASS does not establish ground-truth perceptual-quality acceptance。 |
Unresolved evidence requirements
Portrait v2 formal quality acceptance 仍未建立:已有 VAE 报告没有 ground truth 或正式 blind panel,且不能由 runtime/media gate 自动替代。v2 cold-start speedup 也未建立。完整 NCU 的 tensor-pipe、warp-stall、achieved-occupancy 与 duration counters 仍受权限限制;matched commit/workload/timer 的 BF16↔FP8 performance A/B 仍未建立,因此本报告不给出官方 BF16→FP8 speedup 倍数。
Active Problems and Evidence-Constrained Optimization Space
Portrait v2 已接纳 GPU0 路径的首轮优化,当前 warm KPI 为 34.392314 generated FPS。下一轮不能继续把历史 profiling 占比当作当前精确瓶颈;排序遵循单变量、paired 20 s verification 与质量边界。