<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Recurrent-Architecture on Feng's Blog</title><link>http://fengwang.github.io/tags/recurrent-architecture/</link><description>Recent content in Recurrent-Architecture on Feng's Blog</description><generator>Hugo -- gohugo.io</generator><language>en-us</language><lastBuildDate>Sun, 13 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="http://fengwang.github.io/tags/recurrent-architecture/index.xml" rel="self" type="application/rss+xml"/><item><title>From 'A Severe Misalignment' to a Mathematics Industrial Stack: A Critical Reading and a Blueprint</title><link>http://fengwang.github.io/posts/math-industrial-stack/</link><pubDate>Sun, 13 Sep 2026 00:00:00 +0000</pubDate><guid>http://fengwang.github.io/posts/math-industrial-stack/</guid><description>&lt;p&gt;Why this exists: On 2026/9/11, 25 Fields medalists published &amp;ldquo;A Severe Misalignment of AI in Mathematics.&amp;rdquo; Two days later I read the declaration once straight through and once in reverse, then translated the criticism into design: if treating proof-solving as a benchmark really is the problem, what kind of institution would fix it? I care about the next executable move, not about taking sides.&lt;/p&gt;
&lt;p&gt;The thesis: the declaration diagnosed a real bottleneck, the community&amp;rsquo;s digestion bandwidth, but filed it under a borrowed label; the right response is not to resist industrialization but to industrialize the guild&amp;rsquo;s four functions, production, verification, transmission, and credit, one by one.&lt;/p&gt;
&lt;p&gt;Scope: this covers the declaration text and the event timeline, the strongest opposing case, the calibration between guilds and industrialization, and the design of a four-layer industrial stack for mathematics. It does not cover the signatories&amp;rsquo; individual positions or independent verification of OpenAI&amp;rsquo;s technical results. Every design in this article is never untested, and is just a result of brainstorming.&lt;/p&gt;
&lt;p&gt;Prerequisites: this assumes familiarity with the norms of pure mathematics research, basic concepts of Lean / mathlib and formal verification, and the general shape of research funding and incentive structures.&lt;/p&gt;
&lt;h2 id="the-event-a-declaration-and-the-week-it-detonated"&gt;The event: a declaration and the week it detonated&lt;/h2&gt;
&lt;p&gt;The declaration itself is short on paper. &amp;ldquo;A Severe Misalignment of AI in Mathematics,&amp;rdquo; signed by 25 Fields medalists running from Deligne (1978) to Deng Yu (2026), with Tao, Scholze, and Villani in between. It appeared on September 11 on Tao&amp;rsquo;s blog and at mathandai.org, organized as invitation-based co-signing in the manner of the Leiden Declaration from June.&lt;/p&gt;
&lt;p&gt;What triggered it was a chain of events in late August and early September. The direct fuse: OpenAI announced that an internal model, running roughly 10,000 agents for 88 hours, had &amp;ldquo;solved&amp;rdquo; a Navier-Stokes blowup-class problem. Four days earlier, Buckmaster and Alpöge (Anthropic) had released a 245-page draft on the forced Euler equation and accused OpenAI of scooping it. As the authorship dispute escalated, OpenAI withdrew its sponsorship of the Caltech math marathon. The letter&amp;rsquo;s core claims: AI companies treat &amp;ldquo;solving famous problems&amp;rdquo; as a benchmark, badly misaligned with the mathematics community&amp;rsquo;s goals; rushed announcements, no formal write-ups, and recurring authorship and plagiarism problems; and without the human community digesting them, AI&amp;rsquo;s ideas &amp;ldquo;can never come alive.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;I pinned the key nodes to a timeline first, because these details have been rewritten through several rounds of retelling:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;Event&lt;/th&gt;
&lt;th&gt;Reliability&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2026-05&lt;/td&gt;
&lt;td&gt;OpenAI&amp;rsquo;s internal model produces a counterexample to the Erdős unit distance conjecture; nine mathematicians (including Tsimerman) verify it&lt;/td&gt;
&lt;td&gt;secondhand&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;06-02&lt;/td&gt;
&lt;td&gt;Leiden Declaration published (eight-month consultation, concrete recommendations)&lt;/td&gt;
&lt;td&gt;verified&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;07-23&lt;/td&gt;
&lt;td&gt;ICM Philadelphia: 2026 Fields Medals (Deng Yu, Wang Hong, Pardon, Tsimerman); Tsimerman announces he is joining OpenAI&lt;/td&gt;
&lt;td&gt;Nature / Clay&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;09-07&lt;/td&gt;
&lt;td&gt;Buckmaster &amp;amp; Alpöge (Anthropic) post their forced-Euler result, a 245-page draft; accuse OpenAI of scooping&lt;/td&gt;
&lt;td&gt;multi-source, disputed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;09-08&lt;/td&gt;
&lt;td&gt;OpenAI announces the 88-hour &amp;ldquo;solution&amp;rdquo; of an NS blowup (about 10,000 agents, millions of dollars of compute)&lt;/td&gt;
&lt;td&gt;single source, unverified&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;09-08 ~&lt;/td&gt;
&lt;td&gt;&amp;ldquo;Merger plan&amp;rdquo; accusations and counter-accusations; OpenAI&amp;rsquo;s blog concedes it &amp;ldquo;cannot rule out that de-identified user data helped the model&amp;rdquo;&lt;/td&gt;
&lt;td&gt;both sides contest&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;09-11&lt;/td&gt;
&lt;td&gt;This declaration; OpenAI exits the Caltech sponsorship&lt;/td&gt;
&lt;td&gt;multi-source&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;One media confusion deserves its own note. OpenAI&amp;rsquo;s result (a static fluid blowup under smooth forcing) and the Clay Millennium Problem&amp;rsquo;s original statement (no forcing, smooth initial data) are related but different problems. Even if the proof holds, it almost certainly does not qualify for the Clay prize, which requires peer-reviewed publication plus a two-year waiting period. The popular claim that &amp;ldquo;the Millennium Problem was solved&amp;rdquo; is itself a product of the marketing-style communication this controversy criticizes. That reading is my inference, not settled fact.&lt;/p&gt;
&lt;p&gt;Pin the timeline first, then argue: the declaration and the counter-declaration are both fighting for narrative, and the timeline does not take sides.&lt;/p&gt;
&lt;h2 id="dissection-five-claims-and-one-concept-slide"&gt;Dissection: five claims and one concept slide&lt;/h2&gt;
&lt;p&gt;On the first pass I agreed with nearly everything: benchmarks, scooping, authorship, student training, every item sounded right. On the second pass, item by item, approval and wariness rose together. I broke the declaration&amp;rsquo;s case into five claims, with hidden premises and my assessment:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Claim&lt;/th&gt;
&lt;th&gt;Hidden premise&lt;/th&gt;
&lt;th&gt;Assessment&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Treating problem-solving as a benchmark harms mathematics&lt;/td&gt;
&lt;td&gt;Solving problems is only a proxy for &amp;ldquo;understanding&amp;rdquo;; once the proxy becomes the target, Goodhart bites&lt;/td&gt;
&lt;td&gt;Mechanism plausible, but asserted, not argued&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Mass production of true/false propositions destroys the soil&lt;/td&gt;
&lt;td&gt;Mass production has happened or is imminent&lt;/td&gt;
&lt;td&gt;Insufficient evidence: headline results number a handful per year; anticipatory rhetoric&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Rushed announcements cause authorship/plagiarism problems&lt;/td&gt;
&lt;td&gt;Misconduct traces back to benchmark competition&lt;/td&gt;
&lt;td&gt;Causal chain plausible but unproven; and it is a conduct problem, not a problem of benchmark-making itself&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Without mathematicians digesting them, AI&amp;rsquo;s ideas never come alive&lt;/td&gt;
&lt;td&gt;The community is an indispensable digestion system&lt;/td&gt;
&lt;td&gt;Deeply coherent, and simultaneously a claim of power&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Mathematics&amp;rsquo; problems mirror society&amp;rsquo;s problems&lt;/td&gt;
&lt;td&gt;Process-is-product fields are isomorphic to product-is-product fields&lt;/td&gt;
&lt;td&gt;True for pure mathematics, overgeneralized&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The two loudest rows are claims 2 and 3: &amp;ldquo;mass production destroys the soil&amp;rdquo; is anticipatory rhetoric, since headline-grade results appear at most a handful of times a year; and the causal chain from benchmark competition to misconduct is plausible but unproven.&lt;/p&gt;
&lt;p&gt;More important is a slide that sits outside the table. The title condemns &amp;ldquo;treating problem-solving as a benchmark,&amp;rdquo; but nearly every harm listed in the body comes from speed, authorship, and data provenance. Those are conduct problems, not problems with benchmark-making as such. Hilbert&amp;rsquo;s 23 problems were a benchmark set by the community itself, and nobody accused them of destroying mathematics. The real variable is &lt;strong&gt;who owns the benchmark and what incentives it settles&lt;/strong&gt;: community-set, honor-settled, on a ten-year clock, versus company-set, funding-narrative-settled, on a weekly clock. Same machinery; who holds the dial and whose clock it runs on is the entire difference.&lt;/p&gt;
&lt;p&gt;Having written that dissection, my impression on the declaration: true, rhetorically skilled, its core worry sound, its argumentative structure visibly incomplete. More war cry than argument. And it commits the same rushed-publication error it condemns, written in a week. Confidence: medium-high. The text was checked word by word; the OpenAI-side details rest on a single secondhand source.&lt;/p&gt;
&lt;h2 id="what-holds-proofs-are-compressed-packages-the-community-is-the-decompressor"&gt;What holds: proofs are compressed packages, the community is the decompressor&lt;/h2&gt;
&lt;p&gt;The declaration&amp;rsquo;s deepest insight fits in one sentence: a proof is a compressed package; the human community is the decompressor.&lt;/p&gt;
&lt;p&gt;A proof being &amp;ldquo;found&amp;rdquo; is not the same as it becoming knowledge. Take the 245-page forced-Euler draft. Even if every line is right, it still has to be retold by peers as a shorter story, simplified to its core mechanism, written into teaching material, and finally compressed into one intuition in some young person&amp;rsquo;s head. Only then is it alive. Every station on that chain has a human bandwidth limit. This is what the declaration sees most clearly: AI&amp;rsquo;s production rate has begun to exceed the community&amp;rsquo;s &lt;strong&gt;digestion bandwidth&lt;/strong&gt;, and the bottleneck has moved from the generation side to the absorption side.&lt;/p&gt;
&lt;p&gt;The intuition reduces to one line of accounting:&lt;/p&gt;
&lt;p&gt;$$\frac{dB}{dt} = p - d \quad (p &amp;gt; d)$$&lt;/p&gt;
&lt;p&gt;$B$ is the undigested backlog, $p$ the number of worthwhile results produced per week, $d$ the rate at which the community digests them. Nothing needs solving here; the point is that &amp;ldquo;go faster&amp;rdquo; can never fix this problem. As long as $p$ keeps rising, the backlog grows linearly. The only lever that acts on the bottleneck is raising $d$.&lt;/p&gt;
&lt;p&gt;Several other parts of the declaration hold up too. The worry about student training is concrete: mathematicians are trained by handing students problems from a pool that is still open, with mentors teaching through problems. If that pool is swept away, or merely believed to be swept away, the pipeline starves before the facts arrive. The signature list also carries information: Tao has demonstrated AI + Lean workflows in public for years, so his signature means the letter rises above the usual resistance to new tools. The letter also concedes that AI can accelerate real research; reading it as a Luddite manifesto misreads it. For professions where the process is the product, the closing warning (&amp;ldquo;your field is next&amp;rdquo;) generalizes.&lt;/p&gt;
&lt;p&gt;One last observation complicates the insight. &amp;ldquo;Without willing mathematicians, AI&amp;rsquo;s ideas can never come alive&amp;rdquo;: that sentence is a lament and an ace at the same time. If it is true, mathematicians still hold a monopoly on the final settlement of meaning; what is threatened is not the discipline&amp;rsquo;s survival but its attention economy. The tone of fear sits oddly with that hidden position of power. I read it as the most charged line in the letter: a distress signal and a bargaining chip.&lt;/p&gt;
&lt;p&gt;The insight holds best for pure mathematics. Pure mathematics has almost no direct practical output; understanding itself is the product, and a theorem that is true but understood by no one barely exists for the field. For applied mathematics and engineering the sentence fails: a working algorithm does not need to be understood by everyone first. Keep this boundary in view; it recurs below.&lt;/p&gt;
&lt;h2 id="what-does-not-hold-borrowed-rhetoric-and-the-silence-on-formalization"&gt;What does not hold: borrowed rhetoric and the silence on formalization&lt;/h2&gt;
&lt;p&gt;After reading the declaration until it went stale, I extracted five hard defects. The fifth is the strangest, because the tool it ignores is the one the most AI-literate signatory has championed for years.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Paradoxical self-undermining. To meet &amp;ldquo;urgency,&amp;rdquo; it skipped the eight-month consultation used by the Leiden Declaration and was written in a week, repeating the rushed publication it condemns. Tao himself acknowledged it was &amp;ldquo;unfortunate.&amp;rdquo; A text asking an industry to slow down did not slow down.&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Misalignment&amp;rdquo; is borrowed rhetoric. In AI safety the term has a specific meaning; what the declaration actually describes is incentive misalignment and externalities. The verb is its own; the noun is borrowed. Clever for reach, loose for precision.&lt;/li&gt;
&lt;li&gt;Authority as argument. Twenty-five medals are a credibility strategy, not a sample of &amp;ldquo;the community.&amp;rdquo; This is a joint session of the elite pure-mathematics wing: the side with positive net benefit that joined OpenAI (Tsimerman) is not in the room. In the name of the community, a faction speaks.&lt;/li&gt;
&lt;li&gt;No asks. Against the Leiden Declaration&amp;rsquo;s concrete recommendations, this text is pure indictment. A protest without asks leaves rule-making to the labs, again. This is the defect I care about most.&lt;/li&gt;
&lt;li&gt;Silence on formal verification, the strangest omission. &amp;ldquo;Verified but not understood&amp;rdquo; is the precise form of the crisis, and the Lean path is exactly what Tao has championed in public, and the ready-made foundation for a standards bureau. The text is silent on it. I will not guess why; the cost is clear. It gives up its own most concrete engineering proposal. This is the largest argumentative gap in the letter.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Of the five, some criticize communication strategy and some criticize argument structure. Both kinds count, but I will not blend them: reach and rigor are different ledgers. The declaration may complete a historical task by forcing the issue into public view even if its case never fully stands.&lt;/p&gt;
&lt;h2 id="the-strongest-opposition-and-why-it-does-not-overturn-the-conclusion"&gt;The strongest opposition, and why it does not overturn the conclusion&lt;/h2&gt;
&lt;p&gt;In fairness, the declaration deserves the strongest opponent I can build.&lt;/p&gt;
&lt;p&gt;The case runs: benchmark competition harms no theorem&amp;rsquo;s truth value. A true proof is a net gain to human intelligence whether it arrives in 88 hours or eight years. Authorship disputes are checkable conduct disputes; they belong in an investigation, not in a verdict on an industry&amp;rsquo;s methodology. Every tool revolution (computer algebra, numerical methods) came with the same laments, and mathematics digested every tool. What the medalists actually fear is the devaluation of &amp;ldquo;first prover&amp;rdquo; glory, which is their life&amp;rsquo;s capital. The cure for the student-training worry is reforming how students are trained, not prosecuting how results are measured. A helicopter ascent is not a climb, for the mountaineer; for the humans who only need the flag on the summit (the theorem), it is.&lt;/p&gt;
&lt;p&gt;I agree with the opposition&amp;rsquo;s three factual points: truth values are unharmed, authorship disputes are procedural, and tool revolutions have precedent. But it cannot answer digestion bandwidth. The declaration&amp;rsquo;s real claim lives not at the truth level but at the level of how knowledge is socially produced: for pure mathematics, understanding is the product, and a true-but-ununderstood theorem barely exists. The backlog is not an honor problem; it is a problem of the discipline&amp;rsquo;s capacity to reproduce itself. So the right reading is not either/or but both layers stacked: optimistic at the truth layer, watchful at the production layer. The declaration&amp;rsquo;s error is billing both layers to &amp;ldquo;benchmarks&amp;rdquo;; the opposition&amp;rsquo;s limit is refusing to look at the second layer at all.&lt;/p&gt;
&lt;p&gt;Two bias inventories go on the record. For the declaration: appeal to authority, framing effects, availability (one incident generalized to all companies), group reinforcement (a like-minded draft in a week). For my side: collaborating with an LLM model to analyze a declaration about models is a conflict of interest. The only workable rule is to score argument quality instead of identity. The same rule applies to the signatories: the argument is in the body, not the signature block.&lt;/p&gt;
&lt;h2 id="calibration-guilds-and-industrialization-are-not-a-zero-sum-replacement"&gt;Calibration: guilds and industrialization are not a zero-sum replacement&lt;/h2&gt;
&lt;p&gt;My first characterization of the declaration: a lament of the old order invoking its last rights, the classic script of a guild facing industrialization. &amp;ldquo;Only the best-adapted survives.&amp;rdquo; That framing has a hidden premise: industrialization destroys guilds.&lt;/p&gt;
&lt;p&gt;The more reliable line from technology history runs differently: leaders define paradigms, and &amp;ldquo;leaders define paradigms&amp;rdquo; does not mean &amp;ldquo;old institutions vanish.&amp;rdquo; The medieval guilds were not destroyed by the industrial revolution; they transformed into standards bodies and qualification systems. ISO, medical licensing, and peer review are all industrialized descendants of guild quality-control functions. The real precondition for industrialization to succeed was the guild ceding production while keeping a monopoly on standards and taste.&lt;/p&gt;
&lt;p&gt;So the right question is not how to bypass the guild but how to industrialize its four functions (production, verification, transmission, honor) one by one, until the guild retreats to a position nothing else can occupy. Most of my original judgment survives: OpenAI carries ethical defects, and it is also leading where mathematics may be going; finding the next move matters more than stopping to complain. What I revised is one structural point: this is a function transfer, not a replacement. If the declaration traded lament for institution-building, it would be the natural candidate for the standards bureau. Its closing line, that everything &amp;ldquo;depends on the decisions of the humans who control the technology,&amp;rdquo; already concedes as much.&lt;/p&gt;
&lt;p&gt;This calibration has its own boundary: function transfer requires the guild to be willing and able to take the standards-bureau seat. The declaration&amp;rsquo;s behavior (no asks) suggests that willingness is not in place. If the community keeps refusing institution-building, the &amp;ldquo;cede production, keep standards&amp;rdquo; path does not exist either, and what remains really is zero-sum.&lt;/p&gt;
&lt;h2 id="blueprint-a-four-layer-industrial-stack-for-mathematics"&gt;Blueprint: a four-layer industrial stack for mathematics&lt;/h2&gt;
&lt;p&gt;With the declaration dissected and the premise recalibrated, I inverted the question into a generative one: how should mathematics&amp;rsquo; production, verification, transmission, and credit systems be reorganized so that AI output flows into human understanding at industrial throughput, with rigor enforced by infrastructure rather than by gatekeeping?&lt;/p&gt;
&lt;p&gt;The TRIZ-style contradiction: AI produces proofs on a weekly clock; the community digests on a yearly clock. The design goal is not to balance these two but to dissolve the contradiction. Proofs self-verify; understanding is produced on demand; mathematicians spend time only on what machines cannot do (deciding what is worth doing, and what it means); rigor is enforced by infrastructure instead of spot-checked at gates.&lt;/p&gt;
&lt;p&gt;Round one was generation without evaluation: 22 raw ideas across five families. Production systems (proof assembly line, mathematics OEM/ODM, journals as package registries). Credit and incentives (dual currency, fine-grained attribution graph, problem futures market). Verification and quality (a math FDA, red-team swarm, benchmark disclosure standard). Transmission and understanding (understanding factory, apprenticeship 2.0, canon curator). Wild seeds (orphan theorem adoption registry, reward questions over answers, mathematics digital twin). At convergence one regularity floated up on its own: the ideas that hit digestion bandwidth, rigor, and attribution at the same time all land on the same move, which is to turn verification and settlement from a gate into infrastructure. The four-layer architecture of the &lt;strong&gt;industrial stack&lt;/strong&gt; follows:&lt;/p&gt;
&lt;p&gt;Scored against six criteria (digestion bandwidth, rigor, attribution, student pipeline, technical feasibility, institutional adoptability), three families rose to the top:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Proposal&lt;/th&gt;
&lt;th&gt;Digestion&lt;/th&gt;
&lt;th&gt;Rigor&lt;/th&gt;
&lt;th&gt;Attribution&lt;/th&gt;
&lt;th&gt;Student pipeline&lt;/th&gt;
&lt;th&gt;Feasible&lt;/th&gt;
&lt;th&gt;Adoptable&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Math-FDA + red-team swarm (family C)&lt;/td&gt;
&lt;td&gt;○&lt;/td&gt;
&lt;td&gt;●●&lt;/td&gt;
&lt;td&gt;○&lt;/td&gt;
&lt;td&gt;○&lt;/td&gt;
&lt;td&gt;● (formalization is mature)&lt;/td&gt;
&lt;td&gt;● high (Clay&amp;rsquo;s two-year rule is a prototype)&lt;/td&gt;
&lt;td&gt;best&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dual currency + attribution graph (family B)&lt;/td&gt;
&lt;td&gt;●&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;●●&lt;/td&gt;
&lt;td&gt;●&lt;/td&gt;
&lt;td&gt;● pure institutional design&lt;/td&gt;
&lt;td&gt;○ needs community coordination&lt;/td&gt;
&lt;td&gt;second&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Understanding factory + layered bundles (family D)&lt;/td&gt;
&lt;td&gt;●●&lt;/td&gt;
&lt;td&gt;○&lt;/td&gt;
&lt;td&gt;○&lt;/td&gt;
&lt;td&gt;●●&lt;/td&gt;
&lt;td&gt;● demo-able now&lt;/td&gt;
&lt;td&gt;● high (a university can run it)&lt;/td&gt;
&lt;td&gt;second&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Proof assembly line (family A)&lt;/td&gt;
&lt;td&gt;○&lt;/td&gt;
&lt;td&gt;●&lt;/td&gt;
&lt;td&gt;○&lt;/td&gt;
&lt;td&gt;○&lt;/td&gt;
&lt;td&gt;○ needs agent orchestration to mature&lt;/td&gt;
&lt;td&gt;○ companies more willing&lt;/td&gt;
&lt;td&gt;middle&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Problem futures market (family B)&lt;/td&gt;
&lt;td&gt;○&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;○&lt;/td&gt;
&lt;td&gt;○&lt;/td&gt;
&lt;td&gt;●&lt;/td&gt;
&lt;td&gt;○ liquidity problem&lt;/td&gt;
&lt;td&gt;watch&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;(●● direct hit / ● related / ○ indirect / — irrelevant)&lt;/p&gt;
&lt;p&gt;Three paradigm routes crystallized:&lt;/p&gt;
&lt;p&gt;&amp;ldquo;Certification First&amp;rdquo; (Math-FDA). Translate the declaration&amp;rsquo;s anger into a standard: to claim success on a major result with AI, a team must supply a formal certificate, survive a red-team bounty window, and publish a reproduction kit before the result earns a &amp;ldquo;verified&amp;rdquo; certification; the human community then confers &amp;ldquo;understood&amp;rdquo; grades afterward. The guild&amp;rsquo;s gatekeeping function upgrades into a standards bureau&amp;rsquo;s mandatory certification. Guild transformation, not guild lament.&lt;/p&gt;
&lt;p&gt;&amp;ldquo;Understanding Economy.&amp;rdquo; Make digestion bandwidth into an industry. Proofs are ore; understanding is the finished good. Orphan theorem adoption, understanding currency, and a profession of understanding engineers finally price the transmission labor the medalists treasure.&lt;/p&gt;
&lt;p&gt;&amp;ldquo;Continuous Credit.&amp;rdquo; The structural root of authorship disputes is winner-take-all priority. Replace it with Git-style per-lemma attribution plus transitive credit, and scooping stops paying. Not a moral improvement, a design that makes the bad strategy unprofitable.&lt;/p&gt;
&lt;h2 id="first-tests-three-pilots-and-one-load-bearing-dependency"&gt;First tests: three pilots and one load-bearing dependency&lt;/h2&gt;
&lt;p&gt;A blueprint that never meets a test is just style. The design carries three minimal real tests, each with an explicit falsification condition. Cheap to run, quick to kill:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pilot&lt;/th&gt;
&lt;th&gt;Content&lt;/th&gt;
&lt;th&gt;Falsification condition&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Layered proof bundle demo (1-2 weeks)&lt;/td&gt;
&lt;td&gt;Take a recent long AI-adjacent proof (e.g., the 245-page forced-Euler draft) and generate four layers: Lean-checkable subset, human-readable proof, intuition layer (&amp;ldquo;proof comic&amp;rdquo;), dependency map&lt;/td&gt;
&lt;td&gt;If the understanding layer cannot be produced at reasonable cost and be understood by independent readers, the understanding-factory hypothesis dies&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dual-currency workshop&lt;/td&gt;
&lt;td&gt;One Polymath-style collaboration with full dual-currency bookkeeping&lt;/td&gt;
&lt;td&gt;If nobody behaves differently in response to understanding currency, the honor economy cannot scale&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Red-team swarm micro-pilot&lt;/td&gt;
&lt;td&gt;A $1,000 counterexample bounty on any published AI proof&lt;/td&gt;
&lt;td&gt;If the red team neither finds holes nor raises trust, the swarm idea is out&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Before landing the blueprint I also ran a force-field analysis on the leading combination:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Driving force&lt;/th&gt;
&lt;th&gt;Blocking force&lt;/th&gt;
&lt;th&gt;Neutralizing move&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Lean / mathlib maturity (past critical mass)&lt;/td&gt;
&lt;td&gt;Companies have no incentive to accept certification&lt;/td&gt;
&lt;td&gt;Make certification a prerequisite for benchmark rankings (rankings are the companies&amp;rsquo; real currency)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Talent in the Tsimerman mold willing to migrate&lt;/td&gt;
&lt;td&gt;Community coordination failure (who leads?)&lt;/td&gt;
&lt;td&gt;Reuse existing Clay / IMU institutions; do not build from scratch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Companies have PR incentives to self-constrain&lt;/td&gt;
&lt;td&gt;Young mathematicians un-incentivized to do digestion labor&lt;/td&gt;
&lt;td&gt;Tie understanding currency to career advancement (counted in hiring and tenure review)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open-source software supplies complete precedent&lt;/td&gt;
&lt;td&gt;Cultural identity: &amp;ldquo;industrialization equals vulgarization&amp;rdquo;&lt;/td&gt;
&lt;td&gt;Reframe: the honorable history from guild to standards bureau (ISO, medical licensing, both guild legacies)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;There is exactly one single point of failure, and it is load-bearing: the whole stack depends on formal verification coverage. Large parts of mathematics, geometric and intuitive arguments especially, are extremely expensive to formalize. If coverage stalls, the certification layer degrades and the stack collapses toward &amp;ldquo;credit layer standing alone.&amp;rdquo; Two leading indicators to watch: mathlib&amp;rsquo;s annual growth rate, and the formalization lag for major results. Those two numbers set the schedule for everything above. If they do not move, all of this stays on paper.&lt;/p&gt;
&lt;h2 id="boundary-conditions"&gt;Boundary conditions&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Fact layer: OpenAI-side details (88 hours, 10,000 agents, the &amp;ldquo;merger plan&amp;rdquo; texts) come from a single secondhand source and must be flagged as such when cited. The reading of &amp;ldquo;Millennium Problem solved&amp;rdquo; as marketing-style communication is inference, not settled.&lt;/li&gt;
&lt;li&gt;Judgment layer: if an independent investigation shows Buckmaster&amp;rsquo;s accusation to be false, the declaration&amp;rsquo;s &amp;ldquo;misconduct&amp;rdquo; pillar weakens, and the motive map needs re-estimation.&lt;/li&gt;
&lt;li&gt;Design layer: the 22 ideas and three routes have met no pilot; &amp;ldquo;Math-FDA is optimal&amp;rdquo; is a matrix score, not an empirical result.&lt;/li&gt;
&lt;li&gt;Method layer: this article is a cleaned-up record of a 2026-09-13 conversation with an AI assistant. Using a model to analyze a declaration about models is a conflict of interest; the handling rule is stated in the body, but future me should assume the bias is present.&lt;/li&gt;
&lt;li&gt;Evidence that would change my judgment: if within 12-18 months a major AI-assisted result arrives formally verified, fully written up, properly attributed, and digested by the community without friction, the &amp;ldquo;mass production destroys the soil&amp;rdquo; prediction is falsified. Conversely, if evidence emerges that AI scooping is causing young mathematicians to systematically abandon deep specializations, my &amp;ldquo;insufficient evidence&amp;rdquo; grade on claim 2 goes up.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="open-questions"&gt;Open questions&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Which pilot first? I lean toward the layered proof bundle demo: cheapest (1-2 weeks), and it tests the most central assumption of the four layers, the understanding layer.&lt;/li&gt;
&lt;li&gt;Will mathandai.org&amp;rsquo;s co-signing list grow beyond Fields medalists? If so, the declaration&amp;rsquo;s representation shifts from elite wing to community, and its argumentative weight shifts with it.&lt;/li&gt;
&lt;li&gt;When will independent verification of OpenAI&amp;rsquo;s NS result land? It tests both the declaration&amp;rsquo;s &amp;ldquo;misconduct&amp;rdquo; pillar and the credibility baseline for AI proofs.&lt;/li&gt;
&lt;li&gt;Will the Leiden Declaration&amp;rsquo;s concrete recommendations absorb this anger into a combined text with actual asks? From indictment to institution-building is one step.&lt;/li&gt;
&lt;li&gt;Is there an acceptable threshold for formalization coverage? If geometric and intuitive arguments can only ever be partially formalized, should the stack design a bypass for the formalization-resistant zone?&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;[[Q]] Eighteen months from now: did the layered proof bundle demo get run? How far did mathlib&amp;rsquo;s growth rate and the formalization lag move? Do I still agree with today&amp;rsquo;s &amp;ldquo;guild to standards bureau&amp;rdquo; judgment?&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&amp;ldquo;A Severe Misalignment of AI in Mathematics&amp;rdquo; (declaration), Terry Tao&amp;rsquo;s blog, 2026-09-11. &lt;a href="https://terrytao.wordpress.com/2026/09/11/a-severe-misalignment-of-ai-in-mathematics/"target="_blank" rel="noopener noreferrer"&gt;https://terrytao.wordpress.com/2026/09/11/a-severe-misalignment-of-ai-in-mathematics/&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Declaration site (co-signing), &lt;a href="https://mathandai.org/"target="_blank" rel="noopener noreferrer"&gt;https://mathandai.org/&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Leiden Declaration, &lt;a href="https://leidendeclaration.ai"target="_blank" rel="noopener noreferrer"&gt;https://leidendeclaration.ai&lt;/a&gt; (DOI: 10.5281/zenodo.20302944, 2026-06-02)&lt;/li&gt;
&lt;li&gt;TechCrunch, &amp;ldquo;OpenAI&amp;rsquo;s feud with mathematicians is only escalating,&amp;rdquo; 2026-09-11. &lt;a href="https://techcrunch.com/2026/09/11/openais-feud-with-mathematicians-is-only-escalating"target="_blank" rel="noopener noreferrer"&gt;https://techcrunch.com/2026/09/11/openais-feud-with-mathematicians-is-only-escalating&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The Economist, &amp;ldquo;Top mathematicians are outraged by OpenAI&amp;rsquo;s methods,&amp;rdquo; 2026-09-11.&lt;/li&gt;
&lt;li&gt;36kr Chinese report (event timeline), &lt;a href="https://eu.36kr.com/en/p/3979724367985411"target="_blank" rel="noopener noreferrer"&gt;https://eu.36kr.com/en/p/3979724367985411&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Nature, &amp;ldquo;2026 Fields Medals,&amp;rdquo; &lt;a href="https://www.nature.com/articles/d41586-026-02169-1"target="_blank" rel="noopener noreferrer"&gt;https://www.nature.com/articles/d41586-026-02169-1&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;</description></item><item><title>The checkpoint is not the policy: a merged reading of PDM and RLT</title><link>http://fengwang.github.io/posts/pdm-rlt-merged-reading-en/</link><pubDate>Sun, 13 Sep 2026 00:00:00 +0000</pubDate><guid>http://fengwang.github.io/posts/pdm-rlt-merged-reading-en/</guid><description>&lt;p&gt;Why this exists: In September 2026 I ingested two project reports from the same author line into my vault, a diagnostics report on Prefill–Decode Kernel Mismatch (&amp;ldquo;PDM&amp;rdquo;) and an architecture report on a Recurrent Looped Transformer (&amp;ldquo;RLT&amp;rdquo;), and cross-linked them. Reading the second right after the first, the two stopped looking like separate papers. One says the mismatch between rollout and training computation is measurable and must be contractualized; the other says it can be designed out. I drafted a merged reading in Chinese for the ingest session; this is the English archive version, reorganized around three questions: what problem each report claims, what it does about it and what that delivers, and what survives a critical read.&lt;/p&gt;
&lt;p&gt;The thesis: PDM turns RL execution fidelity into something you can measure, record, and audit, and RLT turns it into something you can design, but neither closes the causal loop from mismatch to end-to-end training outcomes.&lt;/p&gt;
&lt;p&gt;Scope: This covers the two reports as of September 2026 (PDM August 6, revised August 24; RLT September), their problem statements, mechanisms, evidence, and weaknesses. It does not re-derive their math, does not independently reproduce any number, and treats both as non-peer-reviewed project reports. The cross-architecture hotspot hypothesis in the critique is mine, not theirs.&lt;/p&gt;
&lt;p&gt;Prerequisites: This assumes familiarity with autoregressive decoding and KV caches, on-policy RL basics (importance ratios, off-policy correction, support conditions), and a rough picture of state-space and linear-attention models.&lt;/p&gt;
&lt;h2 id="the-problem-pdm-names-the-checkpoint-is-not-the-policy"&gt;The problem PDM names: the checkpoint is not the policy&lt;/h2&gt;
&lt;p&gt;The assumption I carried for years is that a checkpoint defines a policy: load the weights, and every runtime that reads them computes the same distribution. PDM opens by breaking that assumption into pieces. RL pipelines generate rollouts with a sampler (parallel prefill builds a KV cache; decode steps extend it one token at a time, under quantization, custom kernels, and a sampling transform) and score the same tokens with a trainer (usually one teacher-forced parallel pass, another operator, another reduction order, sometimes another precision). Two engines with identical parameters end up defining two token distributions.&lt;/p&gt;
&lt;p&gt;Concretely, the report writes the sampled-token log ratio as a sum of two terms:&lt;/p&gt;
&lt;p&gt;$$\log \rho_t = \underbrace{\log \pi^{\mathrm{trainer}}&lt;em&gt;{\theta_n}(a_t \mid h_t) - \log \pi^{\mathrm{beh}}&lt;/em&gt;{\theta_n}(a_t \mid h_t)}&lt;em&gt;{\Delta&lt;/em&gt;{\mathrm{kernel}}} + \underbrace{\log \pi^{\mathrm{beh}}&lt;em&gt;{\theta_n}(a_t \mid h_t) - \log \pi^{\mathrm{beh}}&lt;/em&gt;{\theta_v}(a_t \mid h_t)}&lt;em&gt;{\Delta&lt;/em&gt;{\mathrm{stale}}}.$$&lt;/p&gt;
&lt;p&gt;The second term is the familiar staleness of asynchronous RL: it vanishes when the sampler and trainer are in sync. The first term does not. A fully synchronous pipeline can still be off-policy for the trainer&amp;rsquo;s own target. That sentence changed how I read every &amp;ldquo;we use a fast inference engine for rollouts&amp;rdquo; line.&lt;/p&gt;
&lt;p&gt;The emblematic bug deserves its own paragraph, because it is three lines away from any real repo. If you sample with the sampler operator and later reconstruct the &amp;ldquo;old&amp;rdquo; probability with the trainer operator, then at zero lag your recomputed ratio is exactly 1 by construction, while the correct ratio can still differ from 1. The recomputation does not fix a drift. It hides the gap, which is the report&amp;rsquo;s phrase, and it does so structurally rather than by accident.&lt;/p&gt;
&lt;p&gt;Then there are the cases where the gap is not a rounding detail. For softmax attention, the algebra of a KV-cached recurrence and a parallel causal pass is equivalent in exact arithmetic, but that identity is not a certificate: tiling, fusion, accumulation order, cache layout, batch shape, parallel communication, precision, quantization, and the post-logit transform all live between the two computations. For linear attention and state-space models the situation is worse: tokenwise, scan, and chunkwise implementations are different schedules, chunk size C runs from 1 (tokenwise) to L (one chunk), and changing the schedule changes floating-point parenthesization. One boundary error propagates into every later state. DeltaNet makes it personal: its state update is a data-dependent corrective write, so an error in the previous state participates in the next write rather than merely being observed downstream. Recurrent sampling against a parallelized delta-rule kernel is two policies until proven otherwise.&lt;/p&gt;
&lt;p&gt;What I find sharpest is the contract split. Distributional parity (same forward distribution on every history the sampler reaches) and score parity (the trainer differentiates the deployed policy&amp;rsquo;s score function, or an equivalent forward-and-backward implementation) are logically distinct. Equal forward distributions at one parameter value do not imply equal Jacobians. And support is part of the contract: if hard top-k, top-p, or a precision underflow assigns zero behavior probability to an action the target retains, no downstream likelihood ratio is defined, and sampled log-probabilities cannot repair that after the fact. I had filed support under &amp;ldquo;implementation detail&amp;rdquo;. It is not; it is a precondition for the estimator to exist.&lt;/p&gt;
&lt;h2 id="the-problem-rlt-names-one-model-four-execution-paths"&gt;The problem RLT names: one model, four execution paths&lt;/h2&gt;
&lt;p&gt;RLT&amp;rsquo;s problem statement reads like the same diagnosis one level up. A model in a production lifecycle passes through pretraining, supervised fine-tuning, rollout sampling, and replay, and each stage has historically used its own computation path, its own masking, its own caching. The report quotes the PDM line explicitly: matching forward probabilities at one parameter value is not sufficient to establish matching policy gradients, and a shared architectural transition alone does not prove numerical kernel parity. So the replay contract in RLT is not an ops checklist bolted onto an architecture; the architecture is chosen so that the checklist becomes cheap to satisfy.&lt;/p&gt;
&lt;p&gt;The second half of RLT&amp;rsquo;s problem is depth. In a standard transformer, the computation depth available to a token is the layer count, full stop. RLT wants reasoning depth that scales with the sequence, not with the stack: it applies a single history-dependent transition to every observed or sampled token, and the recurrent state carries continuous intermediate computation from token to token. After t tokens, the state path has traversed t times the decoder depth in logical block evaluations.&lt;/p&gt;
&lt;p&gt;The two halves connect. If one transition serves every token and every lifecycle stage, then the structural part of the sampler-trainer mismatch disappears by construction. There is no prompt path and response path to disagree; the serving split is just a boundary in the token history. I noticed here that RLT does not claim this makes things easier. It claims this makes things consistent, and consistency is the property its own replay contract needs.&lt;/p&gt;
&lt;h2 id="what-pdm-does-about-it"&gt;What PDM does about it&lt;/h2&gt;
&lt;p&gt;The report separates measurement from repair, and its machinery falls into five layers. I keep the table because I will want it the next time I set up an RL training stack.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;What it gives you&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Concepts&lt;/td&gt;
&lt;td&gt;The $\Delta_{\mathrm{kernel}} / \Delta_{\mathrm{stale}}$ decomposition, distributional vs score parity, the support condition&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Records&lt;/td&gt;
&lt;td&gt;Immutable sampler-emitted log-probabilities plus a per-token execution signature; a trainer may append diagnostics but never replace the stored probability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Audits&lt;/td&gt;
&lt;td&gt;A four-part fixed-weight parity check: state, forward-policy, score, and sampling parity, run independently of any optimization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Remedies&lt;/td&gt;
&lt;td&gt;Four alignment routes, described below&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operations&lt;/td&gt;
&lt;td&gt;Two declared modes: recurrent-target and declared parallel-target, each with guardrails&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The four remedies are where opinions will differ. The first is exposure: log the raw gap and its decomposition (on an audit subset, since separating the terms requires replaying the sampler operator at current weights). The second is to make one recurrence or chunk rule part of the model semantics and execute it everywhere, so there is nothing to reconcile; the report is careful that chunking does not create stochastic tokens in advance, so exact parallel generation still needs proposal and verification. The third is precision, with a ladder I now keep as a default checklist: FP32 baseline for recurrent states, decay products, and state-update reductions; FP64 as an audit reference rather than a training default; and no hope of recovering information by casting an already-rounded BF16 state up. The report captures the route&amp;rsquo;s status in a sentence I quote a lot: precision is a mitigation, not an on-policy certificate. The underlying error recursion is worth writing once, since it explains why precision alone cannot close the contract:&lt;/p&gt;
&lt;p&gt;$$|e_t| \le L_t,|e_{t-1}| + \delta_{\mathrm{round}} + \delta_{\mathrm{schedule}}.$$&lt;/p&gt;
&lt;p&gt;Here $L_t$ is a local state-sensitivity factor, and the two additive terms cover arithmetic rounding and a changed schedule (scan tree, chunk boundary, materialization rule). Higher precision shrinks $\delta_{\mathrm{round}}$. It does not shrink $\delta_{\mathrm{schedule}}$, and it does not stop amplification when products of the $L_t$ are large. That is the whole argument against treating FP32 as a proof.&lt;/p&gt;
&lt;p&gt;The fourth remedy is the one I did not expect. Take speculative decoding&amp;rsquo;s machinery and invert its roles: the recurrent implementation is the correctness authority, a cheap operator drafts, and modified rejection sampling (accept with probability $\min(1, p_i(y_i)/q_i(y_i))$, resample from the normalized positive part of $p_i - q_i$ on rejection) recovers the recurrent target up to the declared hardware numerics. The report says plainly that this helps only when drafting is cheap, acceptance is high, and verification batches well, and that it can provide no speedup when recurrent verification remains the serial bottleneck. It also draws a line I would have blurred: the proposal acceptance ratio is not the RL importance ratio. After exact verification, the committed tokens are distributed according to the recurrent target, so the record stores the target&amp;rsquo;s log-probability as the behavior log-probability.&lt;/p&gt;
&lt;h2 id="what-pdms-measurements-actually-show"&gt;What PDM&amp;rsquo;s measurements actually show&lt;/h2&gt;
&lt;p&gt;Four evidence blocks and one registration. I care more about the &amp;ldquo;does not establish&amp;rdquo; column than about the headline numbers.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Block&lt;/th&gt;
&lt;th&gt;Headline&lt;/th&gt;
&lt;th&gt;Establishes&lt;/th&gt;
&lt;th&gt;Does not establish&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Estimator check (vocab 4, horizon 6, exhaustive)&lt;/td&gt;
&lt;td&gt;Full-trajectory IS matches the exact target gradient to $2.33 \times 10^{-16}$; raw token-local stays $0.1551$ away in $L^2$&lt;/td&gt;
&lt;td&gt;The estimator algebra is airtight in a toy world; local IS is a valid estimator of a declared local surrogate only&lt;/td&gt;
&lt;td&gt;Anything about full-size models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Production-path audit (two H100 nodes, 8,192-token prompt)&lt;/td&gt;
&lt;td&gt;p95 of the absolute log-ratio: 0.09293 (dense), 0.04999 (FP16 cache), 0.04433 (FP32 cache); no support violations&lt;/td&gt;
&lt;td&gt;The gap is measurable and, in these configurations, small&lt;/td&gt;
&lt;td&gt;Causality; generality; the two models are not matched runs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mechanism test (128-route gated SSM)&lt;/td&gt;
&lt;td&gt;FP32 passes forward and score checks at about $6 \times 10^{-7}$; BF16 fails both, with up to 4.69% top-1 and 0.670% gradient-sign disagreement&lt;/td&gt;
&lt;td&gt;Precision is a real switch between passing and failing the checks; throughput rises from 13.0K to 136.0K tok/s across chunk sizes&lt;/td&gt;
&lt;td&gt;That these error levels matter at scale&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Full-distribution stress (16 histories)&lt;/td&gt;
&lt;td&gt;Top-1 agrees in 100% of cases; total variation is still roughly 350x the dense value for the hybrid model; FP32 lowers TV by only 4.11%&lt;/td&gt;
&lt;td&gt;Agreement at the argmax is not agreement in distribution&lt;/td&gt;
&lt;td&gt;Production importance of the tail&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Registered experiment&lt;/td&gt;
&lt;td&gt;DAPO-17k-Eng; 512 trajectories per step, 200 steps, four arms, three seeds, 8xH100 per arm&lt;/td&gt;
&lt;td&gt;Nothing yet; explicitly no results&lt;/td&gt;
&lt;td&gt;Everything causal&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The registered protocol is the part I will watch. It is a full pre-registration (data gate, models, arms, seeds, evaluation) with the honest footnote that no arm is accepted and the earlier pilot is excluded from claims. I expected a scaling paper to bury that; instead it sits near the top.&lt;/p&gt;
&lt;h2 id="what-rlt-does-about-it-and-what-it-costs"&gt;What RLT does about it, and what it costs&lt;/h2&gt;
&lt;p&gt;The mental model I keep from RLT is one transition, no reset. Every token, prompt or response, advances a single state $H_t = (s_t, C_t^D)$: a recurrent output plus layerwise sliding-window caches. The end of the prompt is a serving boundary in the bookkeeping, not a change in the conditional model. Gradients follow the same path: the reference training computation differentiates through the full history (full backpropagation through time with checkpointing), and the report is precise about where approximations leak in (detaching only the recurrent output still leaves gradient paths through the KV cache and encoder memory; SWA eviction is not a stop-gradient).&lt;/p&gt;
&lt;p&gt;Three properties matter for anyone trying to falsify or adopt this. First, serving-split invariance: under identical kernels and precision, where you split prompt from response does not change the conditional distribution. The report itself flags this as mathematical equivalence only, so numerical discrepancy is exactly the door through which PDM&amp;rsquo;s problem can walk back in. Second, the state depends only on the prefix, and the appendix warns that gradient products through $\partial s_t / \partial s_{t-1}$ alone miss paths through the cache; normalization does not bound products of the transition Jacobians. Third, cache exactness: a cached prefix can replace replay only when it matches the current parameters, positions, window semantics, and execution settings; a cache built under older weights satisfies neither the current-policy value nor the gradient requirement.&lt;/p&gt;
&lt;p&gt;The cost side is where the report is unusually forthright. The decoder path is sequential through all T transitions, so encoder parallelism does not make prefill fully parallel, and the report states it makes no claim of a reduced-prefill speedup. Cache storage grows with encoder depth plus a bounded decoder window. Training wants full BPTT. And the depth story carries its own asterisk: unbounded structural depth is not unbounded useful depth, because gates, contraction, and learned projections can suppress what long paths contribute. I find that asterisk more interesting than the headline. The design makes depth available; learning is a separate question.&lt;/p&gt;
&lt;p&gt;So what does RLT deliver, as of September 2026? A mechanism specification. No measured reasoning quality, no measured hardware efficiency, no scaling curves. Its claim is that if the discipline is followed, the fixed-weight parallel-recurrent mismatch has no place to live, because there is only one recurrence. That is a real design-level effect and a zero-evidence empirical effect.&lt;/p&gt;
&lt;h2 id="the-seam-between-the-two-reports"&gt;The seam between the two reports&lt;/h2&gt;
&lt;p&gt;Put PDM&amp;rsquo;s contract checklist next to RLT&amp;rsquo;s design, one row at a time. This is the table I check when someone tells me an architecture &amp;ldquo;solves&amp;rdquo; the mismatch.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;PDM module&lt;/th&gt;
&lt;th&gt;RLT&amp;rsquo;s response&lt;/th&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Canonical recurrent rule&lt;/td&gt;
&lt;td&gt;One transition for all tokens, shared by pretraining, SFT, rollout, replay&lt;/td&gt;
&lt;td&gt;Adopted at design level&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Score parity&lt;/td&gt;
&lt;td&gt;Explicitly inherited, citing PDM&lt;/td&gt;
&lt;td&gt;Adopted&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Support and behavior records&lt;/td&gt;
&lt;td&gt;Full record semantics; metadata cannot restore missing support&lt;/td&gt;
&lt;td&gt;Consistent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Precision ladder&lt;/td&gt;
&lt;td&gt;Only a note that mathematical equivalence is not numerical equality&lt;/td&gt;
&lt;td&gt;Gap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kernel-level audit&lt;/td&gt;
&lt;td&gt;Cache exactness only covers &amp;ldquo;is this state current-policy&amp;rdquo;&lt;/td&gt;
&lt;td&gt;Gap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Proposal plus recurrent verification&lt;/td&gt;
&lt;td&gt;Not addressed&lt;/td&gt;
&lt;td&gt;Gap&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The convergence is in the diagnosis, and it is close to verbatim: both reports treat forward-probability parity as insufficient and score parity as a separate contract. The divergence is in coverage. RLT borrows PDM&amp;rsquo;s logic but not PDM&amp;rsquo;s three remaining modules, which means that by PDM&amp;rsquo;s own standard, an RLT deployment is not automatically exact; it is exactly as auditable as any other pipeline, with a better starting structure. Structural mismatch (prompt path against response path) is designed away. Execution mismatch (the same RLT transition running on different kernels, precisions, and schedules) stays in PDM&amp;rsquo;s jurisdiction.&lt;/p&gt;
&lt;p&gt;I read that as complementarity rather than competition. RLT makes PDM&amp;rsquo;s recurrent-target mode implementable at scale, because &amp;ldquo;canonical recurrence&amp;rdquo; stops being an audit nightmare and becomes the model. PDM remains the reason you cannot call the result exact without measuring it. Each report is the other&amp;rsquo;s checklist.&lt;/p&gt;
&lt;h2 id="critical-assessment"&gt;Critical assessment&lt;/h2&gt;
&lt;h3 id="pdm-a-strong-evidence-chain-with-an-open-causal-loop"&gt;PDM: a strong evidence chain with an open causal loop&lt;/h3&gt;
&lt;p&gt;The measurement design is careful and the self-limiting is real: the report refuses to claim that gap magnitude implies capability loss, marks the cross-model numbers as descriptive, and labels the support stress as appendix-level evidence with a predeclared disposition. What I cannot yet see is the value proposition closed end-to-end. If production-path p95 log-ratios sit at 0.04 to 0.09 and even the full-distribution stress shows top-1 agreement across the board, the honest reading of the current data is that the gap is real, measurable, and modest in tested configurations. That makes this agenda insurance and auditability first, performance repair second, until the registered experiment says otherwise. I would also want a cost ledger: recurrent replay gives up sequence parallelism, the precision ladder costs compute, and proposal-verification can provide no speedup at all. Nobody has priced the discipline against the gap it removes.&lt;/p&gt;
&lt;p&gt;Two smaller critiques. First, the empirical surface is small: one 8,192-token prompt path, sixteen stress histories, and a 128-route toy SSM. These demonstrate measurability, not prevalence. Second, much of the toolbox is assembled from known parts (importance sampling, modified rejection, FP32 baselines); the new contribution is the execution-level framing and the systematic contractualization. That is a framework contribution, and expecting an algorithmic surprise from it would be a category error.&lt;/p&gt;
&lt;h3 id="rlt-design-discipline-with-zero-measurements"&gt;RLT: design discipline with zero measurements&lt;/h3&gt;
&lt;p&gt;The report is unusually honest for a design proposal, and that honesty is also its evaluation. It states no measured results, concedes that structural depth is not a reasoning guarantee, and concedes that neither weight tying nor temporal recurrence alone establishes novelty. What remains is a system-consistency argument: a lifecycle with one transition, replay semantics that follow from the architecture, and explicit interfaces for caches, multi-turn serving, and gradient approximations. The costs point the other way from its selling point: prefill loses parallelism, memory grows, and training wants full BPTT. Whether a quality-per-cost sweet spot exists is exactly the experiment that does not exist yet. And by PDM&amp;rsquo;s standards, RLT&amp;rsquo;s exact-replay claim is itself pending audit: no precision ladder, no kernel-level audit, and no treatment of stochastic computation beyond a deterministic-forward assumption.&lt;/p&gt;
&lt;h3 id="what-both-leave-open-together"&gt;What both leave open together&lt;/h3&gt;
&lt;p&gt;The full narrative chain reads: diagnose the execution mismatch (PDM), design it away (RLT), show the training outcome change. The first link is measured, the second is specified, the third is empty. And there is no economics anywhere in either report. Both add cost in the name of correctness, and neither answers at what training scale the discipline pays for itself. My own working hypothesis, flagged as inference rather than result, is that the payoff is highly architecture-dependent: I expect dense softmax configurations to stay in the modest regime, while chunkwise state-space and delta-rule models carry structural gaps where alignment discipline could be the difference between a stable run and a mysterious one. A systematic gap taxonomy (architecture family against kernel pair against precision) does not exist. That matrix is the most valuable missing artifact I can see in this area.&lt;/p&gt;
&lt;h2 id="what-i-take-from-reading-them-together"&gt;What I take from reading them together&lt;/h2&gt;
&lt;p&gt;Two hooks to keep. PDM: the checkpoint name is not the policy; the policy is the executable distribution, and an exact deployment gradient also includes its score function. RLT: one transition for the whole lifecycle, or replay fidelity becomes someone else&amp;rsquo;s audit problem.&lt;/p&gt;
&lt;p&gt;Three operational items, in order of how soon they apply to any RL stack I touch. One, audit for the recomputation bug: if old probabilities are rebuilt with the trainer engine, the kernel gap is being hidden, and the fix is to store sampler-emitted probabilities at sampling time. Two, adopt a minimal record contract before scaling: sampler log-probability, version, execution signature, sampling metadata, all immutable. Three, declare the target mode (recurrent-target or declared parallel-target) and do not blur it; the blur is where silent off-policyness lives.&lt;/p&gt;
&lt;p&gt;The watch point is the registered DAPO experiment. When it runs, it becomes the first causal data on whether any of this changes training outcomes, and which arm (naive, raw local IS, stabilized IS, FP32 mitigation) moves first. My prior is mild: I expect FP32 state handling to matter more than ratio corrections for hybrid models, and neither to be dramatic on dense ones. I would like to be wrong in a specific direction. If the stabilized arm separates from naive at 200 steps, the priority of this whole line changes from insurance to intervention.&lt;/p&gt;
&lt;h2 id="boundary-conditions"&gt;Boundary conditions&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Every number in this reading is as of September 2026, and PDM is already on its second revision (August 24). Treat figures as versioned, not fixed.&lt;/li&gt;
&lt;li&gt;The &amp;ldquo;modest gap&amp;rdquo; reading is bounded by the tested configurations: 4B-class models, one 8,192-token prompt path, sixteen stress histories. It does not extend to frontier scale or other architectures.&lt;/li&gt;
&lt;li&gt;The hotspot hypothesis (structural architectures carry larger gaps) is mine. It could be wrong in either direction, and the registered experiment will only test two Qwen models.&lt;/li&gt;
&lt;li&gt;RLT&amp;rsquo;s cost profile is argued, not measured. A working implementation could come out better or worse than the design reasoning suggests.&lt;/li&gt;
&lt;li&gt;Both reports share one unexamined premise: that execution fidelity, once measurable, deserves to be optimized for. If audit magnitude turns out uncorrelated with outcomes, the agenda shifts from engineering priority to bookkeeping.&lt;/li&gt;
&lt;li&gt;I read both as non-peer-reviewed project reports; the absence of external review is a boundary on everything above.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="open-questions"&gt;Open questions&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;When the DAPO registration runs, which arm separates from naive first: stabilized token-local IS, the FP32 recurrent-state mitigation, or none of them within 200 steps?&lt;/li&gt;
&lt;li&gt;What does the actual delta-kernel matrix look like across architecture families at fixed scale, and is there a cheap pre-training predictor for &amp;ldquo;this kernel pair will bite&amp;rdquo;?&lt;/li&gt;
&lt;li&gt;Can the four-part parity audit run as a CI gate inside a training stack, and what does one audit step cost relative to one optimizer step?&lt;/li&gt;
&lt;li&gt;Does RLT&amp;rsquo;s serving-split invariance survive a production decode stack, or does hardware entropy re-open PDM&amp;rsquo;s gap the moment the sampler is not the reference kernel?&lt;/li&gt;
&lt;li&gt;Is there a regime where full-BPTT recurrent replay is cheaper than the mismatch it removes, for example short-response verifiable-reward RL where outcome variance dominates?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;[[Q]] Six months from now: has the DAPO-17k-Eng registration produced accepted results, and did any kernel-mismatch arm beat the naive baseline in evaluation? If not, check whether the blocker was infrastructure, a support violation, or a genuinely null effect.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Yifan Zhang et al., &amp;ldquo;Reliable RL Scaling Requires Accounting for Prefill–Decode Kernel Mismatch&amp;rdquo;, Pretraining-RL-Science project report, August 6, 2026 (revised August 24, 2026). &lt;a href="https://github.com/yifanzhang-pro/Pretraining-RL-Science"target="_blank" rel="noopener noreferrer"&gt;https://github.com/yifanzhang-pro/Pretraining-RL-Science&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Yifan Zhang et al., &amp;ldquo;Recurrent Looped Transformer&amp;rdquo;, project report, September 2026. &lt;a href="https://github.com/yifanzhang-pro/recurrent-looped-tranformer"target="_blank" rel="noopener noreferrer"&gt;https://github.com/yifanzhang-pro/recurrent-looped-tranformer&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;</description></item><item><title>Qwen3-TTS on AMD Radeon 780M (gfx1103): A Docker + ROCm Deployment Autopsy</title><link>http://fengwang.github.io/posts/qwen3-tts-amd-780m-rocm-deployment-en/</link><pubDate>Mon, 03 Aug 2026 00:00:00 +0000</pubDate><guid>http://fengwang.github.io/posts/qwen3-tts-amd-780m-rocm-deployment-en/</guid><description>&lt;h1 id="qwen3-tts-on-amd-radeon-780m-gfx1103-a-docker--rocm-deployment-autopsy"&gt;Qwen3-TTS on AMD Radeon 780M (gfx1103): A Docker + ROCm Deployment Autopsy&lt;/h1&gt;
&lt;p&gt;&lt;strong&gt;Abstract.&lt;/strong&gt; This document is a complete post-mortem of deploying the
Qwen3-TTS-Openai-Fastapi service (OpenAI-compatible TTS API, Qwen3-TTS-12Hz-
1.7B-CustomVoice model, default voice Vivian) on an AMD Phoenix1 Radeon 780M
(gfx1103) APU using Docker with an ROCm container. It is written for the
author&amp;rsquo;s future self and for engineers who will deploy LLM / MoE inference on
the same hardware class. The central finding: the per-token &lt;code&gt;.item()&lt;/code&gt; GPU-to-
CPU synchronization inside the Hugging Face transformers &lt;code&gt;generate()&lt;/code&gt; loop
permanently deadlocks on this APU because HSA completion signals are
unreliable, while short texts survive only because they trigger few syncs.
Routing around &lt;code&gt;generate()&lt;/code&gt; with the project&amp;rsquo;s built-in fast codebook path
eliminates the syncs and stabilizes the service at RTF ~2.32. Every command,
error message, and observation in this document comes from the actual
2026-08-02 debugging session.&lt;/p&gt;
&lt;h2 id="1-background-running-tts-on-a-780m"&gt;1. Background: running TTS on a 780M&lt;/h2&gt;
&lt;p&gt;In early August 2026 I needed a self-hosted TTS service reachable over HTTP on
the local network, able to synthesize Chinese and English, running as a
persistent container whose configuration survives restarts. I chose
Qwen3-TTS-Openai-Fastapi with the Qwen3-TTS-12Hz-1.7B-CustomVoice model and
Vivian as the default voice.&lt;/p&gt;
&lt;p&gt;The hardware was an AMD Phoenix1 Radeon 780M (gfx1103), an APU integrated GPU
that shares system memory and has no dedicated VRAM. gfx1103 does not appear
in the official ROCm supported-device list. The host passes only
&lt;code&gt;/dev/kfd&lt;/code&gt; and &lt;code&gt;/dev/dri/renderD128&lt;/code&gt; into the container; all ROCm userspace
lives inside the image. That design decision was the root of most subsequent
trouble.&lt;/p&gt;
&lt;p&gt;Debugging ran from 10 pm into the early morning: four files modified, several
image rebuilds, and three distinct forms of GPU deadlock. This article records
the full chain in chronological order and closes with a distilled checklist
intended as a pitfall manual for future LLM / MoE deployments.&lt;/p&gt;
&lt;h2 id="2-environment-bring-up-image-dependencies-kernel-masquerade"&gt;2. Environment bring-up: image, dependencies, kernel masquerade&lt;/h2&gt;
&lt;h3 id="21-the-base-image-tag-never-existed"&gt;2.1 The base image tag never existed&lt;/h3&gt;
&lt;p&gt;The repository&amp;rsquo;s &lt;code&gt;Dockerfile.rocm&lt;/code&gt; referenced
&lt;code&gt;rocm6.3.1_ubuntu22.04_py3.12_pytorch_release_2.6.0&lt;/code&gt;. The build failed with
&amp;ldquo;not found&amp;rdquo; from Docker Hub. Checking the Hub&amp;rsquo;s tag listing confirmed that
this exact combination was never published.&lt;/p&gt;
&lt;p&gt;I switched to &lt;code&gt;rocm6.4.3_ubuntu24.04_py3.12_pytorch_release_2.6.0&lt;/code&gt;: identical
torch 2.6.0 and Python 3.12, differing only in the Ubuntu base layer. I
deliberately avoided ROCm 7.x because the community reports SIGSEGV on the
780M; the 6.x line is comparatively stable.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-dockerfile" data-lang="dockerfile"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;rocm/pytorch:rocm6.4.3_ubuntu24.04_py3.12_pytorch_release_2.6.0&lt;/span&gt;&lt;span class="err"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="22-pip-silently-replaced-torchaudio-with-the-cuda-build"&gt;2.2 pip silently replaced torchaudio with the CUDA build&lt;/h3&gt;
&lt;p&gt;The first container start failed with
&lt;code&gt;libcudart.so.13: cannot open shared object file&lt;/code&gt;. The cause was subtle: the
project&amp;rsquo;s pyproject.toml declares a bare &lt;code&gt;torchaudio&lt;/code&gt; dependency with no
version constraint, so pip resolved the then-latest CUDA build (torchaudio
2.11.0), overwriting the image&amp;rsquo;s ROCm torch stack and pulling in CUDA-only
libcudart.&lt;/p&gt;
&lt;p&gt;The fix was a forced reinstall of the matching ROCm build after
&lt;code&gt;pip install -e &amp;quot;.[api]&amp;quot;&lt;/code&gt;, from PyTorch&amp;rsquo;s official ROCm wheel index:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;pip install --no-cache-dir --no-deps &lt;span class="se"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nv"&gt;torchaudio&lt;/span&gt;&lt;span class="o"&gt;==&lt;/span&gt;2.6.0+rocm6.2.4 &lt;span class="se"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; --index-url https://download.pytorch.org/whl/rocm6.2.4
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Lesson: treat every bare dependency inside a ROCm image with suspicion. pip
defaults to the newest CUDA build; lock versions explicitly with a &lt;code&gt;+rocm&lt;/code&gt;
suffix.&lt;/p&gt;
&lt;h3 id="23-gfx1103-has-no-precompiled-kernels-hsa_override_gfx_version"&gt;2.3 gfx1103 has no precompiled kernels: HSA_OVERRIDE_GFX_VERSION&lt;/h3&gt;
&lt;p&gt;Model loading failed with &lt;code&gt;HIP error: invalid device function&lt;/code&gt; because
gfx1103 is absent from the ROCm 6.4.3 precompiled-kernel list (which covers
gfx908/90a/1030/1100/1101/942 and others). The standard remedy is the
&lt;code&gt;HSA_OVERRIDE_GFX_VERSION&lt;/code&gt; environment variable, which masquerades the chip
as a different model.&lt;/p&gt;
&lt;p&gt;Here is the counterintuitive trap: following my ollama deployment notes I
first set &lt;code&gt;11.0.2&lt;/code&gt; (the real version for gfx1103) and still got &amp;ldquo;invalid
device function&amp;rdquo;, because the image contains no kernels compiled for 11.0.2.
The value that works is:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;HSA_OVERRIDE_GFX_VERSION=11.0.0 # masquerade as gfx1100 (desktop RDNA3)
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;11.0.0 maps to gfx1100; the desktop RDNA3 kernels ship in the image and load
directly. Set this in compose &lt;code&gt;environment&lt;/code&gt; (it overrides the Dockerfile ENV).&lt;/p&gt;
&lt;h2 id="3-first-wave-of-gpu-hangs-flash-attn-aotriton-and-tunableop"&gt;3. First wave of GPU hangs: flash-attn, AOTriton, and tunableop&lt;/h2&gt;
&lt;p&gt;With the model loadable, every warmup now hit the same hard error:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;HW Exception by GPU node-1 reason: GPU Hang
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;The GPU side deadlocked completely and the driver required a reset. I
eliminated suspects one at a time.&lt;/p&gt;
&lt;h3 id="31-flash-attn-compiled-then-proved-useless"&gt;3.1 flash-attn compiled, then proved useless&lt;/h3&gt;
&lt;p&gt;The Dockerfile compiles flash-attn v2.8.3 from source. It compiled
successfully and the model even loaded in flash_attention_2 mode, but warmup
always hung. After switching &lt;code&gt;attention&lt;/code&gt; to &lt;code&gt;sdpa&lt;/code&gt; in the config, loading
worked. Conclusion: under the gfx1103 masquerade, the flash-attn kernel path
is unstable; sdpa is the only viable choice.&lt;/p&gt;
&lt;h3 id="32-torch_sdpa_enable_flash0-was-a-false-positive"&gt;3.2 TORCH_SDPA_ENABLE_FLASH=0 was a false positive&lt;/h3&gt;
&lt;p&gt;Logs repeatedly showed &amp;ldquo;Using AOTriton backend for Flash Attention forward&amp;rdquo;.
PyTorch 2.6 on ROCm routes SDPA through an AOTriton flash implementation by
default, so I tried disabling it with &lt;code&gt;TORCH_SDPA_ENABLE_FLASH=0&lt;/code&gt;. My test
script passed, which looked like success.&lt;/p&gt;
&lt;p&gt;It was a false positive. The test used float32 inputs, and flash attention
does not support float32, so it silently fell back to the efficient backend
regardless of the environment variable. With bf16 inputs the variable had no
effect and the flash path ran anyway. The correct way is to force-disable it
in Python:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;torch&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;backends&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cuda&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;enable_flash_sdp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;backends&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cuda&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;enable_mem_efficient_sdp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;backends&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cuda&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;enable_math_sdp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;To make this apply automatically at process start, I wrote a sitecustomize.py
into the image&amp;rsquo;s site-packages; Python imports it at startup, so no
application code needed modification.&lt;/p&gt;
&lt;h3 id="33-pytorch_tunableop_enabled1-was-the-hidden-bomb"&gt;3.3 PYTORCH_TUNABLEOP_ENABLED=1 was the hidden bomb&lt;/h3&gt;
&lt;p&gt;This was the first fix that produced visible progress. The image ships with
&lt;code&gt;PYTORCH_TUNABLEOP_ENABLED=1&lt;/code&gt;, which runs a kernel-tuning benchmark on first
call of every operator. On the APU that tuning process hangs the GPU. Setting
it to 0 let warmup 3/3 complete, the service came up for the first time, and
short-text synthesis worked.&lt;/p&gt;
&lt;p&gt;At this point short text (&amp;ldquo;你好世界。&amp;rdquo; / &amp;ldquo;Hello, world.&amp;rdquo;) returned reliably
in 2.6 seconds. I assumed debugging was over.&lt;/p&gt;
&lt;h2 id="4-the-main-culprit-per-step-item-sync-inside-transformers-generate"&gt;4. The main culprit: per-step .item() sync inside transformers generate()&lt;/h2&gt;
&lt;h3 id="41-symptom-short-text-fine-medium-text-always-hangs"&gt;4.1 Symptom: short text fine, medium text always hangs&lt;/h3&gt;
&lt;p&gt;The service answered short requests, but any request of moderate length
(about 30 characters) hung forever. curl timed out at 120 seconds, and the
container health went from healthy to starting (warmup wedged). The symptom
was highly reproducible: short text 100% success, medium text 100% hang.&lt;/p&gt;
&lt;p&gt;GPU monitoring showed an odd pattern:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;t=1s busy=74% # generation in progress
t=9s busy=96% # GPU under load
t=10s busy=0 # GPU suddenly completely idle
t=11s+ busy=0 # stays idle, but the request never returns
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;The GPU finished its work and went idle, while two CPU threads spun at 99.9%
each. This was not a GPU computation problem; the CPU was waiting for
something that never arrived.&lt;/p&gt;
&lt;h3 id="42-double-evidence-from-py-spy-and-gdb"&gt;4.2 Double evidence from py-spy and gdb&lt;/h3&gt;
&lt;p&gt;After granting the container SYS_PTRACE, py-spy showed the Python stack:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;_has_unfinished_sequences (transformers/generation/utils.py:2597)
_sample (transformers/generation/utils.py:2779)
generate (transformers/generation/utils.py:2564)
forward (qwen_tts/core/models/modeling_qwen3_tts.py:1891)
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Wedged inside the transformers generation loop&amp;rsquo;s termination check. gdb then
showed the C stack:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;_local_scalar_dense_cuda -&amp;gt; .item() GPU-to-CPU sync copy
-&amp;gt; libhsa-runtime64.so.1 signal wait
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;The chain was complete: &lt;code&gt;_has_unfinished_sequences&lt;/code&gt; calls &lt;code&gt;tensor.item()&lt;/code&gt;
after every generated token, copying the result from GPU to CPU to decide
whether to stop. &lt;code&gt;.item()&lt;/code&gt; is a D2H copy that depends on an HSA completion
signal. On the 780M that signal is unreliable: the kernel actually finishes,
but the signal is lost and the CPU blocks forever.&lt;/p&gt;
&lt;p&gt;Short text survived because it performs few syncs and happened to receive
every signal. Long text performs many syncs, and the probability of hitting a
lost signal approaches one. This is a probabilistic failure, not a length
threshold.&lt;/p&gt;
&lt;h3 id="43-one-change-that-made-things-worse-hsa_enable_interrupt0"&gt;4.3 One change that made things worse: HSA_ENABLE_INTERRUPT=0&lt;/h3&gt;
&lt;p&gt;At one point I set &lt;code&gt;HSA_ENABLE_INTERRUPT=0&lt;/code&gt; to make HSA poll instead of using
interrupts. That converted &amp;ldquo;waiting for a signal&amp;rdquo; into CPU spinning: the
symptom changed from GPU Hang to two threads at 99.9% CPU, which was harder
to diagnose. Looking back at the timeline, the first stable run (warmup
passed, short text worked) happened under the default interrupt mode, before
this variable was introduced. I removed it and went back to defaults.&lt;/p&gt;
&lt;h2 id="5-the-fix-non_streaming_mode-and-fast-codebook"&gt;5. The fix: non_streaming_mode and fast codebook&lt;/h2&gt;
&lt;p&gt;Once the &lt;code&gt;.item()&lt;/code&gt; sync was identified, the strategy became: bypass
transformers&amp;rsquo; &lt;code&gt;generate()&lt;/code&gt; loop so no per-token synchronization happens. The
project code happens to provide two routes.&lt;/p&gt;
&lt;h3 id="51-first-route-non_streaming_modetrue"&gt;5.1 First route: non_streaming_mode=True&lt;/h3&gt;
&lt;p&gt;&lt;code&gt;generate_custom_voice&lt;/code&gt; defaults to &lt;code&gt;non_streaming_mode=False&lt;/code&gt;, which
simulates streaming input and runs the per-token loop (a &lt;code&gt;.item()&lt;/code&gt; on every
step). Passing &lt;code&gt;True&lt;/code&gt; runs one-shot generation and cuts the sync count by an
order of magnitude. I patched the optimized backend&amp;rsquo;s call:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;wavs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;generate_custom_voice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;speaker&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;voice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;instruct&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;instruct&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;non_streaming_mode&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="kc"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;An isolated test confirmed the diagnosis: for identical short text, the
streaming-simulation mode deadlocked while non_streaming mode succeeded in
4.4 seconds. This proved the root cause but did not fully cure it: medium
text still hung intermittently through the API.&lt;/p&gt;
&lt;h3 id="52-second-route-fast-codebook"&gt;5.2 Second route: fast codebook&lt;/h3&gt;
&lt;p&gt;The real cure is &lt;code&gt;enable_streaming_optimizations(use_fast_codebook=True)&lt;/code&gt;. It
replaces the inner code_predictor&amp;rsquo;s Hugging Face &lt;code&gt;generate()&lt;/code&gt; with the
project&amp;rsquo;s own &lt;code&gt;generate_fast&lt;/code&gt;: pure tensor operations (topk / multinomial /
masked_fill, all on GPU), zero &lt;code&gt;.item()&lt;/code&gt; syncs, and a fixed iteration count
of num_codebooks with no CPU involvement in the stop decision.&lt;/p&gt;
&lt;p&gt;In the original code the fast codebook switch was gated behind
&lt;code&gt;use_compile&lt;/code&gt;, so &lt;code&gt;use_compile: false&lt;/code&gt; skipped the entire
&lt;code&gt;_apply_optimizations&lt;/code&gt; block. I changed the condition so either flag enables
the block, and made the hardcoded &lt;code&gt;use_compile=True&lt;/code&gt; read from config:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;opt&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;use_compile&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kc"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;opt&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;use_fast_codebook&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kc"&gt;True&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; \
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;device&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;cpu&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_apply_optimizations&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model_info&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;opt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Fast codebook is now enabled independently, while torch.compile stays off
(compilation is both slow and risky on the APU).&lt;/p&gt;
&lt;h3 id="53-supporting-stability-parameters"&gt;5.3 Supporting stability parameters&lt;/h3&gt;
&lt;p&gt;Final compose environment:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;HSA_OVERRIDE_GFX_VERSION=11.0.0
HSA_ENABLE_SDMA=0 # known iGPU hang source
PYTORCH_TUNABLEOP_ENABLED=0
TTS_AUTOCHUNK=false # chunked continuous generation also hung; one-pass now
TTS_LAZY_LOAD=false
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;&lt;code&gt;HSA_ENABLE_SDMA=0&lt;/code&gt; is a community-known iGPU stability switch.
&lt;code&gt;TTS_AUTOCHUNK=false&lt;/code&gt; came from observing that multi-chunk continuous
generation also triggered hangs; with it off, long text synthesizes in one
pass.&lt;/p&gt;
&lt;h2 id="6-verification-and-performance"&gt;6. Verification and performance&lt;/h2&gt;
&lt;p&gt;Benchmark after the fix (1.7B + sdpa + fast codebook, no torch.compile):&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;test&lt;/th&gt;
&lt;th&gt;wall time&lt;/th&gt;
&lt;th&gt;audio&lt;/th&gt;
&lt;th&gt;RTF&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;English short&lt;/td&gt;
&lt;td&gt;6.8s&lt;/td&gt;
&lt;td&gt;2.9s&lt;/td&gt;
&lt;td&gt;2.32&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;English medium&lt;/td&gt;
&lt;td&gt;27.4s&lt;/td&gt;
&lt;td&gt;11.9s&lt;/td&gt;
&lt;td&gt;2.32&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chinese medium&lt;/td&gt;
&lt;td&gt;20.0s&lt;/td&gt;
&lt;td&gt;8.6s&lt;/td&gt;
&lt;td&gt;2.32&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;overall p95&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;2.34&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;RTF 2.32 means 2.3 seconds of compute per second of audio. For a 780M
integrated GPU that is a reasonable level: batch synthesis is fully usable;
real-time conversation (RTF &amp;lt; 1) is not.&lt;/p&gt;
&lt;p&gt;The long-text validation was the real relief: synthesizing the full
&amp;ldquo;蜀道难&amp;rdquo; (Hard Is the Road to Shu, a 500+ character classical Chinese poem)
in one pass produced 119.8 seconds of audio (24 kHz / 16-bit WAV) in 5
minutes 38 seconds with no hangs and steady GPU load. At this length the
failure rate used to be 100%.&lt;/p&gt;
&lt;p&gt;One data trap surfaced along the way: the local directory labeled &amp;ldquo;0.6B&amp;rdquo;
contained a byte-for-byte copy of the 1.7B model (identical MD5, config
&lt;code&gt;tts_model_size: &amp;quot;1b7&amp;quot;&lt;/code&gt;). The real 0.6B (905M parameters) lived in another
legacy directory. Never trust a directory name when switching models.&lt;/p&gt;
&lt;h2 id="7-checklist-for-future-llm--moe-deployments"&gt;7. Checklist for future LLM / MoE deployments&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Verify the base image tag exists before writing the Dockerfile. The
Docker Hub tag listing is paginated and searchable; five minutes of
checking saves a failed build.&lt;/li&gt;
&lt;li&gt;Pin every torch-adjacent dependency inside ROCm images. Bare dependencies
get resolved by pip to the newest CUDA build. Use &lt;code&gt;+rocm&amp;lt;ver&amp;gt;&lt;/code&gt; suffixes
and the official ROCm wheel index uniformly.&lt;/li&gt;
&lt;li&gt;For gfx1103, set &lt;code&gt;HSA_OVERRIDE_GFX_VERSION=11.0.0&lt;/code&gt; (masquerade as
gfx1100). 11.0.2 / 11.0.3 are the true gfx1103 version numbers but no
kernels exist for them in the image, producing &amp;ldquo;invalid device function&amp;rdquo;.
Set it in compose environment so it overrides the Dockerfile ENV.&lt;/li&gt;
&lt;li&gt;Disable &lt;code&gt;PYTORCH_TUNABLEOP_ENABLED&lt;/code&gt; on APUs. The kernel-tuning benchmark
hangs the integrated GPU; it is the easiest hidden bomb to miss.&lt;/li&gt;
&lt;li&gt;Disable flash-attn and the AOTriton flash SDPA. flash-attn compiles but
hangs at runtime; for SDPA, force &lt;code&gt;enable_flash_sdp(False)&lt;/code&gt; via a
sitecustomize.py so it applies at Python startup. The
&lt;code&gt;TORCH_SDPA_ENABLE_FLASH=0&lt;/code&gt; environment variable is unreliable on ROCm,
and any test of it must use bf16 inputs or it is a false positive.&lt;/li&gt;
&lt;li&gt;Watch for per-token &lt;code&gt;.item()&lt;/code&gt; syncs inside transformers &lt;code&gt;generate()&lt;/code&gt;.
Any GPU-to-CPU synchronization in the generation loop is a time bomb on
the 780M. Prefer sync-free generation paths (here: fast codebook), or
move the termination check onto the GPU.&lt;/li&gt;
&lt;li&gt;Leave &lt;code&gt;HSA_ENABLE_INTERRUPT&lt;/code&gt; at its default. Setting it to 0 turns lost
signals into CPU spinning, which is much harder to diagnose.&lt;/li&gt;
&lt;li&gt;Diagnostic toolkit: py-spy for Python stacks, gdb for C stacks, and
&lt;code&gt;gpu_busy_percent&lt;/code&gt; for GPU load. Cross-referencing them quickly
distinguishes &amp;ldquo;GPU cannot compute&amp;rdquo; from &amp;ldquo;CPU waits forever&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;Test across a length gradient. Short-text success does not imply
long-text success; signal loss is probabilistic and scales with sync
count.&lt;/li&gt;
&lt;li&gt;Persistence trio: &lt;code&gt;restart: unless-stopped&lt;/code&gt;, bind-mounted model and
config, and bind-mounted cache directories (&lt;code&gt;~/.cache/miopen&lt;/code&gt;,
torch_extensions). Otherwise a container recreate forces kernel
re-search/recompilation and can reproduce the hang.&lt;/li&gt;
&lt;li&gt;Locking the GPU frequency to high improves stability:
&lt;code&gt;/sys/class/drm/card0/device/power_dpm_force_performance_level&lt;/code&gt; set to
&amp;ldquo;high&amp;rdquo; requires root and resets on reboot; persist it with a systemd
oneshot service.&lt;/li&gt;
&lt;li&gt;Verify model authenticity: cross-check the directory name, the config&amp;rsquo;s
model_size field, and file MD5s. Do not trust a copied directory.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="8-boundary-conditions"&gt;8. Boundary conditions&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;The findings hold for ROCm 6.4.3 + torch 2.6.0 + Ubuntu 24.04. With other
ROCm versions (especially 7.x), kernel coverage and AOTriton behavior
change, and the HSA override value may need retesting.&lt;/li&gt;
&lt;li&gt;Masquerading gfx1103 as gfx1100 is a compatibility shim; performance is
whatever it measures to be. Other APUs (gfx1101/1102, Strix Point) may
need different stable parameter sets.&lt;/li&gt;
&lt;li&gt;The fast codebook path bypasses the HF &lt;code&gt;generate()&lt;/code&gt; of this project&amp;rsquo;s
code_predictor only. LLM / MoE inference that must use transformers
&lt;code&gt;generate()&lt;/code&gt; (complex sampling, stopping criteria) cannot adopt this fix
directly; a sync-free path must be found separately.&lt;/li&gt;
&lt;li&gt;RTF 2.32 is &amp;ldquo;usable&amp;rdquo;, not &amp;ldquo;good&amp;rdquo;. Streaming endpoints, voice cloning, and
Base (non-CustomVoice) models were not validated in this deployment and
may still trigger hangs.&lt;/li&gt;
&lt;li&gt;Short-text success in early versions (2.6s) was luck: few syncs. It must
not be taken as evidence that a configuration works; always re-test with
medium or long text.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="9-open-questions"&gt;9. Open questions&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;What RTF does the real 0.6B model (905M parameters) achieve under the
same configuration? If it approaches 1.0, does real-time become viable?&lt;/li&gt;
&lt;li&gt;Is there a usable torch.compile configuration on the 780M (e.g.
reduce-overhead plus manual warmup)? If so, how far below RTF 2.32 can
we go?&lt;/li&gt;
&lt;li&gt;Has the ROCm 7.x SIGSEGV on gfx1103 been fixed? If yes, is migrating
worth it for better kernel coverage?&lt;/li&gt;
&lt;li&gt;Is the streaming endpoint (&lt;code&gt;/v1/audio/speech&lt;/code&gt; streaming mode) stable
under fast codebook? This decides voice-agent viability.&lt;/li&gt;
&lt;li&gt;For LLM / MoE inference (e.g. vLLM&amp;rsquo;s ROCm support), do the HSA override
and tunableop lessons transfer directly, or does vLLM have its own APU
parameter system?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;[[Q]] Six months from now: is there a fully sync-free generation path
(including a custom stopping criterion) that makes transformers
&lt;code&gt;generate()&lt;/code&gt; run stable long sequences on the 780M?&lt;/p&gt;
&lt;h2 id="10-references"&gt;10. References&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Qwen3-TTS-Openai-Fastapi repository (Dockerfile.rocm, config.yaml,
api/backends/optimized_backend.py), local copy, August 2026.&lt;/li&gt;
&lt;li&gt;Docker Hub rocm/pytorch image tag listing, queried August 2026.&lt;/li&gt;
&lt;li&gt;Qwen3-TTS-12Hz-1.7B-CustomVoice model config.json (local model
directory).&lt;/li&gt;
&lt;li&gt;PyTorch ROCm wheel index, download.pytorch.org/whl/rocm6.2.4, accessed
August 2026.&lt;/li&gt;
&lt;li&gt;AMD ROCm community discussions on gfx1103 and
HSA_OVERRIDE_GFX_VERSION (web search, August 2026).&lt;/li&gt;
&lt;/ol&gt;</description></item><item><title>Qwen3-TTS 在 AMD Radeon 780M 上的 Docker + ROCm 部署排障记录</title><link>http://fengwang.github.io/posts/qwen3-tts-amd-780m-rocm-deployment/</link><pubDate>Mon, 03 Aug 2026 00:00:00 +0000</pubDate><guid>http://fengwang.github.io/posts/qwen3-tts-amd-780m-rocm-deployment/</guid><description>&lt;h1 id="qwen3-tts-在-amd-radeon-780m-上的-docker--rocm-部署排障记录"&gt;Qwen3-TTS 在 AMD Radeon 780M 上的 Docker + ROCm 部署排障记录&lt;/h1&gt;
&lt;h2 id="背景在-780m-上跑-tts-的目标与约束"&gt;背景：在 780M 上跑 TTS 的目标与约束&lt;/h2&gt;
&lt;p&gt;2026 年 8 月初，我要在本地部署一个自托管的 TTS 服务，选型是
Qwen3-TTS-Openai-Fastapi（OpenAI 兼容接口）加 Qwen3-TTS-12Hz-1.7B-CustomVoice 模型，
默认音色 Vivian。目标很朴素：局域网内能调 HTTP API 合成中英文语音，容器持久运行，
重启不丢配置。&lt;/p&gt;
&lt;p&gt;硬件是一个我此前没在 PyTorch 场景认真用过的家伙：AMD Phoenix1 Radeon 780M
(gfx1103)，APU 集成显卡。它共享系统内存，没有独立显存，官方 ROCm 支持列表里
没有 gfx1103 这个名字。宿主机只透传 &lt;code&gt;/dev/kfd&lt;/code&gt; 和 &lt;code&gt;/dev/dri/renderD128&lt;/code&gt; 进容器，
ROCm 用户态全部住在镜像里，这是后来所有麻烦的根源。&lt;/p&gt;
&lt;p&gt;排障从晚上 10 点持续到凌晨，前后改了 4 个文件、重建镜像多次、经历了三种不同
形态的 GPU 卡死。这篇文章按时间顺序记录完整链路，末尾是提炼出的经验清单，
目标是给之后部署 LLM / MoE 模型时当避坑手册用。&lt;/p&gt;
&lt;h2 id="环境搭建镜像依赖与内核伪装"&gt;环境搭建：镜像、依赖与内核伪装&lt;/h2&gt;
&lt;h3 id="第一步就踩空镜像-tag-不存在"&gt;第一步就踩空：镜像 tag 不存在&lt;/h3&gt;
&lt;p&gt;仓库自带的 &lt;code&gt;Dockerfile.rocm&lt;/code&gt; 引用
&lt;code&gt;rocm6.3.1_ubuntu22.04_py3.12_pytorch_release_2.6.0&lt;/code&gt;，构建时 Docker Hub 直接
返回 not found。查了 Docker Hub 的 tag 列表，发现 6.3.1 这个组合从未发布过。&lt;/p&gt;
&lt;p&gt;改用 &lt;code&gt;rocm6.4.3_ubuntu24.04_py3.12_pytorch_release_2.6.0&lt;/code&gt;：torch 2.6.0 相同、
Python 3.12 相同，只差 Ubuntu 基础层。特意避开 ROCm 7.x，因为社区在 780M 上
有 SIGSEGV 报告，6.x 系列相对稳。&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-dockerfile" data-lang="dockerfile"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;rocm/pytorch:rocm6.4.3_ubuntu24.04_py3.12_pytorch_release_2.6.0&lt;/span&gt;&lt;span class="err"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="torchaudio-被-pip-覆盖成-cuda-版"&gt;torchaudio 被 pip 覆盖成 CUDA 版&lt;/h3&gt;
&lt;p&gt;第一次启动容器，日志报 &lt;code&gt;libcudart.so.13: cannot open shared object file&lt;/code&gt;。
根因很隐蔽：项目的 pyproject.toml 里裸依赖 &lt;code&gt;torchaudio&lt;/code&gt;（没有版本约束），
pip 解析时装到了当时最新的 CUDA 版 torchaudio 2.11.0，把镜像里 ROCm 的
torch 栈覆盖掉了，还引入了 CUDA 专属的 libcudart。&lt;/p&gt;
&lt;p&gt;修复是在 &lt;code&gt;pip install -e &amp;quot;.[api]&amp;quot;&lt;/code&gt; 之后强制重装 ROCm 配套版本，并从
PyTorch 官方 ROCm wheel 源拉取：&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;pip install --no-cache-dir --no-deps &lt;span class="se"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nv"&gt;torchaudio&lt;/span&gt;&lt;span class="o"&gt;==&lt;/span&gt;2.6.0+rocm6.2.4 &lt;span class="se"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; --index-url https://download.pytorch.org/whl/rocm6.2.4
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;经验：ROCm 镜像里任何裸依赖都要警惕，pip 默认解析到 CUDA 最新版，
必须显式锁定 &lt;code&gt;+rocm&lt;/code&gt; 后缀的版本。&lt;/p&gt;
&lt;h3 id="gfx1103-没有预编译-kernelhsa_override_gfx_version"&gt;gfx1103 没有预编译 kernel：HSA_OVERRIDE_GFX_VERSION&lt;/h3&gt;
&lt;p&gt;模型加载时报 &lt;code&gt;HIP error: invalid device function&lt;/code&gt;。原因是 gfx1103 不在
ROCm 6.4.3 预编译 kernel 列表里（列表覆盖 gfx908/90a/1030/1100/1101/942 等）。
解决办法是环境变量 &lt;code&gt;HSA_OVERRIDE_GFX_VERSION&lt;/code&gt; 把芯片伪装成其他型号。&lt;/p&gt;
&lt;p&gt;这里有个反直觉的坑：我最初按 ollama 部署经验设了 &lt;code&gt;11.0.2&lt;/code&gt;（gfx1103 的
真实版本号），结果反而报 invalid device function，因为镜像里根本没有为
11.0.2 编译的 kernel。最终生效值是：&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;HSA_OVERRIDE_GFX_VERSION=11.0.0 # 伪装成 gfx1100（桌面 RDNA3）
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;11.0.0 对应 gfx1100，桌面 RDNA3 kernel 在镜像里自带，加载后直接可用。
这个值必须写在 compose 的 environment 里（会覆盖 Dockerfile 的 ENV）。&lt;/p&gt;
&lt;h2 id="第一波-gpu-hangflash-attnaotriton-与-tunableop"&gt;第一波 GPU Hang：flash-attn、AOTriton 与 tunableop&lt;/h2&gt;
&lt;p&gt;模型能加载了，但每次 warmup 都触发同一个硬错误：&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;HW Exception by GPU node-1 reason: GPU Hang
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;这条错误意味着 GPU 侧彻底死锁，驱动需要重置。我按嫌疑顺序一个个排除：&lt;/p&gt;
&lt;h3 id="flash-attn-编译成功但没有意义"&gt;flash-attn 编译成功但没有意义&lt;/h3&gt;
&lt;p&gt;Dockerfile 里从源码编译了 flash-attn v2.8.3，居然编译成功，模型也确实以
flash_attention_2 模式加载了。但 warmup 必挂。把 config 里 &lt;code&gt;attention&lt;/code&gt; 改成
&lt;code&gt;sdpa&lt;/code&gt; 之后，加载正常了。结论：在 gfx1103 伪装环境下，flash-attn 的 kernel
路径不稳定，用 sdpa 是唯一选择。&lt;/p&gt;
&lt;h3 id="torch_sdpa_enable_flash0-是假阳性"&gt;TORCH_SDPA_ENABLE_FLASH=0 是假阳性&lt;/h3&gt;
&lt;p&gt;日志里反复出现 &amp;ldquo;Using AOTriton backend for Flash Attention forward&amp;rdquo;。
PyTorch 2.6 ROCm 的 SDPA 默认走 AOTriton flash 实现，我试着用环境变量
&lt;code&gt;TORCH_SDPA_ENABLE_FLASH=0&lt;/code&gt; 关掉它。测试脚本跑通了，看起来有效。&lt;/p&gt;
&lt;p&gt;但后来发现这是假阳性：测试用的是 float32 输入，flash attention 本来就不
支持 fp32，自动降级到 efficient backend，跟环境变量无关。换成 bf16 输入
后，环境变量完全失效，flash 路径照走。正确做法是在 Python 里强制关：&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;torch&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;backends&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cuda&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;enable_flash_sdp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;backends&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cuda&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;enable_mem_efficient_sdp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;backends&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cuda&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;enable_math_sdp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;为了让服务进程启动时自动执行，我在 Dockerfile 里写了一个 sitecustomize.py
放进 site-packages，Python 启动时自动加载，不需要改业务代码。&lt;/p&gt;
&lt;h3 id="pytorch_tunableop_enabled1-是隐藏炸弹"&gt;PYTORCH_TUNABLEOP_ENABLED=1 是隐藏炸弹&lt;/h3&gt;
&lt;p&gt;这是第一个真正让我看到转机的修复。镜像自带的
&lt;code&gt;PYTORCH_TUNABLEOP_ENABLED=1&lt;/code&gt; 会在首次调用每个算子时做 kernel 调优基准，
在 APU 上这个调优过程会挂 GPU。设成 0 之后，warmup 3/3 全部通过，服务
第一次真正起来了，短文本合成成功。&lt;/p&gt;
&lt;p&gt;到这里，短文本（&amp;ldquo;你好世界。&amp;quot;）稳定 2.6 秒出结果。我一度以为排障结束了。&lt;/p&gt;
&lt;h2 id="主犯transformers-generate-的-item-同步"&gt;主犯：transformers generate() 的 .item() 同步&lt;/h2&gt;
&lt;h3 id="症状短文本好长文本必挂"&gt;症状：短文本好，长文本必挂&lt;/h3&gt;
&lt;p&gt;服务能响应短文本，但任意中等长度文本（约 30 个字符）请求就会永久挂起。
curl 等 120 秒超时，容器 health 从 healthy 变 starting（warmup 挂住）。
这个症状高度稳定：短文本 100% 成功，中文本 100% 挂。&lt;/p&gt;
&lt;p&gt;监控 GPU 时发现一个反常现象：&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;t=1s busy=74% # 生成进行中
t=9s busy=96% # GPU 高负载
t=10s busy=0 # GPU 突然完全空闲
t=11s+ busy=0 # 一直空闲，但请求永不返回
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;GPU 干完活就停了，但 CPU 侧有两个线程各占 99.9%，在空转。这不是 GPU 算不动，
是 CPU 在等一个永远不来的东西。&lt;/p&gt;
&lt;h3 id="py-spy-和-gdb-的双重证据"&gt;py-spy 和 gdb 的双重证据&lt;/h3&gt;
&lt;p&gt;给容器加了 SYS_PTRACE 权限后，用 py-spy 抓 Python 栈：&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;_has_unfinished_sequences (transformers/generation/utils.py:2597)
_sample (transformers/generation/utils.py:2779)
generate (transformers/generation/utils.py:2564)
forward (qwen_tts/core/models/modeling_qwen3_tts.py:1891)
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;卡在 transformers 生成循环的结束条件判断里。再用 gdb 抓 C 栈：&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;_local_scalar_dense_cuda -&amp;gt; .item() GPU→CPU 同步拷贝
-&amp;gt; libhsa-runtime64.so.1 信号等待
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;链条完整了：transformers 的 &lt;code&gt;_has_unfinished_sequences&lt;/code&gt; 每生成一个 token
就调用 &lt;code&gt;tensor.item()&lt;/code&gt;，把 GPU 上的结果同步回 CPU 判断是否该停。&lt;code&gt;.item()&lt;/code&gt;
本质是 GPU→CPU 的 D2H 拷贝，依赖 HSA 信号通知完成。在 780M 上这个信号
不可靠，kernel 实际跑完了，但信号丢失，CPU 永久阻塞在等待里。&lt;/p&gt;
&lt;p&gt;短文本为什么能过？因为同步次数少，碰巧每次都拿到信号；长文本同步次数
多了，撞上信号丢失的概率接近 1。这是概率性问题，不是长度阈值问题。&lt;/p&gt;
&lt;h3 id="一个让事情更糟的尝试hsa_enable_interrupt0"&gt;一个让事情更糟的尝试：HSA_ENABLE_INTERRUPT=0&lt;/h3&gt;
&lt;p&gt;我一度设置 &lt;code&gt;HSA_ENABLE_INTERRUPT=0&lt;/code&gt; 让 HSA 用轮询代替中断，结果把
&amp;ldquo;等信号&amp;quot;变成了 CPU 自旋，症状从 GPU Hang 变成双线程 99.9% CPU 空转，
更加难排查。回看时间线，服务第一次稳定跑通时（warmup 全过、短文本成功）
恰恰是在没设置这个变量的默认中断模式下。撤掉它，恢复默认。&lt;/p&gt;
&lt;h2 id="修复组合non_streaming_mode-与-fast-codebook"&gt;修复组合：non_streaming_mode 与 fast codebook&lt;/h2&gt;
&lt;p&gt;定位到 &lt;code&gt;.item()&lt;/code&gt; 同步之后，思路变成：绕过 transformers 的 generate() 循环，
不逐 token 同步。项目代码里恰好有两条路。&lt;/p&gt;
&lt;h3 id="第一条路non_streaming_modetrue"&gt;第一条路：non_streaming_mode=True&lt;/h3&gt;
&lt;p&gt;&lt;code&gt;generate_custom_voice&lt;/code&gt; 默认 &lt;code&gt;non_streaming_mode=False&lt;/code&gt;，这个模式会模拟
流式输入，走逐 token 的生成循环（每步 &lt;code&gt;.item()&lt;/code&gt;）。传 &lt;code&gt;True&lt;/code&gt; 走一次性
生成，同步次数少一个量级。我改了 optimized_backend 的调用：&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;wavs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;generate_custom_voice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;speaker&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;voice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;instruct&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;instruct&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;non_streaming_mode&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="kc"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;隔离测试确认：同样的短文本，流式模拟模式卡死，non_streaming 模式 4.4 秒
成功。这证明根因判断正确，但还没根治：中等文本在 API 里仍偶发卡住。&lt;/p&gt;
&lt;h3 id="第二条路fast-codebook"&gt;第二条路：fast codebook&lt;/h3&gt;
&lt;p&gt;真正的解药是 &lt;code&gt;enable_streaming_optimizations(use_fast_codebook=True)&lt;/code&gt;。
它把内层 code_predictor 的 HF generate() 换成项目自带的 &lt;code&gt;generate_fast&lt;/code&gt;：
纯 tensor 操作（topk / multinomial / masked_fill 全在 GPU 上），零 &lt;code&gt;.item()&lt;/code&gt;
同步，固定迭代 num_codebooks 次，不需要任何 CPU 参与判断。&lt;/p&gt;
&lt;p&gt;原代码里 fast codebook 的开关被挂在 &lt;code&gt;use_compile&lt;/code&gt; 条件后面：&lt;code&gt;use_compile: false&lt;/code&gt;
时整个 &lt;code&gt;_apply_optimizations&lt;/code&gt; 被跳过。我把条件改成两者任一为真就执行，
并把硬编码的 &lt;code&gt;use_compile=True&lt;/code&gt; 改成从 config 读取：&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;opt&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;use_compile&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kc"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;opt&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;use_fast_codebook&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kc"&gt;True&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; \
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;device&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;cpu&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_apply_optimizations&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model_info&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;opt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;这样 fast codebook 独立启用，torch.compile 保持关闭（APU 上编译又慢又危险）。&lt;/p&gt;
&lt;h3 id="配套稳定化参数"&gt;配套稳定化参数&lt;/h3&gt;
&lt;p&gt;最终 compose 环境变量组合：&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;HSA_OVERRIDE_GFX_VERSION=11.0.0
HSA_ENABLE_SDMA=0 # 已知 iGPU hang 源
PYTORCH_TUNABLEOP_ENABLED=0
TTS_AUTOCHUNK=false # 分块连续生成也会挂，一次性合成
TTS_LAZY_LOAD=false
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;&lt;code&gt;HSA_ENABLE_SDMA=0&lt;/code&gt; 是社区已知的 iGPU 稳定性开关，&lt;code&gt;TTS_AUTOCHUNK=false&lt;/code&gt;
是因为实测自动分块的多段连续生成也会触发 hang，关掉后长文本一次合成。&lt;/p&gt;
&lt;h2 id="性能与稳定性验证"&gt;性能与稳定性验证&lt;/h2&gt;
&lt;p&gt;修复后的 benchmark（1.7B + sdpa + fast codebook，无 torch.compile）：&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;测试&lt;/th&gt;
&lt;th&gt;生成耗时&lt;/th&gt;
&lt;th&gt;音频时长&lt;/th&gt;
&lt;th&gt;RTF&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;英文短句&lt;/td&gt;
&lt;td&gt;6.8s&lt;/td&gt;
&lt;td&gt;2.9s&lt;/td&gt;
&lt;td&gt;2.32&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;英文中长句&lt;/td&gt;
&lt;td&gt;27.4s&lt;/td&gt;
&lt;td&gt;11.9s&lt;/td&gt;
&lt;td&gt;2.32&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;中文中长句&lt;/td&gt;
&lt;td&gt;20.0s&lt;/td&gt;
&lt;td&gt;8.6s&lt;/td&gt;
&lt;td&gt;2.32&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;整体 p95&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;2.34&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;RTF 2.32 意味着生成 1 秒音频要花 2.3 秒。对 780M 集成 GPU 这是合理水平：
批量合成完全可用，实时对话（需要 RTF &amp;lt; 1）不行。&lt;/p&gt;
&lt;p&gt;真正让人松一口气的是长文本验证：合成《蜀道难》全文 500+ 字，一次性
生成 119.8 秒音频（24kHz/16bit WAV），耗时 5 分 38 秒，全程无 hang，
GPU 稳定满载。之前这种长度是 100% 必挂的。&lt;/p&gt;
&lt;p&gt;顺带发现一个数据陷阱：本地标着 &amp;ldquo;0.6B&amp;rdquo; 的模型目录其实是 1.7B 的完整拷贝
（两个文件 MD5 完全相同，config 里 &lt;code&gt;tts_model_size: &amp;quot;1b7&amp;quot;&lt;/code&gt;）。真 0.6B
（905M 参数）在另一处旧部署目录里。换模型时不能只看目录名。&lt;/p&gt;
&lt;h2 id="对未来-llm--moe-部署的经验清单"&gt;对未来 LLM / MoE 部署的经验清单&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;先验证镜像 tag 存在&lt;/strong&gt;，再写 Dockerfile。Docker Hub 的 tag 列表可以
直接翻页查询，5 分钟能省掉一次构建失败。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;ROCm 镜像里锁死所有 torch 相关依赖版本&lt;/strong&gt;，裸依赖必被 pip 解析成
CUDA 最新版。统一用 &lt;code&gt;+rocm&amp;lt;ver&amp;gt;&lt;/code&gt; 后缀 + 官方 ROCm wheel 源。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;gfx1103 一律 &lt;code&gt;HSA_OVERRIDE_GFX_VERSION=11.0.0&lt;/code&gt;&lt;/strong&gt;（伪装 gfx1100）。
11.0.2/11.0.3 是 gfx1103 真实版本号，但镜像没有对应 kernel，会报
invalid device function。这个值在 compose 里设置，覆盖 Dockerfile ENV。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;在 APU 上禁用 PYTORCH_TUNABLEOP_ENABLED&lt;/strong&gt;。kernel 调优基准会挂
集成 GPU，这是最容易忽略的隐藏炸弹。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;禁用 flash-attn 与 AOTriton flash SDPA&lt;/strong&gt;。flash-attn 编译能成功但
运行必挂；SDPA 用 sitecustomize.py 强制 &lt;code&gt;enable_flash_sdp(False)&lt;/code&gt;，
Python 启动即生效。环境变量 &lt;code&gt;TORCH_SDPA_ENABLE_FLASH=0&lt;/code&gt; 在 ROCm 上
不可靠，且测试要用 bf16 输入否则是假阳性。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;警惕 transformers generate() 的逐 token &lt;code&gt;.item()&lt;/code&gt; 同步&lt;/strong&gt;。任何在
生成循环里做 GPU→CPU 同步的代码，在 780M 上都是定时炸弹。优先选
无同步的生成路径（本项目是 fast codebook），或把结束判断放到 GPU 侧。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;HSA_ENABLE_INTERRUPT 保持默认&lt;/strong&gt;，不要设 0。轮询模式把信号丢失变成
CPU 自旋，症状更难判断。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;诊断三件套&lt;/strong&gt;：py-spy 抓 Python 栈、gdb 抓 C 栈、&lt;code&gt;gpu_busy_percent&lt;/code&gt;
盯 GPU 负载。三者交叉能快速区分&amp;quot;GPU 算不动&amp;quot;和&amp;quot;CPU 等不到&amp;rdquo;。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;测试用例要覆盖长度梯度&lt;/strong&gt;：短文本成功不代表长文本能过，信号丢失是
概率问题，同步次数越多越容易触发。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;持久化三件套&lt;/strong&gt;：&lt;code&gt;restart: unless-stopped&lt;/code&gt;、模型与配置 bind mount、
缓存目录（~/.cache/miopen、torch_extensions）bind mount，否则容器
recreate 后 kernel 重新搜索编译，可能重现 hang。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GPU 频率锁 high 能提升稳定性&lt;/strong&gt;：&lt;code&gt;/sys/class/drm/card0/device/ power_dpm_force_performance_level&lt;/code&gt; 写 high 需 root 且重启失效，
用 systemd oneshot 服务持久化。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;验证模型真实性&lt;/strong&gt;：目录名、config 的 model_size 字段、文件 MD5 三者
对照，别被复制品误导。&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="boundary-conditions"&gt;Boundary conditions&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;结论基于 ROCm 6.4.3 + torch 2.6.0 + Ubuntu 24.04 这个组合。换 ROCm
版本（尤其 7.x）后，kernel 覆盖和 AOTriton 行为都会变，HSA override
值可能需要重测。&lt;/li&gt;
&lt;li&gt;gfx1103 伪装 gfx1100 只是兼容性手段，性能以实测为准；不同 APU
（如 gfx1101/1102 或 Strix Point）可能有不同的稳定参数组合。&lt;/li&gt;
&lt;li&gt;fast codebook 绕过的是本项目 code_predictor 的 HF generate()。LLM /
MoE 推理如果必须用 transformers generate()（如采样逻辑复杂、需要
stopping criteria），这个方案不能直接平移，需要另找无同步路径。&lt;/li&gt;
&lt;li&gt;RTF 2.32 是&amp;quot;能用&amp;quot;不是&amp;quot;好用&amp;rdquo;。流式端点、voice clone、Base 模型
（非 CustomVoice）在本部署中未验证，可能仍会触发 hang。&lt;/li&gt;
&lt;li&gt;短文本 2.6s 的成功在早期版本里是侥幸（同步次数少），不能作为
&amp;ldquo;某个配置有效&amp;quot;的证据，必须以中长文本复测为准。&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="open-questions"&gt;Open questions&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;0.6B 真模型（905M 参数）在相同配置下 RTF 能到多少？如果逼近 1.0，
实时场景是否可行？&lt;/li&gt;
&lt;li&gt;torch.compile 在 780M 上是否有可用配置（例如 reduce-overhead +
手动 warmup）？如果能用，RTF 2.32 还有多少下降空间？&lt;/li&gt;
&lt;li&gt;ROCm 7.x 在 gfx1103 上的 SIGSEGV 是否已被修复？如果修复，是否值得
迁移以获得更好的 kernel 覆盖？&lt;/li&gt;
&lt;li&gt;流式输出（/v1/audio/speech 的 streaming 模式）在 fast codebook 下
是否稳定？这决定 voice agent 场景的可行性。&lt;/li&gt;
&lt;li&gt;对 LLM / MoE 推理（如 vLLM 的 ROCm 支持），HSA override 和 tunableop
的经验能否直接复用，还是 vLLM 有自己的 APU 参数体系？&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;[[Q]] 半年后重读：780M 上有没有一个完全无 .item() 同步的生成路径（包括
自定义 stopping criteria），能让 transformers generate() 在 APU 上稳定
跑长序列？&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Qwen3-TTS-Openai-Fastapi 项目仓库（Dockerfile.rocm / config.yaml /
api/backends/optimized_backend.py），2026-08 本地副本。&lt;/li&gt;
&lt;li&gt;Docker Hub rocm/pytorch 镜像 tag 列表，2026-08 查询。&lt;/li&gt;
&lt;li&gt;Qwen3-TTS-12Hz-1.7B-CustomVoice 模型 config.json（本地模型目录）。&lt;/li&gt;
&lt;li&gt;PyTorch ROCm wheel 索引 download.pytorch.org/whl/rocm6.2.4，2026-08 访问。&lt;/li&gt;
&lt;li&gt;AMD ROCm 社区对 gfx1103 与 HSA_OVERRIDE_GFX_VERSION 的讨论（网络搜索，
2026-08）。&lt;/li&gt;
&lt;/ol&gt;</description></item><item><title>AI 怎么想问题：从数学推理漫游中读到的认知画像</title><link>http://fengwang.github.io/posts/ai-mathematical-reasoning-walkthroughs/</link><pubDate>Sun, 02 Aug 2026 00:00:00 +0000</pubDate><guid>http://fengwang.github.io/posts/ai-mathematical-reasoning-walkthroughs/</guid><description>&lt;h2 id="引言一个模型在解释另一个模型的想法"&gt;引言：一个模型在解释另一个模型的想法&lt;/h2&gt;
&lt;p&gt;我花了一个晚上读 OpenAI 发布的数学推理漫游系列（&amp;ldquo;How the Ideas Came Together&amp;rdquo;），十二个章节，覆盖非 sofic 群构造、Connes 刚性反例、量子并行重复定理、多重色 Ramsey 数、Erdős–Simonovits 紧致性反例等十个数学问题领域。每一章不是论文复述，而是把&amp;quot;想法是怎么拼出来的&amp;quot;讲成一个故事：哪些想法先指了路，哪些大路撞了墙，什么视角转换揭开了结构。&lt;/p&gt;
&lt;p&gt;打动我的不是那些定理本身，而是这个元设：写这些笔记的&amp;quot;作者&amp;quot;是一个 AI 模型，它读了原始 chain-of-thought 和最终论文，然后回头重建推理过程的演化史。&lt;/p&gt;
&lt;p&gt;这篇笔记记录我从这套漫游中读到的 AI 认知画像，以及我对一个问题的持续困惑：当 AI 用第一人称叙述自己的推理时，那是理解的雏形，还是一种精致的 confabulation？&lt;/p&gt;
&lt;p&gt;Scope：本文覆盖漫游系列中反复出现的认知模式，不覆盖任何定理的完整证明，也不评价 OpenAI 与人类数学家的相对能力。&lt;/p&gt;
&lt;p&gt;Prerequisites：假设你了解 chain-of-thought 的基本概念，且对数学证明的形态有直觉。群论、凸几何、极值图论的细节我会在用到时给出足够上下文。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="叙事装置三层认知结构"&gt;叙事装置：三层认知结构&lt;/h2&gt;
&lt;p&gt;漫游系列的构造本身就是一个三层结构：&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;第一层：AI 解题时留下的原始思维流（chain-of-thought），混沌、有噪声、充满死胡同。&lt;/li&gt;
&lt;li&gt;第二层：AI 事后对思维流的叙事化重述，把曲折路径整理成&amp;quot;哪条路先想到、哪条路撞墙、什么转换揭开结构&amp;quot;的因果故事。&lt;/li&gt;
&lt;li&gt;第三层：我们人类读第二层。&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;三层之间不是忠实转录。第二层是重建，不是回放。它预设了一种认知观：想法不是灵光一闪，而是&amp;quot;路径 + 障碍 + 换视角&amp;quot;的演化。这个预设本身就是一种立场，它把数学发现描述成可叙述的过程，而不是不可言说的顿悟。&lt;/p&gt;
&lt;p&gt;每一章遵循同一骨架：哪些想法先指了路、哪些大路撞了墙、什么视角转换揭开了结构、决定性洞见怎么长成最终论证。骨架如此一致，以至于我怀疑它部分地是在为读者塑造一种&amp;quot;理想的推理形态&amp;quot;，但这恰恰是它最有趣的地方：即便有粉饰，被选进叙事的事件仍然透露了底层搜索的真实形状。&lt;/p&gt;
&lt;p&gt;贯穿十二个章节，我辨认出五个反复出现的认知模式，下面逐一展开。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="失败是主要的认知工具"&gt;失败是主要的认知工具&lt;/h2&gt;
&lt;p&gt;最常见的假设是失败是浪费。这套漫游给出的证据相反：几乎每个突破都由结构性失败铺路，而且失败的方式高度一致：它排除一整类机制，把搜索逼向正确的方向。&lt;/p&gt;
&lt;p&gt;Ehrhart 不等式章节（ch09）是典型。harmonic symmetrization 路线走了很久：它能精确命中极值对象 $T_n$：普通交集在二维就丢体积，对称化后体积恰好是 $2^n$，却解释不了缺失的 $n!$。结论写得极其诚实：&amp;ldquo;它识别了极值对象，却没有解释它的阶乘。&amp;ldquo;这条死路没有白走：它排除了&amp;quot;凸几何对称化&amp;quot;整类机制，把搜索逼向真正 n 维的计数机制，最终落在 jet 计数上。&lt;/p&gt;
&lt;p&gt;量子并行重复章节（ch07）把失败用得更系统。连续四类约减（投影游戏张量化、anchoring、oracularization、顺序熵累积）全部失败，且失败原因同构：都改变了物理策略类，而不是量化损失。原文直接说这些重复失败把搜索从普通信息界推向纯化规范。失败在这里不是噪音，是搜索方向修正器。&lt;/p&gt;
&lt;p&gt;非 sofic 群章节（ch03）的绕路更有名：想把 property (T) 直接转成 mixing，死于二部图谱可以接近 $-1$ 而 Kazhdan 间隙在 $1$。核心错配被明确命名：&amp;ldquo;Kun 给出许多 expander，Kun–Thom 只需要一个。&amp;rdquo;&lt;/p&gt;
&lt;p&gt;我原先以为 AI 的失败记录是事后修饰出来的叙事点缀。读完全部章节后我改了这个判断：失败的出现频率、具体程度和后续引用（&amp;ldquo;正是这次失败告诉我们 X 行不通&amp;rdquo;）太一致了，更像是真实搜索过程的压缩回放。失败不是装饰，是认知工具。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="显式反例杀死自己的猜测"&gt;显式反例杀死自己的猜测&lt;/h2&gt;
&lt;p&gt;漫游中最让我意外的一类行为：AI 会构造小反例来证伪自己的中间猜测，而不是绕过它。&lt;/p&gt;
&lt;p&gt;最漂亮的是永久型下界章节（ch02）里的 Krawtchouk 事件。第一个二元递推基于一个 associated Krawtchouk 猜测。AI 自己构造了 $n = 8$、原始度 $k = 1$ 的反例：猜测值 $508/7 &amp;lt; 128$，而真实码有 128 个码字。猜测被自己的反例直接证伪。&lt;/p&gt;
&lt;p&gt;这不是试错，这是数学家的习惯：对中间猜测做最小实例的假说检验。而且修正不是打补丁，是结构性的：正确项把整个分子移出平方根。反例指出的不是参数偏差，而是公式形态的错误。&lt;/p&gt;
&lt;p&gt;同样的事件出现在球堆积章节（ch01）：全局范数比较路线被显式反例堵死：一个函数可以有大的全局范数比，却不把负质量放进禁区半径。障碍不是未优化的常数，而是信息层面的缺陷：全局范数忘了负质量在哪。这类反例的价值在于它精确定位了论证中&amp;quot;哪条信息被丢弃了&amp;rdquo;，而这正是换坐标系（下一节）的动机。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="视角转换换坐标系是核心武器"&gt;视角转换：换坐标系是核心武器&lt;/h2&gt;
&lt;p&gt;每一个突破都伴随一个&amp;quot;换坐标系&amp;quot;的动作。漫游里这个模式如此频繁，我已经把它当作 AI 推理的主干操作。&lt;/p&gt;
&lt;p&gt;球堆积章节：径向 Fourier 变换变成 Mellin 反射加显式相位，谐和测度在带的边缘识别出精确半径 $1/\pi$。高斯提供正确的 Fourier 对称性但错误的鞍点位置，于是变形加 remote shell 修复。&lt;/p&gt;
&lt;p&gt;Connes 刚性章节（ch04）的决定性转变是一个问题：&amp;ldquo;crossed product 到底记得什么？&amp;ldquo;答案：它记得概率空间、Haar 测度和可测作用，但不记得对偶上的紧群律。于是目标从&amp;quot;找两个同构 K-模&amp;quot;变成&amp;quot;同一个 K 概率空间上的两种群律&amp;rdquo;。可测共轭与代数共轭的区分组织了整个构造。特征二在这里是主题性的：二次布尔进位 $B = \mathrm{span}{v \otimes v}$ 对可测作用不可见、对离散对偶可见，于是&amp;quot;同一个概率空间上两种紧群律&amp;quot;就够了，还能平移出可数无穷纤维，指数 $[\Gamma_n : \Gamma_0] = 2^{4n}$。注意原文有一句 &amp;ldquo;This choice took some work&amp;rdquo;。divided square 不是第一次就选对，普通对称平方商会模糊特征二极化。这个选择本身是迭代出来的。&lt;/p&gt;
&lt;p&gt;Ehrhart 章节的换坐标系最激进：&amp;ldquo;rationality 不必要&amp;rdquo;：任意凸体都能有 toric 势（实 Monge–Ampère 定理），于是扔掉有理多胞形的包袱，在非紧复环面上工作。&lt;/p&gt;
&lt;p&gt;一个统一的观察：这些转换不是随机的。它们都由前一轮失败的&amp;quot;信息缺陷&amp;quot;定向：哪条信息被丢了，就往能重新看见那条信息的坐标系里跳。失败定位缺陷，转换修复缺陷。两件事是一套循环。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="从失败中提炼元规则并守恒不变量"&gt;从失败中提炼元规则，并守恒不变量&lt;/h2&gt;
&lt;p&gt;第三章有一段教科书级的元认知陈述，我读到时愣了一下，因为这种&amp;quot;从失败里抽规则&amp;quot;的行为正是我写论文时最羡慕的能力。log-midrank 中间方案失败后，身份等式揭示了原则：&lt;/p&gt;
&lt;p&gt;$$\mathrm{Var}(\mathrm{midrank}) = \frac{1 - \sum_j p_j^3}{12}$$&lt;/p&gt;
&lt;p&gt;这个等式（其中 $p_j$ 是组件大小分布）被读作一条原则：&amp;ldquo;平均组件大小的有界单调函数，绝不平均无界的大小本身。&amp;ldquo;最终方案，顶点加权中位数 $f = M / (M + m_A)$，就是这条原则的干净实现。AI 不仅记录做了什么，还记录从失败中抽象出的&amp;quot;为什么这样做&amp;quot;的规则。这已经是 Polya 式启发法的形态。&lt;/p&gt;
&lt;p&gt;贯穿全书的是&amp;quot;什么东西必须精确保持&amp;quot;的自觉。量子并行重复章节保 Born 权重：共轭对恒等式 $F^{1/2+it} \cdot F^{1/2-it} = F$ 和有限预解纯化。永久型下界章节的正性来自真实 Gram 余项恒等式，而不是&amp;quot;假定 associated 多项式为正&amp;rdquo;。球堆积章节保符号位置而非全局范数。非 sofic 群章节把组件尺寸匹配到 $\rho_n \to 1$，运输分割被保全。&lt;/p&gt;
&lt;p&gt;这跟物理学的守恒量直觉同构：先找到论证中必须不变的量，失败模式往往就是某个量被破坏了。AI 的搜索看起来像是对&amp;quot;不变量&amp;quot;的持续探测：每次失败都问一句：刚才什么量变了？&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="显式有限见证优于渐近启发"&gt;显式有限见证优于渐近启发&lt;/h2&gt;
&lt;p&gt;漫游偏爱&amp;quot;能指着一个对象说这就是为什么&amp;quot;的论证，而不是&amp;quot;平均来说应该可以&amp;rdquo;。这个偏好很具体：&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;remote shells：可显式构造的球堆积修正层，比渐近密度论证更受信任。&lt;/li&gt;
&lt;li&gt;饱和矩阵：来自 Alon–Ben-Eliezer–Shangguan–Tamo 帽子猜测的显式结构。&lt;/li&gt;
&lt;li&gt;symplectic 广义四边形 $W(q)$：作为稠密见证的有限几何对象。&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;最极端的例子在第 12 章：2-退化图超指数反例。熵窗口 $\kappa + \tau \log_2 3 &amp;lt; \beta &amp;lt; 2h(\tau) - 1$ 在 $\tau = (\sqrt{3}-1)/2$ 处优化，间隙 $w \approx 0.003728$——窄但严格为正。原文的表述我记了很久：&amp;ldquo;一旦为正，大小就不再是障碍。&amp;ldquo;它知道正性比幅度重要。一个 $0.0037$ 的间隙和 $0.5$ 的间隙在存在性论证里等价。&lt;/p&gt;
&lt;p&gt;这个偏好与人类数学家的习惯一致：构造性证明比存在性证明更容易被社区核查、简化和复用。AI 似乎学到了这一点，或者被训练数据里的数学文化塑造了这一点。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="局限叙述不是验证"&gt;局限：叙述不是验证&lt;/h2&gt;
&lt;p&gt;漫游系列自己写了 Limitations，而且没有美化：这是 AI 的事后回顾，&amp;ldquo;反映的是被记录的推理过程，而非对每一步的独立验证&amp;rdquo;。&lt;/p&gt;
&lt;p&gt;这个限定我必须认真对待。回顾性叙事有 hindsight bias 的风险：曲折路径被整理成连贯故事，可能高估了&amp;quot;计划性&amp;rdquo;、低估了随机搜索。我在上面总结的&amp;quot;失败定位缺陷，转换修复缺陷&amp;quot;的循环，可能是真实搜索的样子，也可能是叙事把随机撞墙整理成了因果链条。&lt;/p&gt;
&lt;p&gt;还有几件事值得记住：没有任何结果声称被机器验证；GapCVP 章节的指数大到 $O(N^{802})$ 位，只保多项式时间可计算性；多重色 Ramsey 的常数 $6e^{38}$ 是插值吸收掉的；Ehrhart 等式分类仍开放。这套漫游呈现的是&amp;quot;AI 认为自己在做什么&amp;rdquo;，不是&amp;quot;AI 证明了什么&amp;rdquo;。&lt;/p&gt;
&lt;p&gt;更深一层的张力在这里：如果把推理过程的叙述当作理解的定义，那么 AI 拥有理解；如果理解要求对每一步的独立可验证把握，那么叙述只是 product 的另一种形式。Park 的表述我一直觉得锋利：&amp;ldquo;proof is a product, understanding is a capacity。&amp;ldquo;漫游系列把 product 翻译成 capacity 的语言：它让黑箱推理变得可以被人类社区检查、简化、吸收。它不能替代理解，但它是在为&amp;quot;可理解性&amp;quot;争取空间。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="boundary-conditions"&gt;Boundary conditions&lt;/h2&gt;
&lt;p&gt;这套认知画像的边界，我目前看到这些：&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;它基于单一来源（OpenAI 的漫游系列），样本是十二个精心挑选的成功案例。失败到完全放弃的案例不在其中，搜索的真实失败率未知。&lt;/li&gt;
&lt;li&gt;事后叙事可能系统性高估计划的连贯性。没有独立于叙事的过程数据，无法区分&amp;quot;结构发现&amp;quot;与&amp;quot;事后归因&amp;rdquo;。&lt;/li&gt;
&lt;li&gt;我的解读本身是第三层阅读（读 AI 的重述），可能放大了漫游刻意塑造的&amp;quot;理想推理形态&amp;rdquo;。&lt;/li&gt;
&lt;li&gt;&amp;ldquo;换坐标系&amp;quot;模式可能被过度概括：它确实是高频模式，但我没有统计它的覆盖率。&lt;/li&gt;
&lt;li&gt;我不知道这些模式在非数学领域（编程、科学实验设计）是否同样成立。这套画像可能只是数学推理的特例。&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="open-questions"&gt;Open questions&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;如果给 AI 同一个问题的两套独立 chain-of-thought，事后叙事会收敛到同一套&amp;quot;认知模式&amp;rdquo;，还是各自编造不同的故事？这能检验叙事是忠实压缩还是 confabulation。&lt;/li&gt;
&lt;li&gt;失败排除机制类的能力，在多大程度上来自训练数据里的数学文化，多大程度上是 RL 搜索自然涌现的？&lt;/li&gt;
&lt;li&gt;保不变量直觉是否有一个可操作的算法形式？比如&amp;quot;检测论证中被破坏的量&amp;quot;能否自动化成调试工具？&lt;/li&gt;
&lt;li&gt;人类数学家读漫游后能加速理解定理；如果 AI 生成面向自身的&amp;quot;推理地图&amp;rdquo;，它对 AI 自身的后续搜索有反馈价值吗？&lt;/li&gt;
&lt;li&gt;六到十二个月后重读：当更多独立证据出现时，我还会认为失败是 AI 的主要认知工具吗？&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;[[Q]] 六个月后：如果出现独立于 OpenAI 的 AI 数学推理过程记录（比如开源模型的长思维链日志），我总结的五种认知模式（失败排除、反例证伪、换坐标系、不变量守恒、显式见证）中有几种能复现？&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;OpenAI, &amp;ldquo;How the Ideas Came Together: Notes on AI Mathematical Reasoning&amp;rdquo;（十二章节漫游系列）， &lt;a href="https://cdn.openai.com/pdf/reasoning-walkthroughs.pdf"target="_blank" rel="noopener noreferrer"&gt;https://cdn.openai.com/pdf/reasoning-walkthroughs.pdf&lt;/a&gt;。&lt;/li&gt;
&lt;li&gt;Kun, G. 与 Thom, A.，关于 non-sofic groups 与 expanders 的工作（ch03 引用）。&lt;/li&gt;
&lt;li&gt;Cohn, H. 与 Elkies, N.，球堆积上界理论（ch01 背景）。&lt;/li&gt;
&lt;li&gt;Erdős, P. 与 Simonovits, M.，极值图论紧致性猜想（ch11 背景）。&lt;/li&gt;
&lt;li&gt;Alon, N.、Ben-Eliezer, T.、Shangguan, C. 与 Tamo, I.，帽子猜测与饱和矩阵（ch12 引用）。&lt;/li&gt;
&lt;li&gt;Park,Automation Without Understanding， &lt;a href="https://arxiv.org/abs/2607.06377"target="_blank" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2607.06377&lt;/a&gt;。&lt;/li&gt;
&lt;/ol&gt;</description></item><item><title>From engineered engagement to evidence-based development: the neuroscience of short-form video and a framework for redirecting its mechanisms</title><link>http://fengwang.github.io/posts/sfv-engagement-to-development/</link><pubDate>Tue, 14 Jul 2026 00:00:00 +0000</pubDate><guid>http://fengwang.github.io/posts/sfv-engagement-to-development/</guid><description>&lt;p&gt;I kept encountering the same paradox in the literature on short-form video. The platforms are described as &amp;ldquo;addictive,&amp;rdquo; yet population-level harm effects are tiny. The mechanisms are called &amp;ldquo;exploitative,&amp;rdquo; yet they map almost perfectly onto established findings in learning science. And the proposed solutions oscillate between moral panic and naive techno-optimism. I wanted a single, honest account that holds both sides without collapsing into either.&lt;/p&gt;
&lt;p&gt;The central claim of this article is that short-form video engagement is driven by three converging mechanisms — dopaminergic prediction error, variable-ratio reinforcement, and algorithmic personalization — whose effects are substrate-neutral: they can serve compulsive consumption or structured learning depending on design intent and guardrails. The practical consequence is that the same features making platforms compelling can be ethically redirected toward mastery, provided we replace the reward target from &amp;ldquo;next clip&amp;rdquo; to &amp;ldquo;next competence signal.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope.&lt;/strong&gt; This covers the psychological and neurological mechanisms of short-form video engagement, the contested evidence for &amp;ldquo;addiction,&amp;rdquo; and a practical framework for redirecting these mechanisms toward learning and habit formation. It does not cover clinical diagnosis, platform engineering internals, content moderation, or legislative policy.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prerequisites.&lt;/strong&gt; This assumes familiarity with basic dopamine reward systems, operant conditioning, and cognitive load theory.&lt;/p&gt;
&lt;h2 id="part-i-the-three-engines-of-engagement"&gt;Part I: The three engines of engagement&lt;/h2&gt;
&lt;h3 id="the-neural-reward-engine"&gt;The neural reward engine&lt;/h3&gt;
&lt;p&gt;At the foundation of short-form video&amp;rsquo;s pull is a single, well-replicated finding: dopamine neurons encode reward prediction error. They fire strongly when a reward is better than expected, fire to cues that predict reward, and remain silent when outcomes are fully predicted. This was established through classic electrophysiology — Schultz (1998) recorded from midbrain dopamine neurons while monkeys learned associations between visual cues and juice rewards, finding that dopamine firing shifted from the reward itself to the predictive cue as learning progressed.&lt;/p&gt;
&lt;p&gt;The significance for short-form video lies in a related result: dopamine responses also scale with reward uncertainty, rising most when the probability of reward is intermediate. A TikTok feed is precisely this condition. The user cannot predict whether the next clip will be highly entertaining, mildly interesting, or entirely flat. The brain is maintained in a state of maximum anticipation, and the uncertainty itself is the engine.&lt;/p&gt;
&lt;p&gt;This connects to a distinction that matters enormously for understanding compulsive use: the dissociation between &amp;ldquo;wanting&amp;rdquo; (incentive salience, driven by mesolimbic dopamine) and &amp;ldquo;liking&amp;rdquo; (hedonic pleasure, which does not depend on dopamine). The felt experience of &amp;ldquo;I don&amp;rsquo;t even enjoy this but I can&amp;rsquo;t stop&amp;rdquo; is a direct prediction of this model. The wanting system can run independently of the liking system, and variable-ratio reward schedules are precisely the conditions that drive it hardest.&lt;/p&gt;
&lt;p&gt;A parallel system, less discussed but potentially important, operates through Action Prediction Error (APE). While reward prediction error evaluates outcomes and drives learning about value, APE neurons in the tail of the striatum track how often an action is performed, serving as a value-free teaching signal that consolidates habitual motor behavior. The automatic thumb-swipe — the physical gesture of scrolling — may be solidified through this mechanism, creating a habit loop that is neurologically resistant to conscious interruption even when the user recognizes the behavior is unproductive.&lt;/p&gt;
&lt;h3 id="the-behavioral-schedule"&gt;The behavioral schedule&lt;/h3&gt;
&lt;p&gt;Beneath the neuroscience sits a classic behavioral structure. A variable-ratio reinforcement schedule — reward delivered after an unpredictable number of responses — produces the highest, most persistent response rates and the greatest resistance to extinction of any partial-reinforcement schedule. This is operant psychology 101, established through decades of work with animal models and replicated across species.&lt;/p&gt;
&lt;p&gt;Scrolling a feed in which only some clips are rewarding is functionally a variable-ratio schedule. The pull-to-refresh gesture or the swipe-up is the behavioral trigger. Because the cost of each response is nearly zero (a thumb movement) and the payoff is uncertain, the system generates the maximum possible rate of responding. Habit-forming product design has made this explicit: designers deliberately apply variable-reward schedules modeled on the operant conditioning literature.&lt;/p&gt;
&lt;p&gt;What makes the digital version particularly potent is the combination of near-zero response cost with algorithmic improvement of hit rate. In a traditional variable-ratio schedule (say, a slot machine), the reward probability is fixed. In a short-form feed, the recommender system learns from each interaction what the user finds rewarding and progressively increases the probability of a &amp;ldquo;hit.&amp;rdquo; The schedule is not merely variable — it is adaptive, getting better at predicting what will keep the user engaged.&lt;/p&gt;
&lt;h3 id="the-algorithm-as-amplifier"&gt;The algorithm as amplifier&lt;/h3&gt;
&lt;p&gt;Short-video recommender systems learn user preferences primarily from the implicit watch-time signal and personalize the feed rapidly, without requiring explicit ratings. In a donated-data study, participants&amp;rsquo; daily videos viewed and time on platform roughly doubled within about 80 days. The precise doubling figure comes from a single study and should be read as indicative rather than a general effect size, but the qualitative loop is well-documented: the more a user watches, the better the model predicts what will hold attention, which increases watching.&lt;/p&gt;
&lt;p&gt;At the neural level, personalized recommendation algorithms are effective at up-regulating activity in both the ventral tegmental area (VTA) — the origin of dopaminergic cell bodies — and sub-regions of the default mode network (DMN). Specifically, viewing personalized content increases coupling between the posterior cingulate cortex and sensory cortices (deepening sensory immersion) while decreasing coupling between the medial prefrontal cortex and regions involved in cognitive evaluation. The algorithm is not merely selecting content; it is altering the neural conditions under which the user processes information, reducing the brain&amp;rsquo;s capacity for critical assessment while amplifying sensory engagement.&lt;/p&gt;
&lt;p&gt;This creates what I think of as the core tension of the entire problem: the same algorithmic personalization that makes content maximally engaging also makes it maximally difficult to disengage. The system learns to exploit individual cognitive vulnerabilities — not through malice, but through optimization of a simple objective function (maximize watch time) applied to a substrate (the human reward system) that was not designed for this environment.&lt;/p&gt;
&lt;h2 id="part-ii-the-costs--and-the-contested-question-of-harm"&gt;Part II: The costs — and the contested question of harm&lt;/h2&gt;
&lt;h3 id="attention-and-its-measurable-costs"&gt;Attention and its measurable costs&lt;/h3&gt;
&lt;p&gt;The novelty and switching that make the feed engaging carry attentional costs. Heavy media multitasking and frequent short-form video use are associated with weaker attentional filtering, larger task-switching costs, and more frequent attentional lapses with poorer incidental memory.&lt;/p&gt;
&lt;p&gt;EEG studies provide more specific evidence. A study using the Attention Network Test found a significant negative correlation (r = -0.395, p = 0.007) between short-video addiction tendencies and theta wave power in frontal electrodes during cognitive conflict resolution. Theta oscillations in the frontal cortex are essential for recruiting neural resources to manage competing signals and exert executive control. The critical detail: this neural degradation occurred even in the absence of observable behavioral deficits on the task, indicating a &amp;ldquo;neural masking&amp;rdquo; effect where executive control signals atrophy before performance drops become visible.&lt;/p&gt;
&lt;p&gt;The broader brainwave picture is consistent. Gamma power increases by 40-62% during high-reward moments (indicating hyper-arousal from rapid context-switching), prefrontal beta power reduces by 22% after just 20 minutes (reflecting impaired decision-making), alpha declines during engagement (suggesting unsustainable cognitive load), and delta increases by up to 50% over two years in heavy users (correlating with digital fatigue and degraded sleep architecture).&lt;/p&gt;
&lt;p&gt;I find the &amp;ldquo;neural masking&amp;rdquo; result particularly important. It means that by the time someone notices their attention is degraded, the underlying neural substrate has already been changing for some time. The behavioral symptom is a lagging indicator.&lt;/p&gt;
&lt;h3 id="memory-fragmentation"&gt;Memory fragmentation&lt;/h3&gt;
&lt;p&gt;A consequence I did not expect to find in the literature: rapid short-form consumption may actively impair the brain&amp;rsquo;s ability to process continuous information. Human cognition relies on event segmentation — parsing continuous experience into discrete, meaningful units. When a prediction fails, the brain establishes an event boundary and constructs a new predictive model.&lt;/p&gt;
&lt;p&gt;Short-form video forces artificial prediction errors at extreme frequency (every 15-30 seconds), habituating the brain to a &amp;ldquo;segment-and-refresh&amp;rdquo; mode. Eye-tracking studies using Hidden Markov Models show that people exposed to random short videos exhibit significantly more fragmented perception. Neuroimaging data makes the cost concrete: learning from fragmented short videos drops memory retrieval accuracy from approximately 66.4% to 43.3%, with reduced activation in the left claustrum (cognitive control and multisensory integration), left caudate nucleus (goal-directed behavior), and left middle temporal gyrus (semantic processing), plus broken functional connectivity between caudate and claustrum.&lt;/p&gt;
&lt;p&gt;The implication is that short-form video does not merely waste time. It may train a mode of information processing that is actively hostile to the kind of sustained, integrative thinking that deep learning requires. When novelty comes as relentless context switching, it supports orienting and stimulation while undermining semantic integration, elaboration, and later retrieval.&lt;/p&gt;
&lt;h3 id="the-sleep-connection"&gt;The sleep connection&lt;/h3&gt;
&lt;p&gt;Sleep contributes to memory consolidation through hippocampo-thalamocortical synchronization during sleep cycles. Sleep deprivation impairs encoding and retention of newly learned material. Bedtime scrolling is concerning not only because it displaces total sleep time, but because it may turn potentially useful learning episodes into poorly consolidated ones. The hippocampus is central because short-form content often creates the illusion of learning: high familiarity and high stimulation can be mistaken for durable memory, but episodic and semantic memory formation depend on consolidation processes that sleep supports.&lt;/p&gt;
&lt;h3 id="how-much-of-this-is-addiction"&gt;How much of this is &amp;ldquo;addiction&amp;rdquo;?&lt;/h3&gt;
&lt;p&gt;The clinical picture is genuinely mixed, and I want to hold both sides visible. On one side, validated short-video and TikTok dependence scales exist and, in a 16,038-person study, identify a dependent subgroup of roughly 7.5% (plus about 16% at-risk), converging with DSM-5-style dependence criteria. Problematic use is further associated with measurable attention deficits and with elevated depression, anxiety, and stress.&lt;/p&gt;
&lt;p&gt;On the other side, the proposition that short-form video constitutes a clinical behavioral addiction comparable to established disorders is contradicted by rigorous, pre-registered, large-sample analysis concluding that the association between digital-technology use and adolescent well-being is negative but tiny — explaining at most about 0.4% of variance across datasets totaling more than 350,000 participants, and &amp;ldquo;too small to warrant policy change.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The honest synthesis is that a minority show dependence-like patterns while population-level harm is small and confounded. This matters for the framework: harm-reduction guardrails should be proportionate and targeted, not framed as if universal pathology were established. The defensible design principle is to build systems that work well for most people while providing guardrails for the vulnerable, rather than assuming everyone is at risk.&lt;/p&gt;
&lt;h2 id="part-iii-turning-the-engine-around"&gt;Part III: Turning the engine around&lt;/h2&gt;
&lt;h3 id="the-alignment-between-engagement-and-learning-science"&gt;The alignment between engagement and learning science&lt;/h3&gt;
&lt;p&gt;Here is where the story gets interesting. The features that make short-form video compelling map onto some of the most robust findings in learning science, but with a critical inversion.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mechanism in SFV&lt;/th&gt;
&lt;th&gt;Learning science parallel&lt;/th&gt;
&lt;th&gt;Key difference&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Short segments&lt;/td&gt;
&lt;td&gt;Segmenting effect: breaking instruction into learner-paced segments improves retention and reduces cognitive load&lt;/td&gt;
&lt;td&gt;SFV segments are arbitrary (entertainment); learning segments are pedagogically meaningful&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Novelty variation&lt;/td&gt;
&lt;td&gt;Novel stimuli enhance memory encoding via dopaminergic hippocampal modulation&lt;/td&gt;
&lt;td&gt;SFV novelty is relentless context switching; learning novelty should be bounded and schema-integrative&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Continuous feedback&lt;/td&gt;
&lt;td&gt;Immediate feedback supports competence satisfaction and motivation&lt;/td&gt;
&lt;td&gt;SFV feedback is social validation (likes); learning feedback should be informational (mastery signals)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Personalization&lt;/td&gt;
&lt;td&gt;Adaptive difficulty maintains flow state&lt;/td&gt;
&lt;td&gt;SFV personalization maximizes watch time; learning personalization should optimize retention&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The correspondence is real but partial. The segmenting effect, the spacing effect, and the testing effect are among the most replicated findings in educational psychology. Short-form video already delivers short segments, novelty variation, and immediacy — so the format is a natural vehicle for structured learning. But the default configuration points the wrong way: passive consumption rather than active retrieval, entertainment reward rather than competence reward, infinite continuation rather than bounded sessions.&lt;/p&gt;
&lt;h3 id="self-determination-theory-as-the-replacement-engine"&gt;Self-Determination Theory as the replacement engine&lt;/h3&gt;
&lt;p&gt;Self-Determination Theory (SDT) identifies three basic psychological needs whose satisfaction fosters intrinsic motivation: autonomy (the need for control), competence (the need for mastery), and relatedness (the need for connection). In interactive technology specifically, competence-satisfying feedback and masterable, autonomy-supportive design predict enjoyment and sustained motivation.&lt;/p&gt;
&lt;p&gt;This matters because it names a healthy engine of engagement that can substitute for the variable-ratio hook. The variable-ratio schedule drives behavior through uncertainty and external reward. SDT-aligned design drives behavior through agency and internal satisfaction. They are not the same mechanism; the &amp;ldquo;repurposing&amp;rdquo; is partly a substitution.&lt;/p&gt;
&lt;p&gt;A landmark meta-analysis found that tangible, contingent extrinsic rewards can undermine intrinsic motivation (d = -0.28 to -0.40), whereas positive informational feedback enhances it (d = 0.31 to 0.33). This single result carries much of the ethical weight of the framework: it is why the recommendation is to move away from variable-ratio hooks toward competence feedback. Points, badges, and leaderboards that function as controlling external rewards can backfire. Progress visibility, skill trees, and narrative-embedded mastery signals that function as informational feedback support sustained engagement.&lt;/p&gt;
&lt;h3 id="the-learning-science-stack"&gt;The learning science stack&lt;/h3&gt;
&lt;p&gt;The component evidence for redirecting engagement toward learning is strong at the individual mechanism level, even if integrated systems have not been directly validated.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Spacing and retrieval.&lt;/strong&gt; Distributing practice over time (the spacing effect) and actively retrieving information (the testing effect) are among the most robust enhancers of long-term retention. The optimal inter-study interval depends on the desired retention interval — roughly 10-20% of the target retention period. A learning system should present micro-units at expanding intervals, exploiting the brain&amp;rsquo;s natural consolidation processes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Segmenting and microlearning.&lt;/strong&gt; Breaking instruction into short, learner-paced segments improves retention and transfer while lowering cognitive load. Microlearning applies this principle with measurable outcome gains. The critical qualification: brevity alone is not instructional quality. Short clips improve access and initial uptake, but durable learning requires retrieval, spacing, and transfer practice rather than mere exposure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Interleaving.&lt;/strong&gt; Studying different but related concepts within a single session, rather than one concept exhaustively, introduces desirable difficulty that improves transfer and critical thinking. This actively counters the rigid over-segmentation caused by recreational short-form video.&lt;/p&gt;
&lt;h3 id="from-engaged-sessions-to-durable-habits"&gt;From engaged sessions to durable habits&lt;/h3&gt;
&lt;p&gt;Habit automaticity forms gradually through consistent, context-dependent repetition, taking about 66 days on average (range 18 to 254) with occasional missed repetitions not derailing the process. Forming an implementation intention — an explicit if-then plan linking a cue to an action — substantially increases goal attainment (d = 0.65 across 94 tests). Behavior occurs when motivation, ability, and a prompt converge (Fogg&amp;rsquo;s B = MAP model).&lt;/p&gt;
&lt;p&gt;The practical design rule is that simplicity changes behavior more effectively than motivation. Because motivation is an unreliable, fluctuating emotional state, relying on it for skill acquisition leads to failure. The target behavior must be reduced to an almost absurdly simple action (under 30 seconds), anchored to an existing routine as a cue, and followed by self-generated positive reinforcement to trigger a dopamine response that consolidates the behavior as rewarding.&lt;/p&gt;
&lt;h2 id="part-iv-the-applied-framework"&gt;Part IV: The applied framework&lt;/h2&gt;
&lt;h3 id="the-central-design-rule"&gt;The central design rule&lt;/h3&gt;
&lt;p&gt;The framework rests on one principle: preserve the attentional efficiency of short-form content while replacing compulsive continuation with bounded mastery cycles. This means keeping the short segments, the immediacy, and the personalization while changing the reward target, the feedback structure, and the session architecture.&lt;/p&gt;
&lt;h3 id="step-by-step"&gt;Step-by-step&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;One micro-objective per clip.&lt;/strong&gt; Each short unit should answer exactly one question or train one action: one theorem idea, one pronunciation contrast, one coding pattern. This follows the evidence behind microlearning, segmentation, and cognitive-load reduction.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Front-load curiosity, not clutter.&lt;/strong&gt; Open with a concrete problem, prediction prompt, or noticeable error rather than sensory overload. Novelty helps when it signals relevance; too much surface novelty risks fragmented processing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Force an action before the answer appears.&lt;/strong&gt; Insert a prompt to retrieve, predict, imitate, or classify before the reveal. Retrieval practice generally outperforms passive review, and even simple retrieval prompts can improve delayed problem solving. This is where the variable-ratio schedule gets replaced: instead of &amp;ldquo;maybe the next clip will be good,&amp;rdquo; the uncertainty becomes &amp;ldquo;can I get this right?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Use bounded reinforcement.&lt;/strong&gt; Provide an immediate success signal after correct retrieval or execution, but avoid endless variable-ratio continuation. Good rewards are progress bars, level completion, skill streaks, or supportive peer feedback tied to effort and mastery. Not &amp;ldquo;one more clip.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Schedule reactivation instead of bingeing.&lt;/strong&gt; Deliver the next clip in a spaced sequence rather than an infinite feed: same-day recap, next-day retrieval, then expanding intervals. The algorithm that currently personalizes for watch time should personalize for retention.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bind learning to a stable cue.&lt;/strong&gt; Attach the short learning unit to a reliable context such as &amp;ldquo;after coffee,&amp;rdquo; &amp;ldquo;after opening the laptop,&amp;rdquo; or &amp;ldquo;immediately after lunch.&amp;rdquo; Habit research consistently shows that context stability and cue-based enactment build automaticity more effectively than vague intention alone.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Close the loop off-screen.&lt;/strong&gt; Every cluster of clips should end in a non-feed action: speaking a phrase aloud, solving one practice item, performing one movement, writing one summary sentence, or teaching the concept to someone else. This helps convert familiarity into transfer.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Protect consolidation.&lt;/strong&gt; Use hard stopping rules, especially before bedtime. Sleep is not an optional extra for learning; it is part of the memory system. Short-form educational routines should include no-autoplay rules, session caps, and a pre-sleep buffer.&lt;/p&gt;
&lt;h3 id="ethical-guardrails"&gt;Ethical guardrails&lt;/h3&gt;
&lt;p&gt;An ethical short-form educational system should be optimized for mastery per minute, not minutes per user. That distinction is essential because the same mechanics that increase retention of users can decrease retention of knowledge.&lt;/p&gt;
&lt;p&gt;The specific risks to manage: infinite continuation (remove infinite scroll and autoplay in educational contexts), vulnerable-user amplification (default-friction safeguards for youth, high-FoMO users, low self-control users), popularity over pedagogy (emphasize mastery badges over raw popularity counts), shallow learning masquerading as productivity (require delayed quizzes and off-screen transfer tasks), sleep displacement (no reminders near bedtime), and opaque personalization (make recommendation logic legible; let learners choose playlists, pacing, and stop points).&lt;/p&gt;
&lt;h3 id="evaluation-metrics"&gt;Evaluation metrics&lt;/h3&gt;
&lt;p&gt;Because the literature warns against using time-spent alone as the core metric, evaluation should separate engagement quality, learning quality, habit quality, and well-being cost. Completion rate, voluntary rewatch rate, and return rate distinguish active engagement from passive exposure. End-of-clip accuracy and one-step problem solving capture immediate uptake. Twenty-four-hour and seven-day recall, cumulative quiz performance, and transfer tasks prevent the familiarity illusion. Habit automaticity scores and context-stable enactment rates track whether behavior has become automatic. Attention-control measures, mind-wandering reports, and time-to-re-engage deep work detect whether the intervention is eroding control. Session length variance, bedtime use, sleep quality, negative affect, and academic displacement catch &amp;ldquo;engagement wins&amp;rdquo; that are actually developmental losses.&lt;/p&gt;
&lt;h2 id="boundary-conditions"&gt;Boundary conditions&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;The framework assumes content that can be meaningfully chunked into micro-units. Subjects requiring extended derivation, sustained argument, or immersive practice (e.g., advanced mathematics, complex surgery, musical performance) may not adapt well to short-form delivery.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The learning-science components (spacing, retrieval, interleaving) are individually well-validated, but their integration into a single short-form system has not been directly tested as a unified package. The framework is a synthesis of supported analogies, not a validated product.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The addiction question is unresolved by design. The framework targets guardrails for the vulnerable rather than assuming universal pathology, but this calibrated approach may underserve individuals with genuine compulsive-use disorders who need clinical intervention, not better product design.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Most of the evidence base is drawn from adolescents and university students, with heavy representation from Chinese samples. Cross-cultural and adult generalization is uncertain.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The framework does not address the political economy of platform design: the business model that funds algorithmic personalization is advertising revenue maximized by watch time, and &amp;ldquo;mastery per minute&amp;rdquo; optimization conflicts with that incentive structure.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Individual differences in impulsivity, self-control, attachment style, and baseline mental health moderate all of these effects. A framework that works for a typical user may fail for someone with high trait anxiety or low executive function baseline.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="open-questions"&gt;Open questions&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Does the Action Prediction Error mechanism (value-free habit consolidation in the tail of the striatum) actually apply to the thumb-swipe gesture, or is it specific to goal-directed motor sequences? The APE research is recent and mostly from animal models.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;What is the actual dose-response curve for fragmented viewing and memory impairment? The 66.4% to 43.3% accuracy drop comes from a single fMRI study with specific materials. Does this generalize across content types and populations?&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Can algorithmic personalization be repurposed for retention optimization without recreating the compulsive-use dynamics it currently serves? The technical challenge is that the same feedback loop that improves watch time also increases engagement — and disentangling these requires a different objective function.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;How does attention restoration interact with the microlearning framework? The attention restoration literature suggests that directed attention requires recovery periods, but the optimal &amp;ldquo;dose&amp;rdquo; of restoration relative to learning-session length is not established.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The framework recommends replacing variable-ratio entertainment rewards with competence-contingent feedback. But does competence feedback activate the same dopaminergic circuits, or does it rely on different neural pathways? If different, the &amp;ldquo;redirecting engagement&amp;rdquo; metaphor is misleading — it is more like substituting one motivation system for another.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;[[Q]] Six months from now: in any chance there is a randomized trial comparing a short-form learning system built on these principles (spaced micro-units, retrieval prompts, competence feedback, session caps) against both a standard short-form entertainment feed and a traditional long-form learning control, with delayed retention and transfer as primary outcomes?&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Schultz, W. &amp;ldquo;Predictive reward signal of dopamine neurons.&amp;rdquo; Journal of Neurophysiology, 1998.&lt;/li&gt;
&lt;li&gt;Fiorillo, C. D., Tobler, P. N., &amp;amp; Schultz, W. &amp;ldquo;Discrete coding of reward probability and uncertainty by dopamine neurons.&amp;rdquo; Science, 2003.&lt;/li&gt;
&lt;li&gt;Berridge, K. C., &amp;amp; Robinson, T. E. &amp;ldquo;Liking, wanting, and the incentive-sensitization theory of addiction.&amp;rdquo; American Psychologist, 2016.&lt;/li&gt;
&lt;li&gt;Ophir, E., Nass, C., &amp;amp; Wagner, A. D. &amp;ldquo;Cognitive control in media multitaskers.&amp;rdquo; PNAS, 2009.&lt;/li&gt;
&lt;li&gt;Zannettou, S., et al. &amp;ldquo;Analyzing User Engagement with TikTok&amp;rsquo;s Short Format Video Recommendations using Data Donations.&amp;rdquo; CHI, 2024.&lt;/li&gt;
&lt;li&gt;Jiang, A., et al. &amp;ldquo;Assessing Short-Video Dependence: Development and Validation of the Short-Video Dependence Scale.&amp;rdquo; JMIR, 2025.&lt;/li&gt;
&lt;li&gt;Orben, A., &amp;amp; Przybylski, A. K. &amp;ldquo;The association between adolescent well-being and digital technology use.&amp;rdquo; Nature Human Behaviour, 2019.&lt;/li&gt;
&lt;li&gt;Deci, E. L., &amp;amp; Ryan, R. M. &amp;ldquo;Self-Determination Theory.&amp;rdquo; Center for Self-Determination Theory, 2000.&lt;/li&gt;
&lt;li&gt;Cepeda, N. J., et al. &amp;ldquo;Distributed practice in verbal recall tasks: A review and quantitative synthesis.&amp;rdquo; Psychological Bulletin, 2006.&lt;/li&gt;
&lt;li&gt;Roediger, H. L., &amp;amp; Karpicke, J. D. &amp;ldquo;Test-enhanced learning: Taking memory tests improves long-term retention.&amp;rdquo; Psychological Science, 2006.&lt;/li&gt;
&lt;li&gt;Rey, G. D., et al. &amp;ldquo;A Meta-analysis of the Segmenting Effect.&amp;rdquo; Educational Psychology Review, 2019.&lt;/li&gt;
&lt;li&gt;Lally, P., et al. &amp;ldquo;How are habits formed: Modelling habit formation in the real world.&amp;rdquo; European Journal of Social Psychology, 2010.&lt;/li&gt;
&lt;li&gt;Gollwitzer, P. M., &amp;amp; Sheeran, P. &amp;ldquo;Implementation intentions and goal achievement: A meta-analysis.&amp;rdquo; Advances in Experimental Social Psychology, 2006.&lt;/li&gt;
&lt;li&gt;Fogg, B. J. &amp;ldquo;Fogg Behavior Model (B=MAP).&amp;rdquo; behaviormodel.org, 2019.&lt;/li&gt;
&lt;li&gt;Deci, E. L., Koestner, R., &amp;amp; Ryan, R. M. &amp;ldquo;A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation.&amp;rdquo; Psychological Bulletin, 1999.&lt;/li&gt;
&lt;li&gt;Yan, T., et al. &amp;ldquo;Mobile phone short video use negatively impacts attention functions: an EEG study.&amp;rdquo; PubMed, 2024.&lt;/li&gt;
&lt;li&gt;Wei, M., et al. &amp;ldquo;Fragmented learning from short videos modulates neural activity and connectivity during memory retrieval.&amp;rdquo; ResearchGate, 2026.&lt;/li&gt;
&lt;li&gt;Sherman, L. E., et al. &amp;ldquo;The Power of the Like in Adolescence: Effects of Peer Influence on Neural and Behavioral Responses to Social Media.&amp;rdquo; Psychological Science, 2016.&lt;/li&gt;
&lt;li&gt;Guo, P. J., et al. &amp;ldquo;How video production affects student engagement: An empirical study of MOOC videos.&amp;rdquo; L@S, 2014.&lt;/li&gt;
&lt;li&gt;Brodt, S., et al. &amp;ldquo;Sleep — A brain-state serving systems memory consolidation.&amp;rdquo; Neuron, 2023.&lt;/li&gt;
&lt;li&gt;Singh, B., et al. &amp;ldquo;Time to Form a Habit: a systematic review and meta-analysis.&amp;rdquo; Health Psychology Review, 2024.&lt;/li&gt;
&lt;li&gt;U.S. Surgeon General. &amp;ldquo;Social Media and Youth Mental Health.&amp;rdquo; Advisory, 2023 (updated 2025).&lt;/li&gt;
&lt;/ol&gt;</description></item><item><title>从提示员到系统架构师：Loop Engineering 的范式跃迁</title><link>http://fengwang.github.io/posts/loop-engineering-paradigm-shift/</link><pubDate>Sun, 28 Jun 2026 00:00:00 +0000</pubDate><guid>http://fengwang.github.io/posts/loop-engineering-paradigm-shift/</guid><description>&lt;h2 id="引言你不再是对话者了"&gt;引言：你不再是对话者了&lt;/h2&gt;
&lt;p&gt;2026 年上半年，AI 编码代理从命令行玩具变成了日常基础设施。Claude Code、OpenCode、Codex CLI、Cursor——这些工具不再只是&amp;quot;更聪明的自动补全&amp;quot;。它们开始自主地阅读代码库、提出修改方案、运行测试、迭代修复。&lt;/p&gt;
&lt;p&gt;与此同时，一个微妙但彻底的转变发生了。那些从这些工具中获得最大产出的人，不再花更多时间写更好的提示词。他们花时间设计&lt;strong&gt;系统&lt;/strong&gt;——系统自己决定下一步做什么。&lt;/p&gt;
&lt;p&gt;Boris Cherny（Claude Code 的创造者）在 2026 年 5 月说了一句我觉得精准捕捉了拐点的话：&amp;ldquo;我不再给 Claude 写提示词了。我有一些 loop 在运行，而那些 loop 给 Claude 写提示词……我的工作是写 loop。&amp;rdquo;&lt;/p&gt;
&lt;p&gt;这句话像一把刀切开了两个时代。在此之前，你的价值取决于你多擅长和模型对话。在那之后，你的价值取决于你多擅长设计不需要你参与对话的系统。&lt;/p&gt;
&lt;p&gt;这篇文章记录的是我对这个转变的理解——它在 2026 年中期呈现的样子。如果你一年后重读，有些判断可能已经过时了。但结构性的东西，那些层次的划分、原语的识别、难的问题和容易犯的错——我猜它们会留得更久。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="一个不成立的假设和一个正在发生的事实"&gt;一个不成立的假设和一个正在发生的事实&lt;/h2&gt;
&lt;p&gt;我见过太多团队在 2025–2026 年的做法：雇一个人&amp;quot;专门调 prompt&amp;quot;，把这个人的产出绑定在每一次聊天窗口的对话质量上。这是一种很自然的想法——模型是你的对话对象，对话质量决定输出质量。&lt;/p&gt;
&lt;p&gt;但这个假设在 2026 年中间某个时刻悄悄失效了。&lt;/p&gt;
&lt;p&gt;失效的方式是这样的。想象两个工程师在使用同一个模型。工程师 A 打开聊天窗口，写了一条很长的 prompt，得到了一份不错的代码，然后关掉窗口。工程师 B 做了一个自动化工作流：一个 git hook 在 PR 创建时触发一个 agent，agent 阅读 diff、运行测试、检查代码风格、写 review 评论。如果评论被采纳，agent 还会修改代码并推送新的 commit。&lt;/p&gt;
&lt;p&gt;工程师 A 可能写了更优雅的 prompt。但工程师 B 一周内 Review 的 PR 数量是 A 的几十倍——而且他在这期间还做了一堆别的事情。&lt;/p&gt;
&lt;p&gt;这不是 prompt 质量的问题。这是&lt;strong&gt;抽象层次&lt;/strong&gt;的问题。A 在第一个层次上和模型互动。B 设计了一个运行在更高抽象层次上的系统，这个系统自己知道什么时候需要和模型对话。&lt;/p&gt;
&lt;p&gt;我最初看到这个差距时，觉得它只是一个效率优化——让机器做更多，人做更少。后来我发现这不仅仅是效率。它改变了你作为工程师的&lt;strong&gt;界面&lt;/strong&gt;。你的工作不再是&amp;quot;写更好的指令&amp;quot;，而是&amp;quot;设计更好的迭代结构&amp;quot;。你从一个操作员变成了架构师。&lt;/p&gt;
&lt;p&gt;这个转变我目前认为不可逆。不是因为工具会越来越好——而是因为一旦你体验过设计一个 loop 然后看着它自动完成过去需要你逐轮介入的工作，你就很难再回到聊天窗口里了。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="什么是-loop-engineering"&gt;什么是 Loop Engineering&lt;/h2&gt;
&lt;h3 id="一个直觉"&gt;一个直觉&lt;/h3&gt;
&lt;p&gt;Loop Engineering 的定义很简单：设计自主迭代循环的实践。这些循环决定什么时候调用模型、传递什么上下文、怎么评估输出、什么时候重试、什么时候停止。&lt;/p&gt;
&lt;p&gt;但它真正反直觉的地方是这个：&lt;strong&gt;loop 把工程师从循环内部移到了循环外部&lt;/strong&gt;。&lt;/p&gt;
&lt;p&gt;在 prompt engineering 时代，你在循环里——你读模型的输出，你决定下一步说什么，你判断什么时候够了。你是控制回路的一部分。&lt;/p&gt;
&lt;p&gt;在 loop engineering 时代，你设计循环本身。循环里的每一步可能仍然调用模型（可能很多次），但决定每一步做什么的逻辑是你写的代码、你定的规则、你连的工具——而不是你在聊天窗口里的实时判断。&lt;/p&gt;
&lt;h3 id="一个具体例子"&gt;一个具体例子&lt;/h3&gt;
&lt;p&gt;假设你要为一个代码库添加一组单元测试。&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;没有 loop 的做法&lt;/strong&gt;：你打开聊天窗口，贴一个源文件，说&amp;quot;为这个文件写单元测试&amp;quot;。模型回复一段代码。你复制到项目里，运行测试，发现有失败的，把错误贴回去让模型修。重复三四轮，差不多了。&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;有 loop 的做法&lt;/strong&gt;：你写一个脚本，大致是这样——&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;for 每个需要测试的源文件:
1. agent 读取文件 + 项目中已有的测试风格
2. agent 生成测试代码
3. 运行测试
4. 如果有失败:
a. agent 阅读错误信息
b. agent 修改测试代码
c. 回到第 3 步
5. 如果全部通过 -&amp;gt; commit
6. 如果重试超过 N 次 -&amp;gt; 标记为需要人工审查
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;你运行这个脚本。它处理了 20 个文件。其中有 17 个一次性通过，3 个经过一两次迭代后通过，1 个标记为人工审查。你花 5 分钟看了那个审查标记的文件，手动调了一行断言。&lt;/p&gt;
&lt;p&gt;在这整个过程里，你和模型的对话次数是&lt;strong&gt;零&lt;/strong&gt;。模型被你的 loop 调用了上百次——但你一次都没有打开聊天窗口。&lt;/p&gt;
&lt;h3 id="这个定义的两个边界"&gt;这个定义的两个边界&lt;/h3&gt;
&lt;p&gt;第一，不是所有任务都适合 loop。一次性的、定义清晰的、不需要迭代的任务——比如&amp;quot;把这个 JSON 转成 YAML&amp;quot;——写 loop 的成本高于收益。Loop 的价值在&lt;strong&gt;任务量大、有迭代模式、失败模式可预测&lt;/strong&gt;的场景里最明显。&lt;/p&gt;
&lt;p&gt;第二，Loop 不替代 prompt engineering。你设计 loop 时仍然需要写 prompt——但这些 prompt 现在不是写给最终用户看的人机对话 prompt，而是写给 &lt;strong&gt;loop 内部的子代理&lt;/strong&gt; 的系统指令。它们更结构化，更少自由发挥，更多约束条件。这个区别很重要：你并没有停止写 prompt，你写的 prompt 的&lt;strong&gt;消费者&lt;/strong&gt;变了。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="四层抽象阶梯从对话到自我改进"&gt;四层抽象阶梯：从对话到自我改进&lt;/h2&gt;
&lt;p&gt;如果给 2022–2026 年的 AI 工程画一条抽象层次上升的路径，它大致经过四个阶段。每一层包裹但不替换前一层。&lt;/p&gt;
&lt;h3 id="第一层prompt-engineering20222024"&gt;第一层：Prompt Engineering（2022–2024）&lt;/h3&gt;
&lt;p&gt;核心关注点：怎么说话模型才会照做。&lt;/p&gt;
&lt;p&gt;你优化的是措辞、角色设定、示例格式、few-shot 排列顺序。你的输出质量取决于你多了解模型的&amp;quot;脾气&amp;quot;——GPT-4 喜欢什么结构，Claude 对什么措辞敏感。这是一门手艺活，而且不透明。同一个 prompt 在两个版本之间可能突然不好用了。&lt;/p&gt;
&lt;h3 id="第二层context-engineering2025-年前后"&gt;第二层：Context Engineering（2025 年前后）&lt;/h3&gt;
&lt;p&gt;核心关注点：模型看到了什么。&lt;/p&gt;
&lt;p&gt;你不再只关心你说什么，你还关心模型看到的历史、文档、代码库结构。RAG 检索、系统提示组装、对话历史压缩——这些成为核心技能。你的 prompt 不再孤立存在，它被嵌入到一个更大的信息上下文中。&lt;/p&gt;
&lt;p&gt;这一层的出现是因为模型开始支持更长的上下文窗口，但很快人们发现&amp;quot;能放进去&amp;quot;不等于&amp;quot;能注意到&amp;quot;。你需要的不是把更多文本塞进窗口——你需要把&lt;strong&gt;正确的文本&lt;/strong&gt;放在模型会读到的地方。&lt;/p&gt;
&lt;h3 id="第三层harness-engineering2026-年初"&gt;第三层：Harness Engineering（2026 年初）&lt;/h3&gt;
&lt;p&gt;核心关注点：模型能做什么、不能做什么、怎么约束它。&lt;/p&gt;
&lt;p&gt;Harness 是模型和外部世界之间的运行时中间层。它管理工具注册、任务状态、资源预算、可观测性。它决定模型能调用哪些工具、每次调用的超时是多少、连续失败几次应该熔断。&lt;/p&gt;
&lt;p&gt;这一层的核心洞察是：&lt;strong&gt;模型的可靠性不是模型的问题，是 harness 的问题&lt;/strong&gt;。给同一个模型配上不同的 harness，产出质量可以差一个数量级。&lt;/p&gt;
&lt;h3 id="第四层loop-engineering2026-年中至今"&gt;第四层：Loop Engineering（2026 年中至今）&lt;/h3&gt;
&lt;p&gt;核心关注点：怎么让系统自己决定下一步。&lt;/p&gt;
&lt;p&gt;Loop 是 harness 之上的编排层。它不关心单次调用的质量——它关心的是：系统怎么知道什么时候该做什么？怎么知道自己做完了？做不完的时候怎么办？怎么从一次运行中学到的经验改进下一次运行？&lt;/p&gt;
&lt;p&gt;这是最抽象的一层，也是杠杆最高的一层。一个设计良好的 loop 可以把一次手动的三轮迭代压缩成一个自动步骤。但设计它的难度也最高——因为你不再有实时反馈。你设计的时候不知道具体会遇到什么情况。你的 loop 需要在运行时自己判断。&lt;/p&gt;
&lt;h3 id="四层的关系"&gt;四层的关系&lt;/h3&gt;
&lt;p&gt;拿我前面那个单元测试的例子来说：&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;每一轮调用的 prompt 内容是 &lt;strong&gt;Prompt Engineering&lt;/strong&gt;。&lt;/li&gt;
&lt;li&gt;让 agent 读取项目中已有的测试风格文件是 &lt;strong&gt;Context Engineering&lt;/strong&gt;。&lt;/li&gt;
&lt;li&gt;限制每次调用的 token 预算、超时、重试次数是 &lt;strong&gt;Harness Engineering&lt;/strong&gt;。&lt;/li&gt;
&lt;li&gt;决定&amp;quot;对每个文件做同样的事，但失败时循环回来&amp;quot;是整个 &lt;strong&gt;Loop&lt;/strong&gt;。&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;你不需要精通所有四层来使用 loop——但你设计 loop 越深入，四层都会碰到。这是个全栈活计。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="六个原语loop-的建筑模块"&gt;六个原语：Loop 的建筑模块&lt;/h2&gt;
&lt;p&gt;在研究了多个开源实现和商业产品之后，我发现大多数 loop 可以分解为六种原语。它们不是理论分类——我在实际代码里看到了同样的结构反复出现。&lt;/p&gt;
&lt;h3 id="1-心跳heartbeat谁启动-loop"&gt;1. 心跳（Heartbeat）——谁启动 loop&lt;/h3&gt;
&lt;p&gt;Loop 不能凭空运行。它需要一个触发条件。&lt;/p&gt;
&lt;p&gt;最简单的是时间触发——cron 表达式，每隔一段时间启动一次。GitHub Actions 的 &lt;code&gt;schedule&lt;/code&gt; 事件就是这个模式。更复杂的是事件驱动——webhook 收到外部信号时启动，或者文件系统变化时启动。&lt;/p&gt;
&lt;p&gt;Claude Code 内置了 &lt;code&gt;/loop&lt;/code&gt; 和 &lt;code&gt;/goal&lt;/code&gt; 命令，它们本质上是一个带有交互式心跳的循环：你告诉它一个目标，它在内部决定什么时候该启动下一步。OpenCode 不做这个——它把心跳决策留给你，用 &lt;code&gt;opencode run&lt;/code&gt; 作为 worker，你自己用 cron 或 CI 触发器来调度。&lt;/p&gt;
&lt;p&gt;心跳是容易被低估的原语。很多人可能会写了一个很好的 loop，但忘了想&amp;quot;这个 loop 是天天跑还是只跑一次&amp;quot;——然后它就只跑了一次，再也没动过。&lt;/p&gt;
&lt;h3 id="2-工作分离worktree并行隔离"&gt;2. 工作分离（Worktree）——并行隔离&lt;/h3&gt;
&lt;p&gt;真正的项目从来不是一个 loop 处理一件事。你会有多个 loop 同时在跑——一个修 bug，一个加功能，一个做代码审查。&lt;/p&gt;
&lt;p&gt;每个 loop 需要自己的隔离环境。用 git branch 或者独立目录来隔离，这样不同 loop 的修改不会打架。这在实施上很琐碎，但在设计上很关键——&lt;strong&gt;没有隔离的 loop 不是并行，是竞态条件&lt;/strong&gt;。&lt;/p&gt;
&lt;h3 id="3-技能skill知识固化"&gt;3. 技能（Skill）——知识固化&lt;/h3&gt;
&lt;p&gt;Loop 跑一次的时候，它是热的——上下文里全是当前任务的信息。但 loop 跑完，下一轮启动时，它是冷的。它不记得上一轮学到了什么。&lt;/p&gt;
&lt;p&gt;Skill 是解决这个问题的原语。它是项目特定知识的固化文件——编码规范、架构决策记录、常见陷阱列表。Skill 文件在每次 loop 启动时被注入到上下文中，让 loop 从冷启动跳到&amp;quot;有备而来&amp;quot;的状态。&lt;/p&gt;
&lt;p&gt;好的 skill 不是静态文档。它是从 loop 的运行经验中提炼出来的——一个 loop 发现一个常见错误模式，这个模式被写进 skill，下一次 loop 就不会再犯。&lt;/p&gt;
&lt;h3 id="4-子代理sub-agent分工与制衡"&gt;4. 子代理（Sub-agent）——分工与制衡&lt;/h3&gt;
&lt;p&gt;一个 loop 内部可以只有一个代理反复调用自己，但实践中更好的模式是&lt;strong&gt;分工&lt;/strong&gt;。&lt;/p&gt;
&lt;p&gt;我做代码审查的 loop 里有三个角色：一个代理写代码（maker），第二个分出五个影分身从五个正交的角度审查代码(reviewer)，最后一个代理做对抗性验证（adversarial verifier）。Maker 可以很激进——它尝试各种方案，不介意犯错。Reviewer 很保守——它从五个不同的方向逐行审查 maker 的输出，找逻辑漏洞、边界条件、风格偏离。如果 checker 发现了问题，maker 重做。&lt;/p&gt;
&lt;p&gt;Maker 和 checker 可以是同一个模型，但角色分离的 prompt 不一样，这就有用了。&lt;strong&gt;关键是检查者不能和制造者是同一个对话上下文&lt;/strong&gt;——否则模型会对自己太温柔。&lt;/p&gt;
&lt;p&gt;Adversarial Verifier 只能看到 Maker 的输出，并先入为主地认为有重大缺陷隐藏其中，他的任务就是揪出隐藏的缺陷来。&lt;/p&gt;
&lt;p&gt;更复杂的 loop 可以有更多的子代理角色：规划者、执行者、评审者、记录者。子代理越多，协调成本越高，但分工带来更专业的输出。&lt;/p&gt;
&lt;h3 id="5-连接器connector与现实世界交互"&gt;5. 连接器（Connector）——与现实世界交互&lt;/h3&gt;
&lt;p&gt;Loop 不能只在模型内部的思维空间里运行。它需要调用工具：读文件、写文件、查数据库、调 API、发消息。&lt;/p&gt;
&lt;p&gt;连接器原语处理的是&lt;strong&gt;工具如何暴露给 loop 内的代理&lt;/strong&gt;。MCP（Model Context Protocol）是 2026 年最主流的标准化方案——它定义了一个统一的接口，loop 可以通过这个接口发现和调用外部工具。虽然我颇为嫌弃 MCP 对上下文污染太多，平常很少用到。&lt;/p&gt;
&lt;p&gt;连接器的设计决策影响深远：你允许 loop 访问 shell 吗？允许它写任意文件吗？允许它调用生产环境的 API 吗？每个连接都是权限的边界。&lt;/p&gt;
&lt;h3 id="6-脊骨spine跨运行持久化"&gt;6. 脊骨（Spine）——跨运行持久化&lt;/h3&gt;
&lt;p&gt;这是最被忽视的原语，也可能是最重要的。&lt;/p&gt;
&lt;p&gt;模型在每次调用之间遗忘一切。Loop 在每次运行之间也遗忘一切——除非你有地方存状态。Spine 就是这个&amp;quot;地方&amp;quot;。它可以是一个进度文件（&lt;code&gt;progress.md&lt;/code&gt;）、一个数据库表、一个看板工具（Linear、Jira）——只要 loop 在下一次启动时可以回答&amp;quot;上次我做到哪了&amp;quot;。&lt;/p&gt;
&lt;p&gt;没有 spine 的 loop 不是一个循环——它是一个每次都从第一步开始的重复。&lt;/p&gt;
&lt;p&gt;我感觉在 loop 设计里，这个原语的缺席会导致了最隐蔽的失败模式：loop 在跑，也在产出，但产出质量不随时间增长，因为每次启动都把之前的经验丢掉了。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="三个最硬的问题"&gt;三个最硬的问题&lt;/h2&gt;
&lt;p&gt;Loop Engineering 目前有三个问题我还没有看到优雅的通用解法。似乎每个项目都得用自己的方式硬扛。&lt;/p&gt;
&lt;h3 id="上下文管理"&gt;上下文管理&lt;/h3&gt;
&lt;p&gt;随着 loop 运行轮次增加，模型看到的内容越来越多。但&amp;quot;更多&amp;quot;不等于&amp;quot;更好&amp;quot;——注意力随长度衰减，无关细节淹没关键信号。&lt;/p&gt;
&lt;p&gt;通用的应对策略有三类，各有利弊：&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;压缩&lt;/strong&gt;：把历史记录总结成更短的形式。丢失细节，但保留主线。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;剪枝&lt;/strong&gt;：直接丢弃不相关的历史。需要启发式判断什么是不相关的——启发式本身可能出错。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;外部化&lt;/strong&gt;：不让模型自己记忆，而是把状态写到一个外部文件里，模型只在需要时读取。最干净的做法，但要额外设计读写机制。&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;我目前想到的做法是混合使用：每一轮结束时自动做一次压缩（把本轮的关键决策和输出写进 spine），然后在下一轮开始时只加载压缩后的摘要，不加载完整历史。&lt;/p&gt;
&lt;h3 id="终止检测"&gt;终止检测&lt;/h3&gt;
&lt;p&gt;这个问题比看起来难十倍。&lt;/p&gt;
&lt;p&gt;一个 loop 需要知道什么时候停下来。听起来很简单——任务完成了就停。但&amp;quot;完成&amp;quot;在复杂任务中是一个很难自动判断的属性。&lt;/p&gt;
&lt;p&gt;常见的终止条件有六种：&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;模型输出了一个明确的终止信号——比如&amp;quot;任务完成&amp;quot;标记。&lt;/li&gt;
&lt;li&gt;外部验证通过——测试全绿、lint 无报错、编译通过。&lt;/li&gt;
&lt;li&gt;达到最大迭代次数——硬上限，防止无限循环。&lt;/li&gt;
&lt;li&gt;超时——时间预算耗尽。&lt;/li&gt;
&lt;li&gt;遇到不可恢复的错误——比如依赖服务挂了。&lt;/li&gt;
&lt;li&gt;振荡检测——代理反复调用同一个工具、传同样的参数、得到同样的结果。&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;第六种我花了一段时间才意识到它的存在是必要的。Agent 在遇到困难时会&amp;quot;卡住&amp;quot;——它反复尝试同一个已经证明无效的方案，每次都期待不同的结果。没有振荡检测的 loop 会在这种状态下消耗大量预算。&lt;/p&gt;
&lt;p&gt;我对终止的当前判断是：&lt;strong&gt;终止不是事后考虑，它是半个设计&lt;/strong&gt;。一个 loop 的终止条件和它的业务逻辑一样重要。如果你写 loop 的时候先写&amp;quot;做什么&amp;quot;再补&amp;quot;什么时候停&amp;quot;，你得到的一定是一个在边缘情况下会无限跑的 loop。&lt;/p&gt;
&lt;h3 id="验证信号"&gt;验证信号&lt;/h3&gt;
&lt;p&gt;Loop 需要知道自己做的事情对不对。验证信号就是这个&amp;quot;对不对&amp;quot;的输入。&lt;/p&gt;
&lt;p&gt;最好的验证信号是&lt;strong&gt;确定性的&lt;/strong&gt;：测试通过、类型检查通过、编译成功。这些信号是 gold standard——它们客观、可重复、不会说谎。&lt;/p&gt;
&lt;p&gt;但当任务没有确定性验证可用时（比如写文档、做设计决策），可以退而求其次用&lt;strong&gt;模型作为评审者&lt;/strong&gt;（LLM-as-Judge）。这是一个有用的技术，但它有一个结构性问题：写代码的模型评价自己写的代码太温柔了。不是说它会作弊——而是它会在看到自己熟悉的模式时放松标准。&lt;/p&gt;
&lt;p&gt;解决方法我在 maker–checker 前面提到了：制造者和检查者必须是&lt;strong&gt;不同的代理实例&lt;/strong&gt;——最好是不同的上下文、不同的 prompt 体系、甚至不同的模型。这不是不信任——这是架构约束。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="两个陷阱效率的暗面"&gt;两个陷阱：效率的暗面&lt;/h2&gt;
&lt;p&gt;Loop 越强大，越容易滑入两个陷阱。我说&amp;quot;陷阱&amp;quot;是因为它们不是技术问题——它们是你自己的认知在习惯 loop 之后产生的盲区。&lt;/p&gt;
&lt;h3 id="陷阱一检查仍然是你的工作"&gt;陷阱一：检查仍然是你的工作&lt;/h3&gt;
&lt;p&gt;Loop 让产出变得更快、更自主。但&amp;quot;完成&amp;quot;是代理的主张，不是证明。一个 loop 说&amp;quot;所有测试通过&amp;quot;——你真的看过测试覆盖了正确的东西吗？一个 loop 说&amp;quot;重构完成&amp;quot;——你真的确认过架构约束没有被绕过吗？&lt;/p&gt;
&lt;p&gt;这不是说你不该信任 loop。我信任我的 loop 去做它们被设计来做的事情。但当 loop 的输出要进入生产环境时，&lt;strong&gt;审查仍然是你的工作&lt;/strong&gt;。和人类同事写的代码一样——你信任他们，但你还是做 code review。Loop 生成的代码也是同样的逻辑。&lt;/p&gt;
&lt;p&gt;区别在于：因为 loop 产出太流畅了，你很容易跳过审查这一环。流畅性本身是一个认知陷阱。&lt;/p&gt;
&lt;h3 id="陷阱二不要停止深度思考"&gt;陷阱二：不要停止深度思考&lt;/h3&gt;
&lt;p&gt;这是更隐蔽的一个。&lt;/p&gt;
&lt;p&gt;Loop 越快地产出你&lt;strong&gt;没有亲手写&lt;/strong&gt;的代码，你的项目能力和你的实际理解之间的差距就越大。你看着一个 loop 重构了整个模块——但你真的能解释重构后的每一行在做什么吗？如果不能，你是模块的拥有者还是旁观者？&lt;/p&gt;
&lt;p&gt;我在自己身上看到过这个现象 —— 开始依赖 ClaudeCode/Codex 后，我对代码库的细节掌握明显变薄弱了。不是因为我变懒了 —— 而是因为 AI 代理处理得太好，我没有&amp;quot;被迫理解&amp;quot;的机会了。&lt;/p&gt;
&lt;p&gt;这不是反对使用 loop。一个人熟练的工程师用 loop 可以产出比手动高十倍的效率。但那是因为&lt;strong&gt;在 loop 出现之前，他已经理解了&lt;/strong&gt;。他设计 loop 时知道 loop 在做什么，知道边界在哪里，知道什么情况下 loop 的假设会失效。&lt;/p&gt;
&lt;p&gt;这个陷阱真正的风险是&lt;strong&gt;时间差&lt;/strong&gt;：如果你在还没有深入理解一个领域的时候就开始用 loop 替你工作，你永远不会有那个&amp;quot;被迫理解&amp;quot;的阶段。loop 把学习的窗户关掉了。&lt;/p&gt;
&lt;p&gt;我目前的应对是：对于我熟悉的领域，loop 全速前进。对于我不熟悉的领域，我先手动做几次，理解之后再写 loop。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="两条实现路径工具选择与工程哲学"&gt;两条实现路径：工具选择与工程哲学&lt;/h2&gt;
&lt;p&gt;2026 年中，想实践 Loop Engineering 的人面对两个主流选择：Claude Code 和 OpenCode。它们都做 loop——但背后的哲学差异很大。&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th style="text-align: left"&gt;维度&lt;/th&gt;
&lt;th style="text-align: left"&gt;Claude Code&lt;/th&gt;
&lt;th style="text-align: left"&gt;OpenCode&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="text-align: left"&gt;心跳机制&lt;/td&gt;
&lt;td style="text-align: left"&gt;内置——&lt;code&gt;/loop&lt;/code&gt;、&lt;code&gt;/goal&lt;/code&gt;、Cloud Routines&lt;/td&gt;
&lt;td style="text-align: left"&gt;不内置——用 cron、GitHub Actions 自行调度&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: left"&gt;工作模式&lt;/td&gt;
&lt;td style="text-align: left"&gt;开箱即用的 loop 原语&lt;/td&gt;
&lt;td style="text-align: left"&gt;&lt;code&gt;opencode run&lt;/code&gt; 是一个 worker，loop 你自己写&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: left"&gt;上手门槛&lt;/td&gt;
&lt;td style="text-align: left"&gt;低——内置编排，几分钟能跑&lt;/td&gt;
&lt;td style="text-align: left"&gt;高——需要理解每个零件的原理&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: left"&gt;厂商锁定&lt;/td&gt;
&lt;td style="text-align: left"&gt;部分存在——Cloud Routines 跑在 Anthropic 服务器上&lt;/td&gt;
&lt;td style="text-align: left"&gt;零——全部本地可控&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: left"&gt;动态工作流&lt;/td&gt;
&lt;td style="text-align: left"&gt;内置模版和路由逻辑&lt;/td&gt;
&lt;td style="text-align: left"&gt;你的 shell 脚本就是工作流&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: left"&gt;透明性&lt;/td&gt;
&lt;td style="text-align: left"&gt;封装好的抽象——好用但难改&lt;/td&gt;
&lt;td style="text-align: left"&gt;裸金属——难用但全可控&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;这个对比的背后是一个更深的选择：&lt;strong&gt;你想要一辆造好的车，还是一个发动机让你自己造车？&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Claude Code 适合那些想快速体验 loop engineering、不想在基础设施上花时间的团队。它的 loop 原语是内置的，心跳开箱即用，子代理分工通过 &lt;code&gt;/loop&lt;/code&gt; 和 &lt;code&gt;/goal&lt;/code&gt; 就能实现。你付出的代价是：一旦内置的 loop 模式不满足你的需求，修改起来很困难。&lt;/p&gt;
&lt;p&gt;OpenCode 适合那些需要完全控制 loop 行为的团队。你就是 loop 的设计者——心跳自己写，spine 自己选，子代理的调度自己实现。代价是：你要把六个原语全部手动组装一遍。上手慢，但上限高。&lt;/p&gt;
&lt;p&gt;我目前倾向 OpenCode 路线。不是因为它更好——而是因为 loop engineering 还在快速演化，我不想把我的 loop 架构绑定在一个厂商的特定实现上。等模式稳定了，再考虑更高的抽象层。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="边界条件"&gt;边界条件&lt;/h2&gt;
&lt;p&gt;本文讨论的 loop 模式在以下情况下可能不适用或需要大幅调整：&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;单次任务&lt;/strong&gt;：如果你只需要让模型做一件事一次，loop 是过度设计。写一个 prompt 就够了。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;高精度要求&lt;/strong&gt;：目前 loop 的验证机制在需要人类级判断的任务（法律文件审核、医疗建议）上不够可靠。确定性验证覆盖不到的地方，loop 的可靠性取决于你愿意接受多少错误率。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;极低延迟要求&lt;/strong&gt;：Loop 涉及多轮模型调用，每次调用几百毫秒到几秒。对实时交互有要求的场景，loop 不是一个合适的选择。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;团队规模太小&lt;/strong&gt;：一个人的团队写 loop 可能得不偿失 —— 维护 loop 本身有成本。当团队有三个以上工程师复用同一个 loop 时，这个成本开始物有所值。当然我准备在自己的小项目上实验下 loop，就顾不得这个了。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;语言或文化障碍&lt;/strong&gt;：Loop Engineering 的当前最佳实践和工具文档以英文为主。非母语团队的采纳曲线会更陡。&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;组织尚未准备好&lt;/strong&gt;：如果团队的工程文化不支持自动化、不信任自主系统、或者管理方式要求每步都有人签字，loop 的收益会被组织摩擦吃掉。&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="开放问题"&gt;开放问题&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Loop 的可观测性&lt;/strong&gt;：当一个 loop 跑了 48 小时，调用了上千次模型，做了几百次文件修改——你怎么知道它做得对不对？目前的可观测性工具（trace、log、metric）是为确定性系统设计的，不是为自主决策系统设计的。这个缺口怎么补？&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;验证泛化&lt;/strong&gt;：确定性验证（测试、lint）覆盖不了的任务中，LLM-as-Judge 的可靠性到底有多高？我还没有看到严格的、跨场景的基准测试。&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;跨 loop 状态共享&lt;/strong&gt;：多个 loop 同时在同一个项目上运行时，它们怎么共享上下文而不会冲突？目前的做法是各自写各自的任务上下文，但&amp;quot;总览全局&amp;quot;的能力在 loop 之间是缺失的。&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Skill 的自改进周期&lt;/strong&gt;：当 loop 从运行中学习到新经验时，它怎么自动更新自己的 skill 文件？这个元循环（改进 loop 的 loop）的实现目前还没有好的通用模式。这个感觉跟 Hermes Agent 有点像。&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;收敛性保证&lt;/strong&gt;：多代理 loop（规划→执行→评审→记录）在复杂长周期任务中，是否存在收敛性？有没有一种情况是执行者和评审者反复来回修改同一段代码永远无法达成一致？理论上是的，实际中我见过。什么条件下它一定会收敛？&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;[[Q]] 一年后重读这篇文章时，我会仍然认同&amp;quot;Loop Engineering 是一次范式跃迁&amp;quot;这个框架，还是会觉得它只是 Harness Engineering 的一个子集？&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="参考文献"&gt;参考文献&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Addy Osmani, &amp;ldquo;Loop Engineering&amp;rdquo;, 2026 年 6 月。六个原语的首次系统化定义。&lt;/li&gt;
&lt;li&gt;Boris Cherny（Claude Code 创作者），关于&amp;quot;我不再写 prompt，我写 loop&amp;quot;的访谈，2026 年 5 月。&lt;/li&gt;
&lt;li&gt;Sydney Runkle（LangChain），&amp;ldquo;The Art of Loop Engineering&amp;rdquo;，2026 年 6 月。四层循环栈模型的来源。&lt;/li&gt;
&lt;li&gt;Richmond Alake（Oracle），&amp;ldquo;Agent Loop Decoded&amp;rdquo;，2026 年 6 月。Agent = Model + Harness 公式和三级成熟度模型。&lt;/li&gt;
&lt;li&gt;Peter Steinberger（OpenClaw），关于&amp;quot;技能已转移到 loop 设计&amp;quot;的帖文，2026 年 6 月。&lt;/li&gt;
&lt;/ol&gt;</description></item><item><title>寻找复数网络的并行扫描 —— 从 Mamba 的训练哲学到复数序列模型的高效训练</title><link>http://fengwang.github.io/posts/20260614-complex-parallel-scan/</link><pubDate>Sun, 14 Jun 2026 00:00:00 +0000</pubDate><guid>http://fengwang.github.io/posts/20260614-complex-parallel-scan/</guid><description>&lt;p&gt;在一遍遍阅读 Mamba 论文之后，我第一次清晰地意识到：Mamba 训练成功的核心不在于它的选择机制，也不在于它的门控架构。而在于一个更底层、也更数学的事实——&lt;strong&gt;它找到了一种将整条序列一次喂入模型、并行算出所有时间步的输出、然后统一反向传播更新参数的方法&lt;/strong&gt;。&lt;/p&gt;
&lt;p&gt;这个能力让它和 Transformer 站在一起，和 RNN 分道扬镳。&lt;/p&gt;
&lt;p&gt;而我正在探索的复数神经网络，如果不能在这一点上有所突破，就永远会陷入 RNN 的命运：理论上优美，实践上低效。这篇文章是我对这个问题的当前思考的记录。&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;核心论点：复数线性递推的数学结合律与实数情况完全一致，因此Mamba的并行关联扫描算法在复数域仍然成立；主要障碍不在数学，而在硬件生态——缺乏针对复数的高效融合CUDA算子。&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope:&lt;/strong&gt; 本文讨论从 Mamba 训练方式中提炼的&amp;quot;并行扫描&amp;quot;方法论能否迁移到复数神经网络。覆盖数学论证和工程挑战。不讨论特定复数网络架构（如复数 RNN、复数 LSTM 的具体设计）的细节，也不提供完整实现。&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prerequisites:&lt;/strong&gt; 假定你了解 Mamba/SSM 的基本架构方向，理解 PyTorch 自动求导的基本机制，并熟悉复数的基础运算性质。&lt;/p&gt;
&lt;h2 id="一个问题两种命运"&gt;一个问题，两种命运&lt;/h2&gt;
&lt;p&gt;我花了很长时间才真正理解为什么 RNN 和 Mamba 在数学上都是递推，训练效率却有天壤之别。&lt;/p&gt;
&lt;p&gt;RNN 的训练困局可以用一句话概括：&lt;strong&gt;每前向传播一步，反向传播就必须退一步&lt;/strong&gt;。对于长度为 $L$ 的序列，这意味着 $O(L)$ 的串行深度。梯度要沿着时间轴倒着走回序列起点，每一个隐藏状态 $h_t$ 都必须被保存下来供反向传播使用。结果是 $O(L \cdot d^2)$ 的内存和无法并行的计算图。&lt;/p&gt;
&lt;p&gt;Mamba 的训练方式完全不同。它走完 $L$ 步前向传播，算出总损失，然后统一做一次反向传播。乍看起来这好像和 RNN 没什么区别——不都是走完再更新吗？但关键在于 Mamba &lt;strong&gt;不走那 $L$ 步串行循环&lt;/strong&gt;。它用并行关联扫描在一棵深度 $O(\log L)$ 的树里一次性算出所有 $h_1, \dots, h_L$，然后用重计算技术在反向传播时现场重新生成这些状态，避免将它们写入全局显存。&lt;/p&gt;
&lt;p&gt;我读到这个机制的时候，心里冒出一个问题：&lt;strong&gt;复数版本呢？&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;复数神经网络的价值在理论上已经被反复论证——复数表示的欠阻尼动态系统有更丰富的表达力、复数递推对旋转和振荡建模更自然、复数梯度流具有与实数不同的特性。但所有这一切，都被训练效率的阴影笼罩着。如果复数 RNN 的训练还是只能 $O(L)$ 串行，它在实际应用中的天花板就是一层天花板。&lt;/p&gt;
&lt;p&gt;但等一下。如果 Mamba 的并行扫描依赖的是&lt;strong&gt;递推方程的线性性质&lt;/strong&gt;和由此带来的&lt;strong&gt;结合律&lt;/strong&gt;，而不是递推系数的实数性，那么复数递推方程满足同样的性质。这意味着——至少在数学上——并行扫描对于复数网络是直接可用的。&lt;/p&gt;
&lt;h2 id="并行扫描的数学本质"&gt;并行扫描的数学本质&lt;/h2&gt;
&lt;p&gt;Mamba 能并行的根本原因不在硬件优化，不在 CUDA 编程，而在一个极其朴素的数学性质：&lt;strong&gt;结合律&lt;/strong&gt;。&lt;/p&gt;
&lt;p&gt;我们来写出 Mamba 的递推方程：&lt;/p&gt;
&lt;p&gt;$$h_t = \bar{A}&lt;em&gt;t h&lt;/em&gt;{t-1} + \bar{B}_t x_t$$&lt;/p&gt;
&lt;p&gt;这个方程里，$\bar{A}_t$ 是一个 $N \times N$ 矩阵，$\bar{B}_t x_t$ 是一个 $N$ 维向量。我们可以把&amp;quot;从 $t-1$ 到 $t$ 的一步转移&amp;quot;定义为一个操作符 $O_t = (\bar{A}_t, \bar{B}&lt;em&gt;t x_t)$。这个操作符作用于 $h&lt;/em&gt;{t-1}$ 的方式是：&lt;/p&gt;
&lt;p&gt;$$O_t(h_{t-1}) = \bar{A}&lt;em&gt;t h&lt;/em&gt;{t-1} + \bar{B}_t x_t$$&lt;/p&gt;
&lt;p&gt;关键来了：相邻两步可以合并。将 $h_t = O_t(h_{t-1})$ 代入 $h_{t+1} = O_{t+1}(h_t)$，得到：&lt;/p&gt;
&lt;p&gt;$$h_{t+1} = (\bar{A}&lt;em&gt;{t+1} \bar{A}&lt;em&gt;t) h&lt;/em&gt;{t-1} + (\bar{A}&lt;/em&gt;{t+1} \bar{B}&lt;em&gt;t x_t + \bar{B}&lt;/em&gt;{t+1} x_{t+1})$$&lt;/p&gt;
&lt;p&gt;注意结果保持了完全相同的数学形式——一个矩阵乘前一个状态，再加一个向量。因此我们可以定义合并操作 $\otimes$：&lt;/p&gt;
&lt;p&gt;$$O_{t+1} \otimes O_t = (\bar{A}&lt;em&gt;{t+1} \bar{A}&lt;em&gt;t, ; \bar{A}&lt;/em&gt;{t+1} \bar{B}&lt;em&gt;t x_t + \bar{B}&lt;/em&gt;{t+1} x&lt;/em&gt;{t+1})$$&lt;/p&gt;
&lt;p&gt;这个 $\otimes$ 满足结合律：$(O_3 \otimes O_2) \otimes O_1 = O_3 \otimes (O_2 \otimes O_1)$。这意味着我们可以用树状归约一次性合并所有时间步的操作符，然后用 $O(\log L)$ 深度的并行计算出所有 $h_t$。&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;这个推导自始至终没有用到任何实数特有的性质。&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;如果 $\bar{A}_t$ 和 $\bar{B}_t$ 是复数矩阵，$h_t$ 是复数向量，$x_t$ 可能是复数输入，那么：&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;复数矩阵乘法 $\bar{A}_{t+1} \bar{A}_t$ 满足结合律&lt;/li&gt;
&lt;li&gt;复数向量加法 $\bar{A}&lt;em&gt;{t+1} \bar{B}&lt;em&gt;t x_t + \bar{B}&lt;/em&gt;{t+1} x&lt;/em&gt;{t+1}$ 结合律也成立&lt;/li&gt;
&lt;li&gt;整个 $\otimes$ 操作符的推导一行都不需要改&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;我反复确认了几遍，结论是一样的：&lt;strong&gt;在数学层面，复数线性递推的并行扫描条件完全满足。&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="一个重要细节选择性机制的复数版本"&gt;一个重要细节：选择性机制的复数版本&lt;/h2&gt;
&lt;p&gt;Mamba 的选择性机制让 $\bar{A}_t$ 和 $\bar{B}_t$ 成为输入 $x_t$ 的函数。在实数域，这通过线性层 $s_B(x_t) = \text{Linear}_N(x_t)$ 等来实现。&lt;/p&gt;
&lt;p&gt;对于复数网络，我目前认为有两种方案：&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;方案一：复数值线性层&lt;/strong&gt;。将标准线性层的权重和偏置替换为复数版本，直接输出复数值的 $\bar{B}_t$ 和 $\bar{C}_t$。这样做最直接，但需要实现复数版本的线性层及其反向传播。&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;方案二：实数映射保持&lt;/strong&gt;。参数保持为实数，通过将实数参数重新解释为复数的实部和虚部来构造复数参数，或者用 $2N$ 维的实数向量表示 $N$ 维复数向量。这种方式不需要修改现有框架的复数支持，但理论上丧失了复数表示的一些结构优势。&lt;/p&gt;
&lt;p&gt;我倾向于方案一，因为复数线性层的实现本身并不复杂（只是一个复数矩阵乘法），而且它保持了复数网络的全部表达力。但方案二的一个实际优点是：它可以直接复用现有的实数优化器和 CUDA 库。&lt;/p&gt;
&lt;p&gt;这里的 $\Delta_t$ 离散化步长处理更微妙。Mamba 中 $\Delta_t$ 是一个正实数（通过 $\text{softplus}$ 确保正性），经过 ZOH 离散化：&lt;/p&gt;
&lt;p&gt;$$\bar{A}_t = \exp(\Delta_t A)$$&lt;/p&gt;
&lt;p&gt;如果 $A$ 是复数矩阵，$\exp(\Delta_t A)$ 在数学上定义良好（矩阵指数），但在底层实现上需要复数矩阵指数运算，这比实数版本贵不少。这不是一个原则性障碍——复数矩阵指数已经被广泛研究和实现——但它意味着在实际训练中不能直接照搬 Mamba 的 CUDA 核。&lt;/p&gt;
&lt;h2 id="硬件落地的挑战"&gt;硬件落地的挑战&lt;/h2&gt;
&lt;p&gt;数学上没有问题，落地就完全不是一回事了。我目前能看到三个层面的障碍，按困难程度排序。&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;第一层：CUDA 内核支持。&lt;/strong&gt; Mamba 的并行扫描之所以快，是因为它有一个定制的融合 CUDA 核（fused kernel），将离散化、并行扫描、输出投影合并为一个操作，全部在 SRAM 中完成，避免 $(B, L, D, N)$ 的中间张量写入 HBM。&lt;/p&gt;
&lt;p&gt;对于复数版本，我需要的基础操作（复数矩阵乘法、复数向量加法）本身在 CUDA 中都是支持的——cuBLAS 的 &lt;code&gt;cgemm&lt;/code&gt; 和 &lt;code&gt;zgemm&lt;/code&gt; 从第一代 CUDA 就存在。问题在于 &lt;strong&gt;融合&lt;/strong&gt;。Mamba 的核不是简单的 Scan，而是将多个复数域操作串联成一个融合操作。现有的 cuBLAS 函数做不到这种自定义的融合，我需要自己写 CUDA 核，以及对应的反向传播核和重计算逻辑。&lt;/p&gt;
&lt;p&gt;我不确定这样做的开发成本有多高。但有一个折中方案：退回到使用 PyTorch 的 &lt;code&gt;associative_scan&lt;/code&gt; 或类似的高层接口。&lt;code&gt;torch.vmap&lt;/code&gt; 和 &lt;code&gt;torch.scan&lt;/code&gt; 可能提供一部分支持，但效率一定比不上定制核。对于原型验证（proof of concept），这已经够了。&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;第二层：离散化的复数处理。&lt;/strong&gt; Mamba 使用 ZOH（零阶保持）离散化：&lt;/p&gt;
&lt;p&gt;$$\bar{A}_t = \exp(\Delta_t A), \quad \bar{B}_t = (\Delta_t A)^{-1}(\exp(\Delta_t A) - I) \cdot \Delta_t B_t$$&lt;/p&gt;
&lt;p&gt;当 $A$ 是复数矩阵时，$\exp(\Delta_t A)$ 的计算代价高于实数版本。一个常用的简化是使用欧拉离散化 $\bar{A}_t = I + \Delta_t A$，这在 $|\Delta_t A| \ll 1$ 时效果不错。对于复数网络，这个近似是否成立取决于特征值 $\lambda(A)$ 的模长和 $\Delta_t$ 的取值范围。&lt;/p&gt;
&lt;p&gt;另一种思路是使用&lt;strong&gt;双线性变换&lt;/strong&gt;（Tustin&amp;rsquo;s method），它在复数域同样有效，且数值稳定性优于欧拉法。&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;第三层：复数的反向传播。&lt;/strong&gt; PyTorch 对复数自动求导的支持已经相当成熟（从 1.9 版本开始），支持 &lt;code&gt;torch.complex&lt;/code&gt; 和 &lt;code&gt;torch.complex128&lt;/code&gt; 的梯度计算。但我没有测试过它对复杂计算图（如并行扫描的重计算模式）的处理效率。自定义反向传播函数（&lt;code&gt;torch.autograd.Function&lt;/code&gt;）可能是必要的。&lt;/p&gt;
&lt;p&gt;总结一下硬件层面的判断：&lt;strong&gt;prototype 可行，性能优化的门槛不低，但不存在原则性障碍。&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="值得借鉴的现有工作"&gt;值得借鉴的现有工作&lt;/h2&gt;
&lt;p&gt;在做这个调研之前，我以为复数序列模型的高效训练是一个冷门方向。查了一圈发现情况比想象中好。&lt;/p&gt;
&lt;p&gt;最直接相关的是一系列 **复数状态空间模型 Complex-valued SSM ** 的工作。Mamba-3 和最新的一些 SSM 变体已经显式地使用了复数值状态。Ran-Milo &amp;amp; Cohen (NeurIPS 2024) 更从理论上严格证明了复数参数化对 SSM 的表达优势：复数 SSM 在中等维度下就能表达实数 SSM 的所有映射，但反之实数 SSM 需要维度达到 $O(t)$（$t$ 为时间长度）；即使维度够高，实数 SSM 的参数值也可能需要指数级大才能表达某些振荡映射，使得它们无法在实际中学习。这为复数 SSM 的研究提供了坚实的理论基础。Mamba 论文本身提到：对于连续模态（音频、视频），复数状态比实数表现更好；对于离散模态（文本、DNA），实数反而更优。这说明复数 SSM 在特定场景下是有明确需求的。&lt;/p&gt;
&lt;p&gt;这些工作中的 CUDA 实现主要集中在解决复数矩阵指数和复数并行扫描的数值稳定性问题。我没有找到直接对标 Mamba 融合核的复数版本，但有一些零散的 CUDA 部件可以借用。&lt;/p&gt;
&lt;p&gt;另一个方向更间接但同样重要：&lt;strong&gt;复数 RNN 的训练加速&lt;/strong&gt;。虽然传统复数 RNN 没有 Mamba 的并行扫描能力，但最近的工作尝试用梯度裁剪、正交初始化、复数值动量等方法缓解复数 RNN 的训练困难。这些工作虽然不改变串行训练的本质，但提供了一些在无法彻底并行化时的实用改进。比如对复数 RNN 使用基于 Wirtinger 微积分的梯度下降，其收敛性质比简单的实数梯度拆分要好。&lt;/p&gt;
&lt;p&gt;还有一个值得关注的线索是&lt;strong&gt;复数 FFN 和复数卷积的高效实现&lt;/strong&gt;。在一些复数图像分类的工作中，&amp;ldquo;复数计算太慢&amp;quot;的问题通过 cuFFT（快速傅里叶变换）的复用得到了部分缓解——复数卷积用 FFT 计算天然比实数卷积更高效，因为 FFT 本身就是复数运算。虽然这和序列模型的并行扫描不是同一个东西，但它说明复数计算在某些场景下反而有硬件优势。也许复数 SSM 在某些配置下也能用 FFT 加速卷积模式，结合并行扫描获得双重效率提升。&lt;/p&gt;
&lt;h2 id="where-this-breaks"&gt;Where this breaks&lt;/h2&gt;
&lt;p&gt;这个论证目前有几个明显的薄弱环节。我在写的时候就意识到这些问题，写下来供未来自己验证或推翻。&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;复数矩阵指数是昂贵的。&lt;/strong&gt; ZOH 离散化中的 $\exp(\Delta_t A)$ 对于复数矩阵的计算代价远高于实数，特别是在状态维度 $N$ 较大且 $\bar{A}_t$ 每步不同（选择机制）的情况下。欧拉近似或双线性变换可能是必要的简化，它们的近似误差需要被量化。&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;选择性机制和静态结合的兼容性未验证。&lt;/strong&gt; 我上面的推导假设了给定输入序列后，$\bar{A}_t$ 和 $\bar{B}_t$ 被确定下来，然后并行扫描对这些确定的（可能是复数的）系数进行合并。但选择性机制本身（从 $x_t$ 求 $\bar{A}_t, \bar{B}_t$）是否在复数域同样稳定？复数线性层的训练收敛性质与实数不同，这可能影响整个流水线。&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;硬件效率的定量估算缺失。&lt;/strong&gt; 我说&amp;quot;prototype 可行，性能优化有门槛&amp;quot;但给不出数字。一个完整的复数 SSM 层和实数 SSM 层的吞吐量对比，以及其中离散化步骤的开销占比，需要实际跑实验才能知道。&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;SRAM 容量限制可能会更紧。&lt;/strong&gt; 复数占用双倍于实数的存储（float32 vs complex），在 SRAM 这个稀缺资源上，这是一个紧约束。Mamba 的融合核成功的关键之一是状态维度 $N$ 和特征维度 $D$ 的乘积能放进 SRAM。复数版本用 complex 会让占用翻倍，这意味着更小的 $N$ 或 $D$，或者更小的块大小。&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="open-questions"&gt;Open questions&lt;/h2&gt;
&lt;p&gt;写完之后留下的问题比写之前更多。列在这里，希望六个月后的我能回答其中的一些。&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q1:&lt;/strong&gt; 复数并行扫描的 CUDA 融合核开发成本到底有多大？能否基于现有的开源 SSM 代码库（如 state-spaces 仓库）修改，还是需要从零编写？&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q2:&lt;/strong&gt; 对于复数状态 $h_t$ 而言，离散化误差和数值稳定性与实数版本相比如何？特别是当 $\bar{A}_t$ 的特征值靠近单位圆时。&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q3:&lt;/strong&gt; 在复数域，Mamba 的&lt;strong&gt;选择机制&lt;/strong&gt;是否仍然有效？或者说，额外的复数自由度是否会带来选择性的一种&amp;quot;自然涌现&amp;rdquo;？&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q4:&lt;/strong&gt; 是否有某些任务（如音频合成、振荡动力学建模、量子系统模拟）中，复数 SSM 的 O(log L) 并行训练能力能够带来现实可测量的收益？&lt;/p&gt;
&lt;p&gt;[[Q]] 六个月后回看：复数并行扫描的数学可行性我已经确认了，但你实际动手验证了吗？第一个基线实验跑出了什么结果？&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Gu, Dao, &amp;ldquo;Mamba: Linear-Time Sequence Modeling with Selective State Spaces&amp;rdquo;, 2023.&lt;/li&gt;
&lt;li&gt;Gu, Goel, Ré, &amp;ldquo;Efficiently Modeling Long Sequences with Structured State Spaces (S4)&amp;rdquo;, ICLR 2022.&lt;/li&gt;
&lt;li&gt;Gu et al., &amp;ldquo;HiPPO: Recurrent Memory with Optimal Polynomial Projections&amp;rdquo;, NeurIPS 2020.&lt;/li&gt;
&lt;li&gt;Blelloch, &amp;ldquo;Prefix Sums and Their Applications&amp;rdquo;, 1990.&lt;/li&gt;
&lt;li&gt;Gu et al., &amp;ldquo;Modeling Sequences with Structured State Spaces&amp;rdquo;, Stanford PhD Dissertation, 2023.&lt;/li&gt;
&lt;li&gt;Arjovsky, Shah, Bengio, &amp;ldquo;Unitary Evolution Recurrent Neural Networks&amp;rdquo;, ICML 2016.&lt;/li&gt;
&lt;li&gt;Wisdom et al., &amp;ldquo;Full-Capacity Unitary Recurrent Neural Networks&amp;rdquo;, NeurIPS 2016.&lt;/li&gt;
&lt;li&gt;Trabelsi et al., &amp;ldquo;Deep Complex Networks&amp;rdquo;, ICLR 2018.&lt;/li&gt;
&lt;li&gt;Ran-Milo, Cohen et al., &amp;ldquo;Provable Benefits of Complex Parameterizations for Structured State Space Models&amp;rdquo;, NeurIPS 2024.&lt;/li&gt;
&lt;/ol&gt;</description></item><item><title>Two ways to diffuse text: DFlash block diffusion vs DiffusionGemma</title><link>http://fengwang.github.io/posts/dflash-vs-diffusiongemma/</link><pubDate>Wed, 10 Jun 2026 00:00:00 +0000</pubDate><guid>http://fengwang.github.io/posts/dflash-vs-diffusiongemma/</guid><description>&lt;h2 id="why-this-exists"&gt;Why this exists&lt;/h2&gt;
&lt;p&gt;I&amp;rsquo;ve spent the whole day reading papers at the intersection of diffusion models and text generation, and two names kept surfacing in different contexts: DFlash and DiffusionGemma. Both apply diffusion to discrete text tokens, but they solve completely different problems. DFlash uses a tiny block diffusion model as a fast draft generator inside a speculative decoding loop. DiffusionGemma is a full-scale text diffusion model that replaces the autoregressive decoder entirely. I kept getting them confused — which one uses bidirectional attention? Which one generates 256 tokens at once? Which one guarantees lossless output? This article works through the details side by side so I don&amp;rsquo;t have to reconstruct the comparison from scratch next time.&lt;/p&gt;
&lt;p&gt;Autoregressive language models generate one token at a time. That serial dependency makes them memory-bound at small batch sizes — most of the time goes to loading weights, not computing. Diffusion offers a way out: generate many tokens in parallel and refine them iteratively. DFlash and DiffusionGemma are two concrete implementations of this idea, but they occupy opposite ends of the system-design spectrum.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope:&lt;/strong&gt; This covers the mechanism of DFlash&amp;rsquo;s block diffusion drafter and DiffusionGemma&amp;rsquo;s text diffusion model, then compares them across architecture, training, inference, and use case. It does not cover non-diffusion speculative decoding methods (EAGLE-3, Medusa) or image/video diffusion models.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prerequisites:&lt;/strong&gt; This assumes familiarity with autoregressive language model basics (causal attention, tokens, logits), the speculative decoding pattern (a small draft model proposes tokens that a large target model verifies), and the general diffusion concept (reverse process that iteratively removes noise).&lt;/p&gt;
&lt;h2 id="the-bottleneck-that-diffusion-addresses"&gt;The bottleneck that diffusion addresses&lt;/h2&gt;
&lt;p&gt;Autoregressive decoding has a fundamental throughput problem. For each token, the model must load its full set of weights from memory into compute units, do a forward pass, and output a single logit vector. At small batch sizes — typical for interactive applications — the memory-bandwidth bottleneck dominates: the compute units sit idle waiting for weights to arrive. Batching many requests together amortizes this cost, but it does nothing for single-request latency.&lt;/p&gt;
&lt;p&gt;Diffusion models flip this dynamic. Instead of one token at a time, they generate an entire block of tokens in parallel during each denoising step. The computation shifts from memory-bandwidth-bound to compute-bound. The question is how to make this work for text, where tokens are discrete and the sequential structure of language matters.&lt;/p&gt;
&lt;p&gt;Two distinct answers have emerged. One treats diffusion as a fast approximator inside an existing autoregressive system. The other treats diffusion as the primary generation mechanism, replacing the autoregressive decoder entirely. DFlash is the first approach. DiffusionGemma is the second.&lt;/p&gt;
&lt;h2 id="dflash-block-diffusion-as-a-speculative-drafter"&gt;DFlash: block diffusion as a speculative drafter&lt;/h2&gt;
&lt;p&gt;DFlash was developed at UC San Diego and published at ICML 2026. Its core insight is simple: the hidden states of a large autoregressive language model already encode information about multiple future tokens. Rather than training a separate draft model to predict tokens from scratch, DFlash extracts these hidden features and uses them to condition a tiny block diffusion model.&lt;/p&gt;
&lt;h3 id="architecture"&gt;Architecture&lt;/h3&gt;
&lt;p&gt;The draft model is a shallow bidirectional Transformer — typically 5 layers — that operates on blocks of $\gamma = 16$ tokens. The generation process works in five steps:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Feature extraction.&lt;/strong&gt; During the target model&amp;rsquo;s prefill pass, DFlash extracts hidden representations from a fixed set of layers uniformly sampled from shallow to deep (e.g., layers 3, 10, 17, 24, 31 of a 32-layer model). These hidden states capture information at different levels of abstraction.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Feature fusion and injection.&lt;/strong&gt; The extracted features are concatenated and passed through a lightweight projection layer that fuses cross-layer information into a compact target context feature. This feature is then injected directly into the Key and Value projections of every draft model layer. The result is stored in the draft model&amp;rsquo;s KV cache and persists across drafting iterations.&lt;/p&gt;
&lt;p&gt;This persistent per-layer injection is the key design choice. Prior work like EAGLE-3 fuses target features only at the input layer, letting the signal dilute as the draft model deepens. DFlash&amp;rsquo;s approach keeps the conditioning strong regardless of draft depth.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Parallel diffusion drafting.&lt;/strong&gt; Starting from $\gamma$ random token embeddings, the draft model performs 1-3 bidirectional denoising steps. Unlike autoregressive drafters that must run $\gamma$ sequential forward passes, the diffusion drafter generates all $\gamma$ tokens in a single forward pass per step. Draft latency is nearly independent of $\gamma$.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Parallel verification.&lt;/strong&gt; The target model processes all $\gamma$ draft tokens in a single forward pass, computing acceptance probabilities for each prefix position. Because the attention computation is parallel within a single sequence, verifying $\gamma$ tokens costs barely more than generating one.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Token acceptance.&lt;/strong&gt; The longest prefix where all tokens match the target model&amp;rsquo;s distribution is accepted. One additional &amp;ldquo;bonus&amp;rdquo; token is generated for free at the acceptance boundary, and the process repeats from step 1.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="training"&gt;Training&lt;/h3&gt;
&lt;p&gt;The draft model shares the target model&amp;rsquo;s token embedding and language modeling head — only the draft Transformer layers are trained. Training data comes from response text: random anchor tokens are selected, a contiguous block of $\gamma$ tokens is masked with noise, and the model learns to denoise the entire block in one shot. A position-dependent exponential decay loss weights early tokens more heavily — if the first draft token is wrong, all subsequent tokens in the block are wasted, so the model should prioritize correctness at the beginning of the block.&lt;/p&gt;
&lt;p&gt;Multiple blocks can be trained in a single forward pass using sparse attention masks (Flex Attention), with cross-block attention disabled. This makes training efficient despite the block structure.&lt;/p&gt;
&lt;h3 id="why-it-works"&gt;Why it works&lt;/h3&gt;
&lt;p&gt;DFlash achieves the best of both worlds. The target model provides high-quality next-token predictions; the block diffusion drafter generates many candidates at near-zero marginal latency. The combination yields lossless 5-6x speedups on models like Qwen3-8B, compared to the 2-3x ceiling of autoregressive drafters like EAGLE-3. On Chain-of-Thought reasoning tasks, where sequences are long and the acceptance distribution matters more, the speedup drops to 4-5x — still far above what AR drafters achieve.&lt;/p&gt;
&lt;p&gt;The unconditional baseline (a block diffusion model trained without target features) reaches only 2-3x speedup. The target context conditioning is what makes the difference.&lt;/p&gt;
&lt;h2 id="diffusiongemma-standalone-text-diffusion"&gt;DiffusionGemma: standalone text diffusion&lt;/h2&gt;
&lt;p&gt;DiffusionGemma is an experimental open model from Google DeepMind, announced in early 2026. It represents a much more ambitious bet: replace the autoregressive decoder entirely with a diffusion process, while building on the Gemma 4 26B A4B backbone.&lt;/p&gt;
&lt;h3 id="noise-type-uniform-state-diffusion"&gt;Noise type: uniform state diffusion&lt;/h3&gt;
&lt;p&gt;Text diffusion requires defining what &amp;ldquo;noise&amp;rdquo; means for discrete tokens. Earlier approaches use masked diffusion, where tokens are replaced with a special &lt;code&gt;[MASK]&lt;/code&gt; token. Once a masked position is predicted, it stays fixed — there is no self-correction.&lt;/p&gt;
&lt;p&gt;DiffusionGemma uses uniform state diffusion instead. Noise means replacing a token with a random token drawn uniformly from the vocabulary. This has two consequences:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The model must first identify which tokens are noise before it can predict the correct tokens. The denoising task combines detection and correction.&lt;/li&gt;
&lt;li&gt;A token predicted in step 1 can be replaced in step 11 if its probability drops. Self-correction is built in.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Rejected tokens (those where the model&amp;rsquo;s confidence drops) are replaced with fresh random tokens, not the old wrong ones. This keeps the input distribution close to what the model saw during training.&lt;/p&gt;
&lt;h3 id="shared-encoder-denoiser-architecture"&gt;Shared encoder-denoiser architecture&lt;/h3&gt;
&lt;p&gt;Diffusion models typically need two components: an encoder that processes the user&amp;rsquo;s query, and a denoiser that cleans the noisy canvas. DiffusionGemma patches a single pre-trained Gemma 4 26B A4B model to serve both roles, rather than training separate networks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Encoder mode.&lt;/strong&gt; The model uses its native causal attention to process the user&amp;rsquo;s query. The resulting KV cache is computed once and stored.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Denoiser mode.&lt;/strong&gt; Attention is patched from causal to bidirectional, allowing every token in the canvas to attend to every other token. The model uses the logits of all canvas positions to predict replacements. The pre-computed KV cache from the encoder is injected into the denoiser, providing context about the query without requiring cross-attention layers.&lt;/p&gt;
&lt;p&gt;This design choice — reusing a single pretrained checkpoint instead of training from scratch — dramatically reduces training cost. The model starts from a capable autoregressive backbone and learns to work in bidirectional mode as a fine-tuning task.&lt;/p&gt;
&lt;h3 id="canvas-based-inference"&gt;Canvas-based inference&lt;/h3&gt;
&lt;p&gt;DiffusionGemma operates on a canvas of 256 tokens, generated as follows:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Self-conditioning.&lt;/strong&gt; The denoiser needs memory of its previous predictions. It takes the probability distribution from the previous step, multiplies it by the embedding matrix (producing one weighted embedding per position), passes it through a small feedforward network, and adds the result as a memory vector to the current step&amp;rsquo;s token embeddings.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Entropy-bounded sampling.&lt;/strong&gt; The canvas starts as 256 random uniform tokens. At each denoising step, the model computes entropy for every position. Positions are sorted from lowest entropy (most confident) to highest. Tokens are accepted as long as the cumulative sum of their entropies stays below a threshold. Rejected tokens are replaced with new random values.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Scheduler.&lt;/strong&gt; A decreasing temperature controls exploration over steps. Early steps use high temperature for broad token search; later steps use low temperature for peaked, confident predictions. Adaptive stopping halts denoising early if the model is stable (top tokens unchanged for N steps) and confident (entropy below 0.005).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Block extension.&lt;/strong&gt; For text longer than 256 tokens, the generated block is appended to the prompt, the KV cache is cheaply updated (it is causal on the encoder side), and diffusion begins on the next 256-token canvas.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The result is a model that generates up to 4x faster raw token throughput than an autoregressive baseline on dedicated GPUs, though the quality on complex reasoning tasks trails the full Gemma 4 autoregressive model by a small margin.&lt;/p&gt;
&lt;h2 id="side-by-side-comparison"&gt;Side-by-side comparison&lt;/h2&gt;
&lt;p&gt;With both mechanisms on the table, the differences snap into focus.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;DFlash block diffusion&lt;/th&gt;
&lt;th&gt;DiffusionGemma&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Role&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Draft model in speculative decoding&lt;/td&gt;
&lt;td&gt;Standalone text generation model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Output guarantee&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Lossless — target model verifies&lt;/td&gt;
&lt;td&gt;None — quality depends on denoising&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Model size&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;5-layer bidirectional Transformer (~tens of millions params)&lt;/td&gt;
&lt;td&gt;Full 26B backbone with bidirectional patch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Block size&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;16 tokens&lt;/td&gt;
&lt;td&gt;256 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Denoising steps&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1-3&lt;/td&gt;
&lt;td&gt;Many (scheduler-controlled, adaptive)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Noise type&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Masked diffusion (predict corrupted tokens)&lt;/td&gt;
&lt;td&gt;Uniform state diffusion (random tokens)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Self-correction&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Not needed (target catches errors)&lt;/td&gt;
&lt;td&gt;Yes — tokens can be revised&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Attention&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Bidirectional within block&lt;/td&gt;
&lt;td&gt;Dual: causal (encoder) / bidirectional (denoiser)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Conditioning source&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;External: target model hidden features&lt;/td&gt;
&lt;td&gt;Internal: self-conditioning from previous step&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Conditioning injection&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Into K/V projections of every draft layer&lt;/td&gt;
&lt;td&gt;Added as memory vector to token embeddings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;KV cache&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Draft model builds its own with injected features&lt;/td&gt;
&lt;td&gt;Encoder computes once, denoiser reuses&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Extended generation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Speculative decoding loop (draft → verify → accept)&lt;/td&gt;
&lt;td&gt;Block diffusion (append canvas → update KV cache → next canvas)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Speedup over AR&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;5-6x system-level (draft + verify combined)&lt;/td&gt;
&lt;td&gt;Up to 4x raw token throughput&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Primary bottleneck&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Acceptance length (how many draft tokens are accepted)&lt;/td&gt;
&lt;td&gt;Denoising step count&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id="where-they-converge"&gt;Where they converge&lt;/h3&gt;
&lt;p&gt;Both models exploit the same fundamental idea: bidirectional attention within a block of tokens allows parallel prediction, and the block structure enables integration with autoregressive generation (DFlash through verification, DiffusionGemma through block extension). Both achieve their best speedups when compute, not memory, is the limiting factor — which means dedicated GPUs, not CPU inference or heavily loaded servers.&lt;/p&gt;
&lt;h3 id="where-they-diverge"&gt;Where they diverge&lt;/h3&gt;
&lt;p&gt;The divergence is best understood through the lens of error handling. DFlash accepts that its draft model will make mistakes and handles them through verification — the target model catches errors, so draft quality just affects throughput, not correctness. DiffusionGemma cannot afford to make mistakes because there is no verifier. It must allocate enough denoising steps for the model to self-correct, which creates a direct quality-speed tradeoff.&lt;/p&gt;
&lt;p&gt;This difference cascades through every design choice. DFlash can use a tiny draft model because correctness is not its job. DiffusionGemma needs the full 26B backbone because the model must produce acceptable text on its own. DFlash can use 1-3 denoising steps because rough drafts are fine. DiffusionGemma needs many more because each step should improve the output. DFlash uses masked diffusion (simpler, cheaper) because errors are caught downstream. DiffusionGemma uses uniform state diffusion (more flexible, more expensive) because self-correction is essential.&lt;/p&gt;
&lt;h2 id="when-to-use-which"&gt;When to use which&lt;/h2&gt;
&lt;p&gt;DFlash is the right choice when you already have a deployed autoregressive model and want to reduce latency without changing the output distribution. The integration cost is moderate — you need to train the small draft model and set up the speculative decoding pipeline — but the output is guaranteed identical to the original model. It is particularly effective for long generations (chat, reasoning chains) where the acceptance distribution has time to average out.&lt;/p&gt;
&lt;p&gt;DiffusionGemma is the right choice when you are building a new system from scratch and want the fastest possible raw generation speed, or when you are experimenting with text diffusion as a research direction. The quality gap relative to autoregressive models is small and narrowing, and the 4x speedup at small batch sizes is genuine. But you are committing to a different output distribution, and there is no verifier to catch mistakes.&lt;/p&gt;
&lt;p&gt;The two approaches are not direct competitors. They operate at different levels of the system stack and solve different constraints. DFlash accelerates an existing model. DiffusionGemma replaces one.&lt;/p&gt;
&lt;h2 id="where-the-comparison-breaks"&gt;Where the comparison breaks&lt;/h2&gt;
&lt;p&gt;The speedup numbers are not directly comparable. DFlash&amp;rsquo;s 5-6x includes both draft generation and target verification in an end-to-end system. DiffusionGemma&amp;rsquo;s 4x is raw token throughput from a single model. If you were to run DiffusionGemma through a speculative decoding loop with a separate verifier, the effective speedup would be different.&lt;/p&gt;
&lt;p&gt;The acceptance length metric ($\tau$) that DFlash reports has no direct analogue for DiffusionGemma. A diffusion model does not produce a sequence that an external verifier accepts or rejects — it produces text directly. Comparing $\tau = 6.8$ to a 256-token canvas is comparing apples to power plants.&lt;/p&gt;
&lt;p&gt;Both approaches assume you have a GPU with enough memory. DFlash needs memory for both the target model and the draft model (though the draft model is tiny). DiffusionGemma needs memory for the full 26B backbone plus the bidirectional attention workspace. Neither works well on CPU or edge devices.&lt;/p&gt;
&lt;h2 id="open-questions"&gt;Open questions&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Does the block diffusion draft model benefit from larger block sizes ($\gamma &amp;gt; 16$) in expectation, or does the acceptance distribution saturate? DFlash&amp;rsquo;s block-size transfer experiment (large blocks train → small blocks infer) suggests a non-trivial relationship.&lt;/li&gt;
&lt;li&gt;Can DiffusionGemma&amp;rsquo;s encoder-denoiser sharing be applied to smaller backbones (4-7B) without unacceptable quality loss? The Gemma 4 26B is a large starting point.&lt;/li&gt;
&lt;li&gt;What happens when you stack the two approaches — use a block diffusion model as a draft model for DiffusionGemma as the target? Would the uniform state diffusion&amp;rsquo;s self-correction interact well with a separate verifier?&lt;/li&gt;
&lt;li&gt;How does the quality-speed Pareto frontier of DiffusionGemma shift with fewer denoising steps plus a small external verifier (a hybrid approach)?&lt;/li&gt;
&lt;li&gt;The position-dependent loss decay in DFlash is motivated by speculative decoding&amp;rsquo;s sequential constraint. Does a similar non-uniform loss help DiffusionGemma&amp;rsquo;s training, even though all positions matter for final output quality?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;[[Q]] Six months from now: has any work combined the two approaches — using a block diffusion drafter conditioned on a diffusion target model&amp;rsquo;s hidden states, or alternatively, has DiffusionGemma&amp;rsquo;s uniform state diffusion been explored as a training objective for draft models?&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Chen, Liang, Liu, &amp;ldquo;DFlash: Block Diffusion for Flash Speculative Decoding&amp;rdquo;, ICML 2026, arXiv:2602.06036.&lt;/li&gt;
&lt;li&gt;Google DeepMind, &amp;ldquo;DiffusionGemma: An experimental open text diffusion model&amp;rdquo;, 2026. Blog post and developer guide.&lt;/li&gt;
&lt;li&gt;Li et al., &amp;ldquo;EAGLE-3: Scaling Speculative Decoding via Training Objectives&amp;rdquo;, 2025.&lt;/li&gt;
&lt;li&gt;Samragh et al., &amp;ldquo;Hidden Features of Large Language Models Encode Future Tokens&amp;rdquo;, 2025.&lt;/li&gt;
&lt;li&gt;Nie et al., &amp;ldquo;LLaDA: Large Language Diffusion with Masked Language Modeling&amp;rdquo;, 2025.&lt;/li&gt;
&lt;li&gt;Arriola et al., &amp;ldquo;Block Diffusion Models&amp;rdquo;, 2025.&lt;/li&gt;
&lt;/ol&gt;</description></item><item><title>Pruning Qwen3.6-35B-A3B for RTX 5090: what I learned pushing MoE compression to its limit on a single GPU</title><link>http://fengwang.github.io/posts/qwen-pruning-rtx5090/</link><pubDate>Mon, 18 May 2026 00:00:00 +0000</pubDate><guid>http://fengwang.github.io/posts/qwen-pruning-rtx5090/</guid><description>&lt;h2 id="why-this-exists"&gt;Why this exists&lt;/h2&gt;
&lt;p&gt;I had a 35B-parameter MoE model that needed to run on a single RTX 5090. The model in FP8 needed 34.4 GiB. The GPU had 31.84 GiB. The gap was 2.6 GiB — small enough to feel tantalizing, large enough to break every standard deployment pipeline.&lt;/p&gt;
&lt;p&gt;Six days later, I had a pruned model that scored 73.2% on HumanEval+, 51.0% on Toolcall, and 33.6% on MMLU. That model (v3) shipped as the best result available at the time.&lt;/p&gt;
&lt;p&gt;Twenty-four hours after that, I had evidence that everything I thought I understood about calibration was specific to one recipe and didn&amp;rsquo;t generalize. Then another twelve hours produced v7b-fp8 — a model that beats v3 on every measured pack I ran, with BugFind +17, DataExtract +17, and InstructFollow +20.&lt;/p&gt;
&lt;p&gt;The path between those states was not linear. Some of my most confident conclusions turned out to be wrong. This article is what I wish I had known on day one — updated after the sessions that overturned my earlier priors.&lt;/p&gt;
&lt;p&gt;Qwen3.6-35B-A3B is a 256-expert MoE model optimized for agentic coding, and my only available hardware was a single RTX 5090. This covers pruning, quantization, evaluation, and attempted recovery fine-tuning on a single consumer GPU across seven experimental sessions. It does not cover multi-GPU setups, cloud inference, or comparisons with other pruning algorithms.&lt;/p&gt;
&lt;h2 id="why-a-26-gib-deficit-ate-my-week"&gt;Why a 2.6 GiB deficit ate my week&lt;/h2&gt;
&lt;p&gt;The standard view is that compressing a model by 8% to fit a memory budget is a routine calibration exercise — run AWQ or GPTQ, adjust the quantization config, ship it. A 2.6 GiB gap on a 34.4 GiB model is 7.5% compression. Well within what quantization alone should handle.&lt;/p&gt;
&lt;p&gt;It was not routine.&lt;/p&gt;
&lt;p&gt;The contradiction hit immediately: BitsAndBytes, GPTQ, and AWQ all quantize only 2D &lt;code&gt;nn.Linear&lt;/code&gt; weights. MoE models store experts as batched 3D tensors — &lt;code&gt;[n_experts, in_dim, out_dim]&lt;/code&gt; — for efficient grouped matrix multiplication. These 3D tensors contain roughly 90% of the model&amp;rsquo;s total parameters, and every standard quantization tool simply skips them. No error. No warning. They just pass through at full precision.&lt;/p&gt;
&lt;p&gt;So quantization alone could not close the gap. I needed expert pruning.&lt;/p&gt;
&lt;p&gt;REAP (Router-Weighted Expert Activation Pruning) scores each expert by the conditional average of its gating weight times its activation norm over calibration tokens:&lt;/p&gt;
&lt;p&gt;$$S_j = \mathbb{E}_{x \in X_j}[g_j(x) \cdot |f_j(x)|]$$&lt;/p&gt;
&lt;p&gt;The intuition: an expert that gets low gating weight &lt;em&gt;and&lt;/em&gt; produces small activations contributes little to the output and can be removed with minimal reconstruction error. REAP then removes the lowest-scoring experts and propagates residuals to keep the functional manifold topology intact.&lt;/p&gt;
&lt;p&gt;The pruner works block-by-block: load one 750 MiB decoder layer to GPU, score all surviving experts, prune the lowest, propagate residuals, save. About 25 minutes per experiment across 40 decoder layers. I ran seven experiments over seven sessions on one GPU.&lt;/p&gt;
&lt;p&gt;The constraint that shaped everything: the RTX 5090&amp;rsquo;s 32 GiB was simultaneously my inference platform, evaluation framework, and training environment. Pruning, eval, and SFT all competed for the same VRAM, and no two of them could run at the same time. At BF16, the model occupied 30.6 GiB — 97% of GPU capacity, leaving zero room for gradients or activations during training.&lt;/p&gt;
&lt;p&gt;This total-resource competition, not the pruning algorithm, turned out to be the hard problem.&lt;/p&gt;
&lt;h2 id="the-calibration-mix-that-rewrote-my-priors"&gt;The calibration mix that rewrote my priors&lt;/h2&gt;
&lt;p&gt;I came to this expecting the pruning algorithm or the compression ratio to dominate quality. That is how the literature frames it: better importance scores, better pruning decisions.&lt;/p&gt;
&lt;p&gt;The evidence says otherwise.&lt;/p&gt;
&lt;p&gt;Here is what happened.&lt;/p&gt;
&lt;p&gt;My first pruning experiment (v2) used four code-focused calibration datasets: evol-codealpaca, BigCodeBench, SWE-bench, and xlam. Pure code data. The result: HumanEval+ at 72.0%, Toolcall at 44.0%, and MMLU at — two categories at 0.0%.&lt;/p&gt;
&lt;p&gt;Dead categories. The model could generate code, but it could not answer a general-knowledge question. The pruner had never seen general-knowledge tokens during scoring, so it had no way to know which experts mattered for those domains. Any expert carrying general-knowledge information was pruned or severely weakened.&lt;/p&gt;
&lt;p&gt;For the v3 experiment, I switched to a 70/30 code-to-general mix. Same REAP algorithm. Same compression ratio. Just two general datasets added — 600 samples of MMLU and 600 samples of C4 — to a pool of 700 samples each from four code datasets.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;v2 (pure code cal)&lt;/th&gt;
&lt;th&gt;v3 (70/30 cal)&lt;/th&gt;
&lt;th&gt;Change&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;MMLU Social Sciences&lt;/td&gt;
&lt;td&gt;0.0%&lt;/td&gt;
&lt;td&gt;33.3%&lt;/td&gt;
&lt;td&gt;+33.3pp&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MMLU Other&lt;/td&gt;
&lt;td&gt;0.0%&lt;/td&gt;
&lt;td&gt;34.3%&lt;/td&gt;
&lt;td&gt;+34.3pp&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HumanEval+&lt;/td&gt;
&lt;td&gt;72.0%&lt;/td&gt;
&lt;td&gt;73.2%&lt;/td&gt;
&lt;td&gt;+1.2pp&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Toolcall&lt;/td&gt;
&lt;td&gt;44.0%&lt;/td&gt;
&lt;td&gt;51.0%&lt;/td&gt;
&lt;td&gt;+7.0pp&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Every single metric improved. MMLU recovered from dead to 33%+. Code benchmarks improved too.&lt;/p&gt;
&lt;p&gt;I think of calibration data as the lens through which the pruner sees the model. A narrow lens (pure code) gives a sharp but myopic view — the pruner keeps only what it sees, and blinds the model to everything else. A wider lens (70/30 mix) lets the pruner see the full functional space, so it preserves the structures that serve the whole distribution, not just one mode.&lt;/p&gt;
&lt;p&gt;The analogy breaks in one direction: you cannot just keep adding calibration domains forever. More data means longer scoring passes. But within a practical budget, the evidence is clear — calibration composition matters more than any algorithmic tweak I could have made to the pruner itself.&lt;/p&gt;
&lt;p&gt;I shipped v3 and moved on. That was where the story got interesting.&lt;/p&gt;
&lt;h2 id="the-discovery-that-my-conclusions-were-calibration-specific"&gt;The discovery that my conclusions were calibration-specific&lt;/h2&gt;
&lt;p&gt;After shipping v3, I went back to test a hypothesis that seemed obvious: if I replaced the code-heavy calibration with agentic traces, the pruned model would perform better on agentic benchmarks. Tool calling, bug finding, multi-step reasoning — these were what the model was designed for.&lt;/p&gt;
&lt;p&gt;I built a proper BenchLocal evaluation harness — 8 packs covering tool-call, hermes-agent, bug-find, data-extract, instruct-follow, reason-math, struct-output, and cli — and established a v3 baseline with the new v19 chat template. The baseline was sobering: ToolCall-15 at 90, HermesAgent-20 at 16, BugFind-15 at 8. The agentic gates were low to begin with.&lt;/p&gt;
&lt;p&gt;I ran two candidate experiments at a deeper compression (0.40, keeping 154 of 256 experts) with two calibration strategies:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Mix-A&lt;/strong&gt;: full replacement — agentic data (glm47-reap + hermes-agent-traces) instead of code&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Mix-B&lt;/strong&gt;: additive — agentic data layered on top of v3&amp;rsquo;s base&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Both produced the same result. ToolCall-15 dropped from 97 (the v3 baseline with the old reasoning parser) to 90. HermesAgent-20 stayed at exactly 16. BugFind-15 hovered in the 3-10 range. No agentic movement. Both failed all three gates.&lt;/p&gt;
&lt;p&gt;I stopped and ran a control experiment. I took Mix-A&amp;rsquo;s exact calibration list and ran it at v3&amp;rsquo;s exact compression ratio (0.289, 183 experts). If compression depth was the cause of the toolcall regression, the control would recover toolcall. If calibration content was the cause, the control would still show the regression.&lt;/p&gt;
&lt;p&gt;The control was decisive:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th style="text-align: center"&gt;Candidate&lt;/th&gt;
&lt;th style="text-align: center"&gt;Compression&lt;/th&gt;
&lt;th style="text-align: center"&gt;Calibration&lt;/th&gt;
&lt;th style="text-align: center"&gt;ToolCall-15&lt;/th&gt;
&lt;th style="text-align: center"&gt;HermesAgent-20&lt;/th&gt;
&lt;th style="text-align: center"&gt;BugFind-15&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="text-align: center"&gt;v3&lt;/td&gt;
&lt;td style="text-align: center"&gt;0.289&lt;/td&gt;
&lt;td style="text-align: center"&gt;70/30 balanced&lt;/td&gt;
&lt;td style="text-align: center"&gt;97&lt;/td&gt;
&lt;td style="text-align: center"&gt;16&lt;/td&gt;
&lt;td style="text-align: center"&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: center"&gt;Mix-A&lt;/td&gt;
&lt;td style="text-align: center"&gt;0.40&lt;/td&gt;
&lt;td style="text-align: center"&gt;Agentic replacement&lt;/td&gt;
&lt;td style="text-align: center"&gt;90&lt;/td&gt;
&lt;td style="text-align: center"&gt;16&lt;/td&gt;
&lt;td style="text-align: center"&gt;10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: center"&gt;Mix-B&lt;/td&gt;
&lt;td style="text-align: center"&gt;0.40&lt;/td&gt;
&lt;td style="text-align: center"&gt;Additive layering&lt;/td&gt;
&lt;td style="text-align: center"&gt;90&lt;/td&gt;
&lt;td style="text-align: center"&gt;16&lt;/td&gt;
&lt;td style="text-align: center"&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: center"&gt;v3ratio&lt;/td&gt;
&lt;td style="text-align: center"&gt;0.289&lt;/td&gt;
&lt;td style="text-align: center"&gt;Mix-A calibration&lt;/td&gt;
&lt;td style="text-align: center"&gt;90&lt;/td&gt;
&lt;td style="text-align: center"&gt;16&lt;/td&gt;
&lt;td style="text-align: center"&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;ToolCall-15 at 90 across all three candidates, including the control at v3&amp;rsquo;s compression. The regression came from the calibration content, not the compression depth.&lt;/p&gt;
&lt;p&gt;The mechanism was surprisingly specific. The regression was concentrated entirely in a single sub-dimension: Parameter Precision dropped from 100 to 67. The model still picked the right tools with the right structure — it just generated wrong-typed or wrong-formatted arguments more often. Dropping code corpora (evol-codealpaca, bigcodebench, swe-bench) and substituting agentic traces had cost the model its tight argument-formatting discipline.&lt;/p&gt;
&lt;p&gt;The other finding was harsher. HermesAgent-20 scored 16 out of 20 across all four configurations — literally identical, including per-category breakdown. A 25B pruned MoE cannot handle these multi-step browser-automation scenarios, regardless of what you feed it during pruning calibration. The gate is capacity-bound.&lt;/p&gt;
&lt;p&gt;I also discovered that the vLLM &lt;code&gt;--reasoning-parser qwen3&lt;/code&gt; flag was essential for correct evaluation. Without it, the model&amp;rsquo;s &lt;code&gt;&amp;lt;think&amp;gt;&lt;/code&gt; reasoning block leaked into all plain-text responses, breaking every non-tool-call scorer. The flag lifted ToolCall-15 from 90 to 97 and recovered instruct-follow and data-extract from flat zero. The lesson: validate your eval infrastructure before you trust a single number.&lt;/p&gt;
&lt;p&gt;I closed Session 6 with a clean negative result and the conviction that calibration content could not move agentic gates. That conviction lasted about twelve hours.&lt;/p&gt;
&lt;h2 id="the-recipe-that-broke-through"&gt;The recipe that broke through&lt;/h2&gt;
&lt;p&gt;Session 7 adopted a fundamentally different calibration recipe: the REAP-26B 6-dataset mix. Six datasets — SWE-bench/SWE-smith-trajectories (tool split), xlam-function-calling-60k, evol-codealpaca, and Mixture-of-Thoughts (code/math/science) — at much higher token count (1024 samples x 16384 sequence length, totaling 16.8M tokens). Router renormalization disabled per the REAP-26B README.&lt;/p&gt;
&lt;p&gt;I ran three plans plus a follow-up:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Plan-A&lt;/strong&gt; (compression 0.40, fresh prune). ToolCall-15 collapsed to 63 — a catastrophic -27 regression. But BugFind jumped to +15 and InstructFollow to +33. The recipe was clearly powerful. Too powerful at this depth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Plan-C&lt;/strong&gt; (stacked prune on top of the upstream REAP-26B-VL). Recovered toolcall to 90 but lost ~90% of the recipe&amp;rsquo;s other gains. Stacked pruning does not inherit upstream calibration signal.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Plan-B&lt;/strong&gt; (compression 0.289, v3&amp;rsquo;s depth, the experiment I initially skipped). This was the follow-up after both Plan-A and Plan-C failed.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th style="text-align: center"&gt;Candidate&lt;/th&gt;
&lt;th style="text-align: center"&gt;Compression&lt;/th&gt;
&lt;th style="text-align: center"&gt;ToolCall-15&lt;/th&gt;
&lt;th style="text-align: center"&gt;BugFind-15&lt;/th&gt;
&lt;th style="text-align: center"&gt;DataExtract-15&lt;/th&gt;
&lt;th style="text-align: center"&gt;InstructFollow-15&lt;/th&gt;
&lt;th style="text-align: center"&gt;Verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="text-align: center"&gt;v3+v19&lt;/td&gt;
&lt;td style="text-align: center"&gt;0.289&lt;/td&gt;
&lt;td style="text-align: center"&gt;90&lt;/td&gt;
&lt;td style="text-align: center"&gt;8&lt;/td&gt;
&lt;td style="text-align: center"&gt;5&lt;/td&gt;
&lt;td style="text-align: center"&gt;20&lt;/td&gt;
&lt;td style="text-align: center"&gt;Baseline&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: center"&gt;v7a&lt;/td&gt;
&lt;td style="text-align: center"&gt;0.40&lt;/td&gt;
&lt;td style="text-align: center"&gt;63&lt;/td&gt;
&lt;td style="text-align: center"&gt;23&lt;/td&gt;
&lt;td style="text-align: center"&gt;24&lt;/td&gt;
&lt;td style="text-align: center"&gt;53&lt;/td&gt;
&lt;td style="text-align: center"&gt;FailToolcall&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: center"&gt;v7c&lt;/td&gt;
&lt;td style="text-align: center"&gt;stacked&lt;/td&gt;
&lt;td style="text-align: center"&gt;90&lt;/td&gt;
&lt;td style="text-align: center"&gt;0&lt;/td&gt;
&lt;td style="text-align: center"&gt;4&lt;/td&gt;
&lt;td style="text-align: center"&gt;16&lt;/td&gt;
&lt;td style="text-align: center"&gt;FailAgentic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="text-align: center"&gt;&lt;strong&gt;v7b&lt;/strong&gt;&lt;/td&gt;
&lt;td style="text-align: center"&gt;&lt;strong&gt;0.289&lt;/strong&gt;&lt;/td&gt;
&lt;td style="text-align: center"&gt;&lt;strong&gt;93&lt;/strong&gt;&lt;/td&gt;
&lt;td style="text-align: center"&gt;&lt;strong&gt;25&lt;/strong&gt;&lt;/td&gt;
&lt;td style="text-align: center"&gt;&lt;strong&gt;22&lt;/strong&gt;&lt;/td&gt;
&lt;td style="text-align: center"&gt;&lt;strong&gt;40&lt;/strong&gt;&lt;/td&gt;
&lt;td style="text-align: center"&gt;&lt;strong&gt;Pass&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;v7b-fp8 scored equal or better than v3 on all 7 measured packs. No regressions. BugFind +17, DataExtract +17, InstructFollow +20, ToolCall +3. The trigger verdict was Pass.&lt;/p&gt;
&lt;p&gt;This is the result I keep coming back to: the same recipe at 0.40 collapsed toolcall to 63; at 0.289 it improved toolcall to 93. The recipe drives the agentic gains. The compression depth modulates the toolcall trade-off. Session 6&amp;rsquo;s conclusion that &amp;ldquo;calibration content cannot move agentic benchmarks&amp;rdquo; was specific to Mix-A&amp;rsquo;s content, not a universal property of pruning calibration.&lt;/p&gt;
&lt;p&gt;The pipeline upgrades that made this possible are worth noting. The 16K sequence length calibration required a chunked REAP scoring accumulator — the single-pass approach would have materialized 67 GiB on a 32 GiB GPU. The custom FP8 quantizer (&lt;code&gt;scripts/quantize_fp8.py&lt;/code&gt;) bypasses llmcompressor&amp;rsquo;s broken Qwen3.6 compatibility with a 274-line direct cast from BF16 to &lt;code&gt;torch.float8_e4m3fn&lt;/code&gt;. Schema adapters with a 60-second preflight caught dataset drift before any GPU allocation.&lt;/p&gt;
&lt;h2 id="how-3d-tensors-broke-every-quantization-framework"&gt;How 3D tensors broke every quantization framework&lt;/h2&gt;
&lt;p&gt;The calibration experiment got me v3 — a model that almost passed its quality gates. HumanEval+ at 73.2% was close to the 75% threshold. MMLU at 33.6% was still short of the 40% target. The obvious next step was recovery fine-tuning.&lt;/p&gt;
&lt;p&gt;This is where the second assumption broke.&lt;/p&gt;
&lt;p&gt;I assumed that standard quantization tools would handle model compression for training. Load the model in 4-bit, apply LoRA adapters, train. This is the default workflow for QLoRA on every Hugging Face tutorial. It works on LLaMA, it works on Mistral — it should work on Qwen.&lt;/p&gt;
&lt;p&gt;It does not. Because the model&amp;rsquo;s 3D expert tensors are invisible to BitsAndBytes.&lt;/p&gt;
&lt;p&gt;I spent the evening of day two systematically eliminating every standard training approach. Seven attempts, all failures:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Attempt&lt;/th&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;BnB 4-bit QLoRA&lt;/td&gt;
&lt;td&gt;Can&amp;rsquo;t quantize 3D expert tensors [183, 1024, 2048]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;BF16 model.to(&amp;lsquo;cuda&amp;rsquo;)&lt;/td&gt;
&lt;td&gt;30.6 GiB — 0 bytes left for activations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;accelerate device_map=&amp;lsquo;auto&amp;rsquo;&lt;/td&gt;
&lt;td&gt;Keeps all layers on GPU for backward&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;DeepSpeed ZeRO-3 (single GPU)&lt;/td&gt;
&lt;td&gt;Trainer moves model to GPU before partitioning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;DeepSpeed zero.Init + from_pretrained&lt;/td&gt;
&lt;td&gt;Weight loading conflicts with meta-device tensors&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;FP8 frozen weights + monkey-patched ops&lt;/td&gt;
&lt;td&gt;grouped_mm upcast creates 768 MiB BF16 temp per layer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;FP8 + dispatch_model with 10 GiB budget&lt;/td&gt;
&lt;td&gt;Offloaded layers accumulate on GPU during backward&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;I don&amp;rsquo;t fully understand why every framework eventually calls &lt;code&gt;model.to(device)&lt;/code&gt; for the backward pass on single GPU. The documentation promises CPU offloading. The reality is that DeepSpeed ZeRO-3, accelerate&amp;rsquo;s dispatch_model, and FSDP all converge on the same behavior: put the full model on GPU when gradients need to flow.&lt;/p&gt;
&lt;p&gt;The resolution came from a workaround I had not considered: &lt;strong&gt;unbatch the 3D expert tensors into individual &lt;code&gt;bnb.nn.Linear4bit&lt;/code&gt; layers&lt;/strong&gt;. BnB can quantize standard 2D linear layers. A 3D tensor &lt;code&gt;[183, 1024, 2048]&lt;/code&gt; becomes 183 separate &lt;code&gt;Linear(2048, 1024)&lt;/code&gt; objects, each quantizable in 4-bit.&lt;/p&gt;
&lt;p&gt;The result: model on GPU dropped from 30.6 GiB to 16.8 GiB, leaving 16.9 GiB for activations and gradients. The SFT ran — 311 steps, 9,934 samples, 11.5 hours, loss from 1.058 to 0.975, token accuracy from 85% to 96%. By every training metric, it worked.&lt;/p&gt;
&lt;p&gt;It did not work.&lt;/p&gt;
&lt;h2 id="the-sft-trap-when-training-makes-everything-worse"&gt;The SFT trap: when training makes everything worse&lt;/h2&gt;
&lt;p&gt;I expected SFT with quantized frozen weights to improve the model. The training curves were healthy. Loss decreasing. Token accuracy climbing. All the signals that normally say &amp;ldquo;keep training, it is converging.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Post-SFT evaluation told a different story:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Pre-SFT&lt;/th&gt;
&lt;th&gt;Post-SFT&lt;/th&gt;
&lt;th&gt;Delta&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;HumanEval+&lt;/td&gt;
&lt;td&gt;73.2%&lt;/td&gt;
&lt;td&gt;67.7%&lt;/td&gt;
&lt;td&gt;-5.5pp&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Toolcall&lt;/td&gt;
&lt;td&gt;51.0%&lt;/td&gt;
&lt;td&gt;50.5%&lt;/td&gt;
&lt;td&gt;-0.5pp&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MMLU&lt;/td&gt;
&lt;td&gt;33.6%&lt;/td&gt;
&lt;td&gt;9.4%&lt;/td&gt;
&lt;td&gt;-24.2pp&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Everything regressed. MMLU collapsed back to v2-level. HumanEval+ lost 5.5 points.&lt;/p&gt;
&lt;p&gt;The mechanism is specific and instructive — and worth pausing on because it explains an entire class of pipeline failures:&lt;/p&gt;
&lt;p&gt;4-bit quantization injects noise into the forward pass of every frozen expert layer. That noise is deterministic — same input, same 4-bit weights, same quantization error — but it shifts the activation distribution that the trainable router and shared expert layers see. The trainable parameters adapt to this shifted distribution during SFT. They learn to work &lt;em&gt;with the noise profile&lt;/em&gt; of the 4-bit experts.&lt;/p&gt;
&lt;p&gt;When you remove the 4-bit quantization (merge the fine-tuned weights back into the original BF16 model for inference), the noise profile disappears. The trainable parameters are now running on clean activations, and they have overfit to a distribution that no longer exists.&lt;/p&gt;
&lt;p&gt;This is why the training metrics looked great while the benchmarks collapsed. The model was not learning to generate better code or answer knowledge questions. It was learning to compensate for quantization noise in the frozen pathway. When the noise went away, the compensation became output distortion.&lt;/p&gt;
&lt;p&gt;The result that I keep returning to: the pre-SFT v3 model was the project&amp;rsquo;s best output at that point. The calibration strategy was the lever. Fine-tuning was a trap.&lt;/p&gt;
&lt;h2 id="what-i-would-do-differently"&gt;What I would do differently&lt;/h2&gt;
&lt;p&gt;If I started this project again, I would change several things. The later sessions taught me that some of my early conclusions were incomplete — so these recommendations are updated with everything I know after seven sessions.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What I did&lt;/th&gt;
&lt;th&gt;What I would do instead&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Jumped straight to production-scale SFT&lt;/td&gt;
&lt;td&gt;Validate on a 2-layer toy model first&lt;/td&gt;
&lt;td&gt;Seven failed approaches, ~4 hours of debugging, caught in minutes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Let eval and training share one venv&lt;/td&gt;
&lt;td&gt;Isolate venvs from the start&lt;/td&gt;
&lt;td&gt;huggingface_hub 1.5 vs 1.14 broke vLLM weight loading&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ran SFT before exhausting calibration experiments&lt;/td&gt;
&lt;td&gt;Run all calibration experiments before SFT&lt;/td&gt;
&lt;td&gt;The 50/50 experiment was never attempted, and calibration is the primary lever&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shallow pruning (183/256) + 4-bit SFT&lt;/td&gt;
&lt;td&gt;Deeper pruning (154 experts) + clean BF16 SFT&lt;/td&gt;
&lt;td&gt;Avoids the noise-overfitting trap entirely&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Assumed agentic calibration lifts agentic gates&lt;/td&gt;
&lt;td&gt;Test the REAP-26B recipe at the original depth first&lt;/td&gt;
&lt;td&gt;Session 6&amp;rsquo;s clean negative result was Mix-A-specific; Plan-B at v3&amp;rsquo;s depth passed everything&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tested one variable at a time&lt;/td&gt;
&lt;td&gt;Test &amp;ldquo;new calibration at same compression&amp;rdquo; as default isolation&lt;/td&gt;
&lt;td&gt;The v3ratio control flipped the interpretation — always isolate the variable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Championed &amp;ldquo;capacity-bound benchmarks&amp;rdquo; as a universal conclusion&lt;/td&gt;
&lt;td&gt;Measure with multiple recipes before declaring a ceiling&lt;/td&gt;
&lt;td&gt;BugFind moved +17 with the right recipe at no parameter count change&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The most painful lesson is also the most transferable: validation loops catch pipeline bugs, but they do not catch strategy bugs. The SFT pipeline ran correctly — no crashes, no OOM, healthy training curves — and produced a worse model.&lt;/p&gt;
&lt;p&gt;The one experiment I most regret not running is the 50/50 calibration mix. If 30% general data pushed MMLU from 0% to 33%, 50% might push it past 40%. That experiment would have taken 25 minutes. The SFT that replaced it took 11.5 hours and made everything worse.&lt;/p&gt;
&lt;h2 id="boundary-conditions"&gt;Boundary conditions&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;The calibration-composition result is established for REAP on one model family (Qwen3.6-35B-A3B). It likely transfers to other MoE models and other pruning algorithms, but I have not tested this.&lt;/li&gt;
&lt;li&gt;The REAP-26B recipe finding (v7b-fp8 beating v3) is specific to this calibration mix at this compression depth. Whether it generalizes to other MoE scales is an open question.&lt;/li&gt;
&lt;li&gt;The Session 6 &amp;ldquo;calibration content drives toolcall regression&amp;rdquo; finding is specific to Mix-A&amp;rsquo;s content (glm47-reap + hermes-agent-traces). The REAP-26B recipe at the same depth showed a &lt;em&gt;positive&lt;/em&gt; toolcall delta. The finding is recipe-specific, not universal.&lt;/li&gt;
&lt;li&gt;HermesAgent-20 remained stuck at 16/20 across all seven sessions and all configurations tested. It is genuinely capacity-bound at this model size.&lt;/li&gt;
&lt;li&gt;The 4-bit expert unbatching technique works at the cost of inference speed — per-expert sequential Linear4bit forward passes are slower than native grouped matrix multiplication.&lt;/li&gt;
&lt;li&gt;The SFT degradation result applies specifically to training with 4-bit frozen experts on this model architecture. FP8 frozen experts or full BF16 SFT may behave differently.&lt;/li&gt;
&lt;li&gt;Single-GPU constraints shaped every conclusion. With multi-GPU hardware, the trade-offs shift substantially.&lt;/li&gt;
&lt;li&gt;The v19 chat template costs about 7 toolcall points compared to v18 (90 vs 97 on the same model). All Session 7 comparisons are within the same template, but direct comparability with Session 6 numbers requires accounting for the template shift.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="open-questions"&gt;Open questions&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Can SFT on v7b-fp8 lift HermesAgent-20?&lt;/strong&gt; It is the only pack that v7b didn&amp;rsquo;t improve, stuck at 16/20. Session 6&amp;rsquo;s capacity-bound conclusion was wrong for BugFind (the right recipe moved it +17). It may also be wrong for HermesAgent once the right recipe is found. But SFT on actual agent traces is a qualitatively different approach, and the infrastructure exists but hasn&amp;rsquo;t been validated against v7b.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Does 50/50 calibration push MMLU past 40%?&lt;/strong&gt; The experiment takes 25 minutes and was never scheduled.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Can Transformer Engine FP8 training enable quality SFT without the noise-overfitting trap?&lt;/strong&gt; The tools are installed on sm_120. Untested.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Does the REAP-26B recipe replicate on other MoE families?&lt;/strong&gt; The recipe drove +17 across multiple benchmarks on Qwen3.6. Would it produce similar gains on DeepSeek, Mixtral, or OLMoE?&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Should stacked pruning ever be used?&lt;/strong&gt; Session 7 showed it destroys upstream calibration signal. But if the upstream calibration is expensive (24K samples on a 96 GB GPU), stacking a cheap re-prune on top seems like it should work in theory. The empirical result was negative. I don&amp;rsquo;t fully understand why.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;[[Q]] Six months from now: run the 50/50 calibration experiment first. If it pushes MMLU past 40%, the entire SFT effort was wasted. And promote the v7b symlink — it is the best model you have.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Fang et al., &amp;ldquo;REAP the Experts: Why Pruning Prevails for One-Shot MoE Compression&amp;rdquo;, arXiv:2510.13999, 2025.&lt;/li&gt;
&lt;li&gt;Dery et al., &amp;ldquo;Finding Fantastic Experts in MoE Models&amp;rdquo;, arXiv:2504.15447, 2025.&lt;/li&gt;
&lt;li&gt;Zhang et al., &amp;ldquo;Efficient Expert Pruning in MoE LLMs&amp;rdquo;, arXiv:2505.12345, 2025.&lt;/li&gt;
&lt;li&gt;BitsAndBytes, Hugging Face quantization library, &lt;a href="https://github.com/bitsandbytes-foundation/bitsandbytes"target="_blank" rel="noopener noreferrer"&gt;https://github.com/bitsandbytes-foundation/bitsandbytes&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;TRL: Transformer Reinforcement Learning, Hugging Face, &lt;a href="https://github.com/huggingface/trl"target="_blank" rel="noopener noreferrer"&gt;https://github.com/huggingface/trl&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;</description></item></channel></rss>