{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "Notes by Mouhssine Rifaki",
  "home_page_url": "https://mouhssine.rifaki.me/notes/",
  "feed_url": "https://mouhssine.rifaki.me/feed.json",
  "description": "Essays on deep learning theory, reinforcement learning, and mathematical statistics.",
  "language": "en-US",
  "authors": [
    {
      "name": "Mouhssine Rifaki",
      "url": "https://mouhssine.rifaki.me/"
    }
  ],
  "items": [
    {
      "id": "https://mouhssine.rifaki.me/notes/autonomy-as-coordination/",
      "url": "https://mouhssine.rifaki.me/notes/autonomy-as-coordination/",
      "title": "Why autonomy is a coordination problem",
      "summary": "Autonomy scales, but safety does not come from any single system getting better. When many agents share space and time, decisions collide and hazards compound, so keeping us safe becomes a coordination\u2026",
      "content_html": "<!-- date: 2026-07-07 -->\n<p>Autonomy scales, but safety does not come from any single system getting better. When many agents share space and time, decisions collide and hazards compound, so keeping us safe becomes a coordination problem that compounds with every participant. Nobody in the field disagrees with that. Such wide agreement is what makes the sentence nearly useless on its own. What I want to argue is the part people do disagree with: coordination cannot be added afterwards. The prevailing plan is to build the vehicles now and standardize the coordination between them later as a protocol or a service or a certificate class. My argument is that the plan fails. Where the question admits of measurement I have measured it. A few hours of live air traffic tell us what proximity does in a real sky. A small experiment anyone can rerun in twenty seconds tells us what a coordination convention amounts to inside a learned policy. Each is linked where it is used and each contradicted something I believed before I ran it.</p>\n\n        <h2>Layering</h2>\n\n        <p>The instinct to add coordination later is the layering instinct and it is usually right. Networking works because TCP does not need to know how Ethernet frames a packet. The layer below can fail, the layer above notices, and the recovery is a retransmission costing milliseconds and nothing else. Layering buys independence at the price of one assumption: that failure at the lower layer is recoverable at the higher one.</p>\n\n        <p>The FAA and TSA proposal of 7 August 2025 on beyond-visual-line-of-sight operations [<a href=\"https://www.federalregister.gov/documents/2025/08/07/2025-14992/normalizing-unmanned-aircraft-systems-beyond-visual-line-of-sight-operations\">1</a>] writes that instinct into law. It creates a whole new certificate class called part 146 for automated data service providers. Operators in controlled airspace or over dense population must buy strategic deconfliction from one of them or certificate themselves to self-provide. The proposal certifies the aircraft on one side and the service on the other, and what passes between the two is a request and a response.</p>\n\n        <p>Perrow drew the distinction that governs this. Interactive complexity means a system with more ways to interact than anyone has modeled; tight coupling means a failure reaching its consequence with nothing in between to absorb it. He argued that the dangerous quadrant belongs to systems with both at once [<a href=\"https://press.princeton.edu/books/paperback/9780691004129/normal-accidents\">2</a>]. Layering answers the first of those and has nothing to offer against the second, since the recovery it depends on costs time that a tightly coupled system has already spent. Two hundred aircraft over a city at two hundred feet occupies that quadrant, where the round trip a retransmission needs runs longer than the interval in which a conflict resolves itself one way or the other.</p>\n\n        <p>Leveson's version is more useful still, because she models safety as a hierarchy of controllers and locates accidents in inadequate control at some level [<a href=\"https://direct.mit.edu/books/oa-monograph/2908/Engineering-a-Safer-WorldSystems-Thinking-Applied\">3</a>]. A coordination service is a controller. The question a layering plan has to answer is what it controls and for how long after it stops speaking.</p>\n\n        <h2>Where the layer ends in the rule's own words</h2>\n\n        <p>The proposal answers that question with unusual precision. What it sets down reads as an engineering specification.</p>\n\n        <p>Proposed \u00a7 108.190(c)(1) requires strategic deconfliction to &quot;perform strategic conflict detection and resolution prior to takeoff, and in relation to other unmanned aircraft operations that are discoverable at that time.&quot; Each qualifier carries weight. The guarantee is computed once before departure over whichever intents happened to be visible at that instant. Proposed \u00a7 108.190(d) adds conformance monitoring. Its specified output is an immediate alert to operations personnel when the aircraft leaves its intent, which is to say the layer's response to a violation is to tell a human. The definitions come from ASTM F3548-21 as the rule's own footnote records, and that standard draws its boundaries in clause 1.13. It &quot;does not purport to address tactical conflicts between UAS&quot; and it &quot;does not purport to address operations in locations where persistent connectivity is unavailable&quot; [<a href=\"https://store.astm.org/f3548-21.html\">4</a>]. Those are the two conditions under which the rest of this note is written.</p>\n\n        <p>The rest of the proposal [<a href=\"https://www.federalregister.gov/documents/2025/08/07/2025-14992/normalizing-unmanned-aircraft-systems-beyond-visual-line-of-sight-operations\">1</a>] follows the same shape. Proposed \u00a7 108.195 orders an unmanned aircraft against aircraft arriving at an airport and against aircraft that broadcast their position. Proposed \u00a7 108.825 defines the collision-avoidance design requirement by pointing back to it, so a part 108 aircraft is ranked against manned traffic and unranked against another part 108 aircraft. Proposed \u00a7 108.210(a) fixes the aircraft-to-coordinator ratio at one to one, waivable by &quot;a method acceptable to the Administrator&quot; that nobody has written. Proposed \u00a7 108.815(b) requires the airframe to &quot;execute a safe predetermined action when reaching the link timeout.&quot; Proposed \u00a7 146.325(a) requires the service provider to report an unscheduled outage after the fact.</p>\n\n        <figure class=\"plate\">\n          <img src=\"https://mouhssine.rifaki.me/notes/img/coordination-coverage.png\" alt=\"A chart with four columns for the phases of one flight: before takeoff, in flight on intent, off intent with the link up, and off intent with the link down. Strategic conflict detection under section 108.190(c)(1) covers the first column only. Conformance monitoring under 108.190(d) covers the two middle columns and alerts operations personnel. The right-of-way rule under 108.195 spans all four columns and orders the aircraft toward aircraft that broadcast. The predetermined action under 108.815(b) covers the last column, one aircraft at a time. A fifth row, joint resolution between two unmanned aircraft after takeoff, spans all four columns with a dashed empty outline labeled a method acceptable to the Administrator.\" loading=\"lazy\" decoding=\"async\" width=\"1290\" height=\"644\">\n          <figcaption>This one is a diagram: it maps the proposed rule text onto the phases of a flight, read off 90 FR 38212 [<a href=\"https://www.federalregister.gov/documents/2025/08/07/2025-14992/normalizing-unmanned-aircraft-systems-beyond-visual-line-of-sight-operations\">1</a>]. Every mechanism that reasons about more than one aircraft at once operates before takeoff. Everything after takeoff reasons about one aircraft.</figcaption>\n        </figure>\n\n        <p>I do not read this as an oversight. It is where a layered design has to stop and the standard says so out loud. A service that computes a guarantee over a snapshot and hands it to a vehicle that will act on it for the next forty minutes has produced something with the structure of a database read taken outside a transaction: accurate at the instant of reading and unconstrained from then on.</p>\n\n        <p>The obvious objection is that detect and avoid is the tactical layer and it is on its way. I think the proposal is quietly clear about why that does not rescue the plan. Proposed \u00a7 108.825 states the design requirement as a capability &quot;to avoid aircraft as required in accordance with \u00a7 108.195&quot;. The specification for detect and avoid is therefore the right-of-way rule, and that rule ranks a part 108 aircraft against traffic that broadcasts while saying nothing about two part 108 aircraft meeting each other. The FAA also records that it &quot;decided not to update \u00a7 91.113 based on the BVLOS ARC's proposal related to 'detect-and-avoid' at this time&quot; because doing so would reach into part 91 and legacy aviation [<a href=\"https://www.federalregister.gov/documents/2025/08/07/2025-14992/normalizing-unmanned-aircraft-systems-beyond-visual-line-of-sight-operations\">1</a>]. Set that aside and a deeper problem remains. Detect and avoid is a sensor and a maneuver on one airframe. It improves what one aircraft can do about a conflict it can see, which puts it in the capability column. Two aircraft avoiding competently can still pick maneuvers that conflict with each other.</p>\n\n        <h2>Measuring what the layer is for</h2>\n\n        <p>The case for a coordination service rests on an arithmetic claim: encounters grow faster than fleets do. Put $N$ vehicles in a fixed volume, assume any pair is as likely to meet as any other, and the pair count grows like $N^2$. I have used that argument before. This time I wanted to know whether it survives contact with a real sky. It does so only in the sense that reality turns out to be worse than the model allows for.</p>\n\n        <p>I took 30 snapshots of every transponder-equipped aircraft over the continental United States from the OpenSky Network [<a href=\"https://doi.org/10.1109/IPSN.2014.6846743\">20</a>]. That is 123,520 airborne state vectors in all, sampled with 12,000 discs of 100 km drawn at random across it. In each disc I counted the aircraft present and the pairs within 5, 10, 20 and 40 nautical miles laterally and 1,000 feet vertically. That vertical figure is the standard separation minimum. Then I asked what the same aircraft at the same altitudes would produce if they were thrown at random into the same disc. The null needs no model of its own. It is the same population with the geography taken out. The full set of numbers is <a href=\"https://mouhssine.rifaki.me/notes/experiments/airspace-results.json\">here</a>, along with the <a href=\"https://mouhssine.rifaki.me/notes/experiments/airspace/states.csv.gz\">snapshots</a> they were computed from.</p>\n\n        <p>Real traffic packs 6.74 times more pairs within 5 NM than the scattered version of itself. The fitted slope of the pair count against the aircraft count is 2.58 with a standard error of 0.01, above the 2 that uniform mixing predicts. The excess shrinks as the threshold widens and reaches 1.65 at 40 NM. That is what we would expect if the clustering happens at the scale of terminal areas.</p>\n\n        <p>Each number comes with a qualification. The slope compares different places at one moment. A disc holding more aircraft is usually also a different kind of place, so some of the excess over 2 comes from busier airspace being more concentrated to begin with. The ratio avoids that problem by comparing each disc against itself with the geography removed. It does depend on the size of the disc: the null scatters aircraft across whatever disc we choose, so a wider disc makes random placement look emptier. Repeating the whole measurement at 50 and 200 km gives ratios at 5 NM of 3.89 and 11.34 and slopes of 2.62 and 2.37. The ratio moves with the scale as it should and the slope holds above 2 at every disc size tried.</p>\n\n        <figure class=\"plate\">\n          <img src=\"https://mouhssine.rifaki.me/notes/img/coordination-clustering.png\" alt=\"Two panels from measured air traffic. Panel A plots the mean number of aircraft pairs within 10 nautical miles and 1000 feet against the number of aircraft in a 100 km disc, on log axes, for the measured data and for the same aircraft scattered at random in the same disc. The measured curve has a slope of about 2.7 and stays above both the scattered curve and a dashed slope-2 guide. Panel B plots measured divided by scattered against the aircraft count for four lateral thresholds; the ratio is highest at 5 nautical miles and decays toward 1 at 40 nautical miles, and rises with the aircraft count.\" loading=\"lazy\" decoding=\"async\" width=\"1290\" height=\"744\">\n          <figcaption>Measured against a null with the geography removed. The wedge in panel A is the slope uniform mixing predicts, drawn as a reference so it does not assert an intercept. Panel B is the ratio of the two, and it rises with density, which is why the fitted slope comes out above 2.</figcaption>\n        </figure>\n\n        <p>This is manned aviation. The note concerns uncrewed traffic and I am using manned data on purpose. It is the only large population of vehicles that already shares airspace under a mature coordination regime, so it measures what a century of structure and procedure and separation standards actually achieves. The direction of the answer matters more than its size. All of that has not produced a sky that mixes uniformly. It has produced one that concentrates, and the model I was about to reason from understates the density a coordination service has to handle.</p>\n\n        <p>The second measurement concerns time. For every converging pair within 20 NM and 1,000 feet I computed the time to closest approach from the two velocity vectors. That comes to 26,137 pairs with a median of 158 seconds. The tail is the part that matters: 1,528 of those pairs, 5.8% of them, reach closest approach inside twelve seconds. The first percentile arrives at 1.9 seconds.</p>\n\n        <figure class=\"plate\">\n          <img src=\"https://mouhssine.rifaki.me/notes/img/coordination-cpa.png\" alt=\"A histogram on log-spaced time bins from one second to one hour, showing the time to closest approach for every converging aircraft pair measured within 20 nautical miles and 1000 feet. The distribution peaks in the three to five minute range with a median near 170 seconds, and a dark block of pairs under twelve seconds is visible at the left.\" loading=\"lazy\" decoding=\"async\" width=\"1290\" height=\"748\">\n          <figcaption>Time to closest approach for real converging pairs, on log-spaced bins. Linear bins put every short encounter in one column at the origin, and the short encounters are the ones a staleness budget is about.</figcaption>\n        </figure>\n\n        <p>Twelve seconds is not an arbitrary mark. A failure detector on one-second heartbeats needs three missed beats before it declares a link gone. An agreement about what to do costs at least two more message delays after that. Roughly a third of the budget goes on noticing, and by Fischer, Lynch and Paterson the remainder cannot be bounded at all in the asynchronous model. That leaves the 1,528 pairs in the left-hand block with no interval in which to spend it.</p>\n\n        <p>Which makes the timeout a safety parameter set by the geometry of the corridor it applies to and therefore something to publish per corridor. The harder consequence is that whatever the vehicle does inside those seconds has to work without agreement. Agreement is the one thing the budget cannot buy, and the theory here is unusually blunt about why.</p>\n\n        <h2>What the theory forbids</h2>\n\n        <p>Fischer, Lynch and Paterson showed that no deterministic protocol achieves consensus in an asynchronous system where a single process may fail [<a href=\"https://dl.acm.org/doi/10.1145/3149.214121\">5</a>]. A design that assumes the fleet will agree therefore has a timing model hidden inside it that someone should be made to state. Gilbert and Lynch's proof of Brewer's conjecture says a partitioned system chooses between staying available and staying consistent [<a href=\"https://dl.acm.org/doi/10.1145/564585.564601\">6</a>]. For an aircraft that reads as acting on a belief it can no longer refresh or stopping. Bernstein, Zilberstein and Immerman put the price of the missing channel in complexity terms. The finite-horizon problem for two agents who cannot share their observations is complete for nondeterministic exponential time [<a href=\"https://arxiv.org/abs/1301.3836\">7</a>], against the PSPACE-completeness Papadimitriou and Tsitsiklis had established for the same horizon under a single controller [<a href=\"https://doi.org/10.1287/moor.12.3.441\">8</a>].</p>\n\n        <p>Halpern and Moses is the one I keep coming back to. They formalized what a group of processes can know about what the others know. Their finding on common knowledge, the regress where everyone knows that everyone knows, is that &quot;formally speaking, in practical systems common knowledge cannot be attained&quot; [<a href=\"https://arxiv.org/abs/cs/0006009\">9</a>]. A protocol whose correctness rests on it therefore has no implementation whatever the hardware. That same paper introduces weaker variants that are attainable and those are where a working design has to aim.</p>\n\n        <p>Set that beside a coordination service. The service sends both aircraft the same deconfliction plan and both receive it. Each of them now holds a fact and neither can establish that the other holds it, nor that the other knows they hold it. The regress runs on without terminating. Two parties who each hold a fact are a weaker thing than two parties who can each rely on the other acting on it, and the difference between those is exactly what a coordination service would have to sell. Since no protocol closes that gap, whatever reliance two aircraft place in each other must have been built into them before either took off.</p>\n\n        <h2>The contingency</h2>\n\n        <p>The rule's own contingency provision shows this happening. Proposed \u00a7 108.815(b) [<a href=\"https://www.federalregister.gov/documents/2025/08/07/2025-14992/normalizing-unmanned-aircraft-systems-beyond-visual-line-of-sight-operations\">1</a>] requires each airframe to execute a safe predetermined action at link timeout and each of those actions is safe on its own terms. What no one has checked is how they compose. The only moment they will ever run is the moment they all run together, since the link they lost was in many cases the same link.</p>\n\n        <p>I drew this scenario before I measured it. Five aircraft separated along one corridor, two returning to a pad at one end and three to a pad at the other, their contingency paths crossing six times in a tidy knot in the middle. It was a persuasive picture and it was wrong. I only know that because the same data that produced the two figures above can be asked the same question.</p>\n\n        <p>The question I put to the data was this. Take every configuration of 6 or more aircraft below 10,000 feet inside one of those discs, send each aircraft to the nearest of the 821 large and medium airports in the continental United States [<a href=\"https://ourairports.com/data/\">21</a>], and count the crossings between pairs that were separated at the moment the link went. Across 1,696 configurations holding 29,580 aircraft the count of crossings came to 0. That is a number I refused to publish until the crossing predicate had been checked against an independent orientation test on twenty thousand random segment pairs, a check that now runs every time the script does.</p>\n\n        <p>The zero turns out to be a property of the rule and not of the sky. Sending every aircraft to its nearest field keeps each path inside one cell of the Voronoi diagram of those fields: the cell is an intersection of half-planes and therefore convex [<a href=\"https://doi.org/10.1145/116873.116880\">22</a>], the aircraft is inside it, and the destination is the site that generates it, so the whole path is inside it too. Two aircraft in different cells never meet and two in the same cell meet only at the runway they already share. To make sure that argument was doing the work I recounted the same configurations with the aircraft rotated onto each other's fields. Every position and every destination is unchanged and only the assignment moves, and the count comes to 201,754 crossings in 1,626 of the 1,696 configurations. So the phrase \u201cnearest airport\u201d is carrying a geometric guarantee, and the guarantee holds only for as long as every aircraft resolves that phrase to the same field. Two airframes running different airport databases are two airframes with different Voronoi diagrams.</p>\n\n        <figure class=\"plate\">\n          <img src=\"https://mouhssine.rifaki.me/notes/img/coordination-return.png\" alt=\"Two panels. Panel A is a plan view of one measured configuration: 32 real aircraft below 10,000 feet inside a 100 km disc centered at 35.5 degrees north and 81.4 degrees west, each joined by a line to the nearest of 3 real airports, with 28 of them converging on KCLT in a single funnel. Panel B is a histogram of how many aircraft arrive at one field within two minutes of each other across all 1,696 measured configurations, peaking at two and three and reaching 21.\" loading=\"lazy\" decoding=\"async\" width=\"1290\" height=\"800\">\n          <figcaption>Real positions, real fields, real ground speeds. In this one, 28 of 32 aircraft pick KCLT. The picture I had drawn showed paths crossing. The measurement shows a funnel, and rotating the same aircraft onto each other's fields turns 0 crossings into 201,754.</figcaption>\n        </figure>\n\n        <p>Converging is a worse property than crossing. In 67% of those configurations three or more aircraft reach the same field within two minutes of each other at their current ground speeds, and the worst has 21. Nothing in the pre-takeoff deconfliction set knows about it, because a contingency path is not an operational intent and was never in the set that was checked. Nothing after takeoff can negotiate it either. The link is what failed, so a set of individually compliant aircraft converges on one runway with no mechanism anywhere in the system that would notice.</p>\n\n        <p>Aviation has already run this experiment at full scale. Over \u00dcberlingen on 1 July 2002 a Tupolev Tu-154M and a Boeing 757 collided and 71 people died. The usual telling is that the crews received contradictory instructions. That is true and slightly misses the point. Both aircraft carried a working collision-avoidance system and both were talking to a working controller. Those are two coordination layers. Each was internally correct and each produced a resolution that would have been sufficient on its own. What did not exist was a rank between them. The Tupolev crew followed the controller and the Boeing crew followed the equipment. Among the immediate causes the German federal investigator recorded that the Tupolev crew &quot;followed the ATC instruction to descend and continued to do so even after TCAS advised them to climb&quot;. Among the systemic ones it found that the rules for ACAS and TCAS issued by ICAO and by national authorities and by the manufacturer and by the operators &quot;were not standardised, incomplete and partially contradictory&quot;. Its first recommendation asked ICAO to require pilots to follow a resolution advisory &quot;regardless of whether contrary ATC instruction is given prior to, during, or after&quot; it [<a href=\"https://www.bfu-web.de/EN/Publications/FinalReports/2002/Report_02_AX001-1-2_Ueberlingen_Report.pdf?__blob=publicationFile&amp;v=1\">10</a>]. Adding a second correct coordination layer to a system that already has one makes it worse until the two are ranked against each other, and that ranking has to be settled when the system is designed.</p>\n\n        <h2>The convention is in the weights</h2>\n\n        <p>Everything so far applies to hand-written controllers. Learned ones fail the layering plan in a way that is harder to see and probably harder to fix.</p>\n\n        <p>Coordination between two agents means agreeing on an arbitrary choice: who yields, which side we pass on, which lever we pull. The choice is arbitrary in the strict sense that the alternatives are equally good. What makes one of them correct is only that the other party made it too. A policy trained by self-play has to settle on one and where it settles is decided by its own initialization and sampling noise. Lanctot and colleagues named the resulting co-adaptation joint-policy correlation [<a href=\"https://arxiv.org/abs/1711.00832\">11</a>]. Carroll and colleagues built a coordination task on Overcooked and found that agents from self-play and from population-based training each assume a partner much like itself and converge on protocols that fail against a human or a model of one [<a href=\"https://arxiv.org/abs/1910.05789\">12</a>]. Hu and colleagues framed it as the zero-shot coordination problem in which self-play &quot;can produce agents that establish highly specialized conventions that do not carry over to novel partners&quot;. Their remedy works by exploiting the very symmetries such a convention settles at random [<a href=\"https://arxiv.org/abs/2003.02979\">13</a>].</p>\n\n        <p>I wanted a number for how large the effect is, so I ran the smallest version of it I could. The game is Hu's lever game. Two players independently choose one of ten levers. Matching levers pay that lever's value and mismatching levers pay nothing. One lever pays 1.0 and the other nine pay 0.9. Ten agreements exist and nine are worse by a hair. I trained eight policies by self-play with REINFORCE on a logit vector, one per seed, then scored every policy against every other in the same frame. The whole thing is one file of numpy with a forty-five line training loop and it takes about twenty seconds to run. <a href=\"https://mouhssine.rifaki.me/notes/experiments/results.json\">What it produced</a> is here, seeds and hyperparameters and environment included.</p>\n\n        <figure class=\"plate\">\n          <img src=\"https://mouhssine.rifaki.me/notes/img/coordination-crossplay.png\" alt=\"Two eight by eight heatmaps of expected return for every ordered pair of trained policies. Panel A, self-play: the diagonal is 0.90 and almost every off-diagonal cell is 0.00, with a mean of 0.06; the eight seeds chose 6 distinct levers and none of them chose the payoff-best lever. Panel B, other-play: six of the eight seeds converged on the best lever and score 1.0 against each other, while seeds 4 and 5 score near zero against everything, giving a diagonal of 0.79 and an off-diagonal mean of 0.54.\" loading=\"lazy\" decoding=\"async\" width=\"1290\" height=\"856\">\n          <figcaption>Every trained policy against every other, evaluated in closed form with no sampling. A self-play score reads the diagonal of panel A and reports 0.90. An independently trained partner delivers 0.06, smaller by a factor of 14. The eight seeds picked six distinct levers between them, and not one picked the lever worth more.</figcaption>\n        </figure>\n\n        <p>Every one of those policies is individually optimal and converged and would pass any test that pairs it with itself or with a copy of itself. What none of them carries is a way to tell us which lever it settled on, or to be told to settle on a different one. We cannot publish a standard that says take lever 1 and have these policies comply, because the convention exists only as a pattern spread across the weights and there is no field to write a different one into.</p>\n\n        <p>Which is why the known fix changes the training objective. Other-play trains each policy against a relabeling of itself drawn fresh each episode from the game's symmetry group. A policy that locked onto one of the nine interchangeable levers now meets a partner who locked onto a different one eight times in nine, and the only strategy the gradient can reward is the one every relabeling leaves alone. Panel B is the same experiment with that one change and cross-play goes from 0.06 to 0.54.</p>\n\n        <p>It also stops at 0.54, and the reason turned out to be more interesting than a clean result would have been. Write a policy as mass $q$ on the distinguished lever with the rest spread over the nine. Its expected return under a random relabeling is $$J(q) = q^2 + \\frac{0.9\\,(1-q)^2}{9},$$ a parabola whose interior minimum is at $q = 1/11$. Gradient ascent leaves that point in whichever direction it starts from. Six of my eight seeds initialized above $1/11$ and every one of them converged to $q = 0.9998$. Two initialized below it and both converged to $q = 0.0024$. Ten thousand steps at four times the step size moved neither group.</p>\n\n        <figure class=\"plate\">\n          <img src=\"https://mouhssine.rifaki.me/notes/img/coordination-barrier.png\" alt=\"Two panels. Panel A plots the other-play objective J of q against q from 0 to 1, a parabola with an interior minimum at q equals one eleventh, marked with a dashed line, from which gradient ascent departs in whichever direction it starts. Panel B plots q after training against q at initialization for the eight seeds: the six that started above one eleventh all finished at 1, the two that started below it finished at 0, and the transition is a step at the barrier.\" loading=\"lazy\" decoding=\"async\" width=\"1290\" height=\"660\">\n          <figcaption>Whether other-play works on a given seed is settled before the first gradient step. The barrier at $1/11$ is derived in closed form and then measured; the eight seeds sort onto either side of it with no exceptions.</figcaption>\n        </figure>\n\n        <p>So the fix works and whether it works on any particular run is decided by the initialization. That is a toy and I want to be careful about how far it generalizes. The lever game has an exact symmetry group that I can write down and sample from. Airspace conventions are approximate, the group is unknown in advance, and a real policy has far more ways to settle an arbitrary choice than ten. Every one of those differences makes the problem harder than the one I ran.</p>\n\n        <p>If the convention is in the weights then which weights are flying is a safety-relevant fact about an aircraft, in the same category as its position. Look at what the proposed broadcast actually carries under \u00a7 108.195(a)(2)(ii) [<a href=\"https://www.federalregister.gov/documents/2025/08/07/2025-14992/normalizing-unmanned-aircraft-systems-beyond-visual-line-of-sight-operations\">1</a>]: latitude, longitude, geometric altitude, velocity, an ICAO 24-bit address and three integrity figures. All three integrity figures describe the position source. None of them describes the thing making the decisions. Three policy versions rolling out across eight aircraft in a volume give 6,561 ways to assign versions to airframes and 45 distinct mixes even when the airframes are interchangeable. Certification evidence covers the three configurations where every airframe is on the same build.</p>\n\n        <h2>What it would take to train the contract in</h2>\n\n        <p>If the contract has to be in the policy then the co-player distribution and the partition profile have to be in the training environment, which makes the environment the thing we are designing. That is the premise of unsupervised environment design. Dennis and colleagues make the generator an agent rewarded for regret. They define regret as the gap between what an antagonist achieves in a configuration and what the learner achieves, which pushes the generator toward configurations at the frontier of what the learner can handle, since an unsolvable one earns it no regret at all [<a href=\"https://arxiv.org/abs/2012.02096\">14</a>]. Jiang and colleagues later showed that prioritized replay of randomly generated levels belongs to the same family and earned the approach a robustness guarantee at Nash equilibria [<a href=\"https://arxiv.org/abs/2110.02439\">15</a>]. The configuration covers density, version mix, link profile and geometry. That is what the search ranges over.</p>\n\n        <p>The reported numbers then have to be joint or the exercise reverts to per-agent testing with extra steps: conflicts per flight hour at each density, the worst operator's detour alongside the mean, and the version mix that was airborne. All of it against a baseline of fixed structure and no learning, since a coordination layer that cannot beat assigned altitudes and a precedence rule has not earned its complexity. Melting Pot exists to score exactly this kind of generalization to unfamiliar co-players [<a href=\"https://arxiv.org/abs/2107.06857\">16</a>]. Mixed-autonomy traffic supplies the nearest fielded analog. Vinitsky and colleagues found that their learned controllers held a strategy that worked across penetration rates from 5% to 40% where a hand-tuned feedback controller &quot;degrade[d] immediately upon penetration rate variation&quot; [<a href=\"https://arxiv.org/abs/2011.00120\">17</a>]. A controller tuned for one composition of the population stops working when the composition moves. That is the same difficulty the version mix presents in another setting.</p>\n\n        <p>I do not want to oversell this. A regret objective needs a best response to compare against and in a population the best response is itself a joint quantity that nobody can compute. Every practical estimator substitutes something cheaper and the substitution is where the guarantee leaks. I do not know how to close that gap. What I am confident of is the direction: the co-player distribution is a design input and treating it as one costs nothing at training time and cannot be recovered afterwards.</p>\n\n        <h2>The narrow window</h2>\n\n        <p>Whether a coordination service should exist is settled by the pair count. The question this note raises is whether one can be sufficient, and the answer is that the service cannot be the whole of the coordination. The layer's guarantee expires at takeoff and common knowledge is not something it can deliver. The conventions that would let two vehicles agree without it are set during training and unreachable after.</p>\n\n        <p>Some of what follows costs almost nothing. Contingency envelopes can go into the pre-takeoff deconfliction set alongside nominal intents, since the moment they are needed is the moment nothing can be negotiated. Precedence between two unmanned aircraft can be written into the rule, which is what the BFU asked ICAO to do after \u00dcberlingen. A precedence rule needs no round trip. A policy version can be a field in a broadcast. Wurman, D'Andrea and Mountz describe a Kiva installation for a large distribution center needing &quot;500 or more vehicles&quot;. They move on a weighted grid. System-wide resource allocation is centralized in one job manager and each vehicle plans its own path across that grid [<a href=\"https://aaai.org/ojs/index.php/aimagazine/article/view/2082\">18</a>].</p>\n\n        <p>The training is the expensive part. Retraining a fleet's policies against a co-player distribution that includes the other fleets is beyond what any one operator can do, and the current interface between operators has no way to express it even as a request. Kuchar and Yang separated the pairwise from the global in their survey of conflict detection twenty-six years ago [<a href=\"https://dl.acm.org/doi/10.1109/6979.898217\">19</a>]. Almost everything fielded since has been pairwise because pairwise is what one vehicle can reason about with what it can see.</p>\n\n        <p>The window in which the contract can still be trained in is the window before the fleets exist, and the fleets are being built now. Whether saying so changes any plan is not something I can affect. What I can do is set the argument out so it can be checked. The data and every measured number are linked where they are used.</p>\n\n        <h2>Further reading</h2>\n        <ul class=\"further\">\n          <li><a href=\"https://store.astm.org/f3548-21.html\">ASTM F3548-21, UAS Traffic Management USS Interoperability, 2021. The scope section is the shortest honest statement of what strategic deconfliction does and does not cover</a></li>\n          <li><a href=\"https://arxiv.org/abs/2202.10450\">R. Mirsky, I. Carlucho, A. Rahman, E. Fosong, W. Macke, M. Sridharan, P. Stone, and S. V. Albrecht. A survey of ad hoc teamwork research. arxiv 2202.10450, 2022</a></li>\n          <li><a href=\"https://arxiv.org/abs/1902.00506\">N. Bard et al. The Hanabi challenge: a new frontier for AI research. arxiv 1902.00506, 2019</a></li>\n          <li><a href=\"https://pubsonline.informs.org/doi/10.1287/trsc.1050.0127\">D. Braess, A. Nagurney, and T. Wakolbinger. On a paradox of traffic planning. Transportation Science, 2005. Adding a corridor can make every user worse off, which is worth knowing before opening one</a></li>\n          <li><a href=\"https://dl.acm.org/doi/10.1145/359545.359563\">L. Lamport. Time, clocks, and the ordering of events in a distributed system. CACM, 1978. A shared track built on wall clocks from independent receivers reorders events under load</a></li>\n          <li><a href=\"https://www.dcs.gla.ac.uk/~johnson/Eurocontrol/Ueberlingen/Ueberlingen_Final_Report.PDF\">C. Johnson et al. Review of the BFU \u00dcberlingen accident report. Eurocontrol, 2004</a></li>\n        </ul>\n\n<h2>References</h2>\n\n        <ul class=\"refs\">\n          <li>[<a href=\"https://www.federalregister.gov/documents/2025/08/07/2025-14992/normalizing-unmanned-aircraft-systems-beyond-visual-line-of-sight-operations\">1</a>] FAA and TSA. Normalizing unmanned aircraft systems beyond visual line of sight operations. Notice of proposed rulemaking, 90 FR 38212, 7 August 2025. Docket FAA-2025-1908, RIN 2120-AL82 and 1652-AA80.</li>\n          <li>[<a href=\"https://press.princeton.edu/books/paperback/9780691004129/normal-accidents\">2</a>] C. Perrow. Normal Accidents: Living with High-Risk Technologies. Basic Books, 1984. Princeton University Press edition, 1999.</li>\n          <li>[<a href=\"https://direct.mit.edu/books/oa-monograph/2908/Engineering-a-Safer-WorldSystems-Thinking-Applied\">3</a>] N. G. Leveson. Engineering a Safer World: Systems Thinking Applied to Safety. MIT Press, 2011. Open access.</li>\n          <li>[<a href=\"https://store.astm.org/f3548-21.html\">4</a>] ASTM International. F3548-21, Standard Specification for UAS Traffic Management (UTM) UAS Service Supplier (USS) Interoperability, 2021.</li>\n          <li>[<a href=\"https://dl.acm.org/doi/10.1145/3149.214121\">5</a>] M. J. Fischer, N. A. Lynch, and M. S. Paterson. Impossibility of distributed consensus with one faulty process. Journal of the ACM, 32(2):374-382, 1985.</li>\n          <li>[<a href=\"https://dl.acm.org/doi/10.1145/564585.564601\">6</a>] S. Gilbert and N. Lynch. Brewer's conjecture and the feasibility of consistent, available, partition-tolerant web services. ACM SIGACT News, 33(2):51-59, 2002.</li>\n          <li>[<a href=\"https://arxiv.org/abs/1301.3836\">7</a>] D. S. Bernstein, S. Zilberstein, and N. Immerman. The complexity of decentralized control of Markov decision processes. arxiv 1301.3836, 2000. Journal version, with R. Givan: Mathematics of Operations Research, 27(4):819-840, 2002.</li>\n          <li>[<a href=\"https://doi.org/10.1287/moor.12.3.441\">8</a>] C. H. Papadimitriou and J. N. Tsitsiklis. The complexity of Markov decision processes. Mathematics of Operations Research, 12(3):441-450, 1987.</li>\n          <li>[<a href=\"https://arxiv.org/abs/cs/0006009\">9</a>] J. Y. Halpern and Y. Moses. Knowledge and common knowledge in a distributed environment. Journal of the ACM, 37(3):549-587, 1990. arxiv cs/0006009.</li>\n          <li>[<a href=\"https://www.bfu-web.de/EN/Publications/FinalReports/2002/Report_02_AX001-1-2_Ueberlingen_Report.pdf?__blob=publicationFile&amp;v=1\">10</a>] Bundesstelle f\u00fcr Flugunfalluntersuchung. Investigation report AX001-1-2/02, collision near \u00dcberlingen, 1 July 2002.</li>\n          <li>[<a href=\"https://arxiv.org/abs/1711.00832\">11</a>] M. Lanctot, V. Zambaldi, A. Gruslys, A. Lazaridou, K. Tuyls, J. Perolat, D. Silver, and T. Graepel. A unified game-theoretic approach to multiagent reinforcement learning. arxiv 1711.00832, 2017.</li>\n          <li>[<a href=\"https://arxiv.org/abs/1910.05789\">12</a>] M. Carroll, R. Shah, M. K. Ho, T. L. Griffiths, S. A. Seshia, P. Abbeel, and A. Dragan. On the utility of learning about humans for human-AI coordination. arxiv 1910.05789, 2019.</li>\n          <li>[<a href=\"https://arxiv.org/abs/2003.02979\">13</a>] H. Hu, A. Lerer, A. Peysakhovich, and J. Foerster. \"Other-play\" for zero-shot coordination. arxiv 2003.02979, 2020. The lever game and the other-play objective used in this note are theirs; the barrier measurement is mine.</li>\n          <li>[<a href=\"https://arxiv.org/abs/2012.02096\">14</a>] M. Dennis, N. Jaques, E. Vinitsky, A. Bayen, S. Russell, A. Critch, and S. Levine. Emergent complexity and zero-shot transfer via unsupervised environment design. arxiv 2012.02096, 2020.</li>\n          <li>[<a href=\"https://arxiv.org/abs/2110.02439\">15</a>] M. Jiang, M. Dennis, J. Parker-Holder, J. Foerster, E. Grefenstette, and T. Rockt\u00e4schel. Replay-guided adversarial environment design. arxiv 2110.02439, 2021.</li>\n          <li>[<a href=\"https://arxiv.org/abs/2107.06857\">16</a>] J. Z. Leibo, E. Du\u00e9\u00f1ez-Guzm\u00e1n, A. S. Vezhnevets, J. P. Agapiou, P. Sunehag, R. Koster, et al. Scalable evaluation of multi-agent reinforcement learning with Melting Pot. arxiv 2107.06857, 2021.</li>\n          <li>[<a href=\"https://arxiv.org/abs/2011.00120\">17</a>] E. Vinitsky, N. Lichtle, K. Parvate, and A. Bayen. Optimizing mixed autonomy traffic flow with decentralized autonomous vehicles and multi-agent RL. arxiv 2011.00120, 2020.</li>\n          <li>[<a href=\"https://aaai.org/ojs/index.php/aimagazine/article/view/2082\">18</a>] P. R. Wurman, R. D'Andrea, and M. Mountz. Coordinating hundreds of cooperative, autonomous vehicles in warehouses. AI Magazine, 29(1):9-20, 2008.</li>\n          <li>[<a href=\"https://dl.acm.org/doi/10.1109/6979.898217\">19</a>] J. K. Kuchar and L. C. Yang. A review of conflict detection and resolution modeling methods. IEEE Transactions on Intelligent Transportation Systems, 1(4):179-189, 2000.</li>\n          <li>[<a href=\"https://doi.org/10.1109/IPSN.2014.6846743\">20</a>] M. Sch\u00e4fer, M. Strohmeier, V. Lenders, I. Martinovic, and M. Wilhelm. Bringing up OpenSky: a large-scale ADS-B sensor network for research. IPSN 2014, pages 83-94. The state vectors measured here came from its live API.</li>\n          <li>[<a href=\"https://ourairports.com/data/\">21</a>] OurAirports. Public-domain airport database. The 821 continental US large and medium fields used for the return measurement.</li>\n          <li>[<a href=\"https://doi.org/10.1145/116873.116880\">22</a>] F. Aurenhammer. Voronoi diagrams: a survey of a fundamental geometric data structure. ACM Computing Surveys, 23(3):345-405, 1991. Each cell is an intersection of half-planes and therefore convex, which is the whole of the argument about the contingency paths.</li>\n        </ul>",
      "date_published": "2026-07-07T12:00:00Z",
      "date_modified": "2026-08-24T12:00:00Z",
      "tags": [
        "autonomy",
        "coordination"
      ]
    },
    {
      "id": "https://mouhssine.rifaki.me/notes/diffusion-models/",
      "url": "https://mouhssine.rifaki.me/notes/diffusion-models/",
      "title": "Score matching and diffusion",
      "summary": "The setup goes back to Sohl-Dickstein et al. [1]. The forward chain takes data x0 and adds Gaussian noise according to a fixed schedule \u03b1\u00aft, producing intermediate samples x1,\u2026,xT via xt=\u03b1\u00aftx0+1\u2212\u03b1\u00aft\u03f5 with\u2026",
      "content_html": "<!-- date: 2026-04-22 -->\n<p>The setup goes back to Sohl-Dickstein et al. [<a href=\"https://arxiv.org/abs/1503.03585\">1</a>]. The forward chain takes data $x_0$ and adds Gaussian noise according to a fixed schedule $\\bar\\alpha_t$, producing intermediate samples $x_1, \\ldots, x_T$ via \\[x_t = \\sqrt{\\bar\\alpha_t}\\, x_0 + \\sqrt{1-\\bar\\alpha_t}\\,\\epsilon\\] with $\\epsilon \\sim \\mathcal{N}(0, I)$. The schedule is chosen so that $x_T$ is essentially a standard Gaussian. The reverse chain undoes the corruption: sampling from the data distribution reduces to learning the conditional $p(x_{t-1} \\mid x_t)$ at each step.</p>\n\n        <h2>Score matching</h2>\n\n        <p>The DDPM forward chain has a clean dual under score matching, and once the two are placed side by side they are not separate ideas. The score function $s(x,t) = \\nabla_x \\log p_t(x)$ is the gradient of the log-density of the noisy distribution at noise level $t$ evaluated at $x$; it points in the direction along which the log-density rises most steeply. Hyv\u00e4rinen's original score-matching objective is the square norm of the score plus the trace of the Hessian of the log-density. The Hessian-trace term is expensive in high dimensions. Vincent's denoising score matching gets around this: with Gaussian noise of variance $\\sigma^2$, the optimal MMSE denoiser $D^*(\\widetilde{x}) = \\mathbb{E}[x \\mid \\widetilde{x}]$ satisfies $$\\nabla_{\\widetilde{x}} \\log p_\\sigma(\\widetilde{x}) = \\frac{D^*(\\widetilde{x}) - \\widetilde{x}}{\\sigma^2},$$ so predicting the noise (equivalently, predicting the clean image) is the same problem as estimating the score of the noisy distribution, with mean-squared error as the loss.</p>\n\n        <p>Song and Ermon [<a href=\"https://arxiv.org/abs/1907.05600\">2</a>] estimated these scores at a sweep of noise levels and used annealed Langevin dynamics to draw samples from the estimates. Ho, Jain, and Abbeel showed that DDPMs trained with a weighted variational bound reduce, in a particular weighting limit, to denoising score matching at multiple noise levels. The point shared by both: learning to predict the noise that was added at a given level is enough to recover the score of the noisy distribution at that level, and sampling is then iterated denoising along the noise schedule.</p>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/diffusion-score-field.png\" alt=\"Sliced score matching loss trajectories on a toy problem: the unconstrained estimator diverges to large negative values while the noise-conditioned variant stays bounded over training iterations.\" loading=\"lazy\" decoding=\"async\" width=\"2678\" height=\"978\">\n          <figcaption>Training-stability comparison from the score-matching literature (Song et al. [<a href=\"https://arxiv.org/abs/1907.05600\">2</a>] and surrounding work). The score-matching loss is unbounded below without noise conditioning; injecting noise restores stable training.</figcaption>\n        </figure>\n\n        <h2>SDE view</h2>\n\n        <p>Song et al. unified the discrete and score-based views inside stochastic differential equations. The forward process is an It\u00f4 SDE $$dx_t=f\\left( x_t,t\\right)\\, dt+g\\left( t\\right)\\, dw_t$$ that continuously corrupts data into noise. Anderson's reverse-time formula gives a backward SDE driven by the score $$dx_t=\\left[ f\\left( x_t,t\\right)-g^{2}\\left( t\\right)\\nabla _{x}\\log p_{t}\\left( x_{t}\\right) \\right] dt+g\\left( t\\right)\\, d\\bar {w}_{t},$$ and a deterministic probability-flow ODE with the same marginals $$dx_{t}=\\left[f\\left( x_{t},t\\right)-\\frac{1}{2}g^{2}\\left( t\\right)\\nabla _{x}\\log p_{t}\\left( x_{t}\\right)\\right] dt.$$ The ODE is useful for likelihood evaluation and fast sampling; the SDE is what most early samplers used. DDPMs, score-based models, and probability-flow ODE samplers are different discretizations of the same underlying dynamics.</p>\n\n        <p>The SDE view also separates modeling from numerics. The noise schedule, solver, parameterization, and preconditioner can all be changed without touching the underlying problem of learning the score.</p>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/diffusion-sde-ode-map.png\" alt=\"Figure 1 of Song et al. 2020. Forward SDE turning data into noise (top) and the score-driven reverse SDE turning noise back into data (bottom), with intermediate sample crops at increasing noise levels.\" loading=\"lazy\" decoding=\"async\" width=\"2582\" height=\"1112\">\n          <figcaption>Figure 1 of Song et al. [<a href=\"https://arxiv.org/abs/2011.13456\">4</a>]. The learned score is the object shared by the stochastic reverse process and (in the same paper) the deterministic probability-flow ODE.</figcaption>\n        </figure>\n\n        <figure class=\"tweet-embed\">\n          <blockquote class=\"twitter-tweet\" data-dnt=\"true\"><p lang=\"en\" dir=\"ltr\">There is a lot of great writing on flow matching out there all of a sudden! This post clarifies the connection with diffusion models -- they are essentially two different ways to describe the same class of models. <a href=\"https://t.co/lLokMmxxdz\">https://t.co/lLokMmxxdz</a></p>&mdash; Sander Dieleman (@sedielem) <a href=\"https://twitter.com/sedielem/status/1863661809355362538?ref_src=twsrc%5Etfw\">December 2, 2024</a></blockquote>\n        </figure>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/diffusion-forward-reverse.png\" alt=\"Figure 2 of Ho, Jain, Abbeel 2020 (DDPM). Directed graphical model of the forward and reverse chains between x_T and x_0.\" loading=\"lazy\" decoding=\"async\" width=\"2858\" height=\"484\">\n          <figcaption>Figure 2 of Ho et al. [<a href=\"https://arxiv.org/abs/2006.11239\">3</a>]. The forward chain corrupts data into noise; the learned reverse chain inverts it step by step.</figcaption>\n        </figure>\n\n\n        <p>Lu et al. [<a href=\"https://arxiv.org/abs/2206.00927\">6</a>] built DPM-Solver out of the semi-linear structure of the probability-flow ODE and cut sampling from thousands of steps to tens, with no retraining required.</p>\n\n        <h2>Karras et al.</h2>\n\n        <p>Karras et al. [<a href=\"https://arxiv.org/abs/2206.00364\">5</a>] argued that diffusion practice had become unnecessarily entangled: sampling schedule, loss weighting, noise parameterization, network preconditioning, and solver choice were all bundled together. Pulling each design decision apart showed that most of the published performance gains came from untangling the design space, not from a new generative principle. Their EDM recipe (continuous noise levels indexed by $\\sigma$, $\\sigma$-conditioned network preconditioning, and a second-order Heun ODE solver) has since become a common baseline.</p>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/diffusion-design-space.png\" alt=\"Table 1 of Karras et al. 2022 (EDM). Explicit design-space tabulation of noise schedules, prediction targets, preconditioning, and samplers across DDPM, NCSN, EDM, and related variants.\" loading=\"lazy\" decoding=\"async\" width=\"1822\" height=\"1078\">\n          <figcaption>Table 1 of Karras et al. [<a href=\"https://arxiv.org/abs/2206.00364\">5</a>]. The score-estimation problem stays fixed while schedules, targets, preconditioning, and solvers move around it.</figcaption>\n        </figure>\n\n        <p>The implementation history runs through three papers. Dhariwal and Nichol showed that diffusion models beat GANs on ImageNet, by combining classifier guidance with architecture and training changes. Ho and Salimans then dropped the auxiliary classifier in favor of classifier-free guidance, training the conditional and unconditional scores jointly and combining them at sampling time. Rombach et al. moved the diffusion process inside the latent space of a pre-trained autoencoder, which is what made high-resolution diffusion feasible at academic compute and what Stable Diffusion is built on.</p>\n\n        <p>Natural data sits near low-dimensional manifolds where direct density modeling is hard. Adding white noise thickens the manifold: at high noise levels the distribution is smooth, at low noise levels it is detailed but local. Diffusion replaces a single full-density estimation problem with a sequence of denoising problems indexed by noise level. Compare with autoregressive models, which generate sequentially conditioned on prior tokens, and with flows, which accept restrictions on the transforms they can express in exchange for tractable likelihoods. Diffusion's training objective is stable in a way GAN training is not, while paying for it with iterative sampling that DPM-Solver, consistency models, and distillation have largely clawed back.</p>\n\n        <p>Semantic abstraction inside the model is still theoretically incomplete. The score tells you how to move from a noisy sample toward a clean one. It does not say why text prompts, latent-space guidance, or multimodal conditioning organize concepts the way they do. Those mechanisms are built on top of a fixed score-matching core, but the core does not force any of them.</p>\n\n        <p>Whether diffusion is the right parameterization is also unclear to me. Flow matching frames generation as learning a vector field that transports a simple prior to the data distribution along possibly non-straight paths, with a regression objective that does not require an SDE. Rectified flow constrains the paths to be nearly straight, which keeps few-step sampling accurate. Consistency models compress the iterative denoising sampler into a single-step generator while preserving sample quality. None of these compete with score-based modeling; they are alternative parameterizations of the same vector-field-across-noise-levels problem.</p>\n        <h2>Further reading</h2>\n        <ul class=\"further\">\n          <li><a href=\"https://arxiv.org/abs/2105.05233\">P. Dhariwal and A. Nichol. Diffusion models beat GANs on image synthesis. arxiv 2105.05233, 2021</a></li>\n          <li><a href=\"https://arxiv.org/abs/2207.12598\">J. Ho and T. Salimans. Classifier-free diffusion guidance. arxiv 2207.12598, 2022</a></li>\n          <li><a href=\"https://arxiv.org/abs/2112.10752\">R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. arxiv 2112.10752, 2021</a></li>\n          <li><a href=\"https://arxiv.org/abs/2210.02747\">Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative modeling. arxiv 2210.02747, 2022</a></li>\n          <li><a href=\"https://arxiv.org/abs/2209.03003\">X. Liu, C. Gong, and Q. Liu. Flow straight and fast: learning to generate and transfer data with rectified flow. arxiv 2209.03003, 2022</a></li>\n          <li><a href=\"https://arxiv.org/abs/2303.01469\">Y. Song, P. Dhariwal, M. Chen, and I. Sutskever. Consistency models. arxiv 2303.01469, 2023</a></li>\n          <li><a href=\"https://www.jmlr.org/papers/v6/hyvarinen05a.html\">A. Hyv&auml;rinen. Estimation of non-normalized statistical models by score matching</a></li>\n          <li><a href=\"https://www.iro.umontreal.ca/~vincentp/Publications/smdae_techreport.pdf\">P. Vincent. A connection between score matching and denoising autoencoders</a></li>\n        </ul>\n\n        \n\n<h2>References</h2>\n\n        <ul class=\"refs\">\n          <li>[<a href=\"https://arxiv.org/abs/1503.03585\">1</a>] J. Sohl-Dickstein, E. A. Weiss, N. Maheswaranathan, and S. Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. arxiv 1503.03585, 2015.</li>\n          <li>[<a href=\"https://arxiv.org/abs/1907.05600\">2</a>] Y. Song and S. Ermon. Generative modeling by estimating gradients of the data distribution. arxiv 1907.05600, 2019.</li>\n          <li>[<a href=\"https://arxiv.org/abs/2006.11239\">3</a>] J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. arxiv 2006.11239, 2020.</li>\n          <li>[<a href=\"https://arxiv.org/abs/2011.13456\">4</a>] Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole. Score-based generative modeling through stochastic differential equations. arxiv 2011.13456, 2020.</li>\n          <li>[<a href=\"https://arxiv.org/abs/2206.00364\">5</a>] T. Karras, M. Aittala, T. Aila, and S. Laine. Elucidating the design space of diffusion-based generative models. arxiv 2206.00364, 2022.</li>\n          <li>[<a href=\"https://arxiv.org/abs/2206.00927\">6</a>] C. Lu, Y. Zhou, F. Bao, J. Chen, C. Li, and J. Zhu. DPM-Solver: a fast ODE solver for diffusion probabilistic model sampling in around 10 steps. arxiv 2206.00927, 2022.</li>\n        </ul>",
      "date_published": "2026-04-22T12:00:00Z",
      "date_modified": "2026-08-21T12:00:00Z",
      "tags": [
        "generative models",
        "theory"
      ]
    },
    {
      "id": "https://mouhssine.rifaki.me/notes/adversarial-examples/",
      "url": "https://mouhssine.rifaki.me/notes/adversarial-examples/",
      "title": "Adversarial examples",
      "summary": "Adversarial examples initially seemed an oddity. Szegedy et al. [1] demonstrated that a minuscule perturbation, meaningless to human eyes, could confidently flip a neural net's prediction. My first\u2026",
      "content_html": "<!-- date: 2026-02-07 -->\n<p>Adversarial examples initially seemed an oddity. Szegedy et al. [<a href=\"https://arxiv.org/abs/1312.6199\">1</a>] demonstrated that a minuscule perturbation, meaningless to human eyes, could confidently flip a neural net's prediction. My first instinct on reading it was to blame weird nonlinearities or overfitting. It turned out to be neither. Goodfellow et al. [<a href=\"https://arxiv.org/abs/1412.6572\">2</a>] gave a much simpler explanation. A linear model with weight $w \\in \\mathbb{R}^d$ will have logit shift $w^\\top \\delta$ under perturbation $\\delta$, and the worst-case $\\ell_\\infty$ behavior inside of $\\|\\delta\\|_\\infty \\le \\epsilon$ is $\\epsilon \\|w\\|_1$, growing like $\\epsilon\\,d$ for dense weights, a tiny per-pixel perturbation accumulated across a high-dimensional input.</p>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/adversarial-fgsm.png\" alt=\"Figure 1 of Goodfellow, Shlens, Szegedy 2015. Panda image plus epsilon times sign of the gradient produces an imperceptible perturbation that flips the classifier's prediction to gibbon.\" loading=\"lazy\" decoding=\"async\" width=\"2550\" height=\"996\">\n          <figcaption>Figure 1 of Goodfellow et al. [<a href=\"https://arxiv.org/abs/1412.6572\">2</a>]. The perturbation is small under the threat model, but it is aligned with the classifier's loss gradient.</figcaption>\n        </figure>\n\n        <p>In a thousand-dimensional input space, even an imperceptible $\\delta$ can produce a logit shift that flips the prediction. Adversarial perturbations are then not a symptom of nonlinear extrema; they are a generic property of high-dimensional linear decision rules. FGSM is then a one-step linearized attack on the inner maximum $\\max_{\\|\\delta\\|_\\infty \\le \\epsilon} L(\\theta, x+\\delta, y)$. The fact that one step works at all is the diagnostic: the model is sensitive to a direction the data distribution does not mark as human-meaningful. Carlini and Wagner later showed that more carefully tuned attack objectives produce much smaller-norm perturbations than FGSM, breaking many defenses whose only evaluation had been against single-step attacks.</p>\n\n        <p>Madry et al. wrote down the saddle-point formulation: define a defense by the worst-case inner-max loss it can resist within a fixed threat model, and treat any defense that fails a stronger attack inside that threat model as broken. The split separates two questions that were previously tangled together: the inner problem defines what the attacker can do, and the outer problem defines what the model has to optimize against. Projected gradient descent then becomes both the canonical attack and the canonical training procedure. It is not perfect, but it made robustness measurable enough that defenses could be compared honestly. Many proposed defenses then turned out not to be robust; they just break weak attacks. Athalye, Carlini, and Wagner cataloged the failure modes under one heading - obfuscated gradients. Stochastic preprocessing, non-differentiable transforms, exploding or vanishing gradients, and gradient shattering each produce attacks that fail without producing classifiers that survive a stronger attack.</p>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/adversarial-minmax-loop.png\" alt=\"Figure 1 of Madry et al. 2018. PGD attack-loss curves over inner-maximization iterations on standard- and adversarially-trained MNIST and CIFAR10 networks: the standard models reach high attack loss easily, the robust models cap the attainable loss.\" loading=\"lazy\" decoding=\"async\" width=\"2888\" height=\"796\">\n          <figcaption>Figure 1 of Madry et al. [<a href=\"https://arxiv.org/abs/1706.06083\">3</a>]. PGD finds many high-loss perturbations on standard networks; on the adversarially-trained networks it caps out near a small bounded value.</figcaption>\n        </figure>\n\n        <p>They broke six of the nine ICLR-2018 defenses completely and a seventh partially, just by replacing the attack with a stronger one inside the same threat model. AutoAttack later turned that lesson into a parameter-free ensemble: a single attack you can run against a defense without per-defense tuning, which exposes inflated robustness numbers automatically. The other branch of progress is certified rather than empirical robustness. Cohen, Rosenfeld, and Kolter produce randomized-smoothing certificates: convolve the classifier with isotropic Gaussian noise of variance $\\sigma^2$ and the smoothed classifier $g(x) = \\arg\\max_c \\mathbb{P}_{\\eta \\sim \\mathcal{N}(0,\\sigma^2 I)}[f(x+\\eta) = c]$ is provably robust within an $\\ell_2$ ball of radius $\\sigma\\,\\Phi^{-1}(p_A)$, where $p_A$ is the lower confidence bound on the top-class probability.</p>\n\n        <figure class=\"tweet-embed\">\n          <blockquote class=\"twitter-tweet\" data-dnt=\"true\"><p lang=\"en\" dir=\"ltr\">The definition of &quot;adversarial examples&quot; I prefer these days is &quot;Adversarial examples are inputs to machine learning models that an attacker has intentionally designed to cause the model to make a mistake&quot; <a href=\"https://t.co/GiXiQBCp5L\">https://t.co/GiXiQBCp5L</a></p>&mdash; Ian Goodfellow (@goodfellow_ian) <a href=\"https://twitter.com/goodfellow_ian/status/984518755546906624?ref_src=twsrc%5Etfw\">April 12, 2018</a></blockquote>\n        </figure>\n\n\n        <p>The certificate is provable. The cost is that the radius is small in practice and only natural in $\\ell_2$. For a less paper-indexed entry point, the Gradient Science adversarial robustness page is still one of the better maps of attacks, defenses, and the evaluation traps. RobustBench is the practical scoreboard, once the question becomes: does this defense survive standard attacks?</p>\n\n        <h2>The accuracy trade-off</h2>\n\n        <p>Tsipras et al. [<a href=\"https://arxiv.org/abs/1805.12152\">4</a>] made the uncomfortable point that robustness can conflict with standard accuracy. In their construction, the standard classifier uses weak but highly predictive features that are not robust. The robust classifier has to ignore them and therefore loses ordinary accuracy. The empirical version of this trade-off is more complicated, but the conceptual point survived: robustness is not just accuracy with additional caution.</p>\n\n        <p>A robust classifier may have to learn different features altogether. Ilyas et al. [<a href=\"https://arxiv.org/abs/1905.02175\">5</a>] put it bluntly: adversarial examples are not bugs, they are features. Their claim was not that every attack direction is semantically relevant to humans. It was that standard datasets contain predictive signal that models can use and that humans do not recognize as robust evidence. Adversarial perturbations take advantage of those signals. Robust training suppresses them. Schmidt et al. [<a href=\"https://arxiv.org/abs/1804.11285\">6</a>] sharpened the price metric in statistical terms: the sample complexity of robust learning can be polynomially bigger than the sample complexity of standard learning, an information-theoretic gap that holds irrespective of the training algorithm or the model family. In their Gaussian-mixture model, standard generalization needs only constant sample complexity while robust generalization at $\\ell_\\infty$ radius $\\epsilon$ requires a polynomial-in-$d$ number of samples.</p>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/adversarial-feature-taxonomy.png\" alt=\"Figure 1 of Ilyas et al. 2019. Robust versus non-robust feature decomposition: standard models exploit non-robust features that are predictive but human-imperceptible.\" loading=\"lazy\" decoding=\"async\" width=\"2824\" height=\"832\">\n          <figcaption>Figure 1 of Ilyas et al. [<a href=\"https://arxiv.org/abs/1905.02175\">5</a>]. A feature can be genuinely predictive and still fail the invariance demanded by the threat model.</figcaption>\n        </figure>\n\n        <p>The gap comes from the structure of the learning problem, not the algorithm. Robustness is paying for invariance, and invariance costs in terms of sample complexity. Engstrom et al. then ran the experiment in the other direction: representations from robust classifiers transfer better than standard ones on a range of downstream tasks, look more semantically aligned in feature visualization, and yield gradients that resemble human-perceptible objects. Robustness, in this reading, is also a representation-learning prior. Whether the prior helps or hurts depends on the downstream task, but it is not free of structure.</p>\n\n        <p>The adversarial-examples literature forced a distinction between predictive validity and human-aligned validity. A feature can be statistically real, useful for test accuracy, and still unacceptable under a robustness constraint. That is a deeper issue than security: the supervised-learning objective does not fully specify the invariances I care about. To the extent that human perception is itself a strong inductive bias, models that do more representation learning end up closer to my geometry; robust models produce gradients and saliency maps that line up with what a person would identify as the object. In medical imaging, robotics, or safety-critical perception that trade-off is worth the cost. In low-stakes classification it often is not.</p>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/adversarial-robust-optimization.png\" alt=\"Adversarial-training loss curves of Madry et al. 2018: PGD-adversarial training loss decays from the initial saddle-point value over 100k MNIST iterations and 75k CIFAR10 iterations.\" loading=\"lazy\" decoding=\"async\" width=\"2500\" height=\"1002\">\n          <figcaption>Adversarial-training convergence in Madry et al. [<a href=\"https://arxiv.org/abs/1706.06083\">3</a>]. The robust optimization objective is solvable: the inner-max-then-outer-min loss curve descends to a bounded plateau under PGD adversarial training.</figcaption>\n        </figure>\n\n        <p>The \"non-robust feature\" label renames an older statistical fact. Predictive validity in distribution is not causal structure, and a model that maximizes the former will exploit signals the latter does not endorse. Adversarial training folds a robustness constraint into the objective. The cleaner long-term fix is on the data side: collect or augment so that the equivalence classes the human cares about are the equivalence classes the dataset enforces.</p>\n        <h2>Further reading</h2>\n        <ul class=\"further\">\n          <li><a href=\"https://arxiv.org/abs/2003.01690\">F. Croce and M. Hein. Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. arxiv 2003.01690, 2020</a></li>\n          <li><a href=\"https://arxiv.org/abs/1608.04644\">N. Carlini and D. Wagner. Towards evaluating the robustness of neural networks. arxiv 1608.04644, 2016</a></li>\n          <li><a href=\"https://arxiv.org/abs/1802.00420\">A. Athalye, N. Carlini, and D. Wagner. Obfuscated gradients give a false sense of security: circumventing defenses to adversarial examples. arxiv 1802.00420, 2018</a></li>\n          <li><a href=\"https://arxiv.org/abs/1902.02918\">J. M. Cohen, E. Rosenfeld, and J. Z. Kolter. Certified adversarial robustness via randomized smoothing. arxiv 1902.02918, 2019</a></li>\n          <li><a href=\"https://arxiv.org/abs/1906.00945\">L. Engstrom, A. Ilyas, S. Santurkar, D. Tsipras, B. Tran, and A. Madry. Adversarial robustness as a prior for learned representations. arxiv 1906.00945, 2019</a></li>\n        </ul>\n\n        \n\n<h2>References</h2>\n\n        <ul class=\"refs\">\n          <li>[<a href=\"https://arxiv.org/abs/1312.6199\">1</a>] C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus. Intriguing properties of neural networks. arxiv 1312.6199, 2013.</li>\n          <li>[<a href=\"https://arxiv.org/abs/1412.6572\">2</a>] I. J. Goodfellow, J. Shlens, and C. Szegedy. Explaining and harnessing adversarial examples. arxiv 1412.6572, 2014.</li>\n          <li>[<a href=\"https://arxiv.org/abs/1706.06083\">3</a>] A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu. Towards deep learning models resistant to adversarial attacks. arxiv 1706.06083, 2017.</li>\n          <li>[<a href=\"https://arxiv.org/abs/1805.12152\">4</a>] D. Tsipras, S. Santurkar, L. Engstrom, A. Turner, and A. Madry. Robustness may be at odds with accuracy. arxiv 1805.12152, 2018.</li>\n          <li>[<a href=\"https://arxiv.org/abs/1905.02175\">5</a>] A. Ilyas, S. Santurkar, D. Tsipras, L. Engstrom, B. Tran, and A. Madry. Adversarial examples are not bugs, they are features. arxiv 1905.02175, 2019.</li>\n          <li>[<a href=\"https://arxiv.org/abs/1804.11285\">6</a>] L. Schmidt, S. Santurkar, D. Tsipras, K. Talwar, and A. Madry. Adversarially robust generalization requires more data. arxiv 1804.11285, 2018.</li>\n        </ul>",
      "date_published": "2026-02-07T12:00:00Z",
      "date_modified": "2026-08-21T12:00:00Z",
      "tags": [
        "robustness",
        "theory"
      ]
    },
    {
      "id": "https://mouhssine.rifaki.me/notes/mode-connectivity/",
      "url": "https://mouhssine.rifaki.me/notes/mode-connectivity/",
      "title": "Mode connectivity",
      "summary": "The old picture of the loss surface as many isolated basins, one per initialization, has not held up. Freeman and Bruna [6] suggested early on that low-loss level sets stay connected. Garipov et al. [1]\u2026",
      "content_html": "<!-- date: 2025-11-04 -->\n<p>The old picture of the loss surface as many isolated basins, one per initialization, has not held up. Freeman and Bruna [<a href=\"https://arxiv.org/abs/1611.01540\">6</a>] suggested early on that low-loss level sets stay connected. Garipov et al. [<a href=\"https://arxiv.org/abs/1802.10026\">1</a>] and Draxler et al. [<a href=\"https://arxiv.org/abs/1803.00885\">2</a>] then made it concrete: two independently trained models can be joined by a smooth low-loss path. The minima are not isolated points; they are reachable from each other through a connected high-dimensional region.</p>\n\n <p>The experiment is direct. Train two models \\(\\theta_1\\) and \\(\\theta_2\\) to low loss, and interpolate linearly between them to define a one-parameter family \\[\\theta(\\alpha) = (1-\\alpha)\\,\\theta_1 + \\alpha\\,\\theta_2\\] for \\(\\alpha \\in [0,1]\\). The straight segment usually crosses a high-loss barrier \\[B = \\max_{\\alpha \\in [0,1]} L(\\theta(\\alpha)) - \\tfrac{1}{2}\\big(L(\\theta_1) + L(\\theta_2)\\big).\\] In other words, the loss jumps up substantially as soon as one steps off either endpoint, and the peak in the middle is typically far above either endpoint loss. Garipov and Draxler's contribution was to show that this barrier exists only along the straight line: if you allow curved paths, you can connect \\(\\theta_1\\) and \\(\\theta_2\\) with a path that stays at low loss throughout. The endpoints are not separated by an insurmountable barrier; the straight line is just the wrong path through parameter space.</p>\n\n <h2>Curves before lines</h2>\n\n <p>The first demonstrations of mode connectivity used non-linear paths. Garipov et al. parametrized the path as polygonal chains and B\u00e9zier curves; Draxler et al. used continuous non-linear paths produced by the Nudged Elastic Band method. Both findings weakened the previous picture in which SGD ends up in sharply-separated basins: if a low-loss path exists between two trained models, the connected low-loss region they both lie in is substantially larger than what a local Hessian analysis at either endpoint would suggest.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/mode-connectivity-paths.png\" alt=\"Figure 1 of Garipov et al. 2018. Loss-landscape view of a low-loss curve connecting two independently trained SGD minima, while the straight segment between them rises through a high-loss barrier.\" loading=\"lazy\" decoding=\"async\" width=\"2894\" height=\"600\">\n <figcaption>Figure 1 of Garipov et al. [<a href=\"https://arxiv.org/abs/1802.10026\">1</a>]. The straight segment between two minima crosses a barrier; the learned curve stays in low-loss territory.</figcaption>\n </figure>\n\n <p>Linear mode connectivity is strictly stronger than the curve-based version: it asks the loss to stay low along the straight segment between two minima, not along an arbitrary path. For independently-trained networks from distinct initializations this typically fails. The linearly connected case is essentially confined to one setup. Frankle, Dziugaite, Roy, and Carbin [<a href=\"https://arxiv.org/abs/1912.05671\">3</a>] formalized it as \"spawning\": train a model \\(\\theta_0\\), fork two copies after \\(k\\) steps using different SGD noise, and train each to convergence. Once \\(k\\) crosses a stability threshold (around 1000-2000 iterations on standard CIFAR networks, and a few percent of training on ImageNet), the two descendants end up linearly connected. They argue this is exactly the lottery-ticket basin: the connected region is the one that the matching sparse sub-network corresponds to, so linear-mode-connectivity becomes a practical test for whether two runs landed in the same effective basin and can therefore be merged without loss.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/mode-connectivity-spawning.png\" alt=\"Figure 3 of Frankle, Dziugaite, Roy, Carbin 2020. Linear interpolation curves for spawn-then-fork SGD pairs at varying late-rewinding points: late enough forks remain linearly connected.\" loading=\"lazy\" decoding=\"async\" width=\"2868\" height=\"582\">\n <figcaption>Figure 3 of Frankle et al. [<a href=\"https://arxiv.org/abs/1912.05671\">3</a>]. Fork late enough and the two descendants stay linearly connected.</figcaption>\n </figure>\n\n\n <h2>The permutation turn</h2>\n\n <p>Entezari et al. [<a href=\"https://arxiv.org/abs/2110.06296\">4</a>] reframed the geometry. Neural networks have permutation symmetries: swapping units in a hidden layer and unswapping them in the next layer leaves the function unchanged. Two models that look far apart in raw parameter coordinates might just be using different unit orderings of the same function. Quotient by those permutations and many independently trained models turn out to be connected by a simple low-loss curve.</p>\n\n <p>Git Re-Basin turns the idea into a concrete algorithm. Given two trained models $\\theta_1, \\theta_2$ with hidden-layer widths $\\{n_\\ell\\}$, it searches over permutation matrices $P_\\ell \\in S_{n_\\ell}$ for the alignment that minimizes $\\lVert \\theta_1 - P \\cdot \\theta_2 \\rVert$ under a weights or activations metric, then checks whether the aligned models can be merged in weight space. Singh and Jaggi give an optimal-transport version of the same idea: a soft assignment between units that reduces to permutation matching when the widths agree, and that handles mismatched widths when they don't. Neither paper proves that all minima sit in one basin, but together they make a strong empirical case that raw parameter-space interpolation overstates the separation. A meaningful fraction of the apparent barrier is just a bad coordinate system.</p>\n\n <figure class=\"tweet-embed\">\n <blockquote class=\"twitter-tweet\" data-dnt=\"true\"><p lang=\"en\" dir=\"ltr\">Say you train Model A. <br><br>Independently, your friend trains Model B, possibly on different data. <br><br>With Git Re-Basin, you can merge models A+B in weight space at _no cost to the loss_</p>&mdash; Samuel &quot;curry-howard fanboi&quot; Ainsworth (@SamuelAinsworth) <a href=\"https://twitter.com/SamuelAinsworth/status/1569719499263471616?ref_src=twsrc%5Etfw\">September 13, 2022</a></blockquote>\n </figure>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/mode-connectivity-simplex.png\" alt=\"Figure 1 of Ainsworth, Hayase, Srinivasa 2023 (Git Re-Basin). Linear-interpolation barrier between two independently trained networks before and after permutation alignment.\" loading=\"lazy\" decoding=\"async\" width=\"1782\" height=\"1368\">\n <figcaption>Figure 1 of Ainsworth et al. [<a href=\"https://arxiv.org/abs/2209.04836\">5</a>]. Aligning hidden units by permutation collapses most of the apparent linear-interpolation barrier.</figcaption>\n </figure>\n\n <h2>Empirical evidence ties together alignment and merging</h2>\n\n <p>Tatro et al. [<a href=\"https://arxiv.org/abs/2009.02439\">7</a>] showed empirically that aligning models before fitting a connecting curve produces shorter curves with lower loss along them, which is the consistency check the permutation story predicts: correcting for symmetries should give a simpler geometry than the raw view. Benton et al. [<a href=\"https://arxiv.org/abs/2102.13042\">8</a>] extended this beyond curves to higher-dimensional simplexes of solutions: once symmetries are corrected, low-loss volumes contain many independently trained checkpoints.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/mode-connectivity-rebasin.png\" alt=\"Figure 2 of Ainsworth, Hayase, Srinivasa 2023. Linear interpolation barriers between two independently trained networks across MNIST/CIFAR-10/ImageNet under naive, activation-matching, weight-matching, and STE-matching alignment schemes.\" loading=\"lazy\" decoding=\"async\" width=\"2886\" height=\"518\">\n <figcaption>Figure 2 of Ainsworth et al. [<a href=\"https://arxiv.org/abs/2209.04836\">5</a>]. Aligning hidden units before interpolation collapses most of the apparent barrier across architectures and datasets.</figcaption>\n </figure>\n\n <h2>Practical applications</h2>\n\n <p>Model merging is what gets built on top. SWA averages late-training checkpoints and works because the trajectory it averages over stays inside one connected low-loss region. Model Soups average independently fine-tuned models from a shared pretraining initialization, which puts every fine-tune inside the same connected component and close to the others. Git Re-Basin generalizes this further by aligning unit permutations so models with no shared initialization can be merged at all. Across architectures, weight space behaves like a workspace where related models can be moved between while preserving function \u2014 once symmetries and shared training histories have been accounted for.</p>\n\n <p>It does not explain generalization. A connected low-training-loss region can contain many bad predictors on held-out data, and showing that two solutions are connected says nothing about how either performs on unseen examples. It also does not guarantee that every architecture, dataset, or training recipe lives in a single basin; the broader single-basin claims have counter-examples. I read \"one wide basin\" as rhetorically appealing but over-reaching the evidence. The established claim is weaker: solutions reachable from a fixed initialization, or from independent runs once permutations are aligned, lie in a single connected low-loss region. That is enough to explain why SWA, model soups, and weight-space ensembling work. I would not push the geometry harder than that.</p>\n        <h2>Further reading</h2>\n        <ul class=\"further\">\n          <li><a href=\"https://arxiv.org/abs/1910.05653\">S. P. Singh and M. Jaggi. Model fusion via optimal transport. arxiv 1910.05653, 2020</a></li>\n          <li><a href=\"https://arxiv.org/abs/1803.05407\">P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. G. Wilson. Averaging weights leads to wider optima and better generalization. arxiv 1803.05407, 2018</a></li>\n          <li><a href=\"https://arxiv.org/abs/2203.05482\">M. Wortsman et al. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. arxiv 2203.05482, 2022</a></li>\n        </ul>\n\n        \n\n<h2>References</h2>\n\n <ul class=\"refs\">\n <li>[<a href=\"https://arxiv.org/abs/1802.10026\">1</a>] T. Garipov, P. Izmailov, D. Podoprikhin, D. Vetrov, and A. G. Wilson. Loss surfaces, mode connectivity, and fast ensembling of DNNs. arxiv 1802.10026, 2018.</li>\n <li>[<a href=\"https://arxiv.org/abs/1803.00885\">2</a>] F. Draxler, K. Veschgini, M. Salmhofer, and F. A. Hamprecht. Essentially no barriers in neural network energy landscape. arxiv 1803.00885, 2018.</li>\n <li>[<a href=\"https://arxiv.org/abs/1912.05671\">3</a>] J. Frankle, G. K. Dziugaite, D. M. Roy, and M. Carbin. Linear mode connectivity and the lottery ticket hypothesis. arxiv 1912.05671, 2019.</li>\n <li>[<a href=\"https://arxiv.org/abs/2110.06296\">4</a>] R. Entezari, H. Sedghi, O. Saukh, and B. Neyshabur. The role of permutation invariance in linear mode connectivity of neural networks. arxiv 2110.06296, 2021.</li>\n <li>[<a href=\"https://arxiv.org/abs/2209.04836\">5</a>] S. K. Ainsworth, J. Hayase, and S. Srinivasa. Git Re-Basin: merging models modulo permutation symmetries. arxiv 2209.04836, 2022.</li>\n <li>[<a href=\"https://arxiv.org/abs/1611.01540\">6</a>] C. D. Freeman and J. Bruna. Topology and geometry of half-rectified network optimization. arxiv 1611.01540, 2016.</li>\n <li>[<a href=\"https://arxiv.org/abs/2009.02439\">7</a>] N. Tatro, P.-Y. Chen, P. Das, I. Melnyk, P. Sattigeri, and R. Lai. Optimizing mode connectivity via neuron alignment. arxiv 2009.02439, 2020.</li>\n <li>[<a href=\"https://arxiv.org/abs/2102.13042\">8</a>] G. Benton, W. J. Maddox, S. Lotfi, and A. G. Wilson. Loss surface simplexes for mode connecting volumes and fast ensembling. arxiv 2102.13042, 2021.</li>\n </ul>",
      "date_published": "2025-11-04T12:00:00Z",
      "date_modified": "2026-08-21T12:00:00Z",
      "tags": [
        "optimization",
        "generalization"
      ]
    },
    {
      "id": "https://mouhssine.rifaki.me/notes/neural-tangent-kernel/",
      "url": "https://mouhssine.rifaki.me/notes/neural-tangent-kernel/",
      "title": "The neural tangent kernel",
      "summary": "The neural tangent kernel was one of the few deep-learning theory ideas that were useful before they became a concept. It doesn't solve generalization, but it makes a very stubborn object analyzable.\u2026",
      "content_html": "<!-- date: 2025-08-09 -->\n<p>The neural tangent kernel was one of the few deep-learning theory ideas that were useful before they became a concept. It doesn't solve generalization, but it makes a very stubborn object analyzable. Jacot, Gabriel, and Hongler [<a href=\"https://arxiv.org/abs/1806.07572\">1</a>] found that if you take a network to infinite width under the right scaling, gradient descent on the parameters becomes kernel gradient descent in function space. The kernel isn't chosen by hand, it's induced by the network at initialization. For a model $f_\\theta$, the tangent kernel is $K_\\theta(x, x') = \\nabla_\\theta f_\\theta(x)^\\top \\nabla_\\theta f_\\theta(x')$. In finite networks this kernel changes as you train. In the infinite width limit under the standard parameterization, it converges to some deterministic kernel $K_\\infty$ and $\\|K_{\\theta_t} - K_\\infty\\|$ approaches $O(1/\\sqrt{n})$ in width $n$.</p>\n\n        <p>Function-space gradient descent then satisfies $\\dot{f}_t = -K_\\infty (f_t - y)$ on the training set and integrates to $f_t = y + e^{-K_\\infty t}(f_0 - y)$. Parameter-space non-convexity stops mattering: the function-space dynamics are linear and driven by a positive semidefinite kernel.</p>\n\n        <h2>NTK as baseline</h2>\n\n        <p>The first thing the NTK explained is why very wide networks optimize so easily. If the kernel is well conditioned on the training data, gradient descent has an easy road to interpolation. The complicated nonconvex path is, at leading order, kernel regression with some specific architecture induced kernel. Du et al. [<a href=\"https://arxiv.org/abs/1810.02054\">4</a>] and Lee et al. [<a href=\"https://arxiv.org/abs/1902.06720\">2</a>] pushed this picture further and showed that wide networks of any depth evolve like their first-order Taylor expansion around initialization.</p>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/ntk-linearization.png\" alt=\"Figure 2 of Lee et al. 2019. Predictions from the linearized infinite-width model match the trajectory of the actual wide finite network during gradient-descent training.\" loading=\"lazy\" decoding=\"async\" width=\"2870\" height=\"814\">\n          <figcaption>Figure 2 of Lee et al. [<a href=\"https://arxiv.org/abs/1902.06720\">2</a>]. The linearization at initialization tracks the wide-network trajectory closely under gradient descent.</figcaption>\n        </figure>\n\n        <p>Du et al. extend the same machinery to a clean proof that overparameterized networks reach zero training loss with a polynomial-width requirement and a global-convergence guarantee that the nonconvex landscape never gave you. But the spectrum of $K_\\infty$ does more than set the speed of convergence. Its eigendecomposition $K_\\infty = \\sum_k \\lambda_k \\phi_k \\phi_k^\\top$ implies that the residual along eigenmode $\\phi_k$ shrinks like $e^{-\\lambda_k t}$, so large-eigenvalue modes are learned quickly and small-eigenvalue modes slowly, or not at all under early stopping. Early stopping, the frequency principle, and the spectral bias of MLPs all become statements about $\\{\\lambda_k\\}$.</p>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/ntk-spectrum.png\" alt=\"Figure 1 of Cao et al. 2019 (spectral bias of deep learning). Projection lengths along the lowest few eigenmodes of the NTK as a function of training step: low-frequency (small k) components are fit much faster than higher-frequency components.\" loading=\"lazy\" decoding=\"async\" width=\"2656\" height=\"1130\">\n          <figcaption>Figure 1 of Cao et al. [<a href=\"https://arxiv.org/abs/1912.01198\">arxiv 1912.01198</a>]. Low-frequency components of the target are absorbed by the network long before higher-frequency components, in the order predicted by the NTK spectrum.</figcaption>\n        </figure>\n\n        <p>Arora et al. write down an exact algorithm for computing $K_\\infty$ for fully connected and convolutional networks of arbitrary depth, making these spectral predictions empirically testable on real datasets.</p>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/ntk-spectrum-timeline.png\" alt=\"Figure 2 of Bordelon, Canatar, Pehlevan 2020. Spectrum-dependent generalization-error scaling: per-mode learning curves $E_k(p)/E_k(0)$ versus number of training samples for varying eigenmode index, input dimension, and depth, all approaching the predicted $1/p^\\alpha$ envelope.\" loading=\"lazy\" decoding=\"async\" width=\"2728\" height=\"746\">\n          <figcaption>Figure 2 of Bordelon et al. [<a href=\"https://arxiv.org/abs/2002.02561\">arxiv 2002.02561</a>]. Generalization on each NTK eigenmode follows a spectrum-dependent power law in sample count.</figcaption>\n        </figure>\n\n        <h2>The lazy-training caveat</h2>\n\n        <p>The catch is that the same condition that makes the theory clean also removes one of the main things deep networks seem to be doing. In the NTK limit, features do not move. Parameters drift by $\\|\\theta_t - \\theta_0\\| = O(1/\\sqrt{n})$ in width $n$ while the function changes by $O(1)$, so the network is producing its outputs by reweighting an almost fixed collection of random features. Chizat, Oyallon, and Bach call this the lazy-training regime and emphasize that it is a property of the scaling, not a universal description of neural networks.</p>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/ntk-lazy-vs-feature-learning.png\" alt=\"Figure 1 of Chizat, Oyallon, Bach 2019. Lazy regime versus feature-learning regime trajectories on a 2D classification problem: the lazy regime stays near initialization while feature learning moves substantially.\" loading=\"lazy\" decoding=\"async\" width=\"2626\" height=\"864\">\n          <figcaption>Figure 1 of Chizat, Oyallon, and Bach [<a href=\"https://arxiv.org/abs/1812.07956\">3</a>]. The NTK limit is powerful because it freezes feature movement; that is also what it cannot explain.</figcaption>\n        </figure>\n\n\n        <p>That caveat matters; a convolutional network trained in the lazy regime can optimize while failing to learn the representations that make convolutional networks useful. A transformer that looks like a fixed random-feature model is not the object that in-context learning, induction-head formation, and abstraction discussions are pointing at. The NTK gives a rigorous theory of one limit. The question is whether that limit keeps the right phenomena. Geiger et al [<a href=\"https://arxiv.org/abs/1906.08034\">5</a>] report a sharp empirical separation: at moderate width and standard initialization scale networks operate near the lazy regime; at lower initialization scale (or if explicit feature-learning parameterizations are used) the same architecture enters a regime where features evolve and test error improves.</p>\n\n        <p>The transition is controlled by initialization scale and width, not by anything intrinsic to the architecture. That is a slightly disappointing answer if you wanted neural networks to be feature learners by default.</p>\n\n        <p>The NTK is not wrong; it is a baseline. A phenomenon that already appears in the NTK limit can be attributed to width, interpolation, and fixed random features, with no representation learning needed. A phenomenon that disappears in the NTK limit is the work of feature learning, finite-width fluctuation, architecture-specific structure, or nonlinearity in the optimization. That makes the NTK a useful negative control for theoretical claims about deep learning, and the simpler question to put to such a claim is: would it still hold if the features were frozen? For most optimization claims, yes. For most generalization and capability claims, no. The Distill circuits thread is the non-theorem-shaped version of what feature learning looks like when someone manages to pry a model open. Greg Yang's $\\mu$P writeup is the practical entry into Tensor Programs, and the microsoft/mup repo is what to grab when the goal is hyperparameter transfer and not the theory.</p>\n\n        <figure class=\"tweet-embed\">\n          <blockquote class=\"twitter-tweet\" data-dnt=\"true\"><p lang=\"en\" dir=\"ltr\">Excited to share our new <a href=\"https://twitter.com/hashtag/neurips2020?src=hash&amp;ref_src=twsrc%5Etfw\">#neurips2020</a> paper /Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel/ (<a href=\"https://t.co/4jmrNfOE6H\">https://t.co/4jmrNfOE6H</a>) with @KDziugaite, Mansheej, <a href=\"https://twitter.com/SKharaghani?ref_src=twsrc%5Etfw\">@SKharaghani</a>, <a href=\"https://twitter.com/roydanroy?ref_src=twsrc%5Etfw\">@roydanroy</a>, <a href=\"https://twitter.com/SuryaGanguli?ref_src=twsrc%5Etfw\">@SuryaGanguli</a> 1/6 <a href=\"https://t.co/iPPP3HmNgm\">pic.twitter.com/iPPP3HmNgm</a></p>&mdash; Stanislav Fort (@stanislavfort) <a href=\"https://twitter.com/stanislavfort/status/1322246600320757760?ref_src=twsrc%5Etfw\">October 30, 2020</a></blockquote>\n        </figure>\n\n        <h2>Feature learning is the missing term</h2>\n\n        <p>The frontier after the NTK was to build infinite-width limits in which features actually move. Mean-field limits treat each unit as a particle in a measure and study gradient flow on that measure; in this scaling features evolve and the kernel is no longer constant. Tensor-program analyses catalogue the parameterizations that produce sensible infinite-width limits at all. The maximal-update parameterization $\\mu$P, introduced by Yang and Hu in Tensor Programs IV (<a href=\"https://arxiv.org/abs/2011.14522\">arxiv 2011.14522</a>), is the one that keeps both feature learning and stable optimization in the limit. The follow-up Tensor Programs V (<a href=\"https://arxiv.org/abs/2203.03466\">arxiv 2203.03466</a>) derives the $\\mu$Transfer hyperparameter-transfer rules from that analysis: tune at small width, scale to large width, and the learning-rate schedule transfers without retuning.</p>\n\n        <p>Feature learning means the tangent kernel is moving substantively: $K_{\\theta_t} - K_{\\theta_0}$ is a structured rotation of the features toward the data, not a small perturbation. The parameter-gradients at the end of training are not the same object as at initialization, and the network has in effect changed the basis it works in. That change is exactly what the pure NTK limit suppresses. Fort et al. ran one of the clearest empirical comparisons: kernel learning matches a finite network early in training but the two diverge later, and the divergence is the gap between lazy convergence to a fixed kernel and feature-driven re-shaping of it.</p>\n\n        <p>I use the NTK as a falsifier, not a model. If a proposed mechanism for a deep-learning phenomenon is already trivially in the NTK regime, then \"the network is wide and the features are random\" suffices, and the explanation has not earned the depth of its hypothesis. The interesting predictions are the ones that disagree with the kernel: where width, lazy init, and architecture-induced spectra are not enough, and where representation change must be doing the work. That includes in-context learning, induction-head formation, and the parts of scaling laws that depend on where compute is spent.</p>\n        <h2>Further reading</h2>\n        <ul class=\"further\">\n          <li><a href=\"https://arxiv.org/abs/1904.11955\">S. Arora, S. S. Du, W. Hu, Z. Li, R. Salakhutdinov, and R. Wang. On exact computation with an infinitely wide neural net. arxiv 1904.11955, 2019</a></li>\n          <li><a href=\"https://arxiv.org/abs/1812.11118\">M. Belkin, D. Hsu, S. Ma, and S. Mandal. Reconciling modern machine learning practice and the bias-variance trade-off. arxiv 1812.11118, 2018</a></li>\n          <li><a href=\"https://arxiv.org/abs/1804.06561\">S. Mei, A. Montanari, and P.-M. Nguyen. A mean field view of the landscape of two-layer neural networks. arxiv 1804.06561, 2018</a></li>\n          <li><a href=\"https://arxiv.org/abs/2011.14522\">G. Yang and E. J. Hu. Tensor Programs IV: feature learning in infinite-width neural networks. arxiv 2011.14522, 2020</a></li>\n          <li><a href=\"https://arxiv.org/abs/1912.01198\">Y. Cao, Z. Fang, Y. Wu, D.-X. Zhou, and Q. Gu. Towards understanding the spectral bias of deep learning. arxiv 1912.01198, 2019</a></li>\n          <li><a href=\"https://arxiv.org/abs/2002.02561\">B. Bordelon, A. Canatar, and C. Pehlevan. Spectrum dependent learning curves in kernel regression and wide neural networks. arxiv 2002.02561, 2020</a></li>\n        </ul>\n\n        \n\n<h2>References</h2>\n\n        <ul class=\"refs\">\n          <li>[<a href=\"https://arxiv.org/abs/1806.07572\">1</a>] A. Jacot, F. Gabriel, and C. Hongler. Neural tangent kernel: convergence and generalization in neural networks. arxiv 1806.07572, 2018.</li>\n          <li>[<a href=\"https://arxiv.org/abs/1902.06720\">2</a>] J. Lee, L. Xiao, S. S. Schoenholz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington. Wide neural networks of any depth evolve as linear models under gradient descent. arxiv 1902.06720, 2019.</li>\n          <li>[<a href=\"https://arxiv.org/abs/1812.07956\">3</a>] L. Chizat, E. Oyallon, and F. Bach. On lazy training in differentiable programming. arxiv 1812.07956, 2018.</li>\n          <li>[<a href=\"https://arxiv.org/abs/1810.02054\">4</a>] S. S. Du, X. Zhai, B. P&oacute;czos, and A. Singh. Gradient descent provably optimizes over-parameterized neural networks. arxiv 1810.02054, 2018.</li>\n          <li>[<a href=\"https://arxiv.org/abs/1906.08034\">5</a>] M. Geiger, S. Spigler, A. Jacot, and M. Wyart. Disentangling feature and lazy training in deep neural networks. arxiv 1906.08034, 2019.</li>\n          <li>[<a href=\"https://arxiv.org/abs/2010.15110\">6</a>] S. Fort, G. K. Dziugaite, M. Paul, S. Kharaghani, D. M. Roy, and S. Ganguli. Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the neural tangent kernel. arxiv 2010.15110, 2020.</li>\n        </ul>",
      "date_published": "2025-08-09T12:00:00Z",
      "date_modified": "2026-08-21T12:00:00Z",
      "tags": [
        "theory",
        "optimization"
      ]
    },
    {
      "id": "https://mouhssine.rifaki.me/notes/implicit-bias/",
      "url": "https://mouhssine.rifaki.me/notes/implicit-bias/",
      "title": "The implicit-bias program",
      "summary": "Why does training a model without an explicit regularizer, with the loss driven nearly to zero, still produce a solution that generalizes? The classical answer is that the objective has to carry the\u2026",
      "content_html": "<!-- date: 2025-05-14 -->\n<p>Why does training a model without an explicit regularizer, with the loss driven nearly to zero, still produce a solution that generalizes? The classical answer is that the objective has to carry the regularizer somewhere. The implicit-bias answer is more subtle: even when the objective has many minima, gradient descent does not choose among them neutrally; the algorithm itself selects a particular kind of solution.</p>\n\n <p>Zhang et al. [<a href=\"https://arxiv.org/abs/1611.03530\">2</a>] sharpened the puzzle that most of this literature now opens with. Standard image classifiers fit random labels as easily as they fit true ones, which kills the simplest capacity-based explanation of generalization: the hypothesis class is large enough to memorize anything. Whatever is producing the generalization, then, must come from the optimizer, the data, or the parameterization, not from the loss function itself.</p>\n\n <p>The clearest version of this story is not actually about neural networks. It is about logistic regression on linearly separable data. Once the classifier separates the data, the empirical classification error is already 0% and the logistic loss continues to decrease as the norm of the weights grows. There is no finite minimizer. What is more surprising is that the direction of the weights still converges. And it converges to the maximum-margin SVM solution.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/implicit-bias-margin.png\" alt=\"Figure 1 of Soudry et al. 2018. Five-panel layout: (A) 2D separable data with the converged separator, (B) normalized weight norm growing logarithmically, (C) logistic loss decaying, (D) angle gap to max-margin direction shrinking, (E) margin gap closing.\" loading=\"lazy\" decoding=\"async\" width=\"2282\" height=\"766\">\n <figcaption>Figure 1 of Soudry et al. [<a href=\"https://arxiv.org/abs/1710.10345\">1</a>]. The loss has no finite minimizer on separable data, the norm grows without bound, and the normalized direction converges to the hard-margin separator.</figcaption>\n </figure>\n\n\n <h2>The linear theorem</h2>\n\n <p>Let \\(x_i \\in \\mathbb{R}^d\\) denote an input vector with binary label \\(y_i \\in \\{-1,+1\\}\\), and let \\(w_t \\in \\mathbb{R}^d\\) denote the weight vector at iteration \\(t\\). For linearly separable data \\(\\{(x_i, y_i)\\}_{i=1}^{n}\\), standard gradient descent on the logistic loss sends \\(\\lVert w_t \\rVert \\to \\infty\\), but the normalized direction \\(w_t / \\lVert w_t \\rVert\\) converges to the L2 max-margin direction \\(\\hat{w} / \\lVert \\hat{w} \\rVert\\), where\n \\[\\hat{w} \\;=\\; \\arg\\min_{w \\in \\mathbb{R}^d} \\tfrac{1}{2}\\lVert w\\rVert^2 \\quad \\text{s.t.}\\quad y_i\\,w^{\\top} x_i \\ge 1 \\quad \\forall i \\in [n].\\]\n </p>\n\n <p>In words: once the classifier has separated the data, gradient descent keeps reducing the logistic loss as long as the separation is preserved, except that unlike a traditional SVM it does so while continuing to grow the norm of the weights without bound. The norm grows only logarithmically with \\(t\\), and the angle gap between \\(w_t/\\lVert w_t\\rVert\\) and \\(\\hat w/\\lVert\\hat w\\rVert\\) closes at the same logarithmic rate. This is why the asymptotic regime takes so many iterations to become visible.</p>\n\n <p>A few things follow from the theorem. Gradient descent behaves as if it had been regularized toward the Euclidean max-margin classifier without anyone writing that regularizer down. It also explains why early stopping helps: because the weight norm diverges over time, stopping early caps the implicit penalty before the iterate gets too far. Ji and Telgarsky extended the result to the non-separable case, showing the iterate still tracks a unique ray defined by the data when no separating hyperplane exists. Nacson et al. [<a href=\"https://arxiv.org/abs/1803.01905\">4</a>] showed that aggressive learning-rate schedules accelerate convergence to the max-margin direction by polynomial factors over plain GD. The norm divergence itself is robust to dataset size and dimension: as long as the data is linearly separable, the iterate keeps moving away from the origin and never lands at a finite minimum.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/implicit-bias-margin-flow.png\" alt=\"Figure 2 of Soudry et al. 2018. Three panels on a real classification dataset: training/validation objective loss, classification error, and L2 norm of the final layer growing as training progresses.\" loading=\"lazy\" decoding=\"async\" width=\"2728\" height=\"746\">\n <figcaption>Figure 2 of Soudry et al. [<a href=\"https://arxiv.org/abs/1710.10345\">1</a>]. On real data, classification error plateaus near zero while the L2 norm of the final layer keeps growing - the asymptotic regime described in the linear theorem.</figcaption>\n </figure>\n\n <h2>The geometry enters</h2>\n\n <p>Gunasekar et al. [<a href=\"https://arxiv.org/abs/1802.08246\">3</a>] generalized the result: different optimization geometries select different implicit regularizers. Steepest descent under the \\(\\ell_p\\) norm minimizes the \\(\\ell_p\\) margin instead of the \\(\\ell_2\\) margin. Mirror descent with respect to a convex potential \\(\\Phi\\) minimizes the \\(\\Phi\\)-min-norm interpolant. Natural gradient and adaptive methods land at interpolants determined by the geometry of their step.</p>\n\n <p>The linear-convolutional-network result is the cautionary case. For fully connected linear predictors, gradient descent picks out the familiar $\\ell_2$ margin geometry. For full-width linear convolutional networks of depth $L$, Gunasekar et al. show that gradient descent instead selects the predictor minimizing the $2/L$-bridge penalty in the discrete Fourier domain. The architecture changes which parameters are being optimized, and that change shifts both the trajectory and the preferred solution. \"Gradient descent likes simple solutions\" is too vague to be a theorem. The more honest statement is that gradient descent likes simple solutions in whichever coordinate system the architecture imposes.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/implicit-bias-optimizer-geometry.png\" alt=\"Three-panel figure from Gunasekar, Lee, Soudry, Srebro 2018: (a) mirror descent with primal momentum, (b) natural gradient descent at varying step sizes, (c) steepest descent under the 4/3 norm. Each optimizer trajectory lands on a different interpolating solution along the same zero-loss line.\" loading=\"lazy\" decoding=\"async\" width=\"2534\" height=\"956\">\n <figcaption>From Gunasekar et al. [<a href=\"https://arxiv.org/abs/1802.08246\">3</a>]. Implicit bias is a selection rule over interpolants determined by the geometry of the optimizer, not a single universal preference.</figcaption>\n </figure>\n\n <h2>Margin in homogeneous networks</h2>\n\n <p>Lyu and Li push the result past linear predictors. If \\(f_\\theta\\) is positively homogeneous in \\(\\theta\\) with order \\(L\\) (which holds for ReLU networks without bias, with \\(L\\) equal to depth), then gradient flow on exponential or logistic loss drives \\(\\theta_t / \\lVert \\theta_t \\rVert\\) to a KKT point of the parameter-space margin program \\(\\max_{\\lVert \\theta \\rVert \\leq 1}\\, \\min_i\\, y_i f_\\theta(x_i)\\). The norm still diverges; the normalized direction still converges; the new content is that even on a non-convex parameter landscape, gradient flow lands on points satisfying first-order optimality conditions for the margin program.</p>\n\n <p>Chizat and Bach prove a parallel mean-field result for two-layer networks with vanishing initializations. The implicit bias there is \\(F_1\\)-norm minimization in function space, which is a different object from parameter-space margin maximization and interacts differently with the data. The state of the field is that implicit regularization in deep networks has several reasonable descriptions, none of which generalize cleanly past shallow or homogeneous models.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/implicit-bias-dynamics.png\" alt=\"Figure 1 of Lyu and Li 2020. Training loss and normalized margin trajectories for homogeneous networks under fixed and loss-based learning rates: the loss collapses while the normalized margin keeps rising toward a KKT point of the parameter-space margin program.\" loading=\"lazy\" decoding=\"async\" width=\"2860\" height=\"756\">\n <figcaption>Figure 1 of Lyu and Li [<a href=\"https://arxiv.org/abs/1906.05890\">arxiv 1906.05890</a>]. The loss keeps shrinking, the weight norm grows, and the useful object is the normalized direction.</figcaption>\n </figure>\n\n <figure class=\"tweet-embed\">\n <blockquote class=\"twitter-tweet\" data-dnt=\"true\"><p lang=\"en\" dir=\"ltr\">It is widely thought that neural networks generalize because of implicit regularization of gradient descent. Today at <a href=\"https://twitter.com/hashtag/ICLR2023?src=hash&amp;ref_src=twsrc%5Etfw\">#ICLR2023</a> we show new evidence to the contrary. We train with gradient-free optimizers and observe generalization competitive with SGD.<a href=\"https://t.co/8Vo9rFI9FY\">https://t.co/8Vo9rFI9FY</a></p>&mdash; Tom Goldstein (@tomgoldsteincs) <a href=\"https://twitter.com/tomgoldsteincs/status/1653284772314005505?ref_src=twsrc%5Etfw\">May 2, 2023</a></blockquote>\n </figure>\n\n\n\n <p>My own reading is that the implicit-bias program is one of the few cases I can point to of a research direction being vindicated and outgrown at the same time. Soudry et al. [<a href=\"https://arxiv.org/abs/1710.10345\">1</a>] is true; the mechanism is real; the linear case is the only setting where I can prove anything I trust. What is unclear is whether the same phenomenon is the dominant explanation for why large feature-learning networks generalize, or whether at scale the data distribution and the architecture have already done so much of the work that the optimizer's preference is a small correction. I currently believe the second, but I do not have a falsifier I trust, which is exactly the position the field is in.</p>\n <h2>Further reading</h2>\n <ul class=\"further\">\n <li><a href=\"https://arxiv.org/abs/1806.00468\">S. Gunasekar, J. Lee, D. Soudry, and N. Srebro. Implicit bias of gradient descent on linear convolutional networks. arxiv 1806.00468, 2018</a></li>\n <li><a href=\"https://arxiv.org/abs/1803.07300\">Z. Ji and M. Telgarsky. Risk and parameter convergence of logistic regression. arxiv 1803.07300, 2018</a></li>\n <li><a href=\"https://arxiv.org/abs/1906.05890\">K. Lyu and J. Li. Gradient descent maximizes the margin of homogeneous neural networks. arxiv 1906.05890, 2019</a></li>\n <li><a href=\"https://arxiv.org/abs/2002.04486\">L. Chizat and F. Bach. Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss. arxiv 2002.04486, 2020</a></li>\n <li><a href=\"https://arxiv.org/abs/2106.09524\">S. Pesme, L. Pillaud-Vivien, and N. Flammarion. Implicit bias of SGD for diagonal linear networks: a provable benefit of stochasticity. arxiv 2106.09524, 2021</a></li>\n <li><a href=\"https://arxiv.org/abs/2007.06738\">E. Moroshko, S. Gunasekar, B. Woodworth, J. D. Lee, N. Srebro, and D. Soudry. Implicit bias in deep linear classification: initialization scale vs training accuracy. arxiv 2007.06738, 2020</a></li>\n </ul>\n\n \n\n<h2>References</h2>\n\n <ul class=\"refs\">\n <li>[<a href=\"https://arxiv.org/abs/1710.10345\">1</a>] D. Soudry, E. Hoffer, M. S. Nacson, S. Gunasekar, and N. Srebro. The implicit bias of gradient descent on separable data. arxiv 1710.10345, 2017.</li>\n <li>[<a href=\"https://arxiv.org/abs/1611.03530\">2</a>] C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals. Understanding deep learning requires rethinking generalization. arxiv 1611.03530, 2016.</li>\n <li>[<a href=\"https://arxiv.org/abs/1802.08246\">3</a>] S. Gunasekar, J. Lee, D. Soudry, and N. Srebro. Characterizing implicit bias in terms of optimization geometry. arxiv 1802.08246, 2018.</li>\n <li>[<a href=\"https://arxiv.org/abs/1803.01905\">4</a>] M. S. Nacson, J. Lee, S. Gunasekar, P. H. P. Savarese, N. Srebro, and D. Soudry. Convergence of gradient descent on separable data. arxiv 1803.01905, 2019.</li>\n </ul>",
      "date_published": "2025-05-14T12:00:00Z",
      "date_modified": "2026-08-21T12:00:00Z",
      "tags": [
        "optimization",
        "generalization"
      ]
    },
    {
      "id": "https://mouhssine.rifaki.me/notes/edge-of-stability/",
      "url": "https://mouhssine.rifaki.me/notes/edge-of-stability/",
      "title": "The edge of stability",
      "summary": "Cohen et al. [1] observed that gradient descent on neural networks spends most of training in a regime where the top Hessian eigenvalue \u03bbmax is above the classical stability threshold 2/\u03b7. The step is\u2026",
      "content_html": "<!-- date: 2025-02-27 -->\n<p>Cohen et al. [<a href=\"https://arxiv.org/abs/2103.00065\">1</a>] observed that gradient descent on neural networks spends most of training in a regime where the top Hessian eigenvalue $\\lambda_{\\max}$ is above the classical stability threshold $2/\\eta$. The step is formally unstable ($\\eta \\lambda_{\\max} > 2$), but the loss does not diverge: the trajectory oscillates along the unstable direction while continuing to make progress on the rest. They called this the edge of stability, and the point is that it is not an edge case but the ordinary regime of neural-network training.</p>\n\n        <h2>Progressive sharpening</h2>\n\n        <p>The first phase Cohen identified is progressive sharpening. From random initialization, gradient descent reliably drives the sharpness ($\\lambda_{\\max}(H)$) of the loss landscape upward during training.</p>\n\n        <p>That direction is already counterintuitive. Classical optimization theory says you want to avoid sharp regions, since each step there costs more. Neural-network training does the opposite: it walks into sharper regions until the classical step size is no longer stable. Progressive sharpening itself is not well understood from first principles. Damian et al. [<a href=\"https://arxiv.org/abs/2209.15594\">2</a>] give a self-stabilization argument for what happens after the threshold is reached, and Ahn et al. [<a href=\"https://arxiv.org/abs/2204.01050\">4</a>] build on it, but neither predicts the sharpening from initialization. The observational picture is clean; the theoretical one is not.</p>\n\n        <h2>Above the threshold</h2>\n\n        <p>Once sharpness crosses $2/\\eta$, the textbook prediction is divergence; instead the trajectory oscillates along the top eigendirection of the Hessian. The component of the iterate along that direction swings back and forth, and the loss still falls because optimization keeps making progress on the better-conditioned directions. The net result is a trajectory that reduces loss while sitting in a region of the landscape that classical theory says it should not occupy. The first theoretical account of why this need not be a pathology comes from Arora et al. [<a href=\"https://arxiv.org/abs/2205.09745\">3</a>]. Their analysis works on a smoothed version of the loss, where the Hessian is treated as locally fixed, and in effect tracks the trajectory you would follow from a given starting point under that fixed Hessian.</p>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/cohen-fig1.png\" alt=\"Figure 1 of Cohen et al. 2021. Train loss (top row) and Hessian sharpness (bottom row) over training steps for a fully-connected net on a CIFAR-10 5k subset, VGG on CIFAR-10, and ResNet on CIFAR-10. In every case, sharpness rises until it hits the $2/\\eta$ threshold (dashed) and oscillates along it.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"441\">\n          <figcaption>Figure 1 of Cohen et al. [<a href=\"https://arxiv.org/abs/2103.00065\">1</a>]. Across architectures, the Hessian's top eigenvalue rises during a progressive-sharpening phase and then sits near the $2/\\eta$ stability threshold for the remainder of training.</figcaption>\n        </figure>\n\n\n        <p>The mechanism is that the oscillation $\\Delta \\theta_t$ across the unstable direction averages to zero, so the effective dynamics is slower and looks like gradient descent on a loss with the steepest direction clipped. The account is formal enough to be checked against real training runs, and in most common settings it holds up.</p>\n\n        <h2>And flat minima</h2>\n\n        <p>The earlier flat-minima story started with Hochreiter and Schmidhuber and continued with Keskar et al. [<a href=\"https://arxiv.org/abs/1609.04836\">7</a>] and the later sharpness aware minimization literature. Broadly, the flat-minima story ran as follows: SGD with a small batch size produces gradient estimates with some noise $\\xi$.</p>\n\n        <p>That noise looks like a random walk and tends to leave sharp minima more often than flat ones, so SGD ends up biased toward flat minima, and that bias was meant to be why neural networks generalize. Edge-of-stability does not contradict the story, but it reshapes it. The learning rate itself caps how sharp a reachable minimum can be: anything with $\\lambda_{\\max} > 2/\\eta$ is unstable for GD, so the trajectory cannot stay there. Gradient descent finds flat minima not because it has noise, but because sharp minima are unstable fixed points under its own dynamics. The original explanation identified the phenomenon and pinned it to the wrong cause.</p>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/cohen-fig3.png\" alt=\"Figure 3 of Cohen et al. 2103.00065. Progressive sharpening isolated; sharpness rises before reaching the 2/eta threshold.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"249\">\n          <figcaption>Figure 3 of Cohen et al. [<a href=\"https://arxiv.org/abs/2103.00065\">1</a>]. Progressive sharpening isolated: Hessian sharpness rises monotonically during the early phase, long before the $2/\\eta$ threshold is reached.</figcaption>\n        </figure>\n\n\n        <p>The reframing lives mostly outside the papers themselves. Off Convex has a few posts on implicit bias, trajectory analysis, and why the classical descent lemma is genuinely misleading for neural networks instead of merely approximate. Ben Recht's ArgMin is the complementary skeptical take for once you have left convex optimization theory $\\mu I \\preceq \\nabla^2 L \\preceq L\\,I$. Clare Lyle's tutorial walks through the $2/\\eta$ arithmetic and ties the phenomenon to warmup (rising $\\eta(t)$) and catapult (loss spike then decay) in one frame.</p>\n\n        <p>Andreyev and Beneventano (<a href=\"https://arxiv.org/abs/2412.20553\">arxiv 2412.20553</a>) extended the story to the mini-batch setting Cohen did not analyze, introducing an \"edge of stochastic stability\" where the quantity that pins at $2/\\eta$ is the expected directional curvature of mini-batch Hessians, not the full-Hessian's top eigenvalue.</p>\n\n        <p>A few previously folklore-level phenomena become intelligible from this picture. Warmup schedules $\\eta(t)$, which start small and increase $\\eta$ over several thousand iterations, let the network settle into edge-of-stability before $\\eta$ reaches its final value; without warmup, the early transient at the full $\\eta$ would hit a too-sharp region and diverge. Decreasing-$\\eta$ schedules at the end of training raise $2/\\eta$, so the trajectory can fine-tune in sharper local minima inside the broader flat region already reached, which empirically pushes training loss down further.</p>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/arora-fig1.png\" alt=\"Figure 1 of Arora et al. 2205.09745. Smoothed-loss analysis of the edge-of-stability oscillations.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"525\">\n          <figcaption>Figure 1 of Arora et al. [<a href=\"https://arxiv.org/abs/2205.09745\">3</a>]. The smoothed-loss analysis makes explicit why the oscillations across the unstable direction do not destroy progress; averaging over a few steps yields an effective slow dynamics on a clipped loss.</figcaption>\n        </figure>\n\n\n        <p>Both schedules had been used empirically for years before any principled account existed. Lewkowycz et al.'s catapult mechanism [<a href=\"https://arxiv.org/abs/2003.02218\">5</a>], where an initial loss spike sometimes precedes a better final solution, is the same dynamics at a larger scale: a large learning rate pushes the trajectory through a briefly very sharp region, the loss spikes, and the trajectory then settles into a different basin from the one it would have reached at a smaller step size.</p>\n\n        <h2>The generalization gap is still open</h2>\n\n        <p>Edge-of-stability gives a clean account of why SGD ends up in flat minima, but it says nothing about why flat minima generalize. Those are distinct questions, and the second one is still open.</p>\n\n        <p>Dinh et al. [<a href=\"https://arxiv.org/abs/1703.04933\">6</a>] showed that the Hessian-based notion of sharpness is not reparameterization invariant, so sharpness in that form cannot directly control generalization.</p>\n\n        <figure class=\"tweet-embed\">\n          <blockquote class=\"twitter-tweet\" data-dnt=\"true\"><p lang=\"en\" dir=\"ltr\">Even with full-batch gradients, DL optimizers defy classical optimization theory, as they operate at the *edge of stability.*<br><br>With <a href=\"https://twitter.com/alex_damian_?ref_src=twsrc%5Etfw\">@alex_damian_</a>, we introduce &quot;central flows&quot;: a theoretical tool to analyze these dynamics that makes accurate quantitative predictions on real NNs. <a href=\"https://t.co/pvvfwoQcOy\">pic.twitter.com/pvvfwoQcOy</a></p>&mdash; Jeremy Cohen (@deepcohen) <a href=\"https://twitter.com/deepcohen/status/1973191790602887544?ref_src=twsrc%5Etfw\">October 1, 2025</a></blockquote>\n          <figcaption>Jeremy Cohen, lead author of [<a href=\"https://arxiv.org/abs/2103.00065\">1</a>], announcing the <a href=\"https://arxiv.org/abs/2410.24206\">central-flows follow-up</a> that makes the 2021 observation a quantitative prediction tool.</figcaption>\n        </figure>\n\n        <h2>Further reading</h2>\n        <ul class=\"further\">\n          <li><a href=\"https://clarelyle.com/posts/2023-10-15-edge.html\">deep dive into the edge of stability</a></li>\n          <li><a href=\"https://arxiv.org/abs/1912.05671\">J. Frankle, G. K. Dziugaite, D. M. Roy, and M. Carbin. Linear mode connectivity and the lottery ticket hypothesis. arxiv 1912.05671, 2019</a></li>\n          <li><a href=\"https://jmlr.org/papers/v25/23-1285.html\">P. M. Long and P. L. Bartlett. Sharpness-aware minimization and the edge of stability. Journal of Machine Learning Research, 2024</a></li>\n          <li><a href=\"https://arxiv.org/abs/2410.24206\">J. M. Cohen, A. Damian, A. Talwalkar, J. Z. Kolter, and J. D. Lee. Understanding optimization in deep learning with central flows. arxiv 2410.24206, 2024</a></li>\n          <li><a href=\"https://direct.mit.edu/neco/article/9/1/1/6027/Flat-Minima\">S. Hochreiter and J. Schmidhuber. Flat minima. Neural Computation, 9(1):1-42, 1997</a></li>\n        </ul>\n\n<h2>References</h2>\n\n        <ul class=\"refs\">\n          <li>[<a href=\"https://arxiv.org/abs/2103.00065\">1</a>] J. M. Cohen, S. Kaur, Y. Li, J. Z. Kolter, and A. Talwalkar. Gradient descent on neural networks typically occurs at the edge of stability. arxiv 2103.00065, 2021.</li>\n          <li>[<a href=\"https://arxiv.org/abs/2209.15594\">2</a>] A. Damian, E. Nichani, and J. D. Lee. Self-stabilization: The implicit bias of gradient descent at the edge of stability. arxiv 2209.15594, 2022.</li>\n          <li>[<a href=\"https://arxiv.org/abs/2205.09745\">3</a>] S. Arora, Z. Li, and A. Panigrahi. Understanding gradient descent on edge of stability in deep learning. arxiv 2205.09745, 2022.</li>\n          <li>[<a href=\"https://arxiv.org/abs/2204.01050\">4</a>] K. Ahn, J. Zhang, and S. Sra. Understanding the unstable convergence of gradient descent. arxiv 2204.01050, 2022.</li>\n          <li>[<a href=\"https://arxiv.org/abs/2003.02218\">5</a>] A. Lewkowycz, Y. Bahri, E. Dyer, J. Sohl-Dickstein, and G. Gur-Ari. The large learning rate phase of deep learning: The catapult mechanism. arxiv 2003.02218, 2020.</li>\n          <li>[<a href=\"https://arxiv.org/abs/1703.04933\">6</a>] L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio. Sharp minima can generalize for deep nets. arxiv 1703.04933, 2017.</li>\n          <li>[<a href=\"https://arxiv.org/abs/1609.04836\">7</a>] N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang. On large-batch training for deep learning: generalization gap and sharp minima. arxiv 1609.04836, 2016.</li>\n        </ul>",
      "date_published": "2025-02-27T12:00:00Z",
      "date_modified": "2026-08-21T12:00:00Z",
      "tags": [
        "optimization",
        "theory"
      ]
    },
    {
      "id": "https://mouhssine.rifaki.me/notes/neural-scaling-laws/",
      "url": "https://mouhssine.rifaki.me/notes/neural-scaling-laws/",
      "title": "Kaplan, Chinchilla, and broken laws",
      "summary": "Kaplan et al. [1] set the baseline picture: test loss falls as a power law L(C)=(C/C0)\u2212\u03b1C in each of model size N, dataset size D, and compute C, with clean exponents \u03b1 that hold over many orders of\u2026",
      "content_html": "<!-- date: 2025-01-06 -->\n<p>Kaplan et al. [<a href=\"https://arxiv.org/abs/2001.08361\">1</a>] set the baseline picture: test loss falls as a power law $L(C) = (C/C_0)^{-\\alpha_C}$ in each of model size $N$, dataset size $D$, and compute $C$, with clean exponents $\\alpha$ that hold over many orders of magnitude. They also claimed that for a given compute budget, there is an optimal allocation between $N$ and $D$, and specifically that $N$ should grow faster than $D$ as $C$ grows.</p>\n\n <p>That paper shaped how labs designed pre-training experiments for the next two years, and the eventual \"Chinchilla\" effort grew out of trying to reproduce and extend its recommendations. It also turned out to be wrong about the optimal $D/N$ ratio, which is the part the field then had to revise.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/kaplan-fig1.png\" alt=\"Figure 1 of arxiv 2001.08361. Test loss plotted against compute on log-log axes, showing a power-law fit over many orders of magnitude.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"493\">\n <figcaption>Figure 1 of Kaplan et al. [<a href=\"https://arxiv.org/abs/2001.08361\">1</a>]. Log-log plot of test loss against compute. The power-law fit is tight over seven orders of magnitude which is what made the result so persuasive.</figcaption>\n </figure>\n\n\n <h2>Chinchilla</h2>\n\n <p>Hoffmann et al. [<a href=\"https://arxiv.org/abs/2203.15556\">2</a>] re-asked Kaplan's question with a larger, better-controlled experiment: over 400 language models from 70 million to over 16 billion parameters, trained on 5 to 500 billion tokens, across a sweep of compute budgets. Fitting a joint regression over model size and tokens gave them per-budget optima $N^{\\star}(C), D^{\\star}(C)$ that were not where Kaplan put them. Their headline rule is that for each doubling of model size, dataset size should also double, which lands at roughly 20 tokens per parameter at the compute-optimal point.</p>\n\n <p>The shift in conclusions came from a methodological gap. Hoffmann et al. argue that Kaplan's runs were too short: they ended well before each model had seen enough tokens to bottom out its loss. Kaplan's reported optima were therefore extrapolations from incomplete training curves, while Hoffmann's were drawn from runs that continued long enough to actually locate the per-size minimum. The two sets of optima ended up substantially different.</p>\n\n <h2>The effects of correcting Kaplan</h2>\n\n <p>GPT-3 and similarly Kaplan-trained models were trained on substantially too little data for their size. Chinchilla shows that for a fixed compute budget, a smaller model trained on more tokens reaches a lower loss than a bigger model trained on fewer. Concretely, Chinchilla-70B is roughly 2.5x smaller than GPT-3 (175 billion parameters) and outperforms it on nearly every benchmark in the original paper. Many factors contribute to that gap, but the dominant one is that Chinchilla-70B is configured at the Chinchilla-optimal point for its compute budget while GPT-3 sits at the Kaplan-optimal point. Subsequent scaling-law work that calibrates against the Chinchilla target accordingly emphasizes data scaling, longer training, and less aggressive model-size growth.</p>\n\n <p>Llama 2 was trained on 2 trillion tokens, well past the Chinchilla-optimal breakpoint and into a regime where the relevant trade-off is no longer training compute but inference cost. Once it became clear that smaller models are much cheaper to run at inference, scaling-law work shifted target from training-compute optimality to deployment optimality.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/hoffmann-fig3.png\" alt=\"Figure 3 of Hoffmann et al. 2022 (Chinchilla). Left: training loss versus parameter count for fixed FLOP budgets from 6e18 up to 3e21, each forming a U-shape with a clear minimum. Middle: optimal parameters versus FLOPs, extrapolating to ~63B parameters at ~1e23 FLOPs. Right: optimal training tokens versus FLOPs, extrapolating to ~1.4T tokens.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"461\">\n <figcaption>Figure 3 of Hoffmann et al. [<a href=\"https://arxiv.org/abs/2203.15556\">2</a>]. The IsoFLOPs decomposition: each line is a fixed compute budget, and the locus of minima defines the Chinchilla parameters/tokens scaling law.</figcaption>\n </figure>\n\n\n <h2>Broken laws</h2>\n\n <p>Caballero et al. [<a href=\"https://arxiv.org/abs/2210.14891\">3</a>] argue that the single-power-law narrative is wrong. Loss versus compute, in their fit, is a smoothly-broken power law: continuous everywhere, with several breaks where the slope changes. Kaplan and Chinchilla's single-exponent fits are then averaging across regimes with different exponents and producing a number that matches none of them.</p>\n\n <p>Whether the broken-laws view is useful depends on what you want the fit for. For coarse extrapolation across orders of magnitude in compute, a single power law still works fine. For predicting at what scale a particular capability shows up, it does not, and the Caballero breakpoints line up with the \"emergent\" capabilities $\\mathbb{1}[L &lt; L_{\\text{threshold}}]$ Wei et al. [<a href=\"https://arxiv.org/abs/2206.07682\">4</a>] documented, which Schaeffer et al. [<a href=\"https://arxiv.org/abs/2304.15004\">5</a>] then argued are largely an artifact of how the underlying metric is discretized.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/caballero-fig1.png\" alt=\"Figure 1 of Caballero et al. 2022. Annotated example of a Broken Neural Scaling Law (BNSL) functional form, marking three break points and four slope regimes between them as the performance metric is plotted against the quantity being scaled (log-log).\" style=\"max-width: 75%;\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"1005\">\n <figcaption>Figure 1 of Caballero et al. [<a href=\"https://arxiv.org/abs/2210.14891\">3</a>]. The BNSL form is a piecewise power law with explicit breaks; a single power law averages across these regimes and misses inflections the data actually shows.</figcaption>\n </figure>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/caballero-fig2.png\" alt=\"Figure 2 of Caballero et al. 2022. Two real-task BNSL fits: top panel ImageNet 25-shot test error versus training-dataset size; bottom panel TriviaQA few-shot test accuracy versus number of model parameters. Red curve is the BNSL fit; green points extend the fit beyond the training range.\" style=\"max-width: 75%;\" loading=\"lazy\" decoding=\"async\" width=\"968\" height=\"1370\">\n <figcaption>Figure 2 of Caballero et al. [<a href=\"https://arxiv.org/abs/2210.14891\">3</a>]. Two real-task examples: ImageNet error versus dataset size (top) and TriviaQA accuracy versus parameter count (bottom). The BNSL form tracks the data through visible breaks where a single power law would not.</figcaption>\n </figure>\n\n <p>Gwern's Scaling Hypotheses essay and the Revisited follow-up are the strongest non-specialist treatments of the underlying premise that capabilities come out of scale. Jacob Steinhardt's Bounded Regret is the blog I send people to when they want a careful read of what scaling laws actually let you predict. Beyond Chinchilla-Optimal is the most direct argument that the 20-tokens-per-parameter ratio is not a universal constant.</p>\n\n <p>Kaplan, Chinchilla, and Caballero all show that with a reasonably behaved architecture, a reasonable data mixture, and enough compute to get past the early warmup phase of pre-training, you can extrapolate loss from small runs to larger ones with usable accuracy. Error bars widen as the extrapolation gets more aggressive but not catastrophically. That predictive ability is what labs use to decide whether an expensive training run is worth doing.</p>\n\n <p>None of the papers here support any particular exponent or ratio as universal. Exponents depend on architecture, data mixture, and optimizer. The 20-tokens-per-parameter Chinchilla figure is one point estimate for one such combination, not a law of physics.</p>\n\n <p>Almost the entire scaling-law literature addresses test loss on the training distribution, and nothing else. It provides zero guidance on what data mixture to choose, what architecture will surpass dense transformers $f_\\theta$, what capabilities will appear at what scale, whether the resulting model will be safe, or how optimization will interact with the learning-rate schedule $\\eta(t)$. All of those are properties of individual training runs and live outside the fitted relationship. Treating scaling laws as though they answered them is a common failure mode, and much of the broken-laws literature is about that failure. Scaling laws tell you how much loss you will incur. Almost everything interesting about a model is invariant to that single number.</p>\n\n <h2>Further reading</h2>\n <ul class=\"further\">\n <li><a href=\"https://simons.berkeley.edu/talks/sasha-rush-cornell-university-hugging-face-2023-08-15\">Scaling Data-Constrained Language Models</a></li>\n <li><a href=\"https://simons.berkeley.edu/talks/when-scale-enough\">When is Scale Enough?</a></li>\n <li><a href=\"https://simons.berkeley.edu/talks/yasaman-bahri-google-deepmind-2023-08-15\">Simons talk on the theoretical side of scaling</a></li>\n <li><a href=\"https://arxiv.org/abs/2010.14701\">T. Henighan et al. Scaling laws for autoregressive generative modeling. arxiv 2010.14701, 2021</a></li>\n <li><a href=\"https://arxiv.org/abs/2406.12907\">T. Pearce and J. Song. Reconciling Kaplan and Chinchilla scaling laws. arxiv 2406.12907, 2024</a></li>\n <li><a href=\"https://proceedings.mlr.press/v235/sardana24a.html\">N. Sardana, J. Portes, S. Doubov, and J. Frankle. Beyond Chinchilla-Optimal: Accounting for inference in language model scaling laws. ICML, 2024</a></li>\n <li><a href=\"https://arxiv.org/abs/2602.07488\">Deriving neural scaling laws from the statistics of natural language. arxiv 2602.07488, 2026</a></li>\n </ul>\n\n<h2>References</h2>\n\n <ul class=\"refs\">\n <li>[<a href=\"https://arxiv.org/abs/2001.08361\">1</a>] J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models. arxiv 2001.08361, 2020.</li>\n <li>[<a href=\"https://arxiv.org/abs/2203.15556\">2</a>] J. Hoffmann et al. Training compute-optimal large language models. arxiv 2203.15556, 2022.</li>\n <li>[<a href=\"https://arxiv.org/abs/2210.14891\">3</a>] E. Caballero, K. Gupta, I. Rish, and D. Krueger. Broken neural scaling laws. arxiv 2210.14891, 2022.</li>\n <li>[<a href=\"https://arxiv.org/abs/2206.07682\">4</a>] J. Wei et al. Emergent abilities of large language models. arxiv 2206.07682, 2022.</li>\n <li>[<a href=\"https://arxiv.org/abs/2304.15004\">5</a>] R. Schaeffer, B. Miranda, and S. Koyejo. Are emergent abilities of large language models a mirage? arxiv 2304.15004, 2023.</li>\n </ul>",
      "date_published": "2025-01-06T12:00:00Z",
      "date_modified": "2026-08-21T12:00:00Z",
      "tags": [
        "scaling",
        "generalization"
      ]
    },
    {
      "id": "https://mouhssine.rifaki.me/notes/lottery-tickets/",
      "url": "https://mouhssine.rifaki.me/notes/lottery-tickets/",
      "title": "Lottery ticket hypothesis",
      "summary": "The lottery-ticket hypothesis of Frankle and Carbin [1] proposes that a randomly initialized dense network already contains a much sparser subnetwork (the \"winning ticket\") which, trained in isolation\u2026",
      "content_html": "<!-- date: 2024-11-16 -->\n<p>The lottery-ticket hypothesis of Frankle and Carbin [<a href=\"https://arxiv.org/abs/1803.03635\">1</a>] proposes that a randomly initialized dense network already contains a much sparser subnetwork (the \"winning ticket\") which, trained in isolation from its original initialization, matches the dense network's accuracy. If true, this is a strong claim about how deep networks represent functions: optimization would be selecting structure that was already present at initialization, not creating new structure. The pruning literature had circled this idea before; Frankle and Carbin's contribution was an actionable procedure for finding these subnetworks.</p>\n\n <h2>In practice, what Frankle and Carbin actually did</h2>\n\n <p>Frankle and Carbin defined a simple process to find these \"winning\" tickets. They called it iterative magnitude pruning. First train the full model until it converges. Remove the lowest-magnitude weights and freeze them. Take the remaining weights and reset them to their original initialization values. Train again. Continue this process several times, removing a portion of weights each time, until you cannot remove any more without losing performance relative to the full model. The resulting subnetwork is considered to be the \"winning\" ticket for that particular initial condition and dataset. The \"rewinding\" variant, resetting to weights from a few steps after initialization rather than to initialization itself, came later in Frankle's follow-up [<a href=\"https://arxiv.org/abs/1903.01611\">8</a>].</p>\n\n <p>On MNIST and small CIFAR architectures, the results are clean. Sparse subnetworks at sparsities of a few percent match the dense network's accuracy. The winning tickets are also tied to a specific initialization: a ticket found from one random init does not transfer to a different one, which is why the procedure is read as discovering structure already present at initialization rather than creating it during training.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/frankle-fig3.png\" alt=\"Figure 3 of Frankle and Carbin 2019. Test accuracy versus training iterations on Lenet-MNIST for lottery tickets at sparsity levels 100%, 51.3%, 21.1%, 7.0%, 3.6%, 1.9% remaining weights, plus reinitialized 51.3% and 21.1% baselines. Winning tickets reach the dense baseline; randomly-reinitialized counterparts plateau lower.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"522\">\n <figcaption>Figure 3 of Frankle and Carbin [<a href=\"https://arxiv.org/abs/1803.03635\">1</a>]. Lottery-ticket subnetworks recover the dense baseline down to a few percent of the original weights; the same masks with random reinitialization do not.</figcaption>\n </figure>\n\n\n <h2>Liu [<a href=\"https://arxiv.org/abs/1810.05270\">2</a>]'s rebuttal</h2>\n\n <p>Liu et al. ran the same procedure at larger scale and found that the gap between a Frankle winning ticket and a fresh random initialization of the same architecture closes once the network is big enough. At ImageNet scale, the winning-ticket effect basically disappears. Read narrowly this refutes Frankle and Carbin's strongest claim, but read alongside the original paper it is more usefully a scaling result: the winning-ticket structure is real on small networks and dissolves as scale grows.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/liu-fig2.png\" alt=\"Figure 2 of Liu et al. 2019. Schematic distinguishing predefined pruning (uniform x% per layer) from automatic pruning (per-layer percentages a%, b%, c%, d% chosen by the algorithm) on a 4-layer model.\" style=\"max-width: 60%;\" loading=\"lazy\" decoding=\"async\" width=\"1518\" height=\"1136\">\n <figcaption>Figure 2 of Liu et al. [<a href=\"https://arxiv.org/abs/1810.05270\">2</a>]. The two pruning regimes the paper distinguishes: predefined per-layer ratios versus automatically-discovered per-layer ratios. The \"rethinking\" results separate which regime the lottery-ticket conclusion survives in.</figcaption>\n </figure>\n\n <h2>Rewinding fixes</h2>\n\n <p>Frankle's response to Liu was a follow-up paper [<a href=\"https://arxiv.org/abs/1903.01611\">8</a>] introducing a small change that brought the result back at scale: instead of resetting the surviving weights to their original initialization, reset them to the values they had a short distance into training ($w_{t+1}=w_t-\\eta\\nabla L$, $\\eta$: step-size, $t$: current iteration). The rewind distance is a few hundred iterations on small models and a few epochs on ImageNet (around epoch 4 of 90 for ResNet-50, about 20k iterations). With rewinding, IMP works at ImageNet scale; rewind past that point and the matching subnetwork stops appearing. Renda et al. [<a href=\"https://arxiv.org/abs/2003.02389\">3</a>] tightened the recipe by comparing rewinding to plain fine-tuning across architectures and showing that rewinding only the learning-rate schedule matches or beats fine-tuning at fixed sparsity.</p>\n\n <h2>Softening the claims</h2>\n\n <p>The revised hypothesis is weaker than the original. The earlier claim was that the winning ticket exists at random initialization. The follow-up softens this: a short window of training (a few hundred iterations on small models, a few epochs on ImageNet) is enough to locate one. By the end of that window, something about the loss landscape $L(\\theta)$ has been fixed that determines the rest of training, and from that point on a sparse subnetwork pulled out of the surviving weights matches the dense network.</p>\n\n <h2>Why rewinding works</h2>\n\n <p>Frankle's companion paper offers an explanation. Early in training, two runs forked from the same initial weights end up in different basins $\\mathcal{B}$, so linear interpolation between two such checkpoints crosses a high-loss barrier. After a small fraction of training (about 1000-2000 iterations on CIFAR-scale networks, roughly the first 3% of the schedule), two runs forked from the same starting weights end up in the same basin instead, and linear interpolation between them stays at low loss throughout. The crossover point is where rewinding starts to work.</p>\n\n <p>The simplest framing of the lottery-ticket findings, given what is now understood about optimization and loss landscapes, is this: once a training run commits to a specific basin $\\mathcal{B}$, there is a sparse sub-network within that basin that matches the dense network's performance. Frankle and Carbin's strongest claims fail at larger scales, but their weaker claims have so far held up under every replication that has tested them. This places lottery-ticket results in close alignment with the mode-connectivity literature [<a href=\"https://arxiv.org/abs/1802.10026\">6</a>], particularly its linear-mode-connectivity refinement [<a href=\"https://arxiv.org/abs/1912.05671\">4</a>] and the later permutation-based alignment work [<a href=\"https://arxiv.org/abs/2209.04836\">7</a>].</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/frankle-lmc-fig3.png\" alt=\"Figure 3 of Frankle, Dziugaite, Roy, Carbin 2020. Linear-interpolation instability versus fork step k across LeNet (MNIST), ResNet-20 (CIFAR-10), VGG-16 (CIFAR-10), ResNet-50 (ImageNet), Inception-v3 (ImageNet). Instability collapses once k passes a small threshold.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"318\">\n <figcaption>Figure 3 of Frankle et al. [<a href=\"https://arxiv.org/abs/1912.05671\">4</a>]. Pairs of runs forked from a shared pre-rewinding checkpoint stay linearly connected; pairs forked from initialization do not. Linear mode connectivity is the operational test for \"same effective basin\".</figcaption>\n </figure>\n\n\n <p>Davis Blalock's 2020 MLSys retrospective on pruning, together with the ShrinkBench benchmark he built, is the survey I keep returning to on what holds up after the Frankle-to-Liu exchange. Blalock separates the stronger and weaker forms of the hypothesis and argues that \"checkpoint pruning\" is a better name than \"lottery ticket pruning\" for the late-rewinding procedures that actually work at scale. Google's <a href=\"https://research.google/pubs/the-state-of-sparsity-in-deep-neural-networks/\">State of Sparsity</a> and <a href=\"https://research.google/pubs/rigging-the-lottery-making-all-tickets-winners/\">Rigging the Lottery</a> are the other two retrospectives I keep going back to.</p>\n\n <p>The piece I keep coming back to is that mainstream theory has not picked up the rewinding point as an object in its own right. If someone could pin down precisely when basin-membership becomes determined, that would identify a structural feature of the loss landscape current frameworks do not explain. Lewkowycz's catapult-phase work [<a href=\"https://arxiv.org/abs/2003.02218\">5</a>] and the edge-of-stability literature look like they are circling the same phenomenon from different directions, and the lottery-ticket case is the most direct entry point for tying those threads together.</p>\n\n <figure class=\"tweet-embed\">\n <blockquote class=\"twitter-tweet\" data-dnt=\"true\"><p lang=\"en\" dir=\"ltr\">How do the lottery ticket hypothesis and the loss landscape relate? Winning lottery tickets always find the same, linearly-connected optimum. Check out our (@KDziugaite, <a href=\"https://twitter.com/roydanroy?ref_src=twsrc%5Etfw\">@roydanroy</a>, <a href=\"https://twitter.com/mcarbin?ref_src=twsrc%5Etfw\">@mcarbin</a>) poster at the SEDL workshop (West 121) and our new paper <a href=\"https://t.co/V9yKTSrNnh\">https://t.co/V9yKTSrNnh</a> <a href=\"https://t.co/uPwQKifo1W\">pic.twitter.com/uPwQKifo1W</a></p>&mdash; Jonathan Frankle (@jefrankle) <a href=\"https://twitter.com/jefrankle/status/1205902384112848899?ref_src=twsrc%5Etfw\">December 14, 2019</a></blockquote>\n <figcaption>Jonathan Frankle in 2019 pointing at the bridge between the LTH and mode connectivity. The mode-connectivity reading is the softer form of the hypothesis that holds at scale.</figcaption>\n </figure>\n\n <h2>Further reading</h2>\n        <ul class=\"further\">\n <li><a href=\"https://research.google/pubs/the-state-of-sparsity-in-deep-neural-networks/\">The State of Sparsity in DNNs</a></li>\n <li><a href=\"https://research.google/pubs/rigging-the-lottery-making-all-tickets-winners/\">Rigging the Lottery</a></li>\n          <li><a href=\"https://arxiv.org/abs/2009.08576\">J. Frankle, G. K. Dziugaite, D. M. Roy, and M. Carbin. Pruning neural networks at initialization: Why are we missing the mark? arxiv 2009.08576, 2020</a></li>\n          <li><a href=\"https://openreview.net/forum?id=Uzb45nolTb\">T. Kumar, K. Luo, and M. Sellke. No free prune: information-theoretic barriers to pruning at initialization. ICML, 2024</a></li>\n        </ul>\n\n<h2>References</h2>\n\n <ul class=\"refs\">\n <li>[<a href=\"https://arxiv.org/abs/1803.03635\">1</a>] J. Frankle and M. Carbin. The lottery ticket hypothesis: Finding sparse, trainable neural networks. arxiv 1803.03635, 2018.</li>\n <li>[<a href=\"https://arxiv.org/abs/1810.05270\">2</a>] Z. Liu, M. Sun, T. Zhou, G. Huang, and T. Darrell. Rethinking the value of network pruning. arxiv 1810.05270, 2018.</li>\n <li>[<a href=\"https://arxiv.org/abs/2003.02389\">3</a>] A. Renda, J. Frankle, and M. Carbin. Comparing rewinding and fine-tuning in neural network pruning. arxiv 2003.02389, 2020.</li>\n <li>[<a href=\"https://arxiv.org/abs/1912.05671\">4</a>] J. Frankle, G. K. Dziugaite, D. M. Roy, and M. Carbin. Linear mode connectivity and the lottery ticket hypothesis. arxiv 1912.05671, 2019.</li>\n <li>[<a href=\"https://arxiv.org/abs/2003.02218\">5</a>] A. Lewkowycz, Y. Bahri, E. Dyer, J. Sohl-Dickstein, and G. Gur-Ari. The large learning rate phase of deep learning: The catapult mechanism. arxiv 2003.02218, 2020.</li>\n <li>[<a href=\"https://arxiv.org/abs/1802.10026\">6</a>] T. Garipov, P. Izmailov, D. Podoprikhin, D. Vetrov, and A. G. Wilson. Loss surfaces, mode connectivity, and fast ensembling of DNNs. arxiv 1802.10026, 2018.</li>\n <li>[<a href=\"https://arxiv.org/abs/2209.04836\">7</a>] S. K. Ainsworth, J. Hayase, and S. Srinivasa. Git re-basin: Merging models modulo permutation symmetries. arxiv 2209.04836, 2022.</li>\n <li>[<a href=\"https://arxiv.org/abs/1903.01611\">8</a>] J. Frankle, G. K. Dziugaite, D. M. Roy, and M. Carbin. Stabilizing the lottery ticket hypothesis. arxiv 1903.01611, 2019.</li>\n </ul>",
      "date_published": "2024-11-16T12:00:00Z",
      "date_modified": "2026-08-21T12:00:00Z",
      "tags": [
        "generalization",
        "optimization"
      ]
    },
    {
      "id": "https://mouhssine.rifaki.me/notes/double-descent/",
      "url": "https://mouhssine.rifaki.me/notes/double-descent/",
      "title": "Double descent",
      "summary": "The classical U becomes a W with a second descent in the overparameterized (p>n) regime and that second descent often goes below the first minimum.",
      "content_html": "<!-- date: 2024-08-25 -->\n<p>The classical U becomes a W with a second descent in the overparameterized ($p \\gt n$) regime and that second descent often goes below the first minimum.</p>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/belkin-fig1.png\" alt=\"Figure 1 of Belkin et al. 2018. Schematic of test risk as a function of model capacity: the classical U-shape to the left of the interpolation threshold and a second descending branch to its right.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"445\">\n          <figcaption>Figure 1 of Belkin et al. [<a href=\"https://arxiv.org/abs/1812.11118\">1</a>]. Classical bias-variance to the left of the interpolation threshold; a second descent in the overparameterized regime on the right.</figcaption>\n        </figure>\n\n        <h2>Nakkiran et al.</h2>\n\n        <p>Nakkiran et al. [<a href=\"https://arxiv.org/abs/1912.02292\">2</a>] made the picture concrete by showing that the W-shape appears in three different axes: model size $p$, training time, and dataset size $n$. Model-wise double descent varies the width $k$ of a ResNet $f_\\theta$; epoch-wise double descent varies the number of training steps; sample-wise double descent varies dataset size with everything else held fixed. The shape is the same each time: a test-error peak near the interpolation threshold, then a descent once you push past it.</p>\n\n        <p>Epoch-wise is the most surprising of the three. Within one run, test error gets worse before it gets better. The worst test error sits roughly at the iteration where training loss first hits zero; train past it and the test error drops again.</p>\n\n        <h2>Label noise</h2>\n\n        <p>The sharpest versions of the double descent peak in these papers come with label noise. Nakkiran's headline plots use ten to twenty percent corrupted labels. Without label noise, the peak is much weaker and sometimes absent. Label noise inflates the variance contribution of the model at the interpolation threshold because the model is being asked to memorize random labels at exactly the capacity where memorization is possible but not easy.</p>\n\n        <p>Past the threshold, extra capacity absorbs the noise into higher-frequency components without disturbing the underlying signal. This is the hinge that connects the toy phenomenon to actual deep learning. Belkin's linear-regression result holds at all noise levels but the gap is small without noise; Nakkiran's dramatic curves require label noise to be visible. At modern language-model scale, with clean labels and large models, test loss is close to monotone in parameter count and scaling-law papers fit clean $L \\propto C^{-\\alpha}$ decay with no visible second peak. The effect is real and the classical bias-variance picture is wrong in the overparameterized regime, but the large peak that gives double descent its name is specific to the label-noise case.</p>\n\n        <p>Boaz Barak's Windows on Theory post makes a version of this argument: the interesting part of double descent is to the right of the peak, not the peak itself. OpenAI's Deep Double Descent post took Nakkiran to a much wider audience and posed the sharper question: given this effect, what kind of complexity control (if any) actually predicts generalization? Google Research's \"A new lens on understanding generalization in deep learning\" recast double descent in terms of an effective-capacity measure that tracks the empirical curves better than parameter count does. Misha Belkin's Simons Institute talks are the video account I keep sending people who want to see how the picture has changed since 2018.</p>\n\n        <figure class=\"tweet-embed\">\n          <blockquote class=\"twitter-tweet\" data-dnt=\"true\"><p lang=\"en\" dir=\"ltr\">So, &quot;double descent&quot; is happening b/c DF isn&#39;t really the right quantity for the the x-axis: like, the fact that we are choosing the minimum norm least squares fit actually means that the spline with 36 DF is **less** flexible than the spline with 20 DF. <br><br>Crazy, huh?<br><br>19/</p>&mdash; Daniela Witten (@daniela_witten) <a href=\"https://twitter.com/daniela_witten/status/1292293122752262145?ref_src=twsrc%5Etfw\">August 9, 2020</a></blockquote>\n        </figure>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/nakkiran-fig1.png\" alt=\"Figure 1 of Nakkiran et al. 2019. Test error and train error versus ResNet18 width parameter under varying label-noise levels (0%, 5%, 10%, 15%, 20%). Test error peaks near the interpolation threshold and decreases again as width grows.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"760\">\n          <figcaption>Figure 1 of Nakkiran et al. [<a href=\"https://arxiv.org/abs/1912.02292\">2</a>]. Model-wise double descent grows visibly with label-noise level: a peak at the interpolation threshold, then a second descent in the overparameterized regime.</figcaption>\n        </figure>\n\n\n        <p>I'm closer to Barak's reading than to what filtered down to practitioner intros.</p>\n\n        <p>After the label-noise caveat, two results from this line of work hold up. Classical capacity measures like VC dimension do not extend cleanly to the overparameterized regime and cannot be expected to predict generalization there. And overparameterized networks with astronomically large VC dimensions can sit well below smaller networks in test error on the same task \u2014 which implies the loss landscape is doing the selection: from a vast pool of interpolating solutions, the optimizer is picking a small subset that generalizes.</p>\n\n        <figure>\n          <img src=\"https://mouhssine.rifaki.me/notes/img/nakkiran-fig4.png\" alt=\"Figure 4 of Nakkiran et al. 2019. Left panel labels the classical (under-parameterized) and modern (over-parameterized) regimes around the interpolation threshold; right panel overlays test error across many epochs (color = epochs 1 to 1000) versus ResNet18 width, with an optimal early-stopping envelope.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"437\">\n          <figcaption>Figure 4 of Nakkiran et al. [<a href=\"https://arxiv.org/abs/1912.02292\">2</a>]. Pulling apart the canonical double-descent shape: the peak sits at the interpolation threshold, and the epoch-coloured family on the right makes epoch-wise double descent visible alongside model-wise.</figcaption>\n        </figure>\n\n        <p>What practitioners did with this was simpler than what the theory suggested: once you are in the overparameterized regime and you have compute to spend, bigger is usually better. The second descent has no obvious endpoint, which is why Kaplan-style scaling laws can fit clean power-law decay in compute \u2014 they are sitting entirely on the right-hand, log-log-linear side of the W-curve. The dramatic 2019 reading of double descent was that the bias-variance tradeoff is fiction and overfitting no longer exists. The second half of that is trivially untrue (overfitting is easy to produce in any small-data regime). The first half is more delicate: above the interpolation threshold, with implicit min-norm regularization ($\\theta^\\star = \\arg\\min_{f_\\theta(X)=y} \\|\\theta\\|$), larger models tend to generalize better rather than worse. That is a statement about a regime, not a law.</p>\n\n        <figure>\n          <div style=\"padding: 20px 16px; background: #E5DFCB; border: 1px solid #eee; border-radius: 4px; text-align:center;\">\n            $$R(p) = \\sigma^2 \\cdot \\frac{n}{|p-n-1|} + \\|\\beta\\|^2 \\cdot \\max\\!\\left(0,\\, 1 - \\frac{n}{p}\\right)$$\n          </div>\n          <figcaption>\n            Expected test risk of the min-norm ridgeless interpolant with $n$ samples and $p$ features under isotropic covariates, from Hastie et al. [<a href=\"https://arxiv.org/abs/1903.08560\">3</a>]. The first term diverges at $p=n$ and is the interpolation peak. The second shrinks as $p \\to \\infty$ and is the second descent. Double descent is not a deep-learning phenomenon in any strict sense since it falls out of the min-norm solution to an overparameterized least squares problem and holds only because the optimizer is selecting a specific well-behaved interpolant out of the many that fit the data.\n          </figcaption>\n        </figure>\n\n\n        <p>Another claim that does not survive the empirical record is that the peak is always exactly at the interpolation threshold. In practice, the exact location of the peak in Nakkiran's ResNet experiments depends on the effective number of parameters under whatever implicit regularization is in use, not on the total parameter count. The peak does not occur precisely at the width at which training error ($L_{\\text{train}}$) first reaches zero; it sits slightly past that point, where the network can memorize noisy labels without disrupting the underlying signal.</p>\n        <h2>Further reading</h2>\n        <ul class=\"further\">\n          <li><a href=\"https://arxiv.org/abs/2001.08361\">J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models. arxiv 2001.08361, 2020</a></li>\n          <li><a href=\"https://arxiv.org/abs/2003.01054\">S. d'Ascoli, M. Refinetti, G. Biroli, and F. Krzakala. Double trouble in double descent: Bias and variance(s) in the lazy regime. arxiv 2003.01054, 2020</a></li>\n          <li><a href=\"https://arxiv.org/abs/2003.01897\">P. Nakkiran, P. Venkat, S. Kakade, and T. Ma. Optimal regularization can mitigate double descent. arxiv 2003.01897, 2020</a></li>\n          <li><a href=\"https://arxiv.org/abs/2203.03466\">G. Yang, E. J. Hu, I. Babuschkin, S. Sidor, X. Liu, D. Farhi, N. Ryder, J. Pachocki, W. Chen, and J. Gao. Tensor Programs V: tuning large neural networks via zero-shot hyperparameter transfer. arxiv 2203.03466, 2022</a></li>\n          <li><a href=\"https://research.google/pubs/understanding-double-descent-requires-a-fine-grained-bias-variance-decomposition/\">B. Adlam and J. Pennington. Understanding double descent requires a fine-grained bias-variance decomposition. NeurIPS, 2020</a></li>\n          <li><a href=\"https://iclr-blogposts.github.io/2024/blog/double-descent-demystified/\">R. Schaeffer et al. Double descent demystified. ICLR Blogposts, 2024</a></li>\n        </ul>\n\n        \n\n<h2>References</h2>\n\n        <ul class=\"refs\">\n          <li>[<a href=\"https://arxiv.org/abs/1812.11118\">1</a>] M. Belkin, D. Hsu, S. Ma, and S. Mandal. Reconciling modern machine learning practice and the bias-variance trade-off. arxiv 1812.11118, 2018.</li>\n          <li>[<a href=\"https://arxiv.org/abs/1912.02292\">2</a>] P. Nakkiran, G. Kaplun, Y. Bansal, T. Yang, B. Barak, and I. Sutskever. Deep double descent: Where bigger models and more data hurt. arxiv 1912.02292, 2019.</li>\n          <li>[<a href=\"https://arxiv.org/abs/1903.08560\">3</a>] T. Hastie, A. Montanari, S. Rosset, and R. J. Tibshirani. Surprises in high-dimensional ridgeless least squares interpolation. arxiv 1903.08560, 2019.</li>\n        </ul>",
      "date_published": "2024-08-25T12:00:00Z",
      "date_modified": "2026-08-21T12:00:00Z",
      "tags": [
        "generalization",
        "scaling"
      ]
    },
    {
      "id": "https://mouhssine.rifaki.me/notes/information-bottleneck/",
      "url": "https://mouhssine.rifaki.me/notes/information-bottleneck/",
      "title": "Reading Tishby's information bottleneck",
      "summary": "Tishby and Zaslavsky's 2015 paper was, until fairly recently, one of the most-cited papers in deep-learning theory. They described training as two distinct phases. In the first, the \"fitting\" phase, the\u2026",
      "content_html": "<!-- date: 2024-07-09 -->\n<p>Tishby and Zaslavsky's 2015 paper was, until fairly recently, one of the most-cited papers in deep-learning theory. They described training as two distinct phases. In the first, the \"fitting\" phase, the mutual information between a hidden representation $T = f_\\theta(X)$ and the input $X$ rises. In the second, the \"compression\" phase, the network discards parts of that input information that are not useful for the prediction.</p>\n\n <p>The two strongest objections to that picture come from Saxe et al. [<a href=\"https://openreview.net/forum?id=ry_WPG-A-\">2</a>], who argue the compression phase is an artifact of the activation function and the mutual-information estimator, and from Goldfeld et al. [<a href=\"https://arxiv.org/abs/1810.05728\">3</a>], who formalize the estimator critique and re-run the analysis with noisy networks where mutual information is well-defined. The point of this post is to reread Tishby with both of those objections in hand and ask what remains.</p>\n\n <p>Both objections are worth reading in full; the summaries below are my best attempt to render them faithfully.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/tishby-fig2.png\" alt=\"Figure 2 of Tishby and Zaslavsky 2015. Qualitative information plane: optimal IB limit (black), suboptimal bifurcations (blue), finite-sample distortion bound (red), and a possible path of the layers in a typical DNN (green), with shaded regions marking the compression gap and generalization gap.\" style=\"max-width: 75%;\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"1224\">\n <figcaption>Figure 2 of Tishby and Zaslavsky [<a href=\"https://arxiv.org/abs/1503.02406\">1</a>]. The IB rate-distortion bound with deep-network layers placed on it: each layer trades compression of $X$ against retention of information about $Y$.</figcaption>\n </figure>\n\n\n <h2>The actual claim of the original paper</h2>\n\n <p>The claim, in plain language: early in training, a deep network builds up features that are useful for predicting the target $Y$, which shows up as $I(T; Y)$ increasing. During the same period $I(T; X)$ also increases, because the hidden layer is just retaining more about the input. Then a second phase kicks in: the network drops the parts of the input information that do not help with the prediction. That second phase is the \"compression phase.\"</p>\n\n <p>The empirical support was a small tanh network whose information-plane plot showed a clean fitting-then-compression trajectory.</p>\n\n <h2>Saxe et al. [<a href=\"https://openreview.net/forum?id=ry_WPG-A-\">2</a>] examine the type of activation function used</h2>\n\n <p>Saxe and coauthors re-ran the experiments with ReLU instead of tanh. The compression phase did not appear: $\\hat{I}(X;T)$ stayed roughly flat across training. Their explanation: tanh saturation pushes each unit's activations into a small number of values, the binning estimator is sensitive to that quantization, and the combination produces a curve that looks like compression but is really an estimator artifact. In a network without saturating units (or with a less binning-sensitive estimator), the curve is gone.</p>\n\n <p>That is a serious problem for the original story. The theoretical pull of Tishby's framing was that compression looked universal, a property of deep learning itself. If it only shows up for one activation function with one estimator, the universality claim is much weaker.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/saxe-fig1.png\" alt=\"Figure 1 of Saxe et al. ICLR 2018. Four information-plane panels (A, B, C, D). Top row uses a small toy network with the binning estimator: (A) tanh nonlinearity reproduces the Shwartz-Ziv & Tishby fitting-then-compression trajectory; (B) ReLU nonlinearity shows no compression phase. Bottom row uses a 784-1024-20-20-20-10 MNIST network with the Kolchinsky-Tracey KDE estimator: (C) tanh, no compression observed except in the final sigmoidal classification layer; (D) extension under the same KDE setup.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"1297\">\n <figcaption>Figure 1 of Saxe et al., <a href=\"https://openreview.net/forum?id=ry_WPG-A-\">ICLR 2018</a>. Swapping tanh for ReLU removes the compression phase (A vs B), and re-running with a KDE estimator at MNIST scale (C, D) also fails to reproduce it. The two-phase information-plane story is contingent on the nonlinearity and the estimator, not a property of training.</figcaption>\n </figure>\n\n\n <h2>Goldfeld et al. formally quantify issues with estimator selection</h2>\n\n <p>Goldfeld and coauthors formalized what Saxe had observed. For continuous random variables and deterministic maps $T = f(X)$ from inputs to representations, mutual information $I(T;X) = H(T) - H(T|X)$ is not even well-defined: $H(T|X)$ collapses, $I(T;X)$ blows up, and the finite numbers showing up in published plots are entirely coming from the noise injected by the estimator (binning, added Gaussian noise of variance $\\sigma^2$, or KDE). Different choices give different numbers, so the published information-plane trajectories were tracking properties of the estimator at least as much as properties of the network.</p>\n\n <p>One reason Tishby's paper still has value despite the empirical claims being discredited is that it offered a third lens on generalization at a time when the dominant lenses were capacity-based (how restrictive or broad the hypothesis class is) and geometry-based (how smooth or rough the loss landscape is around a minimum). Tishby's lens was sufficient-statistics: representations should retain only the information that matters for the prediction task. The terminology stuck even though the original empirical observation did not, and modern self-supervised methods like infoNCE are essentially information-bottleneck objectives in everything but name.</p>\n\n <p>The paper got most of its public reach through Natalie Wolchover's 2017 Quanta piece, \"New Theory Cracks Open the Black Box of Deep Learning,\" which presented Tishby's claims in their strongest form. Reading that piece today is mostly useful as a reminder of how far ahead of the evidence the rhetoric got.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/goldfeld-fig1.png\" alt=\"Figure 1 of Goldfeld et al. 2019. Estimated $I(X; \\mathrm{Bin}(T_\\ell))$ over training epochs for layers 1-5 at four binning resolutions (bin size 0.0001, 0.001, 0.01, 0.1). The apparent compression phase appears or disappears depending on the bin size.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"236\">\n <figcaption>Figure 1 of Goldfeld et al. [<a href=\"https://arxiv.org/abs/1810.05728\">3</a>]. The same training run produces qualitatively different \"information-plane trajectories\" depending on the bin size used to estimate mutual information - the compression phase is partly an estimator artefact.</figcaption>\n </figure>\n\n\n <p>For readers who want non-paper summaries of this debate, Adrian Colyer's three-part Morning Paper series on Tishby's IB theory and Saxe's reply is the most accessible walkthrough.</p>\n\n <p>A weaker version of the original claim does survive: networks trained with SGD often end up with representations that are sufficient for the labels and roughly insensitive to label-irrelevant input variation. \"Information bottleneck\" is a fine descriptive label for that. The strong version, in which training proceeds through two cleanly separated phases divided by a phase transition in $I(T;X)$, has no empirical support, and the further claim that SGD is implicitly minimizing an information-bottleneck objective remains unproven.</p>\n\n <p>Some incorrect papers end up more useful to a field than correct ones, because they hand it vocabulary it did not have.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/goldfeld-fig2.png\" alt=\"Figure 2 of Goldfeld et al. 2019. Architectural diagram of the noisy DNN: $T_{\\ell-1}$ feeds through $\\sigma(W_\\ell^{(k)} T_{\\ell-1} + b_\\ell^{(k)})$ to produce a pre-noise hidden $S_\\ell(k)$, to which Gaussian noise $Z_\\ell(k) \\sim \\mathcal{N}(0,\\beta^2)$ is added to yield the next-layer hidden $T_\\ell(k)$.\" style=\"max-width: 55%;\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"474\">\n <figcaption>Figure 2 of Goldfeld et al. [<a href=\"https://arxiv.org/abs/1810.05728\">3</a>]. The noisy-network construction: adding Gaussian noise after each layer makes mutual information well-defined and lets the analysis distinguish genuine compression dynamics from estimator artefacts.</figcaption>\n </figure>\n\n\n <p>Saxe's argument alone is not fatal: a noisy version of the network has well-defined mutual information and can be analyzed directly, which gets you out of the estimator trap. Goldfeld et al. did exactly that and found the two-phase trajectory does not hold up across estimator choices once the quantities being plotted are well-defined. After that, the empirical case for Tishby's strong claims is essentially gone.</p>\n\n <h2>Further reading</h2>\n <ul class=\"further\">\n <li><a href=\"https://arxiv.org/abs/1703.00810\">R. Shwartz-Ziv and N. Tishby. Opening the black box of deep neural networks via information. arxiv 1703.00810, 2017</a></li>\n <li><a href=\"https://arxiv.org/abs/1612.00410\">A. A. Alemi, I. Fischer, J. V. Dillon, and K. Murphy. Deep variational information bottleneck. arxiv 1612.00410, 2017</a></li>\n <li><a href=\"https://arxiv.org/abs/1807.03748\">A. van den Oord, Y. Li, and O. Vinyals. Representation learning with contrastive predictive coding. arxiv 1807.03748, 2018</a></li>\n <li><a href=\"https://arxiv.org/abs/2305.18887\">K. Kawaguchi, Z. Deng, X. Ji, and J. Huang. How does information bottleneck help deep learning? arxiv 2305.18887, 2023</a></li>\n </ul>\n\n<h2>References</h2>\n \n <ul class=\"refs\">\n <li>[<a href=\"https://arxiv.org/abs/1503.02406\">1</a>] N. Tishby and N. Zaslavsky. Deep learning and the information bottleneck principle. arxiv 1503.02406, 2015.</li>\n <li>[<a href=\"https://openreview.net/forum?id=ry_WPG-A-\">2</a>] A. M. Saxe, Y. Bansal, J. Dapello, M. Advani, A. Kolchinsky, B. D. Tracey, and D. D. Cox. On the information bottleneck theory of deep learning. ICLR, 2018.</li>\n <li>[<a href=\"https://arxiv.org/abs/1810.05728\">3</a>] Z. Goldfeld, E. van den Berg, K. Greenewald, I. Melnyk, N. Nguyen, B. Kingsbury, and Y. Polyanskiy. Estimating information flow in deep neural networks. arxiv 1810.05728, 2019.</li>\n </ul>",
      "date_published": "2024-07-09T12:00:00Z",
      "date_modified": "2026-08-21T12:00:00Z",
      "tags": [
        "theory",
        "generalization"
      ]
    },
    {
      "id": "https://mouhssine.rifaki.me/notes/flat-minima/",
      "url": "https://mouhssine.rifaki.me/notes/flat-minima/",
      "title": "On flat minima",
      "summary": "Whether flat minima generalize better than sharp ones has been an open question for about seven years. The debate seems to close every year and reopen a year later. Most readers entering the field\u2026",
      "content_html": "<!-- date: 2024-04-07 -->\n<p>Whether flat minima generalize better than sharp ones has been an open question for about seven years. The debate seems to close every year and reopen a year later. Most readers entering the field encounter it as a settled topic in some textbook chapter, in one direction or the other, when in fact it isn't. This is where I think it actually stands.</p>\n\n <h2>Hochreiter, Schmidhuber, and the original intuition</h2>\n\n <p>Hochreiter and Schmidhuber introduced the idea in 1997: a minimum that sits in a broad, low-curvature valley should generalize better than one in a sharp valley, on roughly an MDL ground that the broad solution requires fewer bits to specify and is correspondingly less tied to the noise in any one training set. The intuition lay mostly dormant for the next two decades.</p>\n\n <p>Keskar et al. [<a href=\"https://arxiv.org/abs/1609.04836\">1</a>] reignited the topic by reporting that large-batch SGD with $\\theta_{t+1} = \\theta_t - \\eta g_t$ converges to sharper minima than small-batch SGD on a range of standard benchmarks, with a corresponding gap in test accuracy. The 1D schematic that came with that paper has done a lot of work since: it is the picture more or less every later flat-minima discussion is implicitly arguing about.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/keskar-fig1.png\" alt=\"Figure 1 of Keskar et al. [1]. A 1D schematic of a wide basin around one minimum and a narrow basin around another.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"643\">\n <figcaption>Figure 1 of Keskar et al. [<a href=\"https://arxiv.org/abs/1609.04836\">1</a>]. The picture that started the modern argument.</figcaption>\n </figure>\n\n\n <h2>Dinh et al.'s objection, which should have ended the debate</h2>\n\n <p>Dinh et al. argued that this should have ended the debate. Sharpness, measured as $\\lambda_{\\max}(H)$ with $H = \\nabla^{2} L$, is a property of the parameterization, not of the function the network represents. They show explicitly that for any minimum one can find a reparameterization $\\theta \\to \\psi(\\theta)$ that scales the Hessian eigenvalues $\\lambda_i(H)$ to arbitrary values without changing the input-output map. The clean response to this would have been to drop the flat-minima paradigm. The community instead salvaged it by looking for sharpness measures that are invariant under reparameterization. The simplest of these comes from Dziugaite and Roy, who use a PAC-Bayes lens: define sharpness as the largest weight perturbation $\\xi \\sim \\mathcal{N}(0, I)$ a minimum can absorb while keeping training loss $L(\\theta)$ small. That measure is reparameterization-invariant by construction and correlates with generalization in their experiments.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/dinh-fig1.png\" alt=\"Figure 1 of Dinh et al. 2017. Schematic of an $\\epsilon$-flat minimum: a parabolic loss curve in $(\\theta, L)$ with a horizontal cutoff at level $\\epsilon$ above the minimum, shading the connected set of parameters whose loss stays within $\\epsilon$ of the optimum.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"551\">\n <figcaption>Figure 1 of Dinh et al. [<a href=\"https://arxiv.org/abs/1703.04933\">2</a>]. The width of the shaded $\\epsilon$-flat region is the geometric quantity flatness intuitions are pointing at - but its size depends on the parameterization, which is the heart of the reparameterization argument.</figcaption>\n </figure>\n\n\n <h2>SAM and where flatness wins</h2>\n\n <p>Building on Dziugaite and Roy's framing, Foret et al. [<a href=\"https://arxiv.org/abs/2010.01412\">3</a>] turned a reparameterization-aware sharpness measure into a training objective: minimize the worst-case loss in an $\\ell_2$ ball around the current weights. They called the procedure SAM, and it does deliver consistent test-accuracy gains, particularly on architectures without strong built-in inductive biases (vanilla MLPs, plain ViTs without strong augmentation). Behnam Neyshabur, a co-author, has remained one of the more consistent public advocates for SAM as a generalization tool.</p>\n\n <p>On the side that flatness improves generalization at scale, the main public voices are Behnam Neyshabur and collaborators across several papers and talks, and Boaz Barak's posts at Windows on Theory. Ferenc Husz\u00e1r's inFERENCe is the other blog I keep coming back to on this; he writes carefully about flatness, generalization, and the Bayesian readings sitting under both. The Off-convex blog has good coverage of mode connectivity that puts flatness inside a larger geometric story, which I find more useful than treating it as an independent explanation. From the continuous-time view, the stochastic-diffusion picture of SGD is still the most direct way to see why noisy iterates concentrate near flatter minima. Husz\u00e1r's adjacent essay \"Everything that Works Works Because It Is Bayesian\" is the prior-based reading I find most useful.</p>\n\n <p>Stopping the timeline here, the flat-minima view would look basically validated: the naive Hessian definition was broken, but a reparameterization-invariant version of the phenomenon is real and SAM is a way to act on it. Kaddour et al. [<a href=\"https://arxiv.org/abs/2202.00661\">4</a>] complicate that picture. They sweep SAM and SWA against vanilla Adam across architectures and dataset scales, and report that the SAM gain shrinks as either model size or dataset size grows. In the regimes where generalization is most useful to improve, the advantage over a tuned Adam baseline narrows substantially.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/foret-fig1.png\" alt=\"Figure 1 of Foret et al. 2020 (SAM). Left: percent error reduction from SAM across CIFAR10, CIFAR100, ImageNet, finetuning, SVHN, F-MNIST, and noisy CIFAR. Right: 3D loss landscape of a SAM-trained network (smooth blue valley) compared with a sharp, jagged surface from standard training.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"458\">\n <figcaption>Figure 1 of Foret et al. [<a href=\"https://arxiv.org/abs/2010.01412\">3</a>]. The empirical headline (error reduction across tasks) plus the loss-landscape contrast that motivates the worst-case-in-a-ball SAM objective.</figcaption>\n </figure>\n\n\n <p>I am left with an uneven picture. The naive Hessian definition of sharpness has no causal link to generalization (by Dinh). A reparameterization-invariant version does correlate with generalization at small-to-medium scale, and SAM produces real test-accuracy gains on architectures with weak inductive biases. At very large scale the correlation weakens and the SAM gain fades. I do not have a clean account of why. My guess is that with rich enough data and architectures, the optimizer's trajectory and the data distribution dominate whatever local geometry the final minimum has, and the landscape framing stops being the right description. That is speculation, and I would not put weight on it beyond that.</p>\n <h2>Further reading</h2>\n <ul class=\"further\">\n <li><a href=\"https://arxiv.org/abs/1703.11008\">G. K. Dziugaite and D. M. Roy. Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data. arxiv 1703.11008, 2017</a></li>\n <li><a href=\"https://arxiv.org/abs/2106.01548\">X. Chen, C.-J. Hsieh, and B. Gong. When vision transformers outperform ResNets without pre-training or strong data augmentations. arxiv 2106.01548, 2021</a></li>\n <li><a href=\"https://arxiv.org/abs/1803.05407\">P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. G. Wilson. Averaging weights leads to wider optima and better generalization. arxiv 1803.05407, 2018</a></li>\n <li><a href=\"https://arxiv.org/abs/2203.05482\">M. Wortsman et al. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. arxiv 2203.05482, 2022</a></li>\n <li><a href=\"https://direct.mit.edu/neco/article/9/1/1/6027/Flat-Minima\">S. Hochreiter and J. Schmidhuber. Flat minima</a></li>\n </ul>\n\n \n\n<h2>References</h2>\n \n <ul class=\"refs\">\n <li>[<a href=\"https://arxiv.org/abs/1609.04836\">1</a>] N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang. On large-batch training for deep learning: Generalization gap and sharp minima. arxiv 1609.04836, 2016.</li>\n <li>[<a href=\"https://arxiv.org/abs/1703.04933\">2</a>] L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio. Sharp minima can generalize for deep nets. arxiv 1703.04933, 2017.</li>\n <li>[<a href=\"https://arxiv.org/abs/2010.01412\">3</a>] P. Foret, A. Kleiner, H. Mobahi, and B. Neyshabur. Sharpness-aware minimization for efficiently improving generalization. arxiv 2010.01412, 2020.</li>\n <li>[<a href=\"https://arxiv.org/abs/2202.00661\">4</a>] J. Kaddour, L. Liu, R. Silva, and M. J. Kusner. When do flat minima optimizers work? arxiv 2202.00661, 2022.</li>\n </ul>",
      "date_published": "2024-04-07T12:00:00Z",
      "date_modified": "2026-08-21T12:00:00Z",
      "tags": [
        "optimization",
        "generalization"
      ]
    },
    {
      "id": "https://mouhssine.rifaki.me/notes/grokking/",
      "url": "https://mouhssine.rifaki.me/notes/grokking/",
      "title": "Four explanations for Grokking",
      "summary": "The network has generalized but long after it has already fit the data. The paper is Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets by Power et al. [1]. I came across it maybe a\u2026",
      "content_html": "<!-- date: 2024-02-24 -->\n<p>The network has generalized but long after it has already fit the data. The paper is Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets by Power et al. [<a href=\"https://arxiv.org/abs/2201.02177\">1</a>]. I came across it maybe a week after it was posted and didn't really know what to do with it for a while.</p>\n\n <h2>Why this is a puzzle</h2>\n\n <p>Standard stories about generalization do not predict a long delay. The VC and Rademacher story says generalization either happens or doesn't as a function of how well the hypothesis class matches the data distribution. The implicit-bias stories say SGD hits a minimum that generalizes but that minimum should be found at roughly the same time as convergence of training loss, not thousands of steps later.</p>\n\n <p>Grokking says the loss landscape has a second stage not driven by training loss. Some other variable is moving the weights ($\\theta$) during the stretch where the training loss is already near zero.</p>\n\n <h2>Four explanations</h2>\n\n <p>The first explanation - and the one that the original paper kind of points toward - is that weight decay ($\\lambda \\|\\theta\\|_2^2$) is a slow regularizer. The paper notes that grokking only happens with weight decay turned on. Once training loss is zero, the weight-decay term keeps pulling the norm of the weights down even if that loss gradient is tiny. This is slow drift toward a lower-norm solution which may be the one that generalizes. The second explanation comes via mechanistic interpretability.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/grokking-fig1.png\" alt=\"Figure 1 of arxiv 2201.02177. Training and validation accuracy on modular arithmetic as a function of optimization step on a log scale. The validation curve stays at chance while the training curve saturates, then jumps to one hundred percent much later.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"432\">\n <figcaption>Figure 1 of Power et al. [<a href=\"https://arxiv.org/abs/2201.02177\">1</a>]. Training accuracy saturates early; validation accuracy stays flat for orders of magnitude of additional steps and then jumps.</figcaption>\n </figure>\n\n\n <p>Neel Nanda and collaborators identified specific circuits inside the small grokking networks that implement modular arithmetic via a Fourier ($\\hat{f}(\\omega)$) decomposition. The grokking transition is when those circuits finish being assembled. Before the transition the network has to memorize by brute force. After the transition it can actually compute. The circuits-level framing this work builds on is laid out in Elhage et al., which is where I'd send anyone to understand what it could mean to talk about a 'circuit' inside a transformer.</p>\n\n <p>The third explanation, due to Liu, Michaud, and Tegmark in <a href=\"https://arxiv.org/abs/2210.01117\">Omnigrok</a>, is geometric: the generalizing solution lies in a narrow \"Goldilocks zone\" of weight norms, and grokking is what you see when the optimizer has been started outside that zone and is slowly being walked into it. Weight decay is what does the walking, which is consistent with explanation one. The fourth angle is to step back and read all of this as a single process viewed at different mesh scales: grokking = double descent but resolved over training time rather than over model or dataset size.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/nanda-fig2.png\" alt=\"Figure 2 of Nanda et al. 2023. Left: histogram of fraction-of-variance-explained by degree-2 polynomials over neurons. Right: heatmap of components of $W_L$ corresponding to frequency-14 neurons, showing weight concentrated at the sin/cos basis pair for that frequency.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"411\">\n <figcaption>Figure 2 of Nanda et al. [<a href=\"https://arxiv.org/abs/2301.05217\">2</a>]. The grokked network's neurons are well-explained by degree-2 polynomials (left), and individual neurons read off specific Fourier-basis pairs from the embedding (right) - the Fourier circuit is mechanistically visible.</figcaption>\n </figure>\n\n\n <p>In this reading, grokking is double descent unfolding in time. The most lucid public articulations of this \"double descent over time\" reading come from Preetum Nakkiran's writing, and OpenAI's Deep Double Descent writeup paints the picture visually. On the mechanistic side, Neel Nanda wrote an intuitive walkthrough of the Fourier circuit story and maintains a corresponding paper page. Google PAIR's <a href=\"https://pair.withgoogle.com/explorables/grokking/\">Do Machine Learning Models Memorize or Generalize?</a> poses the same question in nearby visual vocabulary.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/nanda-fig5.png\" alt=\"Figure 5 of arxiv 2301.05217. The Discrete Fourier Transform of the grokked network's input embeddings, showing concentration on a small set of frequencies.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"435\">\n <figcaption>Figure 5 of Nanda et al. [<a href=\"https://arxiv.org/abs/2301.05217\">2</a>]. The DFT of the network's learned embeddings concentrates in a small number of frequencies after the transition which is the Fourier-based modular arithmetic circuit made visible.</figcaption>\n </figure>\n\n\n <p>These four views are looking at the same puzzle from different angles, and together they read as one story at different levels of abstraction. Weight decay is the optimization pressure: it selects a minimum-norm interpolant, and in modular arithmetic that interpolant admits the Fourier circuit because Fourier captures the low-rank ($\\text{rank}(W) \\ll d$) structure of the task. The broader pattern of fast memorization followed by slow compression is the shape double descent takes when you resolve it over time instead of over model size.</p>\n\n <figure>\n <img src=\"https://mouhssine.rifaki.me/notes/img/nanda-fig3.png\" alt=\"Figure 3 of Nanda et al. 2023. Average train accuracy (saturates near 1.0 within ~1k epochs), average test accuracy (stays at chance for ~5k epochs then jumps), and corresponding average train/test log-loss curves over epochs. Faded background lines show individual seeds.\" loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"419\">\n <figcaption>Figure 3 of Nanda et al. [<a href=\"https://arxiv.org/abs/2301.05217\">2</a>]. The grokking pattern made averaged: training accuracy saturates fast, test accuracy lags by orders of magnitude before its sudden rise.</figcaption>\n </figure>\n\n\n <p>The one thing none of these explanations cleanly accounts for is the abruptness of the transition. Smoothly shrinking the norm and smoothly assembling circuits should give smoothly rising validation accuracy, not a near-vertical jump. My read is that the sharpness is largely a measurement artifact: softmax classifiers, $\\sigma(z)_j = e^{z_j}/\\sum_k e^{z_k}$, route through a top-1 argmax, so logits that are evolving continuously map onto a piecewise-constant accuracy curve that flips once the right logit crosses its competitor.</p>\n\n <figure class=\"tweet-embed\">\n <blockquote class=\"twitter-tweet\" data-dnt=\"true\"><p lang=\"en\" dir=\"ltr\">So what&#39;s behind grokking?<br>Three phases of training:<br>1 Memorization<br>2 Circuit formation: It smoothly TRANSITIONS from memorising to generalising<br>3 Cleanup: Removing the memorised solution<br><br>Test performance needs a general circuit AND no memorisation so Grokking occurs at cleanup! <a href=\"https://t.co/zLnP92RXKV\">pic.twitter.com/zLnP92RXKV</a></p>&mdash; Neel Nanda (@NeelNanda5) <a href=\"https://twitter.com/NeelNanda5/status/1616590960066203648?ref_src=twsrc%5Etfw\">January 21, 2023</a></blockquote>\n <figcaption>Neel Nanda summarizing the three-phase mechanistic account of grokking from [<a href=\"https://arxiv.org/abs/2301.05217\">2</a>]. The transition to generalization happens during cleanup, not during circuit formation which is why it looks sudden.</figcaption>\n </figure>\n\n <h2>Further reading</h2>\n <ul class=\"further\">\n <li><a href=\"https://www.neelnanda.io/mechanistic-interpretability/modular-addition-walkthrough\">accessible walkthrough</a></li>\n <li><a href=\"https://www.neelnanda.io/grokking-paper\">dedicated paper page</a></li>\n <li><a href=\"https://transformer-circuits.pub/2021/framework/index.html\">N. Elhage et al. A mathematical framework for transformer circuits. transformer-circuits.pub/2021/framework, 2021</a></li>\n <li><a href=\"https://arxiv.org/abs/2210.01117\">Z. Liu, E. Michaud, and M. Tegmark. Omnigrok: Grokking beyond algorithmic data. arxiv 2210.01117, 2022</a></li>\n <li><a href=\"https://arxiv.org/abs/2206.04817\">V. Thilak, E. Littwin, S. Zhai, O. Saremi, R. Paiss, and J. Susskind. The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon. arxiv 2206.04817, 2022</a></li>\n <li><a href=\"https://arxiv.org/abs/1812.11118\">M. Belkin, D. Hsu, S. Ma, and S. Mandal. Reconciling modern machine learning practice and the bias-variance trade-off. arxiv 1812.11118, 2018</a></li>\n <li><a href=\"https://arxiv.org/abs/1912.02292\">P. Nakkiran, G. Kaplun, Y. Bansal, T. Yang, B. Barak, and I. Sutskever. Deep double descent: Where bigger models and more data hurt. arxiv 1912.02292, 2019</a></li>\n <li><a href=\"https://arxiv.org/abs/2303.06173\">X. Davies, L. Langosco, and D. Krueger. Unifying grokking and double descent. arxiv 2303.06173, 2023</a></li>\n <li><a href=\"https://arxiv.org/abs/2501.04697\">L. Prieto, M. Barsbey, P. Mediano, and T. Birdal. Grokking at the edge of numerical stability. arxiv 2501.04697, 2025</a></li>\n <li><a href=\"https://arxiv.org/abs/2311.18817\">K. Lyu, J. Jin, Z. Li, S. S. Du, J. D. Lee, and W. Hu. Dichotomy of early and late phase implicit biases can provably induce grokking. arxiv 2311.18817, 2023</a></li>\n </ul>\n\n<h2>References</h2>\n \n <ul class=\"refs\">\n <li>[<a href=\"https://arxiv.org/abs/2201.02177\">1</a>] A. Power, Y. Burda, H. Edwards, I. Babuschkin, and V. Misra. Grokking: generalization beyond overfitting on small algorithmic datasets. arxiv 2201.02177, 2022.</li>\n <li>[<a href=\"https://arxiv.org/abs/2301.05217\">2</a>] N. Nanda, L. Chan, T. Lieberum, J. Smith, and J. Steinhardt. Progress measures for grokking via mechanistic interpretability. arxiv 2301.05217, 2023.</li>\n </ul>",
      "date_published": "2024-02-24T12:00:00Z",
      "date_modified": "2026-08-21T12:00:00Z",
      "tags": [
        "generalization",
        "optimization"
      ]
    }
  ]
}
