Why autonomy is a coordination problem

Autonomy scales, but safety does not come from any single system getting better. When many agents share space and time, decisions collide and hazards compound, so keeping us safe becomes a coordination problem that compounds with every participant. Nobody in the field disagrees with that. Such wide agreement is what makes the sentence nearly useless on its own. What I want to argue is the part people do disagree with: coordination cannot be added afterwards. The prevailing plan is to build the vehicles now and standardize the coordination between them later as a protocol or a service or a certificate class. My argument is that the plan fails. Where the question admits of measurement I have measured it. A few hours of live air traffic tell us what proximity does in a real sky. A small experiment anyone can rerun in twenty seconds tells us what a coordination convention amounts to inside a learned policy. Each is linked where it is used and each contradicted something I believed before I ran it.

Layering

The instinct to add coordination later is the layering instinct and it is usually right. Networking works because TCP does not need to know how Ethernet frames a packet. The layer below can fail, the layer above notices, and the recovery is a retransmission costing milliseconds and nothing else. Layering buys independence at the price of one assumption: that failure at the lower layer is recoverable at the higher one.

The FAA and TSA proposal of 7 August 2025 on beyond-visual-line-of-sight operations [1] writes that instinct into law. It creates a whole new certificate class called part 146 for automated data service providers. Operators in controlled airspace or over dense population must buy strategic deconfliction from one of them or certificate themselves to self-provide. The proposal certifies the aircraft on one side and the service on the other, and what passes between the two is a request and a response.

Perrow drew the distinction that governs this. Interactive complexity means a system with more ways to interact than anyone has modeled; tight coupling means a failure reaching its consequence with nothing in between to absorb it. He argued that the dangerous quadrant belongs to systems with both at once [2]. Layering answers the first of those and has nothing to offer against the second, since the recovery it depends on costs time that a tightly coupled system has already spent. Two hundred aircraft over a city at two hundred feet occupies that quadrant, where the round trip a retransmission needs runs longer than the interval in which a conflict resolves itself one way or the other.

Leveson's version is more useful still, because she models safety as a hierarchy of controllers and locates accidents in inadequate control at some level [3]. A coordination service is a controller. The question a layering plan has to answer is what it controls and for how long after it stops speaking.

Where the layer ends, in the rule's own words

The proposal answers that question with unusual precision. The answer is worth reading as an engineering specification rather than as regulation.

Proposed § 108.190(c)(1) requires strategic deconfliction to "perform strategic conflict detection and resolution prior to takeoff, and in relation to other unmanned aircraft operations that are discoverable at that time." Each qualifier carries weight. The guarantee is computed once before departure over whichever intents happened to be visible at that instant. Proposed § 108.190(d) adds conformance monitoring. Its specified output is an immediate alert to operations personnel when the aircraft leaves its intent, which is to say the layer's response to a violation is to tell a human. The definitions come from ASTM F3548-21 as the rule's own footnote records, and that standard draws its boundaries in clause 1.13. It "does not purport to address tactical conflicts between UAS" and it "does not purport to address operations in locations where persistent connectivity is unavailable" [4]. Those are the two conditions under which the rest of this note is written.

The rest of the proposal [1] follows the same shape. Proposed § 108.195 orders an unmanned aircraft against aircraft arriving at an airport and against aircraft that broadcast their position. Proposed § 108.825 defines the collision-avoidance design requirement by pointing back to it, so a part 108 aircraft is ranked against manned traffic and unranked against another part 108 aircraft. Proposed § 108.210(a) fixes the aircraft-to-coordinator ratio at one to one, waivable by "a method acceptable to the Administrator" that nobody has written. Proposed § 108.815(b) requires the airframe to "execute a safe predetermined action when reaching the link timeout." Proposed § 146.325(a) requires the service provider to report an unscheduled outage after the fact.

Which of the proposed mechanisms is awake in which phase of a single flight A chart with four columns for the phases of one flight: before takeoff, in flight on intent, off intent with the link up, and off intent with the link down. Strategic conflict detection under section 108.190(c)(1) covers the first column only. Conformance monitoring under 108.190(d) covers the two middle columns and alerts operations personnel. The right-of-way rule under 108.195 spans all four columns and orders the aircraft toward aircraft that broadcast. The predetermined action under 108.815(b) covers the last column, one aircraft at a time. A fifth row, joint resolution between two unmanned aircraft after takeoff, spans all four columns with a dashed empty outline labeled a method acceptable to the Administrator. when each proposed mechanism is awake the phases of one flight, left to right before takeoff in flight, on intent off intent, link up off intent, link down strategic conflict detection § 108.190(c)(1) against the intents discoverable then conformance monitoring § 108.190(d) alerts operations personnel right of way § 108.195 toward aircraft that broadcast predetermined action § 108.815(b) one aircraft at a time joint resolution between two aircraft, after takeoff a method acceptable to the Administrator takeoff landing Read off the proposed rule text of 7 August 2025, 90 FR 38212. A solid outline is a requirement in that phase. The right-of-way rule orders an unmanned aircraft against aircraft that broadcast their position, which leaves two unmanned aircraft unordered.
This one is a diagram rather than a measurement: it maps the proposed rule text onto the phases of a flight, read off 90 FR 38212 [1]. Every mechanism that reasons about more than one aircraft at once operates before takeoff. Everything after takeoff reasons about one aircraft.

I do not read this as an oversight. It is where a layered design has to stop and the standard says so out loud. A service that computes a guarantee over a snapshot and hands it to a vehicle that will act on it for the next forty minutes has produced something with the structure of a database read taken outside a transaction: accurate at the instant of reading and unconstrained from then on.

The obvious objection is that detect and avoid is the tactical layer and it is on its way. I think the proposal is quietly clear about why that does not rescue the plan. Proposed § 108.825 states the design requirement as a capability "to avoid aircraft as required in accordance with § 108.195". The specification for detect and avoid is therefore the right-of-way rule, and that rule ranks a part 108 aircraft against traffic that broadcasts while saying nothing about two part 108 aircraft meeting each other. The FAA also records that it "decided not to update § 91.113 based on the BVLOS ARC's proposal related to 'detect-and-avoid' at this time" because doing so would reach into part 91 and legacy aviation [1]. Set that aside and a deeper problem remains. Detect and avoid is a sensor and a maneuver on one airframe. It improves what one aircraft can do about a conflict it can see and belongs to the capability column rather than the coordination one. Two aircraft avoiding competently can still pick maneuvers that conflict with each other.

Measuring what the layer is for

The case for a coordination service rests on an arithmetic claim: encounters grow faster than fleets do. Put N vehicles in a fixed volume, assume any pair is as likely to meet as any other, and the pair count grows like N2. I have used that argument before. This time I wanted to know whether it survives contact with a real sky. It does so only in the sense that reality turns out to be worse than the model allows for.

I took 30 snapshots of every transponder-equipped aircraft over the continental United States from the OpenSky Network [20]. That is 123,520 airborne state vectors in all, sampled with 12,000 discs of 100 km drawn at random across it. In each disc I counted the aircraft present and the pairs within 5, 10, 20 and 40 nautical miles laterally and 1,000 feet vertically. That vertical figure is the standard separation minimum. Then I asked what the same aircraft at the same altitudes would produce if they were thrown at random into the same disc. The null needs no model of its own. It is the same population with the geography taken out. The script and the numbers it produced are here along with the collector and the snapshots it pulled.

Real traffic packs 6.74 times more pairs within 5 NM than the scattered version of itself. The fitted slope of the pair count against the aircraft count is 2.58 with a standard error of 0.01, above the 2 that uniform mixing predicts. The excess shrinks as the threshold widens and reaches 1.65 at 40 NM. That is what we would expect if the clustering happens at the scale of terminal areas.

Each number comes with a qualification. The slope compares different places at one moment. A disc holding more aircraft is usually also a different kind of place, so some of the excess over 2 is that busier airspace is also more concentrated airspace rather than the same airspace scaled up. The ratio avoids that problem by comparing each disc against itself with the geography removed. It does depend on the size of the disc: the null scatters aircraft across whatever disc we choose, so a wider disc makes random placement look emptier. Repeating the whole measurement at 50 and 200 km gives ratios at 5 NM of 3.89 and 11.34 and slopes of 2.62 and 2.37. The ratio moves with the scale as it should and the slope holds above 2 at every disc size tried.

Measured proximate pairs against aircraft count in a 100 km disc, compared with the same aircraft scattered at random Two panels from measured air traffic. Panel A plots the mean number of aircraft pairs within 10 nautical miles and 1000 feet against the number of aircraft in a 100 km disc, on log axes, for the measured data and for the same aircraft scattered at random in the same disc. The measured curve has a slope of about 2.7 and sits above both the scattered curve and a dashed slope-2 guide. Panel B plots measured divided by scattered against the aircraft count for four lateral thresholds; the ratio is highest at 5 nautical miles and falls toward 1 at 40 nautical miles, and rises with the aircraft count. real traffic is more clustered than the model that worries about it 12,000 discs of 100 km drawn at random over the continental United States, across 30 OpenSky snapshots A. pairs within 10 NM and 1000 ft 10⁻² 10⁻¹ 10⁰ 10¹ 10² 2 5 10 20 50 aircraft in the disc, N slope 2 measured slope 2.58 scattered same aircraft, random places B. measured divided by scattered 1 2 5 10 20 2 5 10 20 50 aircraft in the disc, N 5 NM 10 NM 20 NM 40 NM Across every disc, traffic within 5 NM and 1000 ft is 6.7 times commoner than the same aircraft thrown at random into the same disc, and the fitted slope of 2.58 ± 0.01 is above the 2 that uniform mixing predicts.
Measured against a null with the geography removed. The wedge in panel A is the slope uniform mixing predicts, drawn as a reference so it does not assert an intercept. Panel B is the ratio of the two, and it rises with density, which is why the fitted slope comes out above 2.

This is manned aviation rather than the uncrewed traffic the note is about and I am using it on purpose. It is the only large population of vehicles that already shares airspace under a mature coordination regime, so it measures what a century of structure and procedure and separation standards actually achieves. The direction of the answer matters more than its size. All of that has not produced a sky that mixes uniformly. It has produced one that concentrates, and the model I was about to reason from understates the density a coordination service has to handle.

The second measurement concerns time. For every converging pair within 20 NM and 1,000 feet I computed the time to closest approach from the two velocity vectors. That comes to 26,137 pairs with a median of 158 seconds. The tail is the part that matters: 1,528 of those pairs, 5.8% of them, reach closest approach inside twelve seconds. The first percentile arrives at 1.9 seconds.

Measured time to closest approach for real converging aircraft pairs, on log-spaced bins A histogram on log-spaced time bins from one second to one hour, showing the time to closest approach for every converging aircraft pair measured within 20 nautical miles and 1000 feet. The distribution peaks in the three to five minute range with a median near 170 seconds, and a visible dark block of pairs falls below twelve seconds at the left. how long a real converging pair has time to closest approach for all 26,137 converging pairs seen within 20 NM and 1000 ft, across 30 snapshots 0 600 1,200 1,800 1 s 10 s 1 min 10 min 1 hour time to closest approach, log scale median 158 s 12 s 1,528 of these pairs, 5.8%, reach closest approach in under twelve seconds, which is the dark block on the left. A failure detector on one-second heartbeats has not finished deciding the link is gone. Closing speed at the median is 36 m/s and the first percentile of this distribution is 1.9 s. Bins are log-spaced, because linear bins hide the short encounters inside one column at the origin, and the short encounters are the ones the budget is about.
Time to closest approach for real converging pairs, on log-spaced bins. Linear bins put every short encounter in one column at the origin, and the short encounters are the ones a staleness budget is about.

Twelve seconds is not an arbitrary mark. A failure detector on one-second heartbeats needs three missed beats before it declares a link gone. An agreement about what to do costs at least two more message delays after that. Roughly a third of the budget goes on noticing, and by Fischer, Lynch and Paterson the remainder cannot be bounded at all in the asynchronous model. That leaves the 1,528 pairs in the left-hand block with no interval in which to spend it.

Which makes the timeout a safety parameter set by the geometry of the corridor it applies to and therefore something to publish per corridor. The harder consequence is that whatever the vehicle does inside those seconds has to work without agreement. Agreement is the one thing the budget cannot buy, and the theory here is unusually blunt about why.

What the theory forbids

Fischer, Lynch and Paterson showed that no deterministic protocol achieves consensus in an asynchronous system where a single process may fail [5]. A design that assumes the fleet will agree therefore has a timing model hidden inside it that someone should be made to state. Gilbert and Lynch's proof of Brewer's conjecture says a partitioned system chooses between staying available and staying consistent [6]. For an aircraft that reads as acting on a belief it can no longer refresh or stopping. Bernstein, Zilberstein and Immerman put the price of the missing channel in complexity terms. The finite-horizon problem for two agents who cannot share their observations is complete for nondeterministic exponential time [7], against the PSPACE-completeness Papadimitriou and Tsitsiklis had established for the same horizon under a single controller [8].

Halpern and Moses is the one I keep coming back to. They formalized what a group of processes can know about what the others know. Their finding on common knowledge, the regress where everyone knows that everyone knows, is that "formally speaking, in practical systems common knowledge cannot be attained" [9]. A protocol whose correctness rests on it therefore has no implementation whatever the hardware. That same paper introduces weaker variants that are attainable and those are where a working design has to aim.

Set that beside a coordination service. The service sends both aircraft the same deconfliction plan and both receive it. Each of them now holds a fact and neither can establish that the other holds it, nor that the other knows they hold it. The regress runs on without terminating. Two parties who each hold a fact are a weaker thing than two parties who can each rely on the other acting on it, and the difference between those is exactly what a coordination service would have to sell. Since no protocol closes that gap, whatever reliance two aircraft place in each other must have been built into them before either took off.

The contingency

The rule's own contingency provision shows this happening. Proposed § 108.815(b) [1] requires each airframe to execute a safe predetermined action at link timeout and each of those actions is safe on its own terms. What no one has checked is how they compose. The only moment they will ever run is the moment they all run together, since the link they lost was in many cases the same link.

I drew this scenario before I measured it. Five aircraft separated along one corridor, two returning to a pad at one end and three to a pad at the other, their contingency paths crossing six times in a tidy knot in the middle. It was a persuasive picture and it was wrong. I only know that because the same data that produced the two figures above can be asked the same question.

The question I put to the data was this. Take every configuration of 6 or more aircraft below 10,000 feet inside one of those discs, send each aircraft to the nearest of the 821 large and medium airports in the continental United States [21], and count the crossings between pairs that were separated at the moment the link went. Across 1,696 configurations holding 29,580 aircraft the count of crossings came to 0. That is a number I refused to publish until the crossing predicate had been checked against an independent orientation test on twenty thousand random segment pairs, a check that now runs every time the script does.

One configuration explains the zero. Aircraft below ten thousand feet are near airports, which is usually why they are low. The nearest field is the same field for many of them at once, so their return paths converge on a single runway instead of crossing each other on the way to different ones.

One measured configuration of aircraft and the fields they would return to, with the distribution of simultaneous arrivals Two panels. Panel A is a plan view of one measured configuration: 32 real aircraft below 10,000 feet inside a 100 km disc centered at 35.5 degrees north and 81.4 degrees west, each joined by a line to the nearest of 3 real airports, with 28 of them converging on KCLT in a single funnel. Panel B is a histogram of how many aircraft arrive at one field within two minutes of each other across all 1,696 measured configurations, peaking at two and three and reaching 21. a simultaneous return funnels, it does not scatter 1,696 configurations of 6+ aircraft below 10,000 ft, each sent to its nearest of 821 real fields A. one measured configuration KCLT KINT KJQF 32 aircraft, 3 fields, centered 35.5N 81.4W. 28 of them pick KCLT. B. arrivals at one field inside two minutes 0 150 300 450 0 5 10 15 20 aircraft converging on one field three or more in 67% of configurations, up to 21 Path crossings between separated pairs: 0. The returns converge instead, which a drawing of this does not show. The crossing test is checked against an independent implementation before the count is taken.
Real positions, real fields, real ground speeds. In this one, 28 of 32 aircraft pick KCLT. The picture I had drawn showed paths crossing. What the measurement shows is a funnel, which is worse and harder to see.

Converging is a worse property than crossing. In 67% of those configurations three or more aircraft reach the same field within two minutes of each other at their current ground speeds, and the worst has 21. Nothing in the pre-takeoff deconfliction set knows about it, because a contingency path is not an operational intent and was never in the set that was checked. Nothing after takeoff can negotiate it either. The link is what failed, so a set of individually compliant aircraft converges on one runway with no mechanism anywhere in the system that would notice.

Aviation has already run this experiment at full scale. Over Überlingen on 1 July 2002 a Tupolev Tu-154M and a Boeing 757 collided and 71 people died. The usual telling is that the crews received contradictory instructions. That is true and slightly misses the point. Both aircraft carried a working collision-avoidance system and both were talking to a working controller. Those are two coordination layers. Each was internally correct and each produced a resolution that would have been sufficient on its own. What did not exist was a rank between them. The Tupolev crew followed the controller and the Boeing crew followed the equipment. Among the immediate causes the German federal investigator recorded that the Tupolev crew "followed the ATC instruction to descend and continued to do so even after TCAS advised them to climb". Among the systemic ones it found that the rules for ACAS and TCAS issued by ICAO and by national authorities and by the manufacturer and by the operators "were not standardised, incomplete and partially contradictory". Its first recommendation asked ICAO to require pilots to follow a resolution advisory "regardless of whether contrary ATC instruction is given prior to, during, or after" it [10]. Adding a second correct coordination layer to a system that already has one makes it worse until the two are ranked against each other, and that ranking has to be settled when the system is designed.

The convention is in the weights

Everything so far applies to hand-written controllers. Learned ones fail the layering plan in a way that is harder to see and probably harder to fix.

Coordination between two agents means agreeing on an arbitrary choice: who yields, which side we pass on, which lever we pull. The choice is arbitrary in the strict sense that the alternatives are equally good. What makes one of them correct is only that the other party made it too. A policy trained by self-play has to settle on one and where it settles is decided by its own initialization and sampling noise. Lanctot and colleagues named the resulting co-adaptation joint-policy correlation [11]. Carroll and colleagues built a coordination task on Overcooked and found that agents from self-play and from population-based training each assume a partner much like itself and converge on protocols that fail against a human or a model of one [12]. Hu and colleagues framed it as the zero-shot coordination problem in which self-play "can produce agents that establish highly specialized conventions that do not carry over to novel partners". Their remedy works by exploiting the very symmetries such a convention settles at random [13].

I wanted a number for how large the effect is, so I ran the smallest version of it I could. The game is Hu's lever game. Two players independently choose one of ten levers. Matching levers pay that lever's value and mismatching levers pay nothing. One lever pays 1.0 and the other nine pay 0.9. Ten agreements exist and nine are worse by a hair. I trained eight policies by self-play with REINFORCE on a logit vector, one per seed, then scored every policy against every other in the same frame. The whole thing is one file of numpy with a forty-five line training loop. It takes about twenty seconds to run and what it produced is here.

Cross-play matrices for eight independently trained lever-game policies, under self-play and under other-play Two eight by eight heatmaps of expected return for every ordered pair of trained policies. Panel A, self-play: the diagonal is 0.90 and almost every off-diagonal cell is 0.00, with a mean of 0.06; the eight seeds chose six distinct levers and none of them chose the payoff-best lever. Panel B, other-play: six of the eight seeds converged on the best lever and score 1.0 against each other, while seeds 4 and 5 score near zero against everything, giving a diagonal of 0.79 and an off-diagonal mean of 0.54. what an independently trained partner is worth expected return for every ordered pair of 8 policies, each trained from its own seed and scored with no relabeling A. trained by self-play .90 .00 .00 .00 .00 .00 .00 .00 .00 .90 .90 .00 .00 .00 .00 .00 .00 .90 .90 .00 .00 .00 .00 .00 .00 .00 .00 .90 .00 .90 .00 .00 .00 .00 .00 .00 .90 .00 .00 .00 .00 .00 .00 .90 .00 .90 .00 .00 .00 .00 .00 .00 .00 .00 .90 .00 .00 .00 .00 .00 .00 .00 .00 .90 0 0 1 1 2 2 3 3 4 4 5 5 6 6 7 7 partner's seed own seed diagonal 0.90 off-diagonal 0.06 6 distinct levers chosen 0 of 8 took the best one B. trained by other-play 1.0 1.0 1.0 1.0 .00 .00 1.0 1.0 1.0 1.0 1.0 1.0 .00 .00 1.0 1.0 1.0 1.0 1.0 1.0 .00 .00 1.0 1.0 1.0 1.0 1.0 1.0 .00 .00 1.0 1.0 .00 .00 .00 .00 .15 .09 .00 .00 .00 .00 .00 .00 .09 .14 .00 .00 1.0 1.0 1.0 1.0 .00 .00 1.0 1.0 1.0 1.0 1.0 1.0 .00 .00 1.0 1.0 0 0 1 1 2 2 3 3 4 4 5 5 6 6 7 7 partner's seed own seed diagonal 0.79 off-diagonal 0.54 3 distinct levers chosen 6 of 8 took the best one A self-play score reads the diagonal. Deployment reads everything else. In panel A the diagonal is worth 14 times the rest.
Every trained policy against every other, evaluated exactly rather than sampled. A self-play score reads the diagonal of panel A and reports 0.90. An independently trained partner delivers 0.06, which is fourteen times less. The eight seeds picked six distinct levers between them, and not one picked the lever worth more.

A factor of fourteen, and every one of those policies is individually optimal and converged and would pass any test that pairs it with itself or with a copy of itself. What none of them carries is a way to tell us which lever it settled on, or to be told to settle on a different one. We cannot publish a standard that says take lever 1 and have these policies comply, because the convention exists only as a pattern spread across the weights and there is no field to write a different one into.

Which is why the known fix is a change to the training objective rather than an interface. Other-play trains each policy against a relabeling of itself drawn fresh each episode from the game's symmetry group. A policy that locked onto one of the nine interchangeable levers now meets a partner who locked onto a different one eight times in nine, and the only strategy the gradient can reward is the one every relabeling leaves alone. Panel B is the same experiment with that one change and cross-play goes from 0.06 to 0.54.

It also goes to 0.54 rather than to 1.0, and the reason turned out to be more interesting than a clean result would have been. Write a policy as mass q on the distinguished lever with the rest spread over the nine. Its expected return under a random relabeling is J(q)=q2+0.9(1q)29, a parabola whose interior minimum is at q=1/11. Gradient ascent leaves that point in whichever direction it starts from. Six of my eight seeds initialized above 1/11 and every one of them converged to q=0.9998. Two initialized below it and both converged to q=0.0024. Ten thousand steps at four times the step size moved neither group.

The other-play objective has an interior minimum, and every seed that started below it converged away from the good lever Two panels. Panel A plots the other-play objective J of q against q from 0 to 1, a parabola with an interior minimum at q equals one eleventh, marked with a dashed line, from which gradient ascent departs in whichever direction it starts. Panel B plots q after training against q at initialization for the eight seeds: the six that started above one eleventh all finished at 1, the two that started below it finished at 0, and the transition is a step at the barrier. other-play fixes the convention and adds a barrier q is the probability the policy puts on the one lever every relabeling leaves alone A. expected return under a random relabeling 0 0.25 0.5 0.75 1 0 0.5 1 q minimum at q = 1/11 gradient ascent leaves it in whichever direction it starts B. where each seed started, and ended 0 0.5 1 0.00 0.05 0.10 0.15 0.20 q at initialization 1/11 q after training 6 reached the good lever, filled 2 did not, hollow Every seed that began with more than 0.0909 of its mass on that lever finished at 1.000, and every seed that began below it finished at 0.002. Ten thousand steps at four times the step size moved neither group.
Whether other-play works on a given seed is settled before the first gradient step. The barrier at 1/11 is derived in closed form and then measured; the eight seeds fall on either side of it with no exceptions.

So the fix works and whether it works on any particular run is decided by the initialization. That is a toy and I want to be careful about how far it generalizes. The lever game has an exact symmetry group that I can write down and sample from. Airspace conventions are approximate, the group is unknown in advance, and a real policy has far more ways to settle an arbitrary choice than ten. Every one of those differences makes the problem harder than the one I ran.

If the convention is in the weights then which weights are flying is a safety-relevant fact about an aircraft, in the same category as its position. Look at what the proposed broadcast actually carries under § 108.195(a)(2)(ii) [1]: latitude, longitude, geometric altitude, velocity, an ICAO 24-bit address and three integrity figures. All three integrity figures describe the position source. None of them describes the thing making the decisions. Three policy versions rolling out across eight aircraft in a volume give 6,561 ways to assign versions to airframes and 45 distinct mixes even when the airframes are interchangeable. Certification evidence covers the three configurations where every airframe is on the same build.

What training instead of layering would mean

If the contract has to be in the policy then the co-player distribution and the partition profile have to be in the training environment, which makes the environment the thing we are designing. That is the premise of unsupervised environment design. Dennis and colleagues make the generator an agent rewarded for regret. They define regret as the gap between what an antagonist achieves in a configuration and what the learner achieves, which pushes the generator toward configurations at the frontier of what the learner can handle instead of ones that are simply unsolvable [14]. Jiang and colleagues later showed that prioritized replay of randomly generated levels belongs to the same family and earned the approach a robustness guarantee at Nash equilibria [15]. The configuration covers density, version mix, link profile and geometry. That is what the search ranges over.

The reported numbers then have to be joint or the exercise reverts to per-agent testing with extra steps: conflicts per flight hour at each density, the worst operator's detour alongside the mean, and the version mix that was airborne. All of it against a baseline of fixed structure and no learning, since a coordination layer that cannot beat assigned altitudes and a precedence rule has not earned its complexity. Melting Pot exists to score exactly this kind of generalization to unfamiliar co-players [16]. Mixed-autonomy traffic supplies the nearest fielded analog. Vinitsky and colleagues found that their learned controllers held a strategy that worked across penetration rates from 5% to 40% where a hand-tuned feedback controller "degrade[d] immediately upon penetration rate variation" [17]. A controller tuned for one composition of the population stops working when the composition moves. That is the same difficulty the version mix presents in another setting.

I do not want to oversell this. A regret objective needs a best response to compare against and in a population the best response is itself a joint quantity that nobody can compute. Every practical estimator substitutes something cheaper and the substitution is where the guarantee leaks. I do not know how to close that gap. What I am confident of is the direction: the co-player distribution is a design input and treating it as one costs nothing at training time and cannot be recovered afterwards.

The narrow window

The argument here runs against a coordination service being sufficient rather than against its existence, which the pair count settles on its own. What it argues is that the service cannot be the whole of the coordination. The layer's guarantee expires at takeoff and common knowledge is not something it can deliver. The conventions that would let two vehicles agree without it are set during training and unreachable after.

Some of what follows costs almost nothing. Contingency envelopes can go into the pre-takeoff deconfliction set alongside nominal intents, since the moment they are needed is the moment nothing can be negotiated. Precedence between two unmanned aircraft can be written into the rule, which is what the BFU asked ICAO to do after Überlingen. A precedence rule needs no round trip. A policy version can be a field in a broadcast. Wurman, D'Andrea and Mountz describe a Kiva installation for a large distribution center needing "500 or more vehicles". They move on a weighted grid. System-wide resource allocation is centralized in one job manager and each vehicle plans its own path across that grid [18].

The training is the expensive part. Retraining a fleet's policies against a co-player distribution that includes the other fleets is beyond what any one operator can do, and the current interface between operators has no way to express it even as a request. Kuchar and Yang separated the pairwise from the global in their survey of conflict detection twenty-six years ago [19]. Almost everything fielded since has been pairwise because pairwise is what one vehicle can reason about with what it can see.

The window in which the contract can still be trained in is the window before the fleets exist, and the fleets are being built now. Whether saying so changes any plan is not something I can affect. What I can do is set the argument out so it can be checked. The scripts and the data and every measured number are linked where they are used.

Further reading

References