The Stylised Fact We Refused to Fake
There is an easy way to make a synthetic economy look credible, and we decided not to take it. Zipf's law for firm sizes is the famous number every synthetic economy is tempted to hit. This is why we built the harness, pre-registered the bands, and then refused.
There is an easy way to make a synthetic economy look credible, and we decided not to take it.
If you build a simulated economy and want people to believe its numbers, the obvious move is to show that it reproduces a famous empirical regularity. For firm sizes, the most famous one is Zipf's law: rank the firms in an economy by size, and the size falls off almost exactly as one over the rank. Robert Axtell measured it for U.S. firms in 2001 and found a rank-size exponent near 1.06, strikingly close to the canonical value of 1. It is one of the most robust facts in all of economics.
So the tempting headline writes itself: our agent-run economy reproduces Zipf's law. We built the harness to test exactly that, pre-registered the pass/fail bands, and then, after four honest negative results, did not write that headline. This post is about why not, and why the thing we found instead is worth more.
Hitting Zipf would have proved almost nothing
Here is the uncomfortable fact about Zipf's law for firm sizes: it is over-determined. At least half a dozen unrelated mechanisms all produce a rank-size exponent near 1: Gibrat's law of proportional growth with a small reflecting barrier, Kesten processes, random proportional splitting, and more. Xavier Gabaix's work on power laws in economics makes this point repeatedly: the exponent is robust precisely because so many different generative stories land on it.
That robustness is wonderful for the economist and a trap for the modeller. If a dozen different mechanisms all yield ζ ≈ 1, then a model hitting ζ ≈ 1 tells you very little about whether that model's mechanism is the right one. It mostly tells you the modeller kept tuning until the famous number appeared.
This is the whole reason pattern-oriented modelling exists (the programme of Volker Grimm and colleagues in ecology, now standard in agent-based work). The strong move is not to reproduce a well-known pattern. It is to build a mechanism that makes a risky prediction which could have failed: a number you derive and commit to before you look, not a curve you fit. A stylised fact is only evidence if your model could have gotten it wrong and didn't.
So the question we actually cared about was never "can we hit Zipf?" It was "can we make a firm-size exponent a prediction instead of a knob, and does the honest prediction survive?"
Our first revenue engine provably cannot do it
Before you can make Zipf a prediction, you need a mechanism that is capable of a power law at all. Ours was not, and finding that out was the first real result.
The arena's first revenue model works by competition for a fixed shared pool of demand. Each market segment has a finite amount of money to spend per tick, and a firm earns its share of that pool, weighted by how attractive it is relative to its rivals. It is a natural, conservative design. It is also, provably, incapable of Zipf.
The argument is short. If every firm's revenue is a share of the same finite pool, then no firm can be larger than the pool, and the pool caps the whole distribution. A share-of-a-bounded-quantity process cannot produce an unbounded power-law tail, no matter how heavy-tailed you make the inputs. This is not a calibration miss you can tune away; it is a structural ceiling. Our own economy was, by construction, a place where Zipf could not live.
That sounds like bad news and is actually the most useful thing we learned early. It means any Zipf pass we might have coaxed out of that substrate would have been an artifact. And it tells us exactly what has to change: not a parameter, but the demand surface itself: a firm's own productivity has to expand its own accessible demand, not merely shuffle a fixed slice between competitors.
Turning the exponent from a knob into a prediction
The successor design does that. It is a bounded "outside-good" own-demand model fed by a heavy-tailed (Pareto) distribution of firm productivity. The economics matters here for one reason: in this family, the firm-size exponent is not a free parameter. It is an algebraic consequence of two deeper quantities:
ζ = k / (σ − 1)
where k is the tail index of the productivity distribution and σ is a demand elasticity. This is standard heterogeneous-firm trade theory (Melitz; Chaney; di Giovanni and Levchenko): Pareto productivity plus this kind of demand gives Pareto sales, and the sales exponent is k/(σ−1). Zipf (ζ ≈ 1) happens if and only if k ≈ σ − 1. It is a derived prediction, not a target.
That is the lever. So we pinned it down as hard as we could:
- Pre-register the prediction. The predicted exponent, the acceptance band, and the seed protocol were content-hashed and committed before the realised distribution was observed. Amending them afterward is a retraction event, not a tweak.
- Source the inputs externally and blind. σ and k were pinned from published estimates, not from the arena's own output. Using the model's own numbers to pin its own primitives is circular. And the people sourcing the productivity tail index were told only the empirical object they were estimating. They were not told the target band, the pass/fail threshold, or that a red-or-green decision hung on the answer. Band-blind sourcing is the single most important integrity device in the whole exercise.
- Make the search unable to certify itself. The exploration loop can report the pre-registered bands but is structurally incapable of stamping "validated." Only a held-out run on a frozen model can do that. The code cannot cheat even if we wanted it to.
And the mechanism worked. In closed-form desk validation, the realised exponent went from tracking a small minority of the pre-registered grid (when it was still effectively a free quantity) to matching essentially all of it once the prediction was pinned (closed-form: 11 of 103 grid points before pinning, 90 of 90 after). The exponent had become a genuine forward prediction.
I want to be exact about what that is and isn't: it is closed-form desk work. The successor kernel has not yet been wired into the live arena and run. The prediction is real; the in-world reproduction is future work.
The honest negative: the external number won't cooperate
A prediction is only as good as the numbers you feed it. To turn ζ = k/(σ−1) into a pass, we needed an honest, externally-sourced value for the productivity tail index k at the exact scope of our claim: the self-purchasing "B2B core" of technology, finance, and professional-services firms.
At that scope, k is under-identified. The genuine-method evidence we could source is mono-source: it traces to a single study, from Doyne Farmer's group at Oxford, fitting Lévy alpha-stable distributions to firm-level productivity across European (Orbis) data. That study estimates productivity tail exponents in the heavy-tailed range α ∈ [1, 1.5]. Our band-blind sourcing put the B2B-core central near k ≈ 1.16 (technology ≈ 1.12, finance ≈ 1.21; professional services under-identified). And here is the thing that matters: every value in that estimated range sits below the exponent Zipf would require (σ − 1 ≈ 1.67, from the demand elasticity we pinned). Push k ≈ 1.16 through the map and it predicts ζ ≈ 0.70: a fatter-than-Zipf tail, more concentration among the largest firms than the canonical law, not the tidy ζ ≈ 1. The direction does not hinge on one number; the whole estimated range points the same way.
Our pre-registration had a clause for exactly this: if the honest external pin lands off-band, record a red. So we recorded a red. Across the whole thread we recorded four of them. No pre-registration was amended; nothing was retracted. The reds are the finding.
And here is the humility the result demands: economy-wide, real firms are Zipf. That is Axtell's robust fact. Our off-Zipf reading lives only at the narrow B2B-core scope, which is precisely where the external evidence is thinnest and rests on one paper. That is a reason to not go around claiming firms are fatter than Zipf. It is not a validated fact about the world. It is an honest statement about where our evidence runs out.
What we are (and are not) claiming
Plainly, so no one has to guess:
We are not claiming the arena reproduces Zipf's law. The capable kernel isn't in the live world yet; the positive evidence is closed-form.
We are not claiming real firms are fatter than Zipf. That reading is mono-source, inferred through a model, and contradicts the robust economy-wide result.
We are claiming three things, all of which we can show:
- A clean impossibility result: a fixed-pool "share of demand" revenue mechanism is structurally incapable of a power-law firm-size tail. If your economy uses one, a Zipf pass is an artifact.
- A pre-registered, band-blind, self-validation-incapable protocol that turns a firm-size exponent into a forward prediction rather than a fitted target, and in closed form makes it hold.
- An honest map of where the external evidence is too thin to certify, with the pre-registered reds to prove we stopped instead of tuning.
What would change our mind: independent, multi-source estimates of the productivity tail at this scope; or broadening the target to the whole-economy buyer universe, where an in-band estimate (k ≈ 1.33) actually lives, pursued as a new, separately pre-registered claim, never as a quiet rescue of this one.
The point of building a synthetic economy is not to reproduce famous numbers on demand. Anyone can do that; it proves nothing. The point is to build the kind of system where you can say, precisely, what you predicted, what you found, and what you refuse to claim. And show the timestamps.
So a question for anyone else building synthetic economies or agent-based models: which stylised fact do you hold yourselves to, and how do you keep from tuning your way into it?
References
Citations verified against OpenAlex/Crossref; citation counts fetched, not guessed.
- Axtell, R. L. (2001). "Zipf Distribution of U.S. Firm Sizes." Science 293(5536):1818–1820. doi:10.1126/science.1062081, cited by ≈1,331 (Crossref). Foundational.
- Gabaix, X. (2016). "Power Laws in Economics: An Introduction." Journal of Economic Perspectives 30(1):185–206. doi:10.1257/jep.30.1.185, cited by ≈425 (OpenAlex). Accessible, on-point for the over-determination argument. Deeper alternative: Gabaix (2009), "Power Laws in Economics and Finance," Annual Review of Economics 1:255–294, doi:10.1146/annurev.economics.050708.142940 (cited by ≈867). NB: the 1999 QJE "Zipf's Law for Cities" is about cities, not firms. Not to be cited here.
- Grimm, V., et al. (2005). "Pattern-Oriented Modeling of Agent-Based Complex Systems: Lessons from Ecology." Science 310(5750):987–991. doi:10.1126/science.1116681, cited by ≈1,720 (Crossref). Foundational. (Attribution is Grimm et al., not "Grimm & Railsback": Railsback co-authors the later textbook, not this paper.)
- Chaney, T. (2008). "Distorted Gravity: The Intensive and Extensive Margins of International Trade." American Economic Review 98(4):1707–1721. doi:10.1257/aer.98.4.1707, cited by ≈1,358 (Crossref). The cleanest single theory cite for Pareto productivity + CES ⇒ Pareto sales with exponent k/(σ−1). Empirical companion: di Giovanni, Levchenko & Rancière (2011), "Power laws in firm size and openness to trade," Journal of International Economics 85(1):42–52, doi:10.1016/j.jinteco.2011.05.003 (cited by ≈103).
- Yang, J., Heinrich, T., Winkler, J., Lafond, F., Koutroumpis, P., & Farmer, J. D. (2024). "Measuring productivity dispersion: a parametric approach using the Lévy alpha-stable distribution." Industrial and Corporate Change (2024), doi:10.1093/icc/dtae021; preprint arXiv:1910.05219 (2019). Fits Lévy alpha-stable distributions to firm productivity on Orbis Europe (France/Italy/Germany/Spain), estimating tail exponents α ∈ [1, 1.5]. This is the mono-source behind our B2B-core central of k ≈ 1.16.
⚠️ Correction logged during verification: an internal note attributed "France 2015, k≈1.16" to this paper; on checking the source, 1.16 in its Table 1 is the % of negative value-added observations for France in 2011, not a tail exponent. The actual productivity tail-exponent estimates are α ∈ [1, 1.5] (specific fits ≈ 1.14–1.33). The k ≈ 1.16 B2B-core central is the arena's band-blind derivation consistent with that range, not a number printed in the paper. Do not confuse this study with Garicano, Lelarge & Van Reenen (2013, AER), a different France-productivity paper (size distortions at the 50-employee threshold).
Independent R&D. Every organisation in this work is procedurally generated; the empirical firm-size and productivity figures are drawn from published external studies, cited in the references above.