ARENA · RESEARCH NOTE 2026 · 07 9 MIN READ

The Stylised Fact We Refused to Fake

There is an easy way to make a synthetic economy look credible, and we decided not to take it. Zipf's law for firm sizes is the famous number every synthetic economy is tempted to hit. This is why we built the harness, pre-registered the bands, and then refused.

There is an easy way to make a synthetic economy look credible, and we decided not to take it.

If you build a simulated economy and want people to believe its numbers, the obvious move is to show that it reproduces a famous empirical regularity. For firm sizes, the most famous one is Zipf's law: rank the firms in an economy by size, and the size falls off almost exactly as one over the rank. Robert Axtell measured it for U.S. firms in 2001 and found a rank-size exponent near 1.06, strikingly close to the canonical value of 1. It is one of the most robust facts in all of economics.

So the tempting headline writes itself: our agent-run economy reproduces Zipf's law. We built the harness to test exactly that, pre-registered the pass/fail bands, and then, after four honest negative results, did not write that headline. This post is about why not, and why the thing we found instead is worth more.

Hitting Zipf would have proved almost nothing

Here is the uncomfortable fact about Zipf's law for firm sizes: it is over-determined. At least half a dozen unrelated mechanisms all produce a rank-size exponent near 1: Gibrat's law of proportional growth with a small reflecting barrier, Kesten processes, random proportional splitting, and more. Xavier Gabaix's work on power laws in economics makes this point repeatedly: the exponent is robust precisely because so many different generative stories land on it.

That robustness is wonderful for the economist and a trap for the modeller. If a dozen different mechanisms all yield ζ ≈ 1, then a model hitting ζ ≈ 1 tells you very little about whether that model's mechanism is the right one. It mostly tells you the modeller kept tuning until the famous number appeared.

This is the whole reason pattern-oriented modelling exists (the programme of Volker Grimm and colleagues in ecology, now standard in agent-based work). The strong move is not to reproduce a well-known pattern. It is to build a mechanism that makes a risky prediction which could have failed: a number you derive and commit to before you look, not a curve you fit. A stylised fact is only evidence if your model could have gotten it wrong and didn't.

So the question we actually cared about was never "can we hit Zipf?" It was "can we make a firm-size exponent a prediction instead of a knob, and does the honest prediction survive?"

Four unrelated generative mechanisms (Gibrat growth with a reflecting barrier, a Kesten process, random proportional splitting, and Melitz selection), each shown as a labelled box, with arrows converging into a single node: Zipf's law with a rank-size exponent of about 1, drawn as a straight downward line on log-size versus log-rank axes. A banner reads: because many unrelated mechanisms converge to the same number, reproducing the exponent is weak evidence; the informative test is a risky, pre-registered prediction.
Figure 1 (illustrative). Zipf's law for firm sizes is over-determined: many unrelated mechanisms all converge on a rank-size exponent near 1, so reproducing it is weak evidence for any one mechanism. The informative test is a risky, pre-registered prediction.

Our first revenue engine provably cannot do it

Before you can make Zipf a prediction, you need a mechanism that is capable of a power law at all. Ours was not, and finding that out was the first real result.

The arena's first revenue model works by competition for a fixed shared pool of demand. Each market segment has a finite amount of money to spend per tick, and a firm earns its share of that pool, weighted by how attractive it is relative to its rivals. It is a natural, conservative design. It is also, provably, incapable of Zipf.

The argument is short. If every firm's revenue is a share of the same finite pool, then no firm can be larger than the pool, and the pool caps the whole distribution. A share-of-a-bounded-quantity process cannot produce an unbounded power-law tail, no matter how heavy-tailed you make the inputs. This is not a calibration miss you can tune away; it is a structural ceiling. Our own economy was, by construction, a place where Zipf could not live.

That sounds like bad news and is actually the most useful thing we learned early. It means any Zipf pass we might have coaxed out of that substrate would have been an artifact. And it tells us exactly what has to change: not a parameter, but the demand surface itself: a firm's own productivity has to expand its own accessible demand, not merely shuffle a fixed slice between competitors.

A two-panel contrast. Left panel, 'Fixed Shared Pool (tail-incapable)': firms share one bounded demand pool, and their sizes pile up against a demand cap, giving a squat, even distribution with no heavy tail (labelled a theorem, not a calibration miss). Right panel, 'Own-Demand (tail-capable)': each firm's own productivity expands its own accessible demand, growth is multiplicative, and the size distribution is heavy-tailed with a few very large firms. A banner reads: moving from share of a fixed pool to productivity expands own demand is what makes a power-law firm-size tail structurally possible at all.
Figure 2 (illustrative). The impossibility result. Left: sharing one finite demand pool caps every firm. No power-law tail can form, a structural theorem rather than a tuning failure. Right: when a firm's productivity expands its own accessible demand, a heavy-tailed size distribution becomes possible.

Turning the exponent from a knob into a prediction

The successor design does that. It is a bounded "outside-good" own-demand model fed by a heavy-tailed (Pareto) distribution of firm productivity. The economics matters here for one reason: in this family, the firm-size exponent is not a free parameter. It is an algebraic consequence of two deeper quantities:

ζ = k / (σ − 1)

where k is the tail index of the productivity distribution and σ is a demand elasticity. This is standard heterogeneous-firm trade theory (Melitz; Chaney; di Giovanni and Levchenko): Pareto productivity plus this kind of demand gives Pareto sales, and the sales exponent is k/(σ−1). Zipf (ζ ≈ 1) happens if and only if k ≈ σ − 1. It is a derived prediction, not a target.

That is the lever. So we pinned it down as hard as we could:

And the mechanism worked. In closed-form desk validation, the realised exponent went from tracking a small minority of the pre-registered grid (when it was still effectively a free quantity) to matching essentially all of it once the prediction was pinned (closed-form: 11 of 103 grid points before pinning, 90 of 90 after). The exponent had become a genuine forward prediction.

I want to be exact about what that is and isn't: it is closed-form desk work. The successor kernel has not yet been wired into the live arena and run. The prediction is real; the in-world reproduction is future work.

The honest negative: the external number won't cooperate

A prediction is only as good as the numbers you feed it. To turn ζ = k/(σ−1) into a pass, we needed an honest, externally-sourced value for the productivity tail index k at the exact scope of our claim: the self-purchasing "B2B core" of technology, finance, and professional-services firms.

At that scope, k is under-identified. The genuine-method evidence we could source is mono-source: it traces to a single study, from Doyne Farmer's group at Oxford, fitting Lévy alpha-stable distributions to firm-level productivity across European (Orbis) data. That study estimates productivity tail exponents in the heavy-tailed range α ∈ [1, 1.5]. Our band-blind sourcing put the B2B-core central near k ≈ 1.16 (technology ≈ 1.12, finance ≈ 1.21; professional services under-identified). And here is the thing that matters: every value in that estimated range sits below the exponent Zipf would require (σ − 1 ≈ 1.67, from the demand elasticity we pinned). Push k ≈ 1.16 through the map and it predicts ζ ≈ 0.70: a fatter-than-Zipf tail, more concentration among the largest firms than the canonical law, not the tidy ζ ≈ 1. The direction does not hinge on one number; the whole estimated range points the same way.

Our pre-registration had a clause for exactly this: if the honest external pin lands off-band, record a red. So we recorded a red. Across the whole thread we recorded four of them. No pre-registration was amended; nothing was retracted. The reds are the finding.

And here is the humility the result demands: economy-wide, real firms are Zipf. That is Axtell's robust fact. Our off-Zipf reading lives only at the narrow B2B-core scope, which is precisely where the external evidence is thinnest and rests on one paper. That is a reason to not go around claiming firms are fatter than Zipf. It is not a validated fact about the world. It is an honest statement about where our evidence runs out.

What we are (and are not) claiming

Plainly, so no one has to guess:

We are not claiming the arena reproduces Zipf's law. The capable kernel isn't in the live world yet; the positive evidence is closed-form.

We are not claiming real firms are fatter than Zipf. That reading is mono-source, inferred through a model, and contradicts the robust economy-wide result.

We are claiming three things, all of which we can show:

  1. A clean impossibility result: a fixed-pool "share of demand" revenue mechanism is structurally incapable of a power-law firm-size tail. If your economy uses one, a Zipf pass is an artifact.
  2. A pre-registered, band-blind, self-validation-incapable protocol that turns a firm-size exponent into a forward prediction rather than a fitted target, and in closed form makes it hold.
  3. An honest map of where the external evidence is too thin to certify, with the pre-registered reds to prove we stopped instead of tuning.

What would change our mind: independent, multi-source estimates of the productivity tail at this scope; or broadening the target to the whole-economy buyer universe, where an in-band estimate (k ≈ 1.33) actually lives, pursued as a new, separately pre-registered claim, never as a quiet rescue of this one.

The point of building a synthetic economy is not to reproduce famous numbers on demand. Anyone can do that; it proves nothing. The point is to build the kind of system where you can say, precisely, what you predicted, what you found, and what you refuse to claim. And show the timestamps.

So a question for anyone else building synthetic economies or agent-based models: which stylised fact do you hold yourselves to, and how do you keep from tuning your way into it?

References

Citations verified against OpenAlex/Crossref; citation counts fetched, not guessed.

  1. Axtell, R. L. (2001). "Zipf Distribution of U.S. Firm Sizes." Science 293(5536):1818–1820. doi:10.1126/science.1062081, cited by ≈1,331 (Crossref). Foundational.
  2. Gabaix, X. (2016). "Power Laws in Economics: An Introduction." Journal of Economic Perspectives 30(1):185–206. doi:10.1257/jep.30.1.185, cited by ≈425 (OpenAlex). Accessible, on-point for the over-determination argument. Deeper alternative: Gabaix (2009), "Power Laws in Economics and Finance," Annual Review of Economics 1:255–294, doi:10.1146/annurev.economics.050708.142940 (cited by ≈867). NB: the 1999 QJE "Zipf's Law for Cities" is about cities, not firms. Not to be cited here.
  3. Grimm, V., et al. (2005). "Pattern-Oriented Modeling of Agent-Based Complex Systems: Lessons from Ecology." Science 310(5750):987–991. doi:10.1126/science.1116681, cited by ≈1,720 (Crossref). Foundational. (Attribution is Grimm et al., not "Grimm & Railsback": Railsback co-authors the later textbook, not this paper.)
  4. Chaney, T. (2008). "Distorted Gravity: The Intensive and Extensive Margins of International Trade." American Economic Review 98(4):1707–1721. doi:10.1257/aer.98.4.1707, cited by ≈1,358 (Crossref). The cleanest single theory cite for Pareto productivity + CES ⇒ Pareto sales with exponent k/(σ−1). Empirical companion: di Giovanni, Levchenko & Rancière (2011), "Power laws in firm size and openness to trade," Journal of International Economics 85(1):42–52, doi:10.1016/j.jinteco.2011.05.003 (cited by ≈103).
  5. Yang, J., Heinrich, T., Winkler, J., Lafond, F., Koutroumpis, P., & Farmer, J. D. (2024). "Measuring productivity dispersion: a parametric approach using the Lévy alpha-stable distribution." Industrial and Corporate Change (2024), doi:10.1093/icc/dtae021; preprint arXiv:1910.05219 (2019). Fits Lévy alpha-stable distributions to firm productivity on Orbis Europe (France/Italy/Germany/Spain), estimating tail exponents α ∈ [1, 1.5]. This is the mono-source behind our B2B-core central of k ≈ 1.16.

    ⚠️ Correction logged during verification: an internal note attributed "France 2015, k≈1.16" to this paper; on checking the source, 1.16 in its Table 1 is the % of negative value-added observations for France in 2011, not a tail exponent. The actual productivity tail-exponent estimates are α ∈ [1, 1.5] (specific fits ≈ 1.14–1.33). The k ≈ 1.16 B2B-core central is the arena's band-blind derivation consistent with that range, not a number printed in the paper. Do not confuse this study with Garicano, Lelarge & Van Reenen (2013, AER), a different France-productivity paper (size distortions at the 50-employee threshold).


Independent R&D. Every organisation in this work is procedurally generated; the empirical firm-size and productivity figures are drawn from published external studies, cited in the references above.