Before I've measured anything, what should I believe about where a vowel will land? Picture a range of first-formant values and nothing in hand. The honest answer is a flat line: every height equally likely, because I don't yet know a thing about them. We have a name for that ignorance — entropy — and of all the distributions, the flat line is the one with the most of it. It commits to nothing at all.
Now suppose I measure a single thing. Not every value, but an average — say, the expected value of some constraint over the candidates. Suddenly I know one fact. And the question gets interesting: what is the smallest deformation I can impose on the flat line to honour that fact, without smuggling in anything I didn't measure?
The answer is curiously tidy:
$$p(F_1)=\frac{e^{-\lambda\,v(F_1)}}{Z}.$$
The exponential is the smoothest shape that still honours the average I measured; $\lambda$ is exactly how much bend that fact demands, no more and no less; and the $Z$ underneath is just the bookkeeping that keeps everything summing to one. Drag $\lambda$ in the panel below and watch the flat line fill out — and, beside it, the entropy leaving its maximum the moment you add the first commitment.
Measure a second fact and we pick up a second $\lambda$; the shape stays the least-committed one consistent with both. Each constraint pushes the distribution only as far as it must, and the rest stays as undecided as before.
Over discrete candidates this is, point for point, a maximum-entropy grammar: weighted constraints become probabilities. That, roughly, is how the idea entered phonology (goldwater2003learning; jager2007maximumEntropyStochasticOT). With the weight at zero the grammar "knows nothing" and splits everything evenly:
compilation failed: spawnSync tectonic ETIMEDOUT
Turn the weight on and it commits to whoever violates least — the probability concentrates and the entropy drops:
compilation failed: spawnSync tectonic ETIMEDOUT
You can check the arithmetic instead of taking my word for it. For a handful of values of $\lambda$, notice how the entropy starts at its maximum (the uniform distribution) and only falls as we turn the constraint on:
import math
# Three candidates with violation counts. A single weight (lambda) bends a flat
# distribution into a committed one; the entropy falls from its maximum.
v = {"a": 0, "e": 1, "i": 2}
def maxent(lam):
w = {k: math.exp(-lam * vi) for k, vi in v.items()}
Z = sum(w.values())
return {k: w[k] / Z for k in v}
def entropy(p): # in bits
return -sum(pi * math.log2(pi) for pi in p.values() if pi > 0)
for lam in (0.0, 0.5, 1.0, 2.0):
p = maxent(lam)
row = " ".join(f"[{k}]={p[k]:.2f}" for k in v)
print(f"lambda={lam:>3} {row} H={entropy(p):.3f} bits")
It's worth insisting on what this isn't. It isn't a claim that the world is exponential, or that vowels "want" to follow a pretty curve. It's more a discipline of humility: assume the least, given what you actually measured. Each $\lambda$ is the price of a fact; the flat line is what's left when you've measured nothing. Maximum entropy, minimum assumption.
Barroso, A. M. (2024). Maximum entropy, minimum assumption. alexandrebarroso.com. https://alexandrebarroso.com/notes/maximum-entropy-minimum-assumption.html
@misc{barroso2024maximumentropyminimumassumption,
author = {Alexandre Menezes Barroso},
title = {Maximum entropy, minimum assumption},
year = {2024},
howpublished = {alexandrebarroso.com},
url = {https://alexandrebarroso.com/notes/maximum-entropy-minimum-assumption.html},
note = {alexandrebarroso.com}
}