A tableau is such a familiar object that we forget to really look at it. Candidates down one column, weighted constraints across the others, and at the end one harmony per candidate — the sum of violations times their weights, with the sign flipped so that "higher" means "better". In a stochastic grammar, though, there is no single winner: every candidate gets a slice of probability.
Where does that slice come from? From a calculation you may have already met in other clothes:
$$P(x)=\frac{e^{H(x)}}{\sum_y e^{H(y)}}.$$
Exponentiate the harmonies, add them up, divide each by the sum. If that rings a bell, it's because it is exactly the softmax — the same function that, in machine learning, turns a handful of scores into a handful of probabilities. The harmonies play the part of the scores; the constraint weights, the part of the coefficients; and the last step, the exponential normalisation, is the softmax. A weighted tableau is a little log-linear model, and the "winner" is nothing more than the argmax.
You can watch the gears turn. In the panel below, move the weights and see the harmonies become probabilities on the spot:
And here is the knob the everyday tableau tends to be missing: temperature. Divide each harmony by a $T$ before the softmax and you set how decisive the grammar is. At low $T$, everything collapses onto the favourite — a categorical rule. At high $T$, the probabilities flatten toward the uniform — free variation. Categorical and variable stop being two worlds and become two settings of one knob.
It helps to anchor this in a real tableau. First just the harmonies, as in a Harmonic Grammar:
And the same with a probability column — the harmonies after they pass through the softmax:
To check the arithmetic, here are those same harmonies turned into probabilities, and what temperature does to them:
import math
# Harmonies of three candidates (higher = better). The softmax turns them into
# probabilities; dividing by a temperature T makes the grammar more or less decisive.
H = {"ta": -1.0, "tan": -2.0, "tra": -3.0}
def softmax(H, T=1.0):
e = {k: math.exp(h / T) for k, h in H.items()}
Z = sum(e.values())
return {k: e[k] / Z for k in H}
for T in (0.3, 1.0, 3.0):
p = softmax(H, T)
row = " ".join(f"[{k}]={p[k]:.2f}" for k in H)
print(f"T={T:>3} {row}")
None of this is a metaphor. The bridge between constraint grammars and statistical or connectionist models is not new (Prince & Smolensky, 1997; Pater, 2018), and the very maximum-entropy machine phonology uses is, underneath, a log-linear model (Goldwater & Johnson, 2003; Jäger, 2007). Still, it is a pretty thing to see a tableau and a softmax turn out to be the same drawing. The stochastic grammar was a classifier all along; the winner is just the argmax; and variation is the same softmax at a slightly warmer temperature.
- Prince, A., & Smolensky, P. (1997). Optimality: From Neural Networks to Universal Grammar. Science.
- Pater, J. (2018). Generative linguistics and neural networks at 60: foundation, friction, and fusion.
- Goldwater, S., & Johnson, M. (2003). Learning OT constraint rankings using a maximum entropy model. Preprint.
- Jäger, G. (2007). Maximum Entropy Models and Stochastic Optimality Theory. Preprint.
Barroso, A. M. (2024). A tableau is a softmax. alexandrebarroso.com. https://alexandrebarroso.com/notes/a-tableau-is-a-softmax.html
@misc{barroso2024atableauisasoftmax,
author = {Alexandre Menezes Barroso},
title = {A tableau is a softmax},
year = {2024},
howpublished = {alexandrebarroso.com},
url = {https://alexandrebarroso.com/notes/a-tableau-is-a-softmax.html},
note = {alexandrebarroso.com}
}