From Optimal Transport to LLMs
After changing tracks three times, I found that what followed me was never the tools — it was one question.
Ten years ago in Paris my subject was full-waveform inversion: given the seismic waves recorded at the surface, work backwards to the velocity structure underground. Ten years later, what I do every day is design workflows for collaborating with LLMs. The two look entirely unrelated, and yet the longer I work the more certain I am that the road between them was never broken — the connector is simply not where you would expect it.
A question about what “alike” means
Inversion is, in the end, an optimization problem: guess a model of the subsurface, simulate the waveform it would produce, compare that against the recorded one, correct the model by the difference, and repeat. Every difficulty hides inside the word compare.
The most naive comparison subtracts point by point and sums the squares. Its failure mode is well known: if your guessed model makes the simulated waveform arrive more than a cycle late, a point-by-point comparison treats two waveforms that ought to line up as completely unrelated things. The optimization then slides into a beautiful local minimum and never climbs out. Your model is wrong, and your objective function tells you it is doing fine.
My doctoral work was a different way to compare: treat the waveform as a distribution of mass, and ask how much mass has to move, and how far, to turn this pile into that one. That is the distance optimal transport gives you — computable in practice only with entropic regularization. It is forgiving of shifts by construction: arriving a cycle late costs “some distance moved,” not “totally unrecognizable.” The landscape of the objective function flattens out, and the optimization can finally walk.
For those years I thought I was learning an inversion technique. Only later did I understand what I had really learned: a question. When you say two things are alike, how exactly is that defined?
The same question in different clothes
Around 2019, out of curiosity, I detoured into natural language processing for a while. What struck me most were word embeddings: put words in a space and let distance carry meaning. Looking back, that was an old friend in new clothes — I just did not recognize it at the time.
Then came LLMs and agents. The laziest way to evaluate an agent’s answer is to match it against a reference answer character by character — which is the sum of squared point differences all over again, exactly the same trap. Rephrase it, reorder it, and the score collapses even though the agent was right. Everyone who takes evaluation seriously eventually arrives at the same step: find a distance that is forgiving of irrelevant variation and sensitive to real error.
The other half of the inversion experience transfers even more directly: this is an ill-posed problem. Observations are never enough, the solution is not unique, so you must use priors and regularization to pull it back into the physically reasonable range. When I now write a task book for an agent, draw its boundaries and set its acceptance criteria, I am doing exactly that — not handing it the answer, but narrowing the solution space until only the reasonable part remains.
So what actually transfers
Ten years ago, while I was learning optimal transport and measure theory, I was already turning over whether that way of thinking could carry over into machine learning — those were the years OT was hot in machine learning, with Wasserstein GAN taking it as a loss function and redefining what “two distributions resemble each other” means, and with the earth mover’s distance doing the same job in finance. It turned out to follow me all the way to today’s large-model agent systems. Looking back now, three things really did come across:
- An alertness to the objective function — ask what is being optimized before you ask how well it is being optimized;
- A calm about ill-posedness — a non-unique solution is not a failure, it is the signal that a prior is missing;
- And the discipline of judging evidence inside noise, which I wrote about in Signal and Noise.
When you change tracks you will grieve for the specific skills you spent years sharpening. What actually needs packing is only the way you ask questions.