His findings have implications for any AI system in which multiple algorithms must compete or cooperate, from game-playing programs to digital markets to the training of large language models.
When an AI program beat top poker players in 2017, it did so in part by finding a kind of mathematical sweet spot—a stable strategy that no player could improve on by switching to something else. The achievement highlighted a broader theoretical question: why do the types of approaches used for this problem work, and how effectively can we scale them to larger and more complex game theory challenges?
Hertz Fellow Noah Golowich spent his MIT doctoral thesis working out the theory behind that question, and other topics related to how AI programs tackle game theory—a field that has implications far beyond poker. His thesis, “Theoretical Foundations for Learning in Games and Dynamic Environments,” was awarded the 2025 Hertz Thesis Prize from the Hertz Foundation. He was advised by Constantinos Daskalakis and Ankur Moitra.
“A lot of machine learning is sort of a black box,” said Golowich. “You try out various things and sometimes things work and you don’t really understand why. I was driven to push harder at understanding why things happen the way they do and see if I could get a deeper understanding of it.”
The Hertz Thesis Prize recognizes fellows with exemplary, transformative doctoral theses that have real-world applications. Golowich joins more than 60 fellows who have been previously recognized with the award. Winners are chosen by a group of Hertz Fellows who serve on the Hertz Thesis Prize Committee and are informed by reviews that come from a wide range of reviewers within the Hertz Community.
Golowich’s thesis spans two large domains of theoretical computer science. The first revolves around what happens when multiple AI agents, each running their own learning algorithm, interact with one another in a competition—like a poker game. Each agent acts in their own interest but the system must ultimately reach a stable outcome. Beyond poker, this kind of computational equilibrium could inform the design of programs to tackle financial markets, online auctions, negotiation and decentralized decision-making.
“The fact that you have your own funding is really nice. It does give you a greater sense that you can take risks and work on things that are maybe higher risk, higher reward, or just more outside the sphere of what you might be doing otherwise.”
Golowich and his colleagues showed that, if the agents use a learning algorithm called Optimistic Multiplicative Weights, they could improve the rate at which collective behavior settles on a stable outcome.
“What my work showed was that it’s possible to converge on an equilibrium much faster than had been thought,” explained Golowich.
The second area that his thesis tackled was the question of how a single agent, like a robot or a language model, can efficiently learn to act well in an environment it is unfamiliar with. This kind of learning, known as reinforcement learning, is the framework that underlies how many AI systems learn through trial and error. A robot placed in a new room, for instance, gradually explores in different directions, trying to create a map of what works and what doesn’t. The problem is that in any realistic setting, the number of possible situations to test is very large. AI systems must find ways to explore efficiently, learning as much as possible from each action they take, without burning through computational resources that grow out of control.
“We don’t often have great theoretical understanding of exactly what’s going on under the hood for these reinforcement learning algorithms,” said Golowich. “And I think understanding some of these ideas about exploration can be useful for improving reinforcement learning in language models.”
He also tackled cases where an AI agent had only partial information, which could model a doctor working from an incomplete patient history, or a self-driving car with imperfect sensors. He identified a way to more efficiently find a near-optimal strategy even under those constraints, and proved that his approach was essentially the best that any algorithm could do.
Golowich said that he didn’t set out to include such a wide variety of problems and applications in his thesis, but that it naturally grew over time.
“Over time in graduate school you hear about problems from other people, you get interested in different things, and it’s very easy to broaden out,” he said. “That’s sort of what happened; I started working on one particular thing and over time my interests brought me somewhere broader.”
He credited the Hertz Fellowship with giving him room to follow those interests without worrying whether they fit into a preset plan.
Golowich also found that the wider Hertz community shaped his thinking in unexpected ways. He collaborated with other Hertz Fellows including Ankur Moitra and Robert Kleinberg during the course of his graduate work.
Golowich completed a postdoctoral research fellowship at Microsoft Research in New York City and has joined the computer science faculty at the University of Texas at Austin as an assistant professor this year. His lab will tackle both theoretical and empirical questions around how generative AI systems such as language models work.
“It’s very easy to focus on the impressive engineering accomplishments behind all these large language models,” he said. “But there’s still a lot of room to uncover very deep and interesting results about how these systems work, and why.”
Honorable Mention Thesis Prize Awards
This year, the Hertz Thesis Prize Committee also awarded two honorable mentions, to Hertz Fellows Alex Cohen and Nina Zubrilina.
Alex Cohen
Alex Cohen, who also completed his graduate degree at MIT, was recognized for his thesis “Higher Dimensional Fractal Uncertainty.” Cohen studied the mathematical behavior of waves—a field of mathematics called harmonic analysis. His thesis looked at the limits of how functions and fractals interact in higher dimensions.
Nina Zubrilina
Nina Zubrilina, who studied at Princeton University, was recognized for her thesis “Convergence and Correlations of Coefficients of Cusp Forms.” The thesis investigates the statistical behavior of a class of mathematical objects called cusp forms, which are deeply connected to some of the most fundamental problems in number theory.
About the Hertz Foundation
The Hertz Foundation is the nation’s preeminent nonprofit organization committed to advancing American scientific and technological leadership. For more than 60 years, it has stood as an unwavering pillar of independent support through the renowned Hertz Fellowship, cultivating a multidisciplinary network of innovators whose work has positively impacted millions of lives. Learn more at hertzfoundation.org.