Feed aggregator
When it comes to predicting people’s preferences, it pays to consider “the power of three”
In his 1927 paper, “A law of comparative judgment,” the American psychologist L. L. Thurstone proposed that when people select one option among multiple alternatives, they are picking the one that has the highest value to them, even though they cannot assign a particular number to that choice.
Thurstone was a pioneer of “psychometrics” — a field built upon the premise that mental processes, which we cannot see, can nevertheless be measured and quantified. His 1927 paper laid the groundwork for what are now called random utility models, which provide a mathematical framework for describing human preferences — information that can be relied upon, in turn, to make predictions about various hypothetical situations.
Random utility models (RUMs) are so named because they assess the “utility,” or benefit, that can be obtained from a given choice — such as deciding which book to read first among the stack of novels you brought back from the library. “These models are inherently random,” explains Gabriele Farina, an assistant professor in MIT’s Department of Electrical Engineering and Computer Science (EECS) and principal investigator at the Laboratory for Information and Decision Systems (LIDS), “because people are different. Everyone has their own preferences, and even those preferences can vary from time to time.” For example, someone who normally picks coffee over tea in the morning, and prefers tea after dinner, may, upon occasion, mix up that order entirely.
RUMs, to be sure, are frequently used within government and industry in situations of far greater consequence than the selection of a hot (or iced) beverage. The models routinely facilitate predictions regarding what people will elect to do in so-called counterfactual (“what-if”) scenarios such as: How will they get to work or school if a major thoroughfare is shut down for construction? What routes and modes of transport will they take? Or, if a city suddenly receives a windfall of $20 million, how should those funds be disbursed to maximize the common good?
Given that RUMs have been with us for almost 100 years, growing in sophistication over time, one might imagine that, at this stage, there would be little room for improvement. That, however, is not the case.
A paper presented in April at the International Conference on Learning Representations in Rio de Janeiro, Brazil, uncovered basic facts that show there is much more to be gleaned from these models than had traditionally been supposed. The paper was authored by Yeshwanth Cherapanamjeri, a former MIT postdoc now based at Nanyang Technological University in Singapore; Farina, also core faculty in MIT’s Operations Research Center (ORC); Constantinos Daskalakis, the Avanessians Professor of Computer Science at MIT and a member of MIT's Computer Science and Artificial Intelligence Laboratory; and Sobhan Mohammadpour, an MIT PhD student in computer science based at LIDS and EECS.
The group’s findings stem, in part, from a deficiency in the way RUMs are commonly estimated in practice, which has persisted since the days of Thurstone. The data upon which the models are estimated have been largely drawn from so-called pairwise-comparisons: In a choice between items A and B — whether it pertains to movies on Netflix, competing products on Amazon.com, news stories posted on Google, and so forth — which one would you pick? One reason this approach has been so pervasive, explains Daskalakis, is that “assigning a precise numerical score, such as 4.37, to the benefit you get from a single item is very hard. Whereas comparing two things, and deciding which one you like better, is cognitively much easier to do.” But therein lies the rub, he adds. “With this way of assessing people’s preferences, looking at just two things at a time, it is impossible to find correlations between the numerous choices.”
The standard way of applying RUMs assumes that the utilities derived from A and B are independent, but they may, in fact, be linked, and that would be important to know. If someone campaigning for elective office finds out that a potential voter favors gun control, for instance, there is a reasonable chance that same person also favors government-sponsored child care. Similarly, a fan of independent movies might also be partial to foreign films, but less enthusiastic about Hollywood action blockbusters. “If a digital platform has a blind eye to the existence of such correlations, it will not be able to estimate preferences very accurately,” Daskalakis notes. “And if Netflix regularly shows you an assortment of movies you don’t care about, you might sign off and cancel your subscription.”
The MIT team proved that it is impossible to get information about correlations from two-way comparisons alone. Correlations can be discerned, however, when large numbers of people rate three alternatives in their order of preference. The same information can also be obtained from a combination of best-of-three and best-of-two choices. In practice, Mohammadpour explains, “you would get a bunch of people to rank three items. You could then utilize the method we developed for merging those individual results into one big model that can provide us with the big picture.”
Their research effort, according to Farina, is focused on the computational side of RUMs, devising algorithms that can extract preference information and figuring out how much data is needed to do so or, equivalently, how many experiments need to be run. The good news, he says, is that efficient algorithms are, indeed, possible for this purpose. The requisite number of experiments does not grow exponentially with the number of items in the catalog or database that’s under review.
“This paper provides a crucial breakthrough,” comments Emma Frejinger, a computer scientist at the University of Montreal. “It mathematically proves why traditional data collection fails and demonstrates that simply asking users for their best-of-three [choices] unlocks the ability to accurately train these powerful models. This finding provides a highly practical roadmap for collecting better data to drive more accurate optimizations.”
“Building utility models is going to remain a very active area,” Daskalakis insists. “Just as RUMs have been critical to the internet economy since the late 1990s, they are, and will remain to be, critical to the alignment of AI models going forward.” More importantly, he adds, “RUMs play a central role in the commercial viability and usefulness of large language models [LLMs].” During the training period, people are typically asked to rank the various candidate outputs of these LLMs, from which the models can gain a better sense as to the kind of text — in terms of tone, style, and content — that is preferred.
Given that we’re constantly “besieged with a vast sea of options in so many different domains,” Daskalakis says, “you cannot possibly ask people to communicate all their personal preferences for all possible scenarios. So what you can do instead is build a model that predicts what people think about the different possible outcomes. And you have to keep improving and updating your model in an iterative process until, hopefully, you can make good predictions.”
‘News’ Site Keeps Hallucinating EFF Staffers
What do EFF staffers Sarah Chen, Javier Morales, Caitlin Chin, Emma Rodriguez, and Mikko Kopponen have in common?
For one thing, they don’t exist.
For another, all have been quoted as EFF experts in articles published in the past two months on a site called News-USA Today, which describes itself as “an independent news publisher focused on clear, accurate, and useful journalism.”
Uh…
(Please don’t confuse this site with USA Today, in which real EFF experts are accurately quoted on a regular basis.)
News-USA Today is hardly the only slagheap that’s hallucinating or fabricating EFF personnel and quotes; as we wrote last September, media companies large and small are using AI to generate news content because it’s cheaper than paying for journalists’ salaries, but that savings can come at the cost of the outlets’ reputations— assuming they care about reputation at all.
But this many fake EFF sources in two months? That’s making a play for the championship title of bogus news content.
News-USA Today’s site proclaims, “Our goal is simple: give readers the facts and the context they need to make informed decisions.” It then defines its mission:
- “Deliver timely, factual reporting grounded in verifiable sources and public documents.”
- “Make complex topics understandable without losing nuance or accuracy.”
- “Serve the public interest by surfacing stories that affect lives, institutions, and communities.”
- “Maintain a clear separation between news, analysis, opinion, and sponsored content.”
Attempts to reach contacts listed on the site went unanswered. In fact, after we reached out to them, they published a story on June 9 with quotes from Electronic Frontier Foundation Executive Director Jared Cohen — who also doesn’t exist.
As we noted last year, EFF is all about having our words spread far and wide. Per our copyright policy, any and all original material on the EFF website may be freely distributed at will under the Creative Commons Attribution 4.0 International License (CC-BY), unless otherwise noted.
However, we don't want disreputable sites making up words (or false identities!) for us, whether or not they’re using AI. False quotations that misstate our positions damage the trust that the public and reputable media outlets have in us.
The best thing a news consumer can do is invest a little time and energy to learn how to discern the real from the fake. It’s unfortunate that it's the public’s burden to put in this much effort, but while we're adjusting to new tools and a new normal, a little effort now can go a long way.
As we’ve noted before in the context of election misinformation, the nonprofit journalism organization ProPublica has published a handy guide about how to tell if what you’re reading is accurate or “fake news,” as has FactCheck.org.
A shot of carbon dioxide rewires how cement sets
One September day, it started to snow inside MIT’s Pierce Laboratory.
Researchers depressurized a tank of liquid carbon dioxide (CO2), instantly freezing it and releasing solid flakes. These were blended into cement paste and pressed into discs roughly the size of a dime, each sealed with a thin layer of vegetable oil to keep water in and air out. The team trained lasers on each, observing for the first time the transient chemical reaction that might explain why CO2-injected cement paste gains its strength faster.
Injecting CO2 into cement products like concrete is one way to store it and keep it out of the atmosphere. The process has attracted commercial interest, with a growing number of companies offering CO2-injected concrete mixes. But until now, the underlying cement chemistry hadn't been directly visualized.
A new open-access paper in the Journal of the American Ceramic Society — led by Associate Professor Admir Masic and first-authored by graduate student Marcin Hajduczek, both of the MIT Concrete Sustainability Hub and MIT Department of Civil and Environmental Engineering — describes the chemical sequence that unfolds after CO2 meets fresh cement paste. Co-authors include MIT colleagues Santiago El Awad and Franz-Josef Ulm, alongside researchers from IIT Jodhpur and CarbonCure Technologies.
Previous studies had pieced together a story about CO2 injection’s chemical impacts from theory and indirect evidence; the key reactions simply moved too fast, and vanished too completely, for conventional techniques to catch them in the act. Raman confocal microscopy could — and it works on a simple principle: Illuminate a molecule with a laser, and the scattered light will reveal its identity. The light interacts with each material’s unique chemical bonds, shifting in energy to produce a distinct spectral “fingerprint.” Even the most fleeting and amorphous phases leave a readable trace.
“We’ve used Raman spectroscopy to better understand some of the most interesting materials in history, from the Dead Sea Scrolls to Ancient Roman concrete,” says Masic. “Cement paste may seem less glamorous in comparison, but pointing a laser at CO2-injected cement paste as it hardens allows us to visualize things that haven’t been seen before.”
What they saw, unfolding during 24 hours of continuous scanning, was a three-act chemical drama.
Act One: Capturing calcium
The moment that CO2 is added to the fresh cement paste, it goes to work. It dissolves into the pore solution and reacts with calcium released by the dissolving clinker, precipitating as various forms of calcium carbonate. Clinker is produced by heating limestone and aluminosilicate materials in a kiln, forming the primary ingredient ground into a fine powder to make cement. This happens within the first hour, temporarily slowing the normal hydration reaction, which requires calcium to proceed.
In contrast, when CO2 is not present, the calcium released by the dissolving clinker remains available locally, supporting the gradual formation of the material’s binding phases as it sets.
Left without calcium, the silicates released by the clinker dissolve into the pore solution and precipitate far from their source, linking together into chains that form an interconnected silica gel network throughout the paste. This amorphous, fleeting gel sets the stage for what follows.
Act Two: The ghostly gel
Once the injected CO2 is fully mineralized — around four to five hours after mixing — normal hydration resumes. Calcium hydroxide begins to precipitate into the pore space, and when it does, it encounters the silica gel network waiting for it.
The reaction between the two phases begins immediately, producing calcium silicate hydrate (C-S-H), the compound that gives cement its binding ability. What makes this form of C-S-H distinct is where and how it forms: not clustered around clinker particles as in conventional hydration, but distributed throughout the entire matrix, wherever the silica gel had spread.
The CO2 had temporarily suppressed the paste’s alkalinity, and that lower pH was the only thing keeping the silica-gel intact. As hydration reasserts itself and produces standard hydration products, namely C-S-H and calcium hydroxide, the latter drives pH back up to typical levels in a self-reinforcing loop; the silica-gel reacts with calcium hydroxide through a so-called pozzolanic reaction. Within eight hours, the silica gel is almost entirely gone — the previously well-distributed gel network turns rapidly into additional C-S-H during this critical early window.
“At first, the fleeting nature of the silica gel looked like a fluke in the Raman data. But it quickly became clear that its sudden disappearance was a consistent, undeniable feature of every CO2-injected sample,” says Hajduczek.
Act Three: A rewired matrix
With the silica gel consumed, the paste settles into conventional hydration, but what it leaves behind is measurably different. Because the new binder was distributed more evenly throughout the cement matrix, the resulting microstructure is stronger and more uniform at an early age. In the study, paste mixed with CO2 at 1 percent by cement weight achieved, on average, 13 percent higher compressive strength at 24 hours, compared to reference mixes.
“We’ve been injecting CO2 into cement products for years without fully understanding what it was doing inside. Now that we can see it and understand the underlying mechanism that leads to improved performance, we can start to control it. And there’s a lot of room to push,” says Masic.
The findings also refine a leading explanation for CO2-injected cement paste’s higher early age strength: the calcium carbonate crystals, previously suspected to seed C-S-H growth, turn out to be passive bystanders embedded in the silica gel template rather than reacting to form C-S-H.
Where the chemistry goes next
Knowing the mechanism gives researchers a more specific set of questions to pursue. The silica gel template explains the distribution of the new C-S-H, but directly measuring its mechanical properties remains a next step.
On the practical side, dosage matters: Flood the system with too much CO2 and calcium gets locked into carbonate before the gel can form and react. If the paste used here forms abundant C-S-H, it could theoretically offset up to 40 percent of the carbon emissions from cement production, excluding emissions associated with the fossil fuels used in the process. In practice, however, the achievable offset is likely to be only a fraction of that value, although still potentially significant.
But even with these open questions, the ghostly gel has been caught. And now that researchers know what to look for, the chemistry that unfolds in those first eight hours is no longer invisible.
