Feed aggregator
MIT engineers develop a magnetic transistor for more energy-efficient electronics
Transistors, the building blocks of modern electronics, are typically made of silicon. Because it’s a semiconductor, this material can control the flow of electricity in a circuit. But silicon has fundamental physical limits that restrict how compact and energy-efficient a transistor can be.
MIT researchers have now replaced silicon with a magnetic semiconductor, creating a magnetic transistor that could enable smaller, faster, and more energy-efficient circuits. The material’s magnetism strongly influences its electronic behavior, leading to more efficient control of the flow of electricity.
The team used a novel magnetic material and an optimization process that reduces the material’s defects, which boosts the transistor’s performance.
The material’s unique magnetic properties also allow for transistors with built-in memory, which would simplify circuit design and unlock new applications for high-performance electronics.
“People have known about magnets for thousands of years, but there are very limited ways to incorporate magnetism into electronics. We have shown a new way to efficiently utilize magnetism that opens up a lot of possibilities for future applications and research,” says Chung-Tao Chou, an MIT graduate student in the departments of Electrical Engineering and Computer Science (EECS) and Physics, and co-lead author of a paper on this advance.
Chou is joined on the paper by co-lead author Eugene Park, a graduate student in the Department of Materials Science and Engineering (DMSE); Julian Klein, a DMSE research scientist; Josep Ingla-Aynes, a postdoc in the MIT Plasma Science and Fusion Center; Jagadeesh S. Moodera, a senior research scientist in the Department of Physics; and senior authors Frances Ross, TDK Professor in DMSE; and Luqiao Liu, an associate professor in EECS, and a member of the Research Laboratory of Electronics; as well as others at the University of Chemistry and Technology in Prague. The paper appears today in Physical Review Letters.
Overcoming the limits
In an electronic device, silicon semiconductor transistors act like tiny light switches that turn a circuit on and off, or amplify weak signals in a communication system. They do this using a small input voltage.
But a fundamental physical limit of silicon semiconductors prevents a transistor from operating below a certain voltage, which hinders its energy efficiency.
To make more efficient electronics, researchers have spent decades working toward magnetic transistors that utilize electron spin to control the flow of electricity. Electron spin is a fundamental property that enables electrons to behave like tiny magnets.
So far, scientists have mostly been limited to using certain magnetic materials. These lack the favorable electronic properties of semiconductors, constraining device performance.
“In this work, we combine magnetism and semiconductor physics to realize useful spintronic devices,” Liu says.
The researchers replace the silicon in the surface layer of a transistor with chromium sulfur bromide, a two-dimensional material that acts as a magnetic semiconductor.
Due to the material’s structure, researchers can switch between two magnetic states very cleanly. This makes it ideal for use in a transistor that smoothly switches between “on” and “off.”
“One of the biggest challenges we faced was finding the right material. We tried many other materials that didn’t work,” Chou says.
They discovered that changing these magnetic states modifies the material’s electronic properties, enabling low-energy operation. And unlike many other 2D materials, chromium sulfur bromide remains stable in air.
To make a transistor, the researchers pattern electrodes onto a silicon substrate, then carefully align and transfer the 2D material on top. They use tape to pick up a tiny piece of material, only a few tens of nanometers thick, and place it onto the substrate.
“A lot of researchers will use solvents or glue to do the transfer, but transistors require a very clean surface. We eliminate all those risks by simplifying this step,” Chou says.
Leveraging magnetism
This lack of contamination enables their device to outperform existing magnetic transistors. Most others can only create a weak magnetic effect, changing the flow of current by a few percent or less. Their new transistor can switch or amplify the electric current by a factor of 10.
They use an external magnetic field to change the magnetic state of the material, switching the transistor using significantly less energy than would usually be required.
The material also allows them to control the magnetic states with electric current. This is important because engineers cannot apply magnetic fields to individual transistors in an electronic device. They need to control each one electrically.
The material’s magnetic properties could also enable transistors with built-in memory, simplifying the design of logic or memory circuits.
A typical memory device has a magnetic cell to store information and a transistor to read it out. Their method can combine both into one magnetic transistor.
“Now, not only are transistors turning on and off, they are also remembering information. And because we can switch the transistor with greater magnitude, the signal is much stronger so we can read out the information faster, and in a much more reliable way,” Liu says.
Building on this demonstration, the researchers plan to further study the use of electrical current to control the device. They are also working to make their method scalable so they can fabricate arrays of transistors.
This research was supported, in part, by the Semiconductor Research Corporation, the U.S. Defense Advanced Research Projects Agency (DARPA), the U.S. National Science Foundation (NSF), the U.S. Department of Energy, the U.S. Army Research Office, and the Czech Ministry of Education, Youth, and Sports. The work was partially carried out at the MIT.nano facilities.
ICE Collecting DNA Samples
ICE collected nearly a million DNA samples last year.
Trump is shifting more responsibility to states. These 7 maps show how.
Firefighting resources are ‘critically low’ in record fire year
New Zealand shields polluters from climate lawsuits
Florida emergency manager Guthrie resigns to join Collins ticket in governor race
Data center gas plants to boost US power emissions by 20 percent
UK, Google to test changes to flight paths to tackle aviation’s climate impact
Looming rock collapse threatens Swiss village as permafrost thaws in Alps
The best fabric for extreme heat? It’s more complicated than you think.
Thermal justice in urban climate change adaptation
Nature Climate Change, Published online: 19 August 2026; doi:10.1038/s41558-026-02727-5
Extreme heat events are increasing, and urban adaptation is urgently needed for humans and non-humans to function safely under heat. Here we present a thermal justice framework and highlight how its use facilitates equitable adaptation and minimizes heat-risk displacement.Bringing justice to informal adaptation to heat stress
Nature Climate Change, Published online: 19 August 2026; doi:10.1038/s41558-026-02728-4
Informal adaptation can augment or replace formal adaptation with flexible and responsive actions, yet it can also reproduce and intensify existing injustices. We propose justice-oriented co-creation between informal and formal actions as a pathway for effective heat adaptation.A climate impact taxonomy operationalizing IPCC physical driver and risk concepts
Nature Climate Change, Published online: 19 August 2026; doi:10.1038/s41558-026-02717-7
Adaptation to climate risks requires integrating knowledge across IPCC working groups. This study presents a climate impact taxonomy that connects climatic impact-drivers from Working Group I to representative key risks from Working Group II and provides more direct guidance for risk assessment and adaptation strategies.Ninth Circuit Ruling Will Force Online Platforms That Host User Speech to Fight Lengthy and Costly Lawsuits Before They Are Dismissed Under Section 230
A federal appeals court just made it harder for online services, big and small, to get lawsuits over user speech dismissed early. In California v. Meta, a Ninth Circuit three-judge panel held that the lower court’s denial of Section 230 immunity to Meta is not immediately appealable. The misguided ruling has the potential to have widespread impact and to threaten the free speech of all internet users.
The ruling is bigger than a loss for Meta, which has the resources to defend itself against these lawsuits. The court’s ruling signals that all online services (and internet users) that host others’ speech—including those without Meta’s deep pockets—must bear the burden and expense of fighting lawsuits that Section 230 ultimately precludes. This will have real consequences, incentivizing online services to take down users’ speech in response to spurious legal threats, filter speech preemptively, or simply stop offering a place for people to speak online. So even though some may think that Meta is not a sympathetic company, the ruling should raise concerns for anyone who cares about an open and free internet.
Immunities from Suit Advance Important Public InterestsA little procedural background is necessary to understand the implications of the Ninth Circuit’s ruling.
Meta had moved to dismiss a group of social media addiction cases brought by state attorneys general, school districts, and local governments. Meta argued that Section 230(c)(1) immunity applies because the plaintiffs’ claims, framed as seeking to hold Meta liable for allegedly harmful platform features, really seek to hold the company liable for publishing decisions related to third-party content. Section 230 is one of the most important laws supporting online free speech, because its protections for online services enable them to distribute users’ speech at an unprecedented scale.
The district court ruled that Section 230 does not apply to certain features (and does apply to others) and so denied the motion to dismiss on the claims related to those features. Meta immediately appealed invoking appellate jurisdiction under 28 U.S.C. § 1291, but the question before the Ninth Circuit was whether the appeal was legally appropriate.
Under Section 1291, U.S. circuit courts generally only have jurisdiction to hear appeals of “final decisions” from the district courts. Final decisions are trial court orders ending a case, or come after a trial on the merits. Section 230 appellate cases often arise from a district court’s grant of a defendant platform’s motion to dismiss the plaintiff’s case based on Section 230. Typically, a district court’s denial of a defendant’s motion to dismiss is not a final order—it simply means that the case may continue to discovery and summary judgment or trial, after which time an appeal would be appropriate.
However, federal law allows for “interlocutory appeals,” which are appeals of orders that do not end a case but nonetheless are allowed because they involve important legal issues. For example, there is an exception to Section 1291 called the “collateral order doctrine”—at issue in this case—allowing for immediate appeal if, as the Ninth Circuit explained here, “holding a trial would imperil a substantial public interest.”
Inherent in the collateral order doctrine is the consideration of whether an immunity like Section 230 provides mere “immunity from liability” or a more robust “immunity from suit.”
An immunity from liability does not require an immediate appeal and so demands that Section 1291’s final order rule be followed. That’s because waiting until the end of a case before an appellate court can consider the trial court’s denial of immunity does not prejudice the defendant. The appellate court may overturn the trial court and grant the immunity, and thus the defendant’s right to be immune from liability would be vindicated on appeal.
Immunity from suit is different. It means that the public interest demands that a defendant be able to get out of a case as early as possible and avoid having to litigate the case to the end. The U.S. Supreme Court has held, for example, that qualified immunity is such an immunity, and that a district court’s denial of qualified immunity for a government official is immediately appealable under Section 1291, notwithstanding the lack of a final order. The idea is that the public interest is served when government officials are free to act without fear of consequences when established rights are not implicated, and so determining as soon as possible whether their acts are immune serves that public interest.
Here, the Ninth Circuit held that the district court’s denial of Section 230 immunity for Meta was not immediately appealable under Section 1291’s collateral order doctrine because the immunity is not from suit, but rather from ultimate liability. The panel’s absurd result contravenes the text of Section 230, the statute’s policy goals, and the court’s own prior rulings.
Treating Section 230 as an Immunity from Suit Protects Online Free SpeechMeta rightly argued that Section 230(e)(3) plainly states, “No cause of action may be brought and no liability may be imposed under any State or local law that is inconsistent with this section.” The panel dismissed this argument, stating that this language likely amounts to “redundancy” reflecting only immunity from liability. The court failed to side with the more reasonable position that statutory language should generally not be interpreted as superfluous.
Meta also reminded the panel that the Ninth Circuit has many times over the past two decades framed Section 230 as both an immunity from liability and an immunity from suit. The panel also dismissed this argument, stating, “It is true that we have used the phrase ‘immunity’ somewhat loosely in our section 230 jurisprudence.”
But “loosely” is a gross mischaracterization—the panel did not discuss a seminal prior ruling, Fair Housing Council of San Fernando Valley v. Roommates.com (2008), in which the entire Ninth Circuit, not just a three-judge panel, explicitly ruled that Section 230 is also an immunity from suit. That court rightly explained that Section 230 “must be interpreted to protect websites not merely from ultimate liability, but from having to fight costly and protracted legal battles.”
Why is it important that social media platforms and other internet intermediaries (and their users) have immunity from suit for engaging in publishing activities related to third-party content—and thus a right to immediately appeal when Section 230 immunity is denied?
The Ninth Circuit panel here, using their own words, failed to “evaluate the interests that would be lost through rigorous application of a final judgment requirement” and failed to consider the “substantial public interest” served by treating Section 230 as an immunity from suit.
Section 230 immunity, contrary to what some argue, is not a gift to Big Tech—it applies to all internet intermediaries, big and small, from the large social media companies to smaller entities like community message boards and local ISPs. It even protects internet users who forward others’ emails or host comments on their blogs. In turn, the law supports the free speech of all internet users.
While it is helpful when an internet intermediary can ultimately benefit from Section 230 immunity, if a trial court’s early denial is not immediately appealable, that means the intermediary must bear the extended logistical and financial burdens of defending itself. Under the Ninth Circuit’s logic, anyone hosting others’ speech online would have to endure the pain and expense of discovery, summary judgment, or trial, before they ultimately can be protected by Section 230.
Congress crafted Section 230 to give internet intermediaries legal breathing room, so that they will be incentivized to facilitate online communication and commerce, allowing the rest of us to go online with minimal barriers to entry, without needing to have loads of money or to know how to code. Congress acknowledged in Section 230 itself, “Increasingly Americans are relying on interactive media for a variety of political, educational, cultural, and entertainment services.”
Yet if platforms, especially smaller platforms, know that they will have to defend themselves for years in court before they can ultimately benefit from Section 230 immunity, this alone will create a perverse incentive, as we have explained, to censor user speech, in order to reduce the platforms’ legal exposure. And this incentive is only exacerbated at scale, where the sheer volume of user-generated content hosted by modern platforms makes legal risk astronomical.
Unfortunately, this opinion seems to be part of larger trend reflecting the Ninth Circuit’s increasing disdain for Section 230, and apparently for free speech rights more broadly. The court similarly held last year in Gopher Media v. Melone (2025)—overruling itself—that a trial court’s denial of a defendant’s anti-SLAPP motion also is not immediately appealable under the collateral order doctrine. This is despite the fact that, similar to Section 230, California’s anti-SLAPP law is intended to allow defendants to get harassing lawsuits meant to silence them dismissed early, lest they be chilled from engaging in lawful speech on public issues due to the risk of being mired in litigation, even if they ultimately win a delayed appeal.
When AI art has no author: Study finds generated images often can’t be traced to training data
When an artificial intelligence image generator produces a portrait, whose work went into it? The question sits at the center of lawsuits, licensing deals, and proposed regulations worldwide. Artists want credit. Companies want clarity. Policymakers want a way to assign responsibility.
New work from a team of researchers at MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL) suggests that for models trained on large datasets, the question may often have no answer. It's not that the tools for finding it are inadequate. The connection itself has disappeared.
The scientists identified a phenomenon they call attribution decay, where the more data a generative model is trained on, the less any individual training example matters to any particular output. It feels counterintuitive, but at sufficiently large scales, they find, you can often remove any single image from the training data, or every image by a given artist, or every photograph of a given person, and the generated sample doesn't change.
And if removing something changes nothing, the researchers argue, it can't be said to be responsible for anything.
"If you take away a piece of data and the output of the model doesn't change, then that piece of data didn't affect the output," says Zheng Dai SM ’21, PhD ’24, former MIT CSAIL researcher and lead author on the work. "So it doesn't make much sense to attribute the output to that piece of data. And if you then do this one at a time for every other piece of data and find that the output doesn’t change for any of them either, then it doesn't make much sense to attribute the output to any one of them."
"All previous methods were approximate," says MIT Professor David Gifford, who is an MIT CSAIL principal investigator. "They really could not absolutely show that deleting individual things did not change the output. This paper introduces the first method that is absolute. You're actually deleting the inputs and deleting all influences of the inputs. This is the first exact method for doing large-scale deletion efficiently and showing that the results don't change."
Dai and Gifford's project is described in an open-access paper published today in Nature Communications.
The retraining problem
Testing this idea directly meant answering a what-if question. What would this model have produced if it had never seen this particular image? Answering it honestly means retraining the model from scratch without that image, then doing it again for the next image, and the next. With millions of training examples, the math quickly becomes prohibitive, which is why prior work in the attribution field has relied on approximations that estimate a training example's influence, rather than actually removing it.
Their workaround is an architecture they built themselves, called a "diffusion ensemble." Instead of one monolithic model, it's made up of many smaller components, each trained on a different slice of the data. Want to know what the model would do without a particular image? Just switch off the parts that saw it. No retraining, no approximation. What's left is a true counterfactual model, not an estimate of one.
Of course, a clever architecture only matters if it still works as a generator. So the team put the ensembles head to head with 24 conventional diffusion models trained on the exact same data. The images came out looking about as good by standard measures.
One nice surprise in the numbers: The more training data, the better the ensembles held up against their single-model counterparts, a hint that they may actually be more data-efficient.
"When you have low amounts of data, they do very poorly," says Dai. "But if you have more data, it actually scales better compared to the vanilla diffusion model."
Exploring a counterfactual universe
With ablation working, the researchers could finally ask their question at scale. Take one generated image, then imagine every alternate version of it, each produced by removing a different piece of the training data. The team calls this the image's counterfactual universe. The distance between the original and its most different alternate, the counterfactual radius, captures the most that any single piece of training data could have mattered.
They trained 24 ensembles on datasets from 256 images to more than 160,000, pulled from seven public collections including CIFAR-10, CelebA, MetFaces, and ArtBench. The pattern was consistent: The bigger the training set, the smaller the radius, shrinking along an inverse power law. It held whether differences were measured pixel by pixel or by semantic meaning, with statistical significance both ways.
The team also stress-tested their own result. Maybe ablation itself was the culprit? They redid it the brute-force way at small scale, training 1,282 separate models, and the decay showed up anyway. Maybe bigger datasets just make each removal proportionally smaller? They pinned the removed fraction in place, and it persisted. Fixed epochs, text-prompted models, class-conditioned models, four similarity metrics — the finding survived everything.
The privacy paradox
The implications run in a direction that surprised the researchers themselves.
Gifford sees the finding as bearing directly on the legal question of whether model outputs are derivative works.
"One way to think about this is that these models are creative. They are not simply copying what they are fed, but creating brand new outputs. If those outputs have nothing to do with any individual piece of training data, that raises questions about fair use, about whether the outputs are themselves copyrightable as novel works, and about how authors get compensated when what comes out of a model isn't attributable to anything on the internet."
Gifford also notes that the work shows how to produce outputs that are guaranteed to be unattributable, a capability he frames as an obligation for the industry, rather than a loophole.
"In order for these companies to claim their outputs aren't derivative of the internet in a copyright-infringing way, they need to revise their models to take advantage of the advances in this work, so they can show they're not creating derivatives of individual people or items."
The work looks at diffusion models, now dominant in generating audiovisual media and prevalent in scientific applications including protein structure modeling and therapeutic discovery. Whether the same decay holds for the large language models at the center of the highest-profile copyright litigation is still an open question.
"If attribution worked, it would reliably tell us whether similarities between a model's output and a copyright-protected work are due to copying or coincidence," says James Grimmelmann, a law professor at Cornell Law School and Cornell Tech. "But this paper provides reason to think that attribution will fail for interesting models. Instead, technologists and courts will need to resort to other methods for assessing copying."
Dai and Gifford's work was supported by Schmidt Futures.
Anthea Coster awarded International Union of Radio Science Appleton Prize
MIT Principal Research Scientist and Haystack Observatory Assistant Director Emerita Anthea J. Coster was awarded the prestigious Appleton Prize at the International Union of Radio Science (URSI) General Assembly and Science Symposium in Krakow, Poland, on Aug. 16.
The Appleton Prize recognizes career achievements and outstanding contributions to studies in ionospheric physics. Appleton awardees are regarded as pillars of the URSI atmospheric science community; the citation for Coster, an URSI Fellow, is for “pioneering research in GNSS [Global Navigation Satellite System] science, developing techniques to provide global-scale view of storm responses in the ionosphere, operationalizing novel algorithms, and providing novel ionospheric products to the community.”
The Appleton Prize honors Sir Edward Victor Appleton, a Nobel Prize–winning physicist and former president of URSI (1934–52) who proved the existence of the ionosphere.
Coster joined MIT in 1984, originally at MIT Lincoln Laboratory, where she worked on satellite tracking applications within the Space Surveillance Complex situated at MIT Haystack Observatory. While at Lincoln, she was introduced to the Global Positioning System (GPS), the first GNSS; her GPS research at Lincoln eventually led to an appointment in Haystack’s geospace and atmospheric science research group. She continued and expanded her Lincoln-based GNSS research, focusing on ionospheric and atmospheric applications. At Haystack, Coster started as a research scientist, becoming an MIT principal research scientist in 2012; she also served as assistant director for the observatory from 2015 until 2024.
Her career research focus spans the physics of the ionosphere, magnetosphere, and thermosphere, covering space weather and storm-time effects and coupling of these atmospheric regions, with particular expertise on GNSS positioning and measurement accuracy. Coster’s breakthrough contributions in GNSS applications to frontier geospace research span many areas, including ionosphere-magnetosphere coupling and mid-latitude ionospheric dynamics. A selected number of her accomplishments include the first real-time GNSS ionospheric monitoring system, as well as pioneering work in monitoring tropospheric water vapor with GNSS signals. She also was responsible for the first GNSS observations of storm-enhanced density, a bright and important feature that can span the heavily populated continental United States, with significant impacts to the Federal Aviation Administration Wide Area Augmentation System, which supplements traditional GPS navigation systems.
MIT Haystack Observatory director Phil Erickson says, "Dr. Coster's award from the International Radio Science Union is most well-deserved, and reflects her substantial international impact on the field of geospace remote sensing. Coster's pioneering application of GNSS signals to global and precise maps of total ionospheric electron density has produced a rich and insightful scientific output that anchors and greatly complements the multi-messenger, sensor fusion techniques at the forefront of the research field in near-Earth space weather dynamics. These areas are of critical importance to our increasingly spacefaring civilization."
Coster’s career also encompasses a lifetime of professional service contributions to the U.S. and international geophysical sciences community, including many leadership positions with the U.S. chapter of the Union of Radio Science, the Institute of Navigation, and the American Geophysical Union. She has served as co-chair of NASA's Living with a Star Program Analysis Group and is a current member of the U.S. National Academies of Science, Medicine, and Engineering Space Weather Roundtable.
She is an author or co-author on more than 200 peer-reviewed publications, and is the principal investigator of numerous federal scientific grants from NASA, the National Science Foundation, the Office of Naval Research, and the Air Force Office of Scientific Research. Prominent results of Coster’s work are heavily used, including scientifically rich GNSS total electron content (TEC) and scintillation data products available to the research community through NSF's CEDAR Madrigal database and the Millstone Hill Geospace Facility.
Coster has also made a number of notable contributions to science outreach, such as deploying radio instrumentation with MIT graduate students in Brazil and Peru, presenting outreach talks to high school and middle school students in Rwanda and Zambia, and installing GNSS receivers in Inuit villages and along the remote Steese Highway in Alaska. For many years, she has taught U.N.-sponsored GNSS workshops aimed at workforce education and career advancement in disadvantaged countries.
Originally from Texas, Coster attended the University of Texas at Austin as an undergraduate and earned her master's and doctorate degrees at Rice University in Houston, where she was involved with ionospheric experiments at the Arecibo Observatory in Puerto Rico. She moved to Massachusetts in 1984 to join MIT Lincoln Laboratory.
"Anthea Coster has made seminal contributions to the state of the profession, enabling the international science community to conduct ionospheric research at spatio-temporal scales that were previously unachievable," says Larisa Goncharenko, assistant director and head of the atmospheric and geospace group at Haystack. "Her pioneering work on introducing and relating GPS measurements to fundamental research has led the community to employ GNSS as an information-rich sensor for ionospheric remote sensing and space weather monitoring. Her effort enabled countless discoveries in the near-Earth space environment that has become increasingly important for human activities in space. I am truly in awe of Anthea's pioneering accomplishments, and incredibly proud of her receiving the Appleton Prize."
With this award, MIT Haystack Observatory is now home to three URSI prize recipients. Former director and research scientist John Evans received the Appleton Prize in 1975 with a citation for "ionospheric physics, including application of the incoherent scatter technique," and research scientist Alan Rogers received the 2008 John Howard Dellinger Gold Medal for outstanding contributions to radio astronomy.
LLMs and Contextual Integrity
I have been thinking a lot about AI and integrity. Part of that is contextual integrity. I recently found two papers on the topic.
“CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs“:
Abstract: Large Language Models (LLMs) increasingly use persistent memory from past interactions to enhance personalization and task performance. However, this memory introduces critical risks when sensitive information is revealed in inappropriate contexts. We present CIMemories, a benchmark for evaluating whether LLMs appropriately control information flow from memory based on task context. CIMemories uses synthetic user profiles with over 100 attributes per user, paired with diverse task contexts in which each attribute may be essential for some tasks but inappropriate for others. Our evaluation reveals that frontier models exhibit up to 69% attribute-level violations (leaking information inappropriately), with lower violation rates often coming at the cost of task utility. Violations accumulate across both tasks and runs: as usage increases from 1 to 40 tasks, GPT-5’s violations rise from 0.1% to 9.6%, reaching 25.1% when the same prompt is executed 5 times, revealing arbitrary and unstable behavior in which models leak different attributes for identical prompts. Privacy-conscious prompting does not solve this—models overgeneralize, sharing everything or nothing rather than making nuanced, context-dependent decisions. These findings reveal fundamental limitations that require contextually aware reasoning capabilities, not just better prompting or scaling...
