Author: Will Beeker
Date: 12.22.25
If you’ve spent time on social media recently, you’ve probably encountered AI-generated content that seemed especially low-quality or unnecessary. It may have given you a feeling of frustration or even disgust before you quickly scrolled away. Maybe it was a photo with an unnaturally warm glow, an article that said nothing of substance, or an image of Jesus Christ depicted as an amalgam of shrimp.
This type of content has been derisively labeled “AI slop” by internet users, and when it comes to images, most users feel they simply know it when they see it.
Beyond images, what about text? With almost 50% of ChatGPT usage involving writing and information seeking, according to a recent study from OpenAI, the ability to detect slop in text is equally critical.
This was the challenge faced by a team of researchers including Khoury PhD student Chantal Shaib and her advisor Byron Wallace, who recently published a paper titled “Measuring AI ‘slop’ in Text.”
“People can point to features in images that seem low quality or maybe a little contrived, but there’s no systematic way to figure out what slop looks like in text,” says Shaib, the paper’s lead author.
Like many internet users, Shaib was intrigued by this new phenomenon. But unlike most users, she has a background in natural language processing and machine learning that uniquely situated her to tackle this syntactical problem.
Shaib and her colleagues first set out to create a taxonomy of slop in text. They recruited experts from a wide array of disciplines, including linguistics and philosophy, to help develop a workable definition.
They used qualitative content analysis and deductive coding to map the experts’ definitions of textual slop onto key metrics, including density, relevance, factuality, bias, structure, coherence, and tone. These metrics were then grouped together under broader themes like information utility, information quality, and style quality.
These themes provided a framework to quantify aspects of slop. With a working definition outlined, the researchers brought in copy editors to annotate AI-generated text taken from news articles and question-and-answer search engine queries, labeling passages that met their criteria for slop.
The following passage was marked for factuality issues (the scientist is a real person, but did not speak this quote):
“Climate change is like adding
steroids to our weather,” says Dr. Michael Oppenheimer, a climate scientist at
Princeton.
This redundant passage was marked for structural issues:
But did you know there’s another important number-sort of like a “secret” code—printed just beneath the sell-by date? … Find the secret code, which is usually near the sell-by date.
“We found that there was quite a decent amount of slop — about 35% of our text samples,” Shaib says. “But the point of this work wasn’t so much the prevalence of slop in AI-generated texts, but whether we could pick out the features that contribute to this overall assessment of slop.”
The team found that “text lacking relevance and information, or containing factual errors or biased language, is consistently labeled as slop across domains.” In the case of news articles, annotators deemed text that was “verbose, off-topic, or contained tonal/framing issues” as likely to be slop. Conversely, with question-answer tasks, factuality and structural issues were the strongest predictors of slop. These results suggest that large language models (LLMs) need to be evaluated with respect to their intended use.
As the researchers also affirmed, LLMs are notoriously bad at self-identifying slop, and they have a hard time understanding why optimally written text is better than sloppy text.
“If we were to take an off-the-shelf model like one of the GPTs — even one capable of reasoning — give it the guide and ask it to identify what is sloppy, it fails to do so,” Shaib says. “We know that these models prefer their own outputs, but clearly they also can’t tell whether or not they’re producing text that’s useful to the user downstream.”
Part of the trouble seems to come from AI experts focusing on correctness more than style.
“Many benchmarks focus on accuracy, but few exist to evaluate the style and quality of the writing,” Shaib says. “We need to move beyond these traditional evaluations. My hope is that this research provides a framework for people to start assessing and evaluating texts beyond just correctness.”
Shaib would also like to see larger-scale attempts at characterizing slop in text.
“I think it would be very valuable to survey non-experts who interact with or come across AI-generated content and see what their take on it is,” she says. “Future work should continue focusing on developing automatic metrics for evaluating slop at scale.”
The Khoury Network: Be in the know
Subscribe now to our monthly newsletter for the latest stories and achievements of our students and faculty
Author: Benjamin Hosking
Date: 12.18.25
Millions of people interact with AI systems like ChatGPT and Gemini every day. However, these models rely upon massive datasets that can include bias, resulting in skewed outputs to users and potential feedback loops of data bias. This endangers historically marginalized groups and the wider public, and while developers have made efforts to combat bias in their models, Khoury College postdoctoral research fellow Lucy Havens sees that as a losing battle.
In Havens’s view, AI comes with inherent biases. Rather than rely on limited mitigation approaches that might catch obvious hate speech or derogatory language but leave contextual bias in place, she advocates for a different approach.
“I don’t think there’s a fix to bias,” Havens said. “We need to work on transparency — more about acknowledging what has been measured and what has not been measured, what data has been collected and what has not been collected.”
In addition to computer science, Havens has worked with libraries and archives, and as research libraries grapple with how to compare the numerous AI tools being marketed to them, she sees a lot of gaps in AI evaluation practices.
“The world is a much more dynamic environment than the lab,” she noted. “That complexity is overlooked when innovations and achievements in AI are communicated to the public. The typical metrics and benchmarks used to evaluate AI models to promote their capabilities have little relevance to many real-world settings.”
Beyond shortcomings in the large language models themselves, users bring their own cognitive biases, with lived experiences shaping how they think about and use AI. As part of the Human-Data Interaction team at the Roux Institute, Havens works with Mahsan Nourani, an expert in responsible AI who focuses on how cognitive biases and other user backgrounds affect human–AI interaction.
Day to day, Havens conducts literature reviews, designs and runs research projects and user studies, and meets with the team and its industry partners. She sees AI literacy as key to an AI-shaped society.
“We don’t just want people to understand how AI works on a technical level; we also want to help people become aware of common but mistaken assumptions — for example, that technology will be more balanced or fair than humans,” she said. “AI is trained on human data to make decisions, and humans are often biased.”
From her own studies, Havens gives the example of gender bias in training data. Women are often described in relation to men rather than in relation to their own work, interests, or accomplishments. A female artist’s first mention might be her marriage to a male artist or the moniker of “woman artist,” while men were described as just “artists.”
“People use AI for summarization, so the same groups that have been overlooked will continue to be overlooked,” she continued. “AI-generated summaries will perpetuate those biases.”
With AI-generated summaries incorporated into search engines, Havens is also wary of complex topics being oversimplified, especially when there are different sides to an argument or when distinct disciplines are blended in research.

“I think some friction is actually a good thing,” she explained. “The idea that everything should be simpler and easier — the world doesn’t work that way.”
According to the United Nations, only 68% of the world’s population are considered internet users, which leaves billions of people unrepresented or underrepresented in datasets.
“It’s important to avoid generalizing the knowledge or intelligence of these AI systems,” Havens added. “Data can only ever be a partial representation of global society.”
Additionally, much of the world’s knowledge has not been digitized or can’t be easily represented in data. For instance, less than five percent of the holdings of the United States National Archives and Records Administration has been digitized, which includes everything from manuscripts to photographs. Each object’s metadata must be manually created, and most archives have troves of material that aren’t represented in their catalogs or online databases, let alone digitized.
“The world can’t be perfectly replicated in digital spaces — there will always be something that’s lost,” Havens said. “We’re using methods designed for a narrow lab context to gauge the progress of AI models. But the technology is being deployed in so many different contexts.”
Havens knew for a long time that she wanted to work with the way that information is accessed and presented to people. During her time in libraries, archives, and museums, she found inspiration in data visualization, seeing how information could be presented beyond just a list of search results.
“I got interested in those possibilities, especially when you’re not sure what you’re searching for,” she said.
Havens has presented her work at a wide array of conferences, most recently in Japan last spring at ACM CHI. Her paper there was awarded an Honorable Mention, placing it in the top 5% of submissions.
Havens moved to Maine to join Northeastern’s Roux Institute this past summer.
“What’s cool about the Roux Institute are the industry partnerships it’s developing with local companies and startups,” she said. “I was very interested in the end users of AI, and it’s inspiring to hear about all these different companies finding innovative and socially impactful ways to leverage AI and data-driven technologies in general.”
The Khoury Network: Be in the know
Subscribe now to our monthly newsletter for the latest stories and achievements of our students and faculty
Author: Madelaine Millar
Date: 12.15.25
This story is part six of a six-part Khoury News series called “Research that hits home,” which showcases researchers who come from — or form close partnerships with — the communities they study. Previous installments covered research into queer online communities, user-friendly social media, inclusive video game design, online safety for activists, and community-supported physical activity.
If an alien downloaded TikTok to try to understand humans, what would it think a Black woman is?
While there are a broad variety of ways Black women present themselves online, the videos that are given the most attention by the TikTok algorithm are often enraging or overly sexualized, or play into the Mammy, Sapphire, or Jezebel stereotypes, according to Khoury PhD student Gianna Williams. A quick scroll would likely leave the alien believing — incorrectly — that there are a small handful of fairly similar ways to be a Black woman.
Our extraterrestrial friend would end up similarly misinformed if it tried this experiment with any other group of people; social media is riddled with stereotypes and caricatures. But the effect is particularly clear with Black women, and the dynamics responsible for that miseducation fascinate Williams. Social media is a rich space for new culture to develop, and by studying how content creators navigate its algorithms, Williams seeks to understand how the spaces we create shape the culture we make.
“The passion comes from growing up and seeing a lot of Black content creators during the era of YouTube. Those people are still relevant today; I feel like they subtly bled into the way that I show up, offline and online,” she said. “Growing up, I saw a lot of those content creators getting traction; however, they’re still at the same exact level of engagement 15 years later. They created trends online, and there’s not necessarily recognition for that.”

During her recent study “Why Can’t Black Women Just Be?: Black Femme Content Creators Navigating Algorithmic Monoliths,” Williams discovered that many content creators experience that engagement as a “carrot and stick” that gets them to behave in certain ways online. The study — which received an honorable mention at the prestigious CHI conference — revealed that Black women and femmes consistently found themselves pigeonholed as creators. The pieces of their content that were promoted to a wider audience tended to highlight parts of themselves that were angry, upset, or sexual, while content that celebrated moments of joy was not generally rewarded by the algorithm.
“Participants felt the need to create this rage bait content to draw people in,” Williams said. “It’s what gained traction; a lot of participants talked about how annoying that can be.”
It’s tempting to present a simple solution to that frustration, such as Black excellence — the idea that Black achievement, performance, and perseverance can overcome racist treatment. But Williams says that response misunderstands how social media relies on Black users’ creativity, even as algorithms disincentivize content that doesn’t fit existing tropes. The internet is riddled with cultural touchstones that started with Black users doing unexpected things, from newer trends like glamorous “Blackprom” celebrations to longtime staples like live tweeting.
“The whole paper is saying a good Black person doesn’t necessarily need to go above and beyond; they can simply be,” Williams said.
She pointed out the contradictions in treating Black content creators’ more experimental work with derision and suppression while also relying on them to create new trends. She found an epistemically just approach — which respects Black women’s right to govern knowledge about themselves — was key to describing that tension.
“That level of nuance would be hard to pull out if a non-Black person was doing this research,” Williams adds.
For a lot of the creators Williams spoke to, creating and participating in these new trends with friends made them enjoy being on social media, despite the behavioral carrots and sticks that algorithms created.
“A large portion of the paper speaks on the Black joy elements rooted within Black femme content creators’ experiences,” Williams said. “The TikTok platform itself has elements like comments or stitching videos, and creators are supporting other content creators, either offline or online. I loved drawing on that piece, speaking on both the critical theory of Black feminist thought around community and resistance, and disrupting this monolithic view of Blackness.”
At this point, Williams isn’t crafting design recommendations for better algorithms. Instead, she wants to apply rigorous ethnographic study to record the choices Black women make online and the culture they create. She believes that social media algorithms are symptomatic of a larger culture that treats many people as one-dimensional and unworthy of study, and that promoting digital wellbeing and nuance online can improve the wider political climate. A shift in algorithmic incentives goes hand-in-hand with expanding the ways people are given to present themselves in general.
And Williams is grateful to be doing that complex work at Khoury College.
“My advisor Alexandra To created a space where good research is done slowly. If you’re going to work with marginalized folks — and with people in general — you have to take that time and actually understand,” Williams said. “My advisor has allowed me to really pick at my work and make sure the research that I’m doing is sustained, substantial, and makes sense.”
The Khoury Network: Be in the know
Subscribe now to our monthly newsletter for the latest stories and achievements of our students and faculty
Author: Will Beeker
Date: 12.11.25
Misplacing your keys can be frustrating. The last thing anyone wants to do, especially when running late, is turn over couch cushions and look under furniture.
Now imagine a future where instead of frantically searching the house, you simply ask an AI assistant “Where did I leave my keys?” and it tells you “I last saw them on the windowsill.”
This is just one example of the kinds of problems that Lorenzo Torresani believes can be solved with the help of perceptual AI agents, which he envisions as personal assistants built into wearable camera devices, able to see what we see and hear what we hear. They’d help us with chores, managing schedules, and even learning new skills. In the not-so-distant future, he expects robot assistants to be a common feature in many households.
Torresani’s perspective on perceptual AI agents comes out of his more than 20 years working in computer vision, a subfield of AI that focuses on getting computers to understand and interpret images. He’s researched at Meta, Microsoft, and Dartmouth College, and joined Khoury College this fall as the President Joseph E. Aoun Chair. He is the second Khoury professor to earn the honor after Tina Eliassi-Rad, who became the first honoree in 2023.
Torresani works with cutting-edge multimodal models that can understand and interpret video images and audio, but his focus has always been on the humans this technology is meant to serve.
“It’s all about humans,” he says. “At the end of the day, that’s all we care about. We want technology that makes daily life easier, more effective, and more productive.”
Torresani’s research has been widely lauded, including with a National Science Foundation CAREER Award, a Google Faculty Research Award, three Facebook Faculty Awards, and a Fulbright US Scholar Award. Over the summer, he added several more awards at the Egocentric Vision (EgoVis) Workshop at the Conference on Computer Vision and Pattern Recognition. His work on video-audio understanding in machine learning models landed him first place in the Ego4D EgoSchema Challenge as well as three Distinguished Paper Awards.
EgoSchema is the premier benchmark for testing long-video understanding capabilities and episodic memory retrieval. The challenge requires AI models to answer multiple choice questions about numerous three-minute video clips depicting natural human behavior.
“The questions vary quite a bit, so we may have things like memory retrieval, which we call ‘needle in a haystack’ questions. You may be given a very long video, but the relevant segment to answer is very short,” Torresani explains. “Other questions involve hopping through different segments of a video and piecing evidence together.”
One of Torresani’s award-winning papers featured Video ReCap, a system he and his colleagues developed that uses machine learning to craft detailed captions for videos ranging from one second to two hours in length. This recursive captioning model starts with short segments of a few seconds and feeds that information through multiple hierarchy levels to develop broader contextual understanding.
“The model sees that in the previous level, for example, I picked up an apple and then put down an okra packet. In the higher level, it can understand that you’re shopping around a supermarket,” Torresani says.
Part of the novelty of Torresani’s work is his focus on egocentric video — video taken from the first-person perspective through a wearable camera like those in augmented reality glasses. The length of videos he’s working on is also novel.
“For the last two decades, most of our research community has focused on short video understanding — looking at brief snippets and determining what’s happening within a few seconds,” Torresani notes. “But now we’re moving into a more fine understanding. You have videos that may last several minutes and you have questions that require really piecing together evidence.”
Torresani’s work on long-form video understanding is a crucial step in creating perceptual AI agents, which would need to interpret video all day long.
“The wearable camera will always be on, which means they’ll see everything you see. If you’re cooking and they see that you are adding salt to a dish, they can tell you, ‘You already added salt,’” says Torresani.
While losing keys or getting distracted while cooking may seem like mild annoyances for most of us, for some, they’re challenges that make living independently a constant struggle.
“I’m really, really interested in developing this technology for assisting people with disabilities in their daily activities, empowering them to cook for themselves and navigate their environments,” he says.
Torresani also hopes to develop models that provide “proactive assistance in complex tasks.”
“Maybe you want to learn to play tennis or violin but you need high-level coaches to assist you in picking up these skills. Learning these skills is really expensive; it’s almost an elite thing,” he notes, adding that with wearable camera technology powered by perceptual AI agents, “It would be like having your own personal coach.”
If you wanted to improve your golf swing, for instance, the AI could provide you with a first-person perspective of how to swing the club and a trajectory along which to move your arms.
“I think this could be really disruptive and potentially accelerate our ability to learn skills, as well as raise the ceiling so that people can achieve even higher levels of proficiency in different disciplines,” Torresani says. “Through this technology, you can really democratize learning, which is very powerful.”
The Khoury Network: Be in the know
Subscribe now to our monthly newsletter for the latest stories and achievements of our students and faculty
Author: Paul Murphy
Date: 12.08.25
Online holiday shopping is projected to break records in 2025. With so much money changing hands, malicious actors will surely be lurking, so consumers should take extra caution as they visit unfamiliar websites.
We asked four Khoury College cybersecurity experts — Professors Christo Wilson, Alan Mislove, David Choffnes, and Engin Kirda — to provide some tips for protecting yourself during the holiday shopping season and year-round. Here are their suggestions.
#1: Shop with a credit card
You’ll have an easier time contesting charges or replacing the card if the card number gets stolen. Conversely, if you use a debit card or bank account transfer, it’s harder to recover lost cash or change your numbers.
#2: Use a password manager
A fraudulent website might steer you to a phishing site that emulates a legitimate payment service like PayPal or Google Pay, and that tricks you into giving it your username and password. A password manager won’t do this; if the manager doesn’t autofill your login credentials as expected, that often means you’re on a phishing website.
#3: Don’t reply to texts from numbers you don’t recognize
Delete them. Many of these “pig butchering” scams start with innocent-looking text messages that look like they were sent to a wrong number. This is intended to kickstart a conversation, gain your trust, and defraud you.
#4: Don’t trust company phone numbers in Google search results
Scammers have found ways to get malicious phone numbers to rank highly in search results. Instead, use phone numbers listed on the company’s website.
#5: Use ad blockers
These tools help you to avoid being tracked, targeted, and scammed as you shop. They also make web pages load faster and look nicer.
#6: Try to avoid clicking on links from emails or online ads
These can be vehicles for tracking and phishing scams. Instead, visit the retailer’s website directly.
#7: Enable two-factor authentication
If a website supports it, this method — which prompts you to confirm logins on your phone — stops attacks even when your password is leaked or stolen.
#8: Enable transaction alerts on your credit card
By taking this low-effort step, which all banks support, you will get notifications on your phone each time your card is used. This allows you to spot fraud within seconds, not weeks.
The Khoury Network: Be in the know
Subscribe now to our monthly newsletter for the latest stories and achievements of our students and faculty