In January 1919, more than 2 million gallons of molasses exploded from inside a pressurized holding tank in Boston. A tsunami of syrup, which some reports said reached 40 feet tall, surged through the city’s North End, covering the waterfront, destroying buildings, and drowning dozens of residents.
As the historian Jill Lepore recently wrote in the New Yorker, the Great Molasses Flood is a morbid if apt metaphor for the effect that artificial intelligence is having on the internet. Gooey, disgusting, saccharine AI slop is flowing everywhere these days, gumming up your LinkedIn feed, oozing through Twitter and Instagram, and filling every square inch of the internet with treacly spam and viscous sham, including fake recipes for “glue-topped pizza” on Google; fake biographies on Amazon; and fake images of Trump and Jesus stuck all over the modal Boomer’s Facebook wall.
The Great AI Flood is inundating even the more respectable corners of the landscape of letters. AI has “plunged the book publishing industry into utter chaos,” the Wall Street Journal reported, with “spectacular implosions of big deals over suspected AI use.” The publisher Hachette recently canceled a massive horror novel deal, when an author’s work was flagged as being significantly AI generated. Last week, billionaire Stanley Druckenmiller published an op-ed in the same newspaper, which was identified as “100 percent” written by AI. It’s not just the billionaires that are copy-pasting ChatGPT into emails they dash off to newspapers editors. An analysis by Semafor estimated that of the 310 op-ed columns by guest writers in the Wall Street Journal, New York Times, and Washington Post in the last month, one in six used AI; and one in ten were fully written by the bots.

Why does this matter? One answer is that it doesn’t, really. Of all the existential problems that AI poses to the economy, cybersecurity, biosecurity, and national security, the human fidelity of the written word seems like a relative trifle.
But I care about the slopification of the written word. It matters for education, as I think it would send a terrible message to young people (and their teachers!) in the trenches of high-school English classes if the most esteemed newspapers and publishers simply didn’t care if essays published under your name were effectively composed by an electrical exchange taking place inside some faraway data center. The act of writing is an act of thinking, and to automate our paragraphs is to populate our screens with words without filling our minds with the ideas necessary to conceive of them. Finally, I do not enjoy the idea of a technology that severs the connection between composition and knowledge, such that I cannot know if an author actually knows anything about what they have published under their name.
For many people on the internet, the best medicine for AI slop is Pangram, a machine-writing detection app. Here’s how it works. You read some piece of writing. You copy the text. You go to Pangram. You paste the text in a box, press a button, and voila: You learn what percent of that writing is artificial. If it’s 100 percent AI, you see this:
In my corner of the world, which is journalism and academia and writing in general, Pangram has become nothing less than a cultural force. A meme, even. I’ve seen prominent politicians, chief executives, and public intellectuals exposed as slop-slingers. Their public humiliation sometimes fills me with righteous indignation. But it also makes my toes curl a bit, as I think about people spending all day policing the internet for slop and attempting to Pangram-shame people with a technology that can make mistakes.
Today’s interview is with Pangram founder and chief executive Max Spero. Among other things, we talk about:
what if AI writing takes over the internet and nobody cares?
how Pangram works—and how and when it makes mistakes
why AI writes “like that”
my fear that AI-writing witch hunts would be an unwelcome addition to online life
Pangram’s relationship with Substack
what if a big tech company, like OpenAI or Google, creates their own AI detector that puts Pangram out of business?
Derek Thompson: What problem was Pangram initially built to solve?
Max Spero: After ChatGPT came out, I was thinking a lot about bad actors and what they could do with AI as a source of an infinite volume of text. Somebody could write hundreds of thousands of fake reviews. They could create fabricated news websites. Political misinformation and disinformation could be supercharged.
If you look at the scale of AI and bot traffic on the internet, it’s gone from a small fraction of human traffic to recently passing 50 percent. I think it’s not long before 99% of internet traffic is bots. We’re at risk of the “dead internet theory,” where the internet is just this echo chamber of bots talking to bots. Obviously humanity is not going to go extinct if we can’t tell AI writing from human writing. But what we are doing is trying to save the internet. We’re trying to save the written word.
Thompson: In the last few weeks, the Wall Street Journal has reported that “AI has plunged the book publishing industry into utter chaos.” The New York Times reported that Spotify, LinkedIn, and others are “trying to dig out of a digital sewage heap full of low-quality content made by artificial intelligence.” The Times cited Pangram, which reported that AI was used to create nearly half of the posts with more than 50 words on X or Twitter. That is crazy.
Spero: I’m not sure if that got misquoted. Our study found that 29 percent of long-form content on Twitter—meaning, 250 words or more—was AI generated. Of short-form content—50 to 250 words—it was about 9 percent.
Thompson: Good to fact-check the New York Times live.
Spero: It’s not just the social platforms. LinkedIn had 41% of its long-form content generated by AI. We did a study with Stanford and the Internet Archive, which found that in May of last year, nearly 40% of the internet was AI-generated. I think this number is just going to keep expanding. There are people trying to pollute the internet for different reasons, whether they want their own narratives in the LLM training data or they’re trying to win at SEO. Now it’s easier than ever to produce AI content for search results.

Thompson: Your original thesis was that you wanted to stop “bad actors” from filling the internet with AI dreck. Four years later, is most AI content really written by bad actors? Or is the problem more complex, [such as] a lot of people with busy days [writing] emails and memos with AI?
Spero: I don’t think I fully understood how the economic incentives would work. It’s not just bad actors. It’s everyone. It’s so cheap to produce the written word.
Why is LinkedIn more full of AI content than Twitter? I think it’s because there’s this incentive people feel: Posting on LinkedIn regularly is going to improve my career prospects. It’s going to make me more likely to get a good job because I’m going to be visible to people who are relevant in my career. If posting gives you an advantage to your career, people are going to find the easiest way to post more, which ends up being using AI.
Today’s algorithms also incentivize volume and quantity over quality. Hopefully we’re going to see that change in the coming years as quantity becomes free, while quality is really the limiting factor.
Thompson: Pangram takes text and comes back with a rating between 100% AI and 100% human. Without spilling state secrets, what can you tell us about how Pangram works?
Spero: Pangram is a classifier model. It’s looking at text and saying, “I believe this is AI-generated, AI-assisted, or human-written.” It’s not doing any generation. It’s not an LLM, like ChatGPT.
We have a training set of writing that we know was written by a human. For example, a five-star Yelp review about Denny’s or a 500-word essay on Moby-Dick. Then in each case, we’re picking an AI model and asking it to write something similar. We’re going to ask ChatGPT to write a five-star review about Denny’s. We’re going to ask Claude to write a 500-word essay on Moby-Dick. Pangram compares the human and AI items, which are similar in topic but different in style, and we’re learning the difference between how AI produces text and how humans do.
If you look at the broad field of AI systems, there are models that discriminate. For example, Waymo has cameras and a pedestrian detection system that takes an image input and says either, “Nope, there are no pedestrians,” or, “Yes, there are pedestrians, here they are.” That’s closer to what Pangram is than ChatGPT. Pangram isn’t producing anything. It’s making a judgment based on an input.
Thompson: Is there a certain kind of writing that you’ve trained Pangram on in abundance versus other kinds of writing that are harder to get or less frequent in your training set?
Spero: Our entire human data set is from pre-2022, before ChatGPT was released. From 2022 to 2026, language has changed a little bit, but it’s mostly the same. One of the big things that has changed is how people talk about AI. There’s very little writing about AI and LLMs, and there are a lot of terms today that didn’t exist four years ago.
A big question for us is how we confirm that Pangram still does well on writing about AI, because we don’t have this in our test set to prove that we have a low false-positive rate. Something we have done is to take writing about different algorithms and replace the word “algorithm” with “AI” or “LLM” to simulate modern text without having access to confirmed human-written modern text.
Thompson: Why should people trust Pangram? What is the best third-party empirical evidence that it actually works without a bunch of false positives or false negatives?
Spero: There’s a really good study by the University of Chicago. They tested a whole bunch of AI detectors—some open source, some commercial—and found that Pangram was the best, by a large margin. They tested almost 2,000 human texts and found zero false positives. They found that Pangram was highly accurate across all these different domains. That was part of the reason we started building credibility.
Before that, we were just asking people to trust us. It wasn’t necessarily working that well because there are so many AI detectors that say, “Trust us,” and then end up sucking. They say the Declaration of Independence is AI, or something like that. Pangram is different because we’ve put so many resources behind training this model and being able to tell you at a high level of granularity what degree of AI is in this text.
Thompson: What would you say are the hallmarks of AI writing? And how have those hallmarks changed in the last four years?
Spero: The early ChatGPT and AI models were trained on this instruction-tuned data set and would overuse some words and phrases. For example, “delve,” “tapestry,” “intricate.” They really loved these individual words. If you saw the word “delve” in an out-of-place setting, that was an immediate red flag to some people.
Maybe a year later, they were able to hammer out these word-level inconsistencies. But there were still patterns. For example, “It’s not just X, but Y” is called a negative parallelism, and LLMs love it because it has a lot of impact on the reader. Sentence constructions that use em dashes were associated with good writing. Now I think they’ve even hammered some of these out. So I’m relying on longer-context signals. It’s at the paragraph or sentence level. I could look at the shape of the text and tell you that looks like ChatGPT or Claude.
Thompson: I love this theory that the scale of AI’s tell has changed over time. AI’s tell used to exist at the scale of the word—“delve”—or the level of punctuation, like the em dash. Then AI got better at writing, and the scale was raised to the level of the sentence. It fell in love with sentences like “It’s not X, but Y,” negative parallelism, which, embarrassingly, I used to use a lot.
Spero: It’s good writing in moderation.
Thompson: Something I’ve noticed is that long pieces of AI writing often try to summarize and re-summarize. The tell is at the scale of the paragraph. They’ll be like, “This is genuinely important.” “This is the main course.” “The bottom line is…” Every sentence is trying to outdo the previous one in summarizing what it’s trying to say better. What’s going on there?
Spero: AI wants to try to make every sentence the most impressive sentence. It’s really trying to impress the reader. Partly this is due to the way these things are tuned.
Typically, AI models are trained in two stages. The first stage is pre-training, when it’s predicting the next token. It’s looking at internet data and saying, “I think this word comes next.”
The second stage is reinforcement learning, where it’s already good at predicting the next token. Instead, we’re giving it a new reward function, like giving it a reward if it’s judged as producing good writing or the correct answer. This reinforcement learning turns up a positive signal up to 11. People like when they read an answer and feel it was impressive. You take that reward signal and ratchet it up. Eventually it’s summarizing itself over and over, because that’s how it knows it can make its answer the most legible to the human. It’s making these sentences impressive and meaningful because it knows it’ll get rated as a better writer if it does this.
Thompson: That’s exactly what it does. Every sentence is trying to stand out taller than the one that came before it. It drives me crazy because that’s not how good writing works at any appropriate length.
When I was writing my first long features for The Atlantic, I was working with an editor named Don Peck. With my first drafts, he would say, “You’re signposting too much early in the essay. You’re trying to bang people over the head with ‘This is why this essay is important.’” But in a 7,000-word essay, readers want to go on a journey. They want to get a little lost. They want you to tell them a story and leave the candy at the end of that story.
That, to me, is what AI is terrible at. AI is so often used to summarize or write short posts. It might be overtuned to excellent synopses and great summaries, where every sentence is saying, “I can sum up the best.” But there’s no sense of the talented novelist or creative-nonfiction writer saying, “I’m going to take you on a story here, and you’ll only understand its importance by the time we get to the end.” That seems lost on the models for now.
Spero: One hundred percent. It’s simply not trained for this sort of task, and maybe even if it is, it doesn’t have enough training data. The number of 7,000-word articles is very small. The number of novels is also vanishingly small compared to internet posts. Even on the scale of novels, what makes this novel good? It’s so long that I’m not entirely sure that LLMs are able to pick up the whys.
Thompson: What makes a novel great is [often] change. The characters change. The stories change. The emotions change. The novelists I like best are masters at creating scenes and characters in the first 10 pages that have undergone something wrenching by the end of the story that moves you. But that requires a confident ability to manage meaning across tens of thousands of words that a summary-making machine is not going to be as good at.
Having said that, I’m going to complicate this thesis by referring to studies suggesting there are readers and even MFA writers who cannot tell the difference between human writing and AI-assisted writing, especially when the AI is told explicitly, “Write like William Faulkner. Write like David Foster Wallace.” What do you make of the fact that even sophisticated readers today, according to some seemingly good studies, truly cannot tell the difference between AI and human writing?
Spero: I think the studies are good. But they’re still looking at localized, short passages of text. They’re comparing, at most, a couple hundred words of AI text versus a couple hundred words of famous-author text. At this scale, AI is probably maybe a little bit superhuman at portraying something clearly and concisely in a way that is impressive to the reader. But the problem is long-context writing. It can’t do this over the course of a novel.
Thompson: What you’re saying is it can write a Cormac McCarthy sentence, but it can’t write Blood Meridian.
Spero: Exactly. I don’t want to say AI is never going to be able to do this, because historically if you say that, you’re going to be proven wrong eventually. But today’s AI is really not even close.
Thompson: How is Pangram able, even at this short-passage level, to detect AI-ness that expert readers and writers themselves cannot?
Spero: There’s a lot of information hidden in natural language, and Pangram is able to pick up on this. If you think about the different ways you might phrase a sentence, you might have this wide distribution of five different ways you could phrase it. AI might have a slightly different distribution where it’s going to prefer ways two and three out of the five. Over the course of a document, we can build up increasingly high confidence that something was written by AI because the decisions are consistent with how an AI makes them.
Thompson: It makes me think that Pangram is able to see a mathematical structure in language that I can’t.
Spero: I think that’s true. It’s trained on so many millions of AI documents that you and I would not be able to read all of the Pangram training data in our lifetimes. You can pick up much greater patterns the more you study and the more examples you see. With machine learning, we’re able to do this at a great scale that makes Pangram essentially superhuman at AI detection. It’s better than most people because it’s seen more data.
Thompson: Let’s talk about Pangram as a cultural force, as a maker of memes. There are stories where prominent pieces of writing will be identified as AI, and people will take that Pangram screenshot of “100% AI” and suddenly it’s everywhere. How do you feel about the Pangram rating becoming an internet meme?
Spero: I think it’s emblematic of this greater backlash against AI content. For a long time, people were automating their LinkedIn or whatever, and it felt to a lot of people like they were just having AI shoved in their faces everywhere they turned. That’s part of why the backlash has been so great.
Pangram is powerful for these people as a way to say, “It’s not just my intuition. There’s an accurate third-party tool thing is also helping me make a judgment and say that this is AI generated.”
Thompson: Everyone wants a referee. Your model is very effective, but it’s not perfect. Some folks really do get publicly shamed by the use of Pangram. Even if Pangram were 99% accurate, if it’s used 100,000 times, one should expect 1,000 errors. Does it bother you that the technology might be used to publicly shame people who don’t deserve it because they did write what is being labeled AI writing?
Spero: Our false-positive rate, which is how often we say something human written is actually AI, is about one in 10,000. So of every 10,000 pieces of human writing we scan, one of them is going to be AI.
Thompson: So I said 99% in my question, but it’s actually 99.99%.
Spero: Correct. With that said, we scan enough text that we’re going to have dozens of false positives every day, and that’s a fact of life. Obviously, we want to drag this down. But Pangram already has a lower false-positive rate than your average person.
What we’ve seen in the art world is people will go witch hunt and be like, “This digital art looks like AI.” They might go really deep into “the hair is wrong” or “the finger’s messed up,” and then it’s actually just that the artist made a mistake. It’s not AI generated. That does really suck.
Thompson: A few months ago, someone identified a column by a Guardian sports journalist that they said sounded like artificial intelligence. The Guardian released a statement saying, “No, this is how the writer has always sounded.”
Then you took about 900 of that journalist’s articles, ran them through Pangram, and published a time series showing that their writing had become “increasingly reliant on AI.” What part of you thinks that’s good behavior we should encourage? And what part of you worries that if everyone behaves like this forever, it’s going to feel a little witch-hunty on the internet?
Spero: If the Guardian is telling people, “This detector’s completely wrong. It’s making a bunch of errors,” I can reasonably defend myself and say, “This is not how the writer has always sounded, because there’s actually a point in time where they started using AI, and you can see it. It’s very visible.”
This writer is still writing with AI. You can look at their recent pieces. If the Guardian is happy with this writer’s output, I’m not going to tell them they can’t be. But I really do think this one is AI-generated. It’s not an isolated incident, and it’s very clear after looking at the data that it’s not a false positive either.
Thompson: I’m of two minds. I’m against public shaming. But I also think your behavior was a bit akin to policing. And policing is good when the rules being violated are worth policing.
I think AI writing reduces the authenticity with which people deal with each other on the internet because suddenly someone’s output can plausibly be that of a bot communicating information they don’t even know in their own heads. I worry about young people who never learn to write because they simply learn how to prompt ChatGPT to write their essays for them. So there are values worth protecting, and values worth protecting require some kind of policing. What values do you think you’re protecting?
Spero: I value human authenticity and human-to-human connection. In a sense, AI threatens to replace a lot of this, especially when combined with capitalism and monetary incentives.
Consider a world where it is taboo to ever call someone out for AI writing. What do you think major news and media organizations are going to do? They’re going to let journalists go and push everyone else to output 10 times as many articles with AI because it’s taboo to call people out on this. That monetary incentive is here to stay, so we need an incentive on the other side, which is people who are vocally anti-AI text and pro-humanity.
Thompson: I want to talk about three scenarios that could threaten your company. The first is that the gap between AI and human writing closes. The obvious way this happens is that AI gets better at writing “like a person.” But another way is that humans get worse at writing like humans. We write more and more like AI, and you have this pincer movement where AI and human writing converge.
Spero: Human language does shift. But it changes on a much longer time horizon than AI text. So we should probably only be looking at the AI text because that’s what’s changing really rapidly.
There is a good argument that AI writing will be as good as human writing, possibly even surpass human writing within the next decade. The question will then become, do we care if something is human-authored or not, and why? My argument is we’re still going to care a ton. The internet is going to be full of bots. Internet traffic is going to be 99% bots and 1% humans. I’m going to care that I can get information that I couldn’t get from Claude or ChatGPT. To do this, I need to be hearing from a real human, not a bot.
Thompson: In many ways, we’re impressed by activities that we know robots can do better than humans, but we’re only impressed when the human can do it. Think about pitching in baseball. Can a robot throw a ball faster than 105 miles per hour? Of course. But it’s not interesting. No one wants to watch a robot pitch for the Milwaukee Brewers or Pittsburgh Pirates. They want to watch Paul Skenes or Misiorowski. They want to watch a human throw the fastest fastball ever thrown by a human. That’s the only thing that has athletic or artistic value.
It’s strange to think about that [quality] invading the purely artistic space of painting and writing. We’re already maybe seeing it in mathematics, where people are going to recognize that AI is a better mathematician than any living human mathematician. But I think there’s still going to be a part of us that will be specifically and uniquely impressed when achievements are human achievements rather than technological achievements.
Spero: We connect with the story, too. Look at the art world when photography was invented. Previously, the best thing you could have for a faithful reproduction was a painting. Then suddenly a photograph could do better along basically every axis.
What happened to art? Painting was not completely supplanted by photography. Instead, artists found different ways to find meaning and express themselves. They didn’t simply go farther down the path of realism. They expressed their humanity in a different way.
Thompson: The second scenario that could threaten your company isn’t that the gap between human and AI writing closes, but that people just stop caring. Are you worried about that?
Spero: At least in the medium term, no. Beyond the backlash against AI, there are societal reasons why we’re going to care about human text. I still find it meaningful that if somebody writes me a personal note, it was written by them and it’s not an AI-personalized note. The more AI proliferates and becomes common, I don’t think that sort of thing is going away. There are reasons to prefer human writing that are not about quality. We’re going to prefer human text more as AI text proliferates.
Thompson: The last threat to the company is more technical. There’s this concept in artificial intelligence called “the bitter lesson,” which says that general models with a lot of compute can often do a task better than a narrow model. That might speak to a world in which ChatGPT or Anthropic builds the best possible Pangram; or Apple or Google creates something on-platform that automates the Pangram function. What are you doing to prepare for the possibility that the big guys recognize that consumers are sick of the internet being drowned in AI slop and come up with their own medicine?
Spero: The biggest takeaway an AI researcher should take from the bitter lesson is that more compute and more data will solve a problem more efficiently than human expertise. It doesn’t matter how much of an expert I am in AI slop. If I have more data than anyone else, I can make a better AI detection model.
We’re trying to be on the right side of history with this lesson by having the biggest model and the most data. So far, I think we are, except for the big labs. Anthropic and OpenAI obviously have infinitely more data than us in terms of AI-generated content.
But there’s still a lot of value in Pangram being an external third party. People will trust us because I’m not Anthropic telling you this was or wasn’t AI generated. I am separate and independent from OpenAI, Anthropic, Google. Because of that, I can make the best judgment rather than over- or under-attributing the AI models.
Thompson: It’s like the referees in the NFL or NBA not belonging to the players’ union. They’re not the owners. They’re not the players. They’re a third entity and therefore can maintain some kind of authoritative independence.
Spero: Exactly. Anthropic recently announced that they’re coming out with watermarking. They’ll be able to tell you if AI touched a piece of text because when it produces its output, the watermark will apply a statistical invisible pattern. Somebody can then read that text, put it through the watermark checker, and say, “This text was generated or came out of Claude.”
However, this has the risk of over-attributing authorship to Claude, especially when there’s a truly mixed-authorship document where somebody spent a lot of time working on it and then worked with Claude. Pangram will be able to do a better job correctly attributing that Claude wrote these parts and a human wrote these parts.
Thompson: You guys are already integrated with Substack, which publishes all of my work on the internet. Are there other similar integrations you’re looking to get started?
Spero: We also have a Chrome extension. Even if we don’t have a deal with an individual platform, Pangram will label posts on Reddit, Twitter, LinkedIn, Medium, and Substack as AI, mixed, or human, which is valuable for somebody navigating an internet that’s increasingly full of AI slop.