Ricardo Hausmann, a Harvard economist, published an essay in the Financial Times against the tariffs. He used AI to shorten it. That broke the FT’s rule against AI in the writing process, a rule he told the Chronicle of Higher Education he had not known about. The FT added a note and began a review. On August 13 a freelance writer posted the piece with a one-line verdict: the worst AI slop he had seen in a major publication. Two days later the political scientist, and long time colleague, Brendan Nyhan quote-tweeted him with screenshots of a scan by Pangram, a program that estimates how much of a text a machine wrote, and no words of his own. The tool put the essay above 70 percent. That post has been seen 146,000 times last I checked. (That’s not even a call out, by the way. When Brendan posts these, he usually has good reason to do so. I’ve trusted his eye for a long time.)
Zvezdelina Stankova teaches math at Berkeley. She published an op-ed asking the University of California to bring back the SAT, partly because applicants now use AI to write their admissions essays. She used AI to edit the piece. The San Francisco Standard’s spokesperson told the Chronicle that AI may assist with writing so long as humans stand behind every article, so she broke no rule. The Daily Californian scored her with the same detector anyway and printed the number, 33 percent. By the end of the week she was defending herself in the Chronicle while nobody discussed her argument about admissions.
Hausmann broke a rule and Stankova did not. Their outlets responded differently. The crowd did not.
I write with AI and say so in the disclosure on here (and to be quite clear, and so far at least, I use it quite differently when writing my articles and/or in other professional pieces). So, sure, this essay defends my own substack machinations in a way, and…with candor, I think you should read my substacks with that in mind.
I continue to maintain that my AI use makes these substack posts better, makes me more efficient, adds production value, etc., but that is probably because (I’d like to think anyway) I am a reasonably advanced user of these tools, and other friends have told me that AI tics still appear now and again and it affects what they read, and how they read it.
So, I’m at a bit of a loss, really.
Obviously, you can also check the score and decide for yourself. Any reader who wants a number on this here essay can have one in about four seconds. Go ahead, I’ll wait. All I ask is that you watch what the number does to your reading psychologically, because that’s really what this piece is going to be about.
For the record (and because I am about to ask other people for this), I wrote the argument of this piece, chose the sources, and checked them; a model helped me cut, suggested improvements, gave me some tightening options, and helped me hunt for logic holes, and then I rewrote the suggestions it gave back (that’s what I mean by “iterating”).
I hope you find value in what’s here. If it’s not for you for whatever reason, I get it.
Alexander Kustov admitted the same conflict when he wrote about detectors in June, and his summary of the technology remains the best I know: a detector estimates the probability that a text is AI, never the probability that it is bad.
Never ever ever, in fact, the probability that it is, you know, actually a bad piece.
So, no, I did not just push a button on this and say go. I use AI as a thinking partner. This post took hours of iterating and reading and iterating (and if I’m honest, I am using AI to learn while I am putting these together, so I am getting something out of it too).
I’m guessing this piece will have a high AI score, probably on the pretty high end, and that number will be high because it has been with all of my longer form pieces, and it will still tell you nothing about what actually went on here other than I used AI a good bit. (I haven’t looked yet and I’m close to hitting the go button.)
But back to the point: yes, I understand that AI is still a new thing that we’re all getting used to, a thing that has been probably the most momentous and the fastest adopted technology in world history. People don’t want to like it because it induces fear, uncertainty, and doubt about the future. I get all that. I really do. I feel it too, believe it or not.
And yet, AI is here, AI is not going anywhere, and I have to constantly remind myself and everyone else that AI will never be worse than it is today. It’s not going to slow down, and it will continue its move towards the asymptote.
We gotta figure this out while also looking to where we will be in six months, twelve months from now.
Although being here does not settle anything by itself. That a technology exists and is not leaving tells you the conditions you are working in, and nothing at all about what students need to learn or what a scholar owes a reader. Those still take an argument. I will try to make one further down, and if you think the argument fails, then inevitability will not rescue it and it should not.
But, in today’s piece I want to start with what we actually empirically know about what these tools do to writing and to the people who use them…it’s a lot less than all the shouting suggests. (If there are things that should be added to this list, I’d ask that you place them in the comments or DM me, because I am trying to keep an updated list going.) Then, we’re going to try to lay out what I’m doing and why. I remain open to the possibility that I am wrong about this, but the more I think about it, the more I land where we get to.
Some of the evidence we know, but even it is sometimes read out of context.
When people learn that AI helped write something, they like it less. Researchers ran sixteen preregistered experiments with more than 27,000 readers, who rated creative writing lower once it carried an AI label because it felt less authentic to them. A PNAS study found that workers who disclose AI use get judged less able and less driven. In a working paper (that has not been through peer review and needs a bigger sample), raters scored an article lower after hearing that AI was involved, though every version they read was the same human-written text.
About a thousand Turkish high schoolers practiced math with GPT-4. The ones with an ordinary chatbot did well in practice, then scored 17 percent below the control group on the exam, where they worked alone. The ones whose chatbot was built like a tutor, holding back answers and prompting steps, matched the control group. The paper’s title says “without guardrails,” and everything hangs on that clause. A student who receives the answer skips the work that would have taught him.
Putting together words has become really cheap. Organization Science scored the abstracts of its own submissions with Pangram and counted a 42 percent rise since ChatGPT, almost all of it from manuscripts with detectable AI text, while the rest declined. Across seven thousand submissions the errors wash out and the trend still survives, while a single score on a single op-ed is the error rate applied to one person’s name.
Words also got more alike, which is the part that matters for detection. Doshi and Hauser found AI-assisted stories rated more creative one at a time and more similar to each other in the aggregate. And a Nature Human Behaviour paper published last week ran more than 880,000 texts through model rewrites: meaning survived, with 87 percent of pairs scoring above 0.95 on similarity, while the variance in writing complexity fell somewhere between 21 and 50 percent. That flattening is a good part of what a detector is actually seeing when it lights up a competent edit. It is a signal about sameness, not about badness.
arXiv stopped taking review papers in computer science last fall because hundreds arrived each month, many of them the sort of annotated bibliography a model can produce on demand. It comes down to the fact that the academic journal industry, that includes, editors, reviewers, and colleagues, now carry costs that they did not a couple of years ago.
In a randomized study, sixteen experienced programmers worked on their own code. They predicted AI would speed them up by 24 percent. Afterward they said it had, by 20. The stopwatch said it slowed them by 19. When the lab tried again this year the number leaned the other way, and then they found a leak in the design and said they no longer trust it.
The flagship writing result runs the other way. Noy and Zhang gave 453 college-educated professionals occupation-specific writing tasks, memos and press releases and short reports, and ChatGPT cut their time about 40 percent while blinded professional graders rated the output about 18 percent better. The weaker writers gained the most, so the spread between people narrowed. Two caveats. The tasks run twenty to thirty minutes, which is nothing like an argument built over weeks, and most participants submitted the model’s output more or less unedited, so the graders were partly rewarding raw machine prose. If the graders were rewarding the machine, then the quality gain is not the writer’s. True enough for that experiment, and it is the reason the study cannot carry my case. It also cuts the other way for the scolds, though. If competent readers, blind to the source, rate the machine’s prose higher than the human’s on the same task, then the thing they are detecting when they detect AI is not badness. That is the whole problem in one result.
Clock and quality moved together there. They came apart for the programmers. That is the same shape as the endoscopists against the cover-letter study.
I should say that I am keeping that study in here anyway (even though I throw out another detector finding a few paragraphs below on similar-sounding grounds). The part I use here survives either way, because it is not the speed that matters; it is the spread between what those programmers predicted, what they reported, and what a clock recorded, and that spread widened rather than closed when the lab went back for a second look.
Measuring all this validly, as we are amidst so much change, is really hard. Asking people how well they performed is mostly useless, though asking them what a task cost them still works. In a security study, programmers with an AI assistant wrote weaker code and felt surer of it.
The documented harms happen to people while they are still learning the thing. For experts the record is thin, with one exception: a Lancet study found that experienced endoscopists, after months of AI assistance, found about a fifth fewer growths when they worked unaided.
That result deserves more attention than it has received really, and it is also the strongest evidence against the argument I am making.
It’s an interesting complication, but we need to think a little about this finding. For example, adenoma detection is perceptual vigilance: scanning a moving image for a faint pattern, thousands of times, under time pressure, where the skill lives in calibrated attention that decays without reps. Turn the alarm on and the eye stops working as hard.
Nothing in argumentative writing runs on that kind of automaticity that I can think of. The closest thing to a test we have, a preregistered experiment on writing with AI, also a working paper, found the opposite: people who practiced with help wrote better afterward, unaided, than people who practiced with a search engine or an editor. Two studies pointing opposite ways, in different faculties, would seem to be a reason to go looking for more evidence, not a reason to increase confidence.
There is one study of professionals that helps my argument, and it is weak. Microsoft and Carnegie Mellon researchers surveyed 319 knowledge workers about 936 real uses of AI at work and found that confidence in the tool predicted less critical thinking, while confidence in one’s own expertise predicted more. So, expertise looked protective.
It is also a survey with no performance measure, taken at one moment, and its authors do not claim otherwise, which makes it pretty thin support, really.
In sum, AI damages learning when it replaces the effort that teaches, it can degrade a trained perceptual skill, and for expert judgment in long-form argument nobody has run the study. The cover-letter experiment above is a writing study, but it is not that one. If somebody runs it and finds writers dulling the way those endoscopists did, then this essay that I’m writing is on the wrong side of the argument.
Still a classroom and an op-ed page (and, and, and) still pose different problems, sure, and this controversy applies the classroom rule to op-eds by professors. But “expert” is doing real work in that sentence, isn’t it? :) I mean somebody deploying a skill they already have, not somebody acquiring one, and plenty of people writing op-eds are still acquiring. The line is a spectrum, not a fence, damn it.
The people who build and study these tools mostly agree with all this, which is the part of the fight that gets the least airtime.
The people selling the scan say the same thing about it. Announcing Substack’s own detection feature in July, Chris Best wrote that Pangram “can only detect whether AI was used to make the text, not whether great human care went into creating it.” That is the whole problem in the vendor’s own words, and it has not stopped anyone from treating a percentage as a verdict three weeks later. (The best comment under his post came from a writer named Monica Hebert: a score “cannot tell you who had the original thought, who shaped the argument, who rejected the bad language, who supplied the life.”)
The presence of AI does not prove the absence of a human.
MIT recently said a version of this out loud in August. Its committee on AI in education told faculty not to rely on AI detectors, for practical reasons: the tools do well on pure machine text and poorly on the real case, a student who used AI to outline or edit, and policing with them breeds distrust between instructors and students. The report came out the same week as the two detector screenshots this essay opened with. (That’s just one of the logics I used in making my course AI policy, by the way…a lot of this research I read over the summer when I was deciding how to proceed.)
Faculty are less sure than either side of this fight. Three researchers at Oklahoma State sat sixteen engineering professors down in focus groups last November and published what they heard four days ago. What they found was not enthusiasm and not resistance but something in between, which they call reluctant engagement: accommodation driven by perceived inevitability rather than principled conviction. The worries are specific and they are not about cheating. “They’re offloading their thinking process itself.” “The students think they know things they don’t.” One professor described what happened to debugging conversations: students used to arrive two or three levels deep into a problem, and now it is “here you go, I wrote this giant 100-line code, it’s not working.” Somebody said “I’m not going to grade Gemini and ChatGPT.” Somebody else said “you can’t prove it because none of the AI checkers are accurate.” And one line that nobody in the group had to editorialize: “my office hours are a lot less frequented now.”
Read that honestly and a good deal of it is evidence for the other side. Cognitive offloading is real, the debugging conversations did get shallower, and the office hours did empty out. None of that is confusion. What the study does not show is anybody catching it with a detector. The professor who noticed the shallow debugging noticed it in a conversation, which is the whole argument of this piece arriving from a room I was not in.
Two things about that study before I lean on it. The authors note that everyone on the research team was in the same college of engineering as the participants and that collegial familiarity may have shaped what got said out loud, which is worth holding onto, because inevitability is the cheap answer in a room full of colleagues. Conviction is a position somebody across the table can argue with. And the paper reports no enthusiasts and no one who had successfully redesigned an assessment. The only changes anyone described were defensive: weight the exams more, weight the homework less.
Matteo Wong had made the point in The Atlantic in May, under a headline that says it all: “America Has a Pangram Problem.” Basically every recent high-profile accusation of passing off AI writing, he found, started with that one tool. A horror novel pulled days before release. Articles in major newspapers. Prize-winning short stories. Significant chunks, in Wong’s words, of Pope Leo XIV’s encyclical warning the world about the dangers of AI. Three years ago ZeroGPT declared the United States Constitution machine-written and OpenAI shut down its own detector as too inaccurate to keep, so the tools have improved enormously in a short time. Wong’s subhead still puts it right: they are getting better, and they still aren’t good enough.
Now some evidence that cuts against the argument.
(I present that because I believe in understanding both sides of this and reasoning out why I am pretty aggressive in how I am thinking about the place of this technology in our society and, well, our epistemology. And to be candid, I really do hate that folks think I’m just putting my finger to the wind…no, it’s not that easy, and the effort I expend to be adaptable and nimble and all that, it is costly cognitively, not gonna lie, but away we go.)
A 2023 study showed detectors flagging 61 percent of essays by non-native English speakers as machine work. That was true of the 2023 tools. This June, researchers ran four current detectors over forty graduate theses that non-native speakers wrote before generative AI existed, and none of the four flagged a single one at the threshold the authors used.
Another study, of almost 80,000 peer reviews, gets summarized as: reviewers now read AI-ish phrasing as a sign of foreign authorship. Read closely, it found suggestive but inconclusive evidence that the old rating penalty against authors from non-English-speaking countries eased after ChatGPT arrived, less than the authors hoped, and the phrasing worry appears in a handful of interviews, where people describe a fear rather than a practice.
Non-native scientists pay a heavy tax. In a survey of 908 researchers, they reported spending up to twice as long reading a paper, half again as long writing one, and drawing revision requests to fix the English about twelve times as often as native speakers. A tool that cuts that tax is worth the most to them, so a rule against the tool costs them the most.
That is an inference. Nobody has really measured it.
Daron Acemoglu, working from the task-level evidence, estimates AI’s whole productivity contribution at well under one percent of total factor productivity across a decade. Goldman Sachs projects seven percent added to the level of global GDP over roughly the same span. Those are different objects, so they do not contradict each other on their face, but they still cannot both be describing the same economy, and no one can test the assumptions that seem to divide them.
A survey of about six thousand executives found nine in ten firms reporting no measurable effect on productivity or jobs, three years in. An effect of the size Acemoglu describes would hide inside the noise of the quarterly statistics, so we will not know for years, which is, yes, yet another conundrum.
Even so, the certainty has run a bit ahead of the evidence, and once it even invented some. MIT withdrew a celebrated paper on AI and materials discovery after announcing it had no confidence the data were real. The number spread because many smart people wanted it to be true, and, well, most people’s (including my own) reading habits work the same way if we are not conscious of them.
Dan Williams published the broad form of all this on August 26: most questions about AI are about labor, law, and institutions, and a technical pedigree does not answer them. His example is Geoffrey Hinton telling the world in 2016 to stop training radiologists. The models got better at the narrow task. Demand for radiologists rose for a decade anyway. Williams ends where this essay does: “nobody knows anything” is close, and we are all ignorant, but some are more ignorant than others, is closer.
Epistemic humility, especially in a time of massive change, is really fucking hard, especially when fear and uncertainty are involved. This is why I’m trying to give everyone making good faith attempts at this some grace.
To my thinking, two of the three studies that would narrow this are really cheap to do, and yet nobody has run them. A semester and a research assistant would test a current detector on native and non-native writers side by side, which the June study came close to doing and could not, because it had no native comparison group. A second could follow experts using AI on arguments they already own, which describes most serious use in the world and every professor in this controversy. The third, whether AI help closes the acceptance-rate gap non-native scholars face, is expensive, needs journals to cooperate, and would take years. I know some underemployed graduate students who could staff them, someone give us some money. :)
Some of that is ordinary friction, of course. Human-subjects review on student writing is (very) slow, grant cycles are slower, and a detector market that iterates every six months makes any result stale by the time it clears peer review, which is exactly what happened to the 2023 detector finding above.
Still, friction does not explain why nobody is even trying, and the rest is (yes, here it comes again) incentives. Nobody in the fight really earns anything by closing the question, and every answer would cost somebody something, and some more than others. Ambiguity is cheap, so ambiguity is what we keep.
(Sigh, it’s been a theme in so many circles of late…and I hate to tell you, but there’s a real cost that we all pay sooner or later.)
Kustov put the same dynamic in one line about himself this month: a few months ago some people called for him to be fired over his views on AI, and others privately messaged him asking for his setup, sometimes the same people. Public condemnation and a private request for the recipe, from one mouth. That is not hypocrisy so much as an accurate read of what the norm currently rewards, as well as how confused everyone is about this new…thing.
That story is easy to tell and hard to test. If you’ve been around here for a while, you know what question I am going to ask next: what falsifies it?
If the studies get run in the next two years by people with stakes in the answer, the incentive account is wrong and friction was the whole story. And if somebody runs the expert version and finds writers dulling the way those crazy endoscopists did, a mea culpa would be in order, not a paragraph explaining why my case is different. (When I wrote about sociology’s audit, the strongest evidence was the questions nobody starts. The same sign points here, at a field that includes me…)
And then there’s the last question, to bring it back full circle: who has caught whom doing what, then?
That’s the weird thing about Type I and Type II errors. We always complain about the possibility of a missed test result that should have signaled red when it gave us green. But what about when we get a red when it should have been green (especially when that red says absolutely nothing about the quality of what’s being tested)?
Currently, a test exists that needs no guess about anyone’s motive, right? So you just start by asking whether the response followed the rule or the person. Hausmann broke a policy, Stankova broke none, and yet the response treated them alike.
Two weeks after Hausmann, Stanley Druckenmiller published a Wall Street Journal op-ed attacking the Treasury Secretary and told a reporter, when asked, “of course I used AI,” adding that he writes everything with it now for the same reason he uses a calculator for arithmetic. Axios covered that as a billionaire flex. The Journal’s opinion editor called AI a fact of modern life and defended running the piece with no disclosure at all, because it reflected the author’s own argument and because the author had, in his words, the standing and credibility to make it.
“Standing and credibility.” He said it out loud, as the reason.
Daniel Drezner, watching both cases, put it plainly: what reads as an ethical no-no for an academic reads as a neat efficiency trick for a billionaire. Same tool, same month, both aimed at the same administration. One of them is a Harvard professor and one of them runs a family office.
The accusation also pushed something aside. A Berkeley Law professor’s post of Stankova’s score drew close to three million views, and as one write-up put it, the story stopped being about calculus at some point. Her op-ed warned that admissions essays are now machine-written, and then the op-ed got tried as machine writing instead.
(I would also invite the people who run detectors over other people’s op-eds to publish their own help: the research assistants, the copyeditors, the journal editor who suggested rewrites to the abstract, the friend or colleague who suggests fixes to their transitions. When are we going to develop a test for that? :) )
Drezner, who noticed that double standard before I did, still thinks it is worth keeping: a professor’s whole claim on public attention is the ability to take a body of research and distill it for people who have not read it, so if that is the job, readers are owed “the work.” Then, however, he also names the thing that complicates his own position, which is that plenty of professors already do not write their own op-eds, and nobody has ever demanded a disclosure line for the research assistant or the comms office, which cuts against his points.
The even harder version of the objection landed on August 22, nine days after the FT note. Adrian Vermeule found an SSRN paper attacking his own book whose footnote disclosed that it was “developed and drafted in extended dialogue with Claude,” which had helped “to structure the argument, to draft and revise the prose, and to locate and verify sources.” Lawrence Solum had already reposted the paper with a “Highly Recommended!” and no mention of any of that. Vermeule ran it through Pangram and got 100 percent.
Then he did the thing almost nobody in this fight does. He argued with the actual paper. Several paragraphs on whether the classical idea of determinatio is smuggled Schmitt, with Finnis and Aquinas brought in to show it is not, and a demonstration that the paper had confused “Schmitt said there is irreducible discretion in specifying legal rules” with “any account of irreducible discretion is Schmittian.” He hedged the detector while he was at it, noting he had heard it sometimes yields false positives and that a 100 percent score need not mean every word came from the machine.
He is also right about the thing that should worry me/us most. Structuring an argument, drafting the prose, and locating and verifying sources is not assistance around the edges of scholarship. It is the scholarship. And if the machine verified the sources, then nobody verified the sources, which is a broken compact and not a style complaint. Disclosure did not prevent that. Disclosure described it, in language vague enough that no reader can tell who did what.
And then, having proven he could take the argument apart on its merits, he announced he would not be responding to machine-generated texts in the future.
That is the whole difficulty in one paragraph isn’t it? He earned the right to dismiss the paper, but then dismissed the category instead. He’s also right that the volume of machine-written stuff is going to just increase and become even more of a glut on the inference system.
So, again, I get it, we all feel like we have to draw lines right now. I am drawing one by writing this. But in fairness to everyone out there, we’re all guessing, we’re all human beings trying to figure what to do with….this.
I think a lot about the arguments surrounding graduate education. The professor who sincerely asks what is lost when a graduate student never writes a bad first draft is defending the student. Lord knows I certainly wrote some absolute crap in grad school. I mean, one of my reaction papers was on the pseudoerotic machinations of those arguing against rational choice in an IR field seminar. Those were the days.
So, how do we deal with the argument that the years of doing a thing badly teach you to do it well? Especially when a machine exists that spares you the bad years may spare you the learning. People at a university produce prose, and they produces the next people able to judge prose, and the second job can suffer while the first one booms.
Seems to me that worry deserves a much more extensive conversation than a screenshot on multiple fronts.
As for the rest of it, a lot of this fails the logic test to my eye, and it’s exactly in the direction the incentives predict, especially the case since academia has its own rather obvious share of egocentrism and other Cluster B-related issues. This is not to say, by the way, that the opponents of AI should all be placed in the status preservation box or receive a diagnosis, that’s not what I am saying.
However, it is a question I am guessing that not all of them will ask, and they should.
It is also the case that, well, I am merely a first generation college student who did ok, perhaps I see these things differently than some of my more well-heeled colleagues.
But even within all that, it’s also the case that most people rarely feel themselves defending turf, and the research on motivated reasoning says the ablest of us hire fine, only the best, lawyers inside ourselves to argue for our conclusions, without ever seeing them for what they are.
A fight this loud, running on evidence this thin, has to be at least somewhat about standing; of course, the people fighting have reasons other than status.
But, is that just the screenshot move in a much nicer Sunday dress? Replacing your text is 70 percent machine with your objection is 70 percent status anxiety uses the same logic I have spent the essay refusing: a claim about where something came from, standing in for a claim about whether it is right.
Run it on me too, please! My own status defense, if I am running one, would look like this: this guy whose Substack writing uses these tools, arguing that the norms should stay loose because he sees it as a force that is a rising tide that raises all boats, and finding the studies that say so more convincing than the ones that don’t.
Nope, can’t rule that out.
What came to mind for me in writing this was remembering how chess dealt with it twenty years ago. In 2005 a tournament let anyone enter in any combination (person, computer, etc.), and the winners were two American amateurs (rated around 1685 and 1400 for the chess nerds) running three ordinary home computers, one of them borrowed from a father. They beat grand masters, even those working with machines.
Kasparov’s summary has since become a law with his name on it: weak human plus machine plus a better process beat a strong computer alone, and beat a strong human plus machine plus an inferior process.
That is status threat (that we all feel) in its purest form. The quality of play went up. When everything changed, what really changed underneath was who was the resulting authority.
Chess responded to this by defining the events. Computers are banned at the board, but permitted in correspondence play. Where the ban applies FIDE screens statistically, with a presumption of cheating above a threshold on Ken Regan’s move-matching. Point is, in chess, they seem to have figured it out, but chess also has something to detect against. Move-matching has an observable ground truth, the machine’s best move, computable for everyone. Pangram, does not have a ground truth for the thing anyone ultimately cares about, which is whether the writing is any good, and its makers say so themselves. Scoping a valid instrument to a bounded event is a different act from firing an invalid one at everything in sight.
Remember that the Financial Times had a rule. The Standard did not. The Journal declined to write one. Three events, three rulebooks, one (black-box, proprietary) instrument fired at all of them.
Further, notice something else chess did not do. It never banned machines from chess. Preparation, analysis, training, all of it stayed wide open.
Of course, professional academic writing is different (and I still think of it differently as well, if I’m honest, but I think there’s an inevitability to AI becoming a part of that too as it gets better and better, at least on the empirical and analytic side of things). But that’s because an academic research paper is written over months, in offices and at kitchen counters, with research assistants and copyeditors and the kind colleague who quietly helps you fix your messes.
Amidst all that, to me, it’s apparent that the answer is not an AI-free university. Not in the least.
Progress through that gauntlet of becoming a certifier of knowledge then, has to be some number of bounded moments where the computer is set aside and someone is asked to perform: the exam, the in-person write, the version history, the defense. After seeing what’s on the page, then test the person and what they know. My experience so far has been that the student who used the machine and learned can usually show you. One who didn’t, usually can’t.
I said earlier that the line we’re dealing with is a spectrum and not a fence, and I really believe that. You may still have to put a fence somewhere, but it is absolutely better to admit you chose the spot, and why you chose that spot, than to pretend the ground chose it for you.
But all this costs so much. Time, especially. I firmly believe that AI is going to increase the time we spend with our students, not reduce it. An oral defense is a performance under time pressure. Somebody should probably measure the efficacy of that too, and, yet, nobody has.
To close this way too long of a piece, which is just an attempt to calibrate all that is going on right now and it might have broken my brain in doing so: a detector can (mostly) tell you if a text was edited by an AI.
Whether the text deserved your attention remains a human judgment. That judgment sits where it always has: with readers willing to do the slow thing and read.
I read a lot. Too much even. When I read, I prefer the best produced pieces, the accurate pieces, the well-integrated pieces, to those less so. I read some AI generated text, and I read non-AI generated text. There’s going to be more of the former and it’s going to get better and better. What we do about that, and how it affects our episteme is probably the question of the next five years.
But, hey, what do I know? I’m just some dude writing a substack.
If you want these to keep coming, you can subscribe, and there’s always the coffee button if you’d rather put a price on it you can see.
All your support is appreciated.
Earlier in the stack:
“Different Ways of Knowing” Means Two Things. Did We Sign Up for the Wrong One?, where the provenance-versus-content argument started, including the detector run on my own work.
An Election Can Defund a Field Like Sociology. It Can’t Replicate One., on what crude auditors do when the careful ones won’t.
Half the Building Is Still Standing (Or, On Nature’s SocSci Replication Piece), the replication numbers, and the move this essay tries to run on itself.
Why Are the Humanities Missing the AI Moment?, where the sophistication amplifier first showed up.



"The presence of AI does not prove the absence of a human." That seems to distill this long post.
From a reader/user perspective, I don't care who said or did something, unless it is based on their status or authority. The question is whether it is correct; whether it conforms to reality (even good fiction conforms to reality in its own way).
But from a credit/reward perspective, it does matter who did it. Machines don't get PhDs; people do. That works the other way, too: if AI detects better, give me the AI, not the PhD. The pay and prestige people get is related to what they get credit for.
The mixing of human and machine muddies the latter, not the former. Business deals with this as an increase in productivity. That is not so easily done where quality matters as much or more than quantity.
Great and thoughtful essay. There's an unstated bias in how writers have responded to AI. Most people with successful Substack accounts (strong presence, lots of well-cited journal articles, books, etc.) are probably both good writers and take to it easily. To them, banning AI is like a shop rule that limits competition. While they point to 'AI slop,' there's an equal or larger amount of human slop. We just don't have to look at it often, and eventually people who aren't gifted writers quit trying to write. AI is reducing the cost of writing well, which threatens those who can do it well without assistance. Another analogy (and admittedly, most analogies are lame) is the autopilot on airplanes. Useful, no one much seems to object that their jet plane flight was almost entirely managed by computers and electronic controls. Most of us also do want our pilot to be able to take over in an emergency. (I wrote this with Grammarly checking my spelling -- I don't touch-type, and I've always been a horrendous speller).