16 Comments
User's avatar
Dawson Eliasen's avatar

I think AI detection is and will remain a thorny issue, because of false positives but also because there is no guarantee whatsoever that Pangram will stay as accurate as it is now long term as models get better. But I think the way Substack implemented the integration was about as good as you can hope for and is a good thing, for Substack and for the internet in general. I don't think people have the right to share something that's AI generated or even AI assisted and pass it off as human written, especially when you're invited to pay for that content. But I have also been wondering how much difference Pangram really makes? Lots of people say they can spot AI writing easily (I certainly feel like I can). But there could be crazy survivor bias at play. Is Pangram actually successfully detecting the especially good AI generated writing, or is it just capturing the slop, which makes up almost all of the volume? Are there going to be writers "exposed" as using AI, where previously no one had suspected it? Does any of this even matter if a week from now Anthropic releases a model that completely fools Pangram? Does Anthropic et al even care about fooling detectors? It's also been interesting to see the folks who admit in their disclosure to using AI assistance heavily in their writing process even if the content isn't completely AI "generated" per se, but I'm not sure if that distinction matters? If you're using an LLM to guide the thinking--even if you're working to make sure the prose is human--I'm not sure if I want to buy in to your stuff, even if the price is only my attention.

galactic-beyond's avatar

It is cool that we can detect text that is obviously entirely AI generated[0], but the usefulness of this depends on how strong the correlation between AI-generated-ness and undesirability remains/fluctuates over the years and decades. My personal fear is that entire writings can be soft-censored by compromising either the detection-tool (e.g. Pangram), or the AI lab itself (e.g. make your models write English in a way that is similar to $groupXYZ, where $groupXYZ is some out-group -- this seems likely given how involved the US government already is with these labs). Long term, what we actually want is better spam-filters and better search-engines (and ideally, their operation/algorithms would be inspectable).

[0]: Basically, this move by substack is the pro-social version of the NY State ban on 3d-printed ghost-guns (an ineffective law, that harasses owners of 3d printers, and potentially criminalizes owners of modified printers). I appreciate the integration of Pangram, more as an expression of values and expectations, than as a long-term solution to the problem.

Louis Dormegnie's avatar

> But I have also been wondering how much difference Pangram really makes?

I think it will make a difference (1) for people who don't feel like they can reliably spot AI writing and want to avoid it; (2) for those who can, but want to see whether the entire post is AI-generated (I don't mind a portion of a post being AI-assisted/-written, but often click away if the thing lacks a human touch after ~1min of reading); (3) for whoever intends to make a decision wrt to the author based on their writing (as an investment analyst I won't contact somebody who writes a finance-related post entirely with AI); (etc) a plethora of other niche needs that I couldn't possibly write out.

> Is Pangram actually successfully detecting the especially good AI generated writing, or is it just capturing the slop, which makes up almost all of the volume?

"Especially good AI generated writing" is probably a very marginal issue, and probably blurs the culpability intensity of the author for most readers. I don't think a fully-AI-written post can be "especially good", and I've read thousands of posts on this platform and have sought to cultivate a high quality feed. What can be especially good is AI-assisted writing, where the right-tail portion of the quality comes from the human idea at the core/onset of the article.

When Substack announced the Pangram partnership on the Substack Post, many commenters raised the idea that AI can help an archetypal person I'll call "great-ideas-but-cannot-write-them-out". I believe that group is either much smaller than those commenters think it is, and/or the few who fit the archetype will use their great-ideas personality trait to ensure the writing doesn't end up reading like a who's who of the Wikipedia page "Signs of AI writing".

> Does any of this even matter if a week from now Anthropic releases a model that completely fools Pangram?

The problem with this argument is that we are now about two years from the time where LLMs became good enough to assist or even write posts in full, and yet their tell-tale signs are still there in most cases. So, although it's definitely possible that AI models start exhibiting enough idiosyncrasy in their default responses that their writing become less distinguishable from human writing, so far it hasn't been the case which makes Pangram a de facto great-to-have.

I'm convinced that most people who write the entirety of their posts with AI either don't see a problem with it or are unwilling to put in the effort to remedy common signs of AI writing (which can be prompted effectively). It all comes back to the idea that fully writing posts with AI is a combination of laziness and a lack of taste: if you're doing it then by definition you are the kind of person who doesn't see a problem with it, but readers always have the final say.

Dawson Eliasen's avatar

> When Substack announced the Pangram partnership on the Substack Post, many commenters raised the idea that AI can help an archetypal person I'll call "great-ideas-but-cannot-write-them-out".

Pretty much every idea seems like a good idea until you try to write it out. That's the problem. Writing is thinking. You find the gaps through the process of writing. You can only be sure if it's a good idea if you've gone through the frustration of writing it out in a way that someone else can understand it and maybe even be convinced of it. If you have an idea and you ask an LLM to help you write it, you completely miss this process. The LLM will just serve you up something that satisfies your poorly justified belief in this idea you can't even articulate to yourself. There is nothing innocent about using an AI to write out your ideas.

Phone Free Will's avatar

This is an excellent point. Excellent.

This is what happens to my ideas.. they change, get better or die in the writing.

Louis Dormegnie's avatar

> Pretty much every idea seems like a good idea until you try to write it out.

I'm not sure that's true, but my counter to this sentence lives in the edge cases, in your ("pretty much every"), so I mostly agree with your comment.

I think it's possible to come to a great idea through rich discussion, when all sides bring whatever is needed for the idea to surface. That doesn't mean it wouldn't be refined through writing, but the act of writing itself isn't a sine qua non for a great idea. I do believe some people fit the "great-ideas-but-cannot-write-them-out" archetype, as I think I know a couple in real life, but it's such a self-congratulatory cop-out that it's misappropriated in almost all cases.

Dawson Eliasen's avatar

I agree with everything you're saying here

> "Especially good AI generated writing" is probably a very marginal issue, and probably blurs the culpability intensity of the author for most readers. I don't think a fully-AI-written post can be "especially good",

To clarify, by "especially good AI writing," I meant the top 1% of AI generated or heavily AI assisted writing, not especially good writing, by normal human standards, that was actually written by AI.

Green-2.99's avatar

All good points. Meanwhile though, I get angry at the way most of my colleagues fail to think critically about the fundamental nature of differentiating human versus AI writing, because it is evident from how they talk about it that they think of the detection like magic, i.e., a black box that can "reveal" the correct answer *even from short samples*, "because advanced technology" end of sentence. Whereas it is actually blindingly obvious to a brain that actually thinks that it is literally impossible for it to work that way on short samples that don't happen to be loaded with stereotypical tells. That's because there is rampant nonspecificity and limited data. Every speck of data that is present in the sample is neither specific to AI nor specific to humans. And yet they think that some software can certainly reveal the true answer Because Technology. These same people must somehow believe that I can type one sentence of highly nonspecific symptoms into a medical diagnosis software and trust which answer (which disease) it spits out. "I'm feeling kinda tired and my nose is running." Aha! I can give you a diagnosis! And it must be true because I used a computer to do it!

Dawson Eliasen's avatar

Yeah, the most pernicious quality of LLMs is the way they will give you an answer confidently no matter what. If they were more likely to say something like “there’s not enough information here to come up with anything useful” they’d be a lot less dangerous

nutrient  poeisis's avatar

Slop in, slop amplified, slop out.

We did it to ourselves by allowing immaturity to control the narratives.

Green-2.99's avatar

Iconoclasm is a scorpion on a frog's back midriver. It can't help itself. The cool rebel who's going to stick it to The Man, the disruptor who's going to lauch a startup and get rich. You could mention Chesterton's fence to them but then you'd just be one of the sheep who follow The Man.

nutrient  poeisis's avatar

It’s not a singularity for popularism, it’s a group of conversation and no leaders.

We have it backwards, the individual is a group not a passport.

SusanA's avatar

Dear Eric, I really, really enjoyed today’s missive. Each of the subjects is very interesting; and you told me just enough, but not too much. I, as many, suffer from information overload (self-inflicted), so I found your 2-3 paragraph summary to be just right. I can look up more if I desire. I particularly enjoyed the ones about number of words spoken today as opposed to years ago and the one about teaching reading at an earlier age. Oh yes, and especially the one about whether or not parents - a thousand years ago - loved their children as much as we do ours. Of course they did!!!

Green-2.99's avatar

Scorpions eating frogs, leopards eating faces ...

On Value in Culture's avatar

Number 5 on fewer words - yes, philosophers and linguists have been saying that for a number of years. Except for maybe people who speak more than one language. Anecdotally, we speak to more people and use more words in each language from the back and forth between the two (or three). Would speaking to oneself count as speech? I know many people who do it.

SusanA's avatar

Dear Eric, I really, really enjoyed today’s missive. Each of the subjects is very interesting; and you told me just enough, but not too much. I, as many, suffer from information overload (self-inflicted), so I found your 2-3 paragraph summary to be just right. I can look up more if I desire. I particularly enjoyed the ones about number of words spoken today as opposed to years ago and the one about teaching reading at an earlier age. Oh yes, and especially the one about whether or not parents - a thousand years ago - loved their children as much as we do ours. Of course they did!!!