Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

You definitely can rule out the general case a priori. If the problem were possible, for every text there would be a unique provenance label “human” or “ai”. But since humans and machines have both written many texts, it is not possible.

As an example, you could imagine a giant lookup table that deterministically mapped every text ever written to “human” or “AI”. You would very quickly run into situations where the labels conflict for the same piece of text.

The data is statistically inseparable which makes it impossible to classify from text alone.

 help



That just proves that perfect classification is impossible. Classification doesn't need to be 100% accurate to be useful.

It’s worse. If the data was separable in this way, you would equally be able to train an AI to mask those signs.

IIRC, the big names in LLMs have no real interest in cloaking the LLM-nature of the text, Google adds deliberate watermarks to text, OpenAI developed a watermark for text but reportedly arent't actually using it.

Considering [1], I’m going to challenge that their techniques are currently even mildly effective. Given the absolute academic malpractice these papers are pushing, I’m calling BS; while they want to watermark it, they clearly aren’t actually able to. For images. Which are drastically easier than text.

Their interest is irrelevant in the face of technical impossibility. And that’s before you get into other people who don’t care and will just build adversarial tools to bypass the attempted watermarks. It’s a losing useless battle. Google and OpenAI engage in it to try to catch competitors when there’s a lawsuit or to try to clean their datasets clean.

But it’s absolutely unusable for something like “did someone cheat”.

[1] https://hackerfactor.com/blog/index.php?/categories/1-Image-...


> Their interest is irrelevant in the face of technical impossibility.

I'm responding to "If the data was separable in this way, you would equally be able to train an AI to mask those signs.": yes, if you wanted to you could, the big names clearly don't consider masking to be a priority.

> But it’s absolutely unusable for something like “did someone cheat”.

This is the one case where I'd most expect it to succeed:

I suspect most of the people who do want to cloak-to-cheat, don't have the skills to do so; I also suspect most of them are so unaware of what they don't know that they won't even ask an LLM to write cloaking software for them.


You’re overthinking it. There will be end products specifically for this and they’ll be trained by their peers / the company.

> But since humans and machines have both written many texts, it is not possible.

Maybe you meant "many humans have used AI when writing texts"? Your stated reason that they can't be separated because there are many texts of each kind is nonsensical, you clearly need to supply more reasons than "there are many".




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: