Take a random essay and add in a bunch of the phrases that LLMs love like “load-bearing,” “crucial,” structural,” and “woven,” and then submit the original and the edited version to an LLM and ask which is better. It will choose the second one virtually every time. They have ingrained biases that associate those words with good writing and arguments.
Sometimes I wonder if there’s just one guy somewhere who loved using the word load-bearing, all his papers got trained on, and now he can’t write anything without being assumed to be Claude.
This is why using other LLMs as scorers for benchmarks and evaluations is such a bad idea, they'll have preferences you can't anticipate and won't understand immediately.
Let me know when AI patches this bug in the wheels on all the vehicles on the highway, AND my box fan when it's next to the light with the bad PWM dimming. Then we can resume our existential doom.
There is no realistic path to human oversight for the vast quantities of data these models are trained on. Imagine the cost of having every Reddit comment ingested human reviewed. Insane.
This is how I proofread. I play a screen reader while I read my work. The brain does some autocorrecting that you don’t even notice and occasionally skips over a typo when you read it without the audio.
I also tracked my clothing for three years, but not to evaluate expenditures. I recorded what I wore while cycling along with the weather conditions that day and notes about how my clothes performed. When I want to go for a ride, I check the weather and then check my spreadsheet for a day with similar conditions.
I virtually never click anything texted to me (other than personal stuff from friends and family). If a bank or a shipping company texts me, I go to their website and look up the information myself.
Like the author of this post, I also have a better than average eye for spotting scams, but it’s foolish to assume you’ll be right 100% of the time.
reply