Hacker Newsnew | past | comments | ask | show | jobs | submit | cainxinth's commentslogin

Take a random essay and add in a bunch of the phrases that LLMs love like “load-bearing,” “crucial,” structural,” and “woven,” and then submit the original and the edited version to an LLM and ask which is better. It will choose the second one virtually every time. They have ingrained biases that associate those words with good writing and arguments.

Sometimes I wonder if there’s just one guy somewhere who loved using the word load-bearing, all his papers got trained on, and now he can’t write anything without being assumed to be Claude.

More likely, it crawled-thru Construction permit listings.

The prose equivalent of Artgerm (a comic cover artist whose style looks to have heavily inspired a lot of AI art).

This is why using other LLMs as scorers for benchmarks and evaluations is such a bad idea, they'll have preferences you can't anticipate and won't understand immediately.

The idea that that LLM reliability or bias can be solved with more LLM is... infuriatingly persistent.

Isn't this basically the mythical man-months LLM edition?

I looked closely. They are going the wrong way. Still the best one I've seen by a lot.

Probably some temporal aliasing on your device. The SVG elements are 100% rotating clockwise.

Still a flaw--the SVG should account for how different client devices will render it! Existential doom averted, for now.

Let me know when AI patches this bug in the wheels on all the vehicles on the highway, AND my box fan when it's next to the light with the bad PWM dimming. Then we can resume our existential doom.

Looks fine on safari iPadOS 26.6.1 wheels rotating forwards. I’m a skeptic but man, this actually has me revising my opinion a bit.

> Pompous things like "Your keys, supercharged" or weird yoda-speak stuff like "searches the app remembers"...

It's copywriting. They fed these models the internet, which is loaded with it.


And turns out, the "frontier" labs have no human oversight of the training data going into these models... Explains so much

There is no realistic path to human oversight for the vast quantities of data these models are trained on. Imagine the cost of having every Reddit comment ingested human reviewed. Insane.

Somewhat related, I’ve long wanted an ereader app that color coded the characters in a novel.

> It's exactly like vibe coding but for computer tasks.

That’s precisely why I don’t use it. LLMs are still not reliable enough for me to trust them with my actual system.


This is how I proofread. I play a screen reader while I read my work. The brain does some autocorrecting that you don’t even notice and occasionally skips over a typo when you read it without the audio.

Has anyone found a useable fake review detector since Fakespot and Review Meta went down?


I also tracked my clothing for three years, but not to evaluate expenditures. I recorded what I wore while cycling along with the weather conditions that day and notes about how my clothes performed. When I want to go for a ride, I check the weather and then check my spreadsheet for a day with similar conditions.


I'm fascinated by people who get Type 1 fun from reading the classics. I get type 2 at best.


I virtually never click anything texted to me (other than personal stuff from friends and family). If a bank or a shipping company texts me, I go to their website and look up the information myself.

Like the author of this post, I also have a better than average eye for spotting scams, but it’s foolish to assume you’ll be right 100% of the time.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: