Hacker Newsnew | past | comments | ask | show | jobs | submit | Springtime's commentslogin

Even Mythos can hallucinate[1] with a codebase to analyze. The argument I often see is that humans are fallible, too, but the issues we're seeing in this topic are about those who should know better—including in senior positions—heeding LLM responses/advice over human and concerningly not actually thinking about things at all. Ie: they're not being treated by many as just useful tools/tentative feedback but as authoritative answers/solutions.

There still needs to be critical thinking involved on the human side, even if there's a high rate of accuracy in certain dimensions.

[1] https://news.ycombinator.com/item?id=48434824


That's a bad example: it did not hallucinate about the codebase (which it had access to), but about a rule outside the codebase (which is not in the context, I assume). Yeah, missing requirements can cause it to rely on things it thinks are true. Same as humans do all the time when they lack context.


Christ, no, humans dont makr the same mistakes. And even mire importantly, those who do similar ones, end up not being trusted.

We dont have such a campaign around people as there is about ai.


It seems a key contention of theirs is the possibility that rare books are being destroyed this way, yet the things they cite don't seem to suggest this (based on their paraphrasing), they just throw the following at the end to make it seem like it's occurring to irreplaceable books:

> You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal.

Is there evidence of this? Since otherwise they could very well be describing what is only occurring to in-print or non-rare books. (This is a genuine question since their post doesn't shed any light on it.)


There is zero reason to shred 18th century books. Any such books are out of copyright.


They are, but using the same process for all books is simpler and cheaper.


Yes, that is what the person you are responding to is saying. That's why they are questioning the assumption that this practice extends to 18th-century books that are solidly in the public domain.

If it is happening, it is an outrage. However, the 404 article doesn't actually provide any evidence of this; it just connects the shredding of digitized books and the digitizing of rare books to the assumed shredding of rare books, which isn't necessarily happening.


Though they're not shredded simply due to copyright, they're also shredded due to cost, speed and the quality of the scans.


Just sociopaths doing sociopath things. Bugger civilization.


Is it the UI per se or the site that is requiring local storage? Since the examples don't load with cookies blocked, nor do any of the links/buttons work without it.

Ordinarily I wouldn't comment on this (eg: for the plethora of React-using sites that haven't tested without it yet work fine when manually bypassed) but in this case it's intended as a drop-in UI.


> with all the GPT hype of late, why havent we seen a single GPT yet capable of building a browser engine from scratch supporting the last 20 years of html, css and js specs?

Macsurf is a browser project for Mac OS 9 leveraging LLMs[1] to do this. Tbh it makes sense for this given how time-consuming (absurd? Though the fact it exists tickles me) it'd be otherwise for essentially a single person to support the scope of web tech it does for such an incredibly niche userbase.

[1] https://news.ycombinator.com/item?id=48339534


I looked at the comments here but it doesn't appear to be an accurate characterization of the discourse. If anyone has links to such comments shaping the discourse from before the parent posted I'd be interested.

There's one top-level comment that's just an opinion of censorship of Qwen models (which might be argued to be politically adjacent if one assumes the censorship is referring to sensitive political subjects) but no specifics and there was a rebuttal reply anyway, while an unrelated top-level comment has a child reply believing China models are commoditizing AI to unseat US based models (could be argued a whiff of politics but also just competition) but it was still a positive impression.

Then I went to archive.org to check the earliest crawl of this page (prior to parent's post), but nothing different. So I enabled showdead in my settings to check dead-marked posts and there's one post near the bottom talking about how Chinese models are rarely ground-up foundational models, while another dead comment had an angle about it's a geopolitical strategy (neither of which had replies). Ironically some of the dead comments are praising Qwen.


The comment timestamps already show how old the posts are[1]

Putting politics aside, technical chatter is scarce in large‑model discussions here.

Take the K3 release as a case study: out of more than 1,200 comments, how many actually delve into topics like LatentMoE, KDA, or attention residuals? And how many are offering guesses about per‑head Muon or SiTU?

Users interested in this should look at imageboards

[1]: https://news.ycombinator.com/item?id=48966211


What's weird is with Firefox for Android it's so difficult erasing a site from being remembered. Once visited, in my experience, even after deleting the history entry and last closed tabs item it still auto completes the domain in the addressbar (when it didn't before) and the only workaround is a full history wipe (since the Android version offers no granular timeframe like the desktop version).

So if accidentally clicking some link from some other app that auto opens the default browser it's a PITA to get FF for Android to forget about it.


Android is a wasteland, not surprised


> The current picture is that there is an emerging class distinction between writing (and writers) that use genai vs. writing that does not. As soon as the "this sounds like an LLM" allergy kicks in, the writing instantly gets relegated to a low-status bucket in the reader's mind.

It's not a purity test, it's as the author is communicating they don't care whether the reader has any signals of what is accurate vs inaccurate information, which puts the burden of investigating how much is accurate on the reader at every step when there's some minimum expectation that should be an author's role (outside of topics where there is some expectation of ulterior motives/biases and one would naturally engage more critically minded).

When people complain here it's more often than not when an article has no disclaimer about AI use or what has been human-reviewed, so the burden again falls on the reader who is now even more skeptical. That is more to ask of a reader than when it's coming from say a known expert and the reader is receptive to engage and learn.

That's the reason tired cliches and turns of phrase (overused by LLMs) have become a heuristic for whether to pay attention, because it's a sign that there's some unknown quantity of of the article that hasn't had human review and it's easier to put in the bucket of 'maybe worthwhile but would need a fully human analysis of this' or just outright rejection (as we've seen from comments).

Edit: I see a sibling comment has raised the same observation.


I'm afraid I don't understand this - can I ask what was the sibling comment? Maybe that will help me triangulate.


This[1] comment, mostly in terms of not knowing which parts are generated and which not when not disclosed, which puts an added burden on the reader to assess.

Like, if some non-controversial article makes a statement about something technical (where one's guard isn't already raised) but you've observed signs of LLM use (without any disclosure of to what degree) then instead of thinking it might be an interesting thing to follow-up on or remember one might be thinking instead 'is this something the author themselves understands and has reviewed for accuracy or slipped in by the LLM' and other such distractions (and legwork if wanting to try and fact-check such things on the spot).

It goes from having a perhaps pleasurable, educational read to questioning and being more skeptical/cynical about the material. HN's guidelines meanwhile encourage good faith engagement, which is challenging.

Edit: corrected permalink (had accidental extra digit).

[1] https://news.ycombinator.com/item?id=48888422


I wasn't aware they operated as baseline Bluetooth headphones by default on other devices, so even having read the readme it wasn't obvious.


How else did you think they worked


I assumed, like Apple's Airtags (which use Bluetooth LE), that they used some proprietary pairing that's incompatible with non-Apple devices.


Airtags can be scanned with your non-Apple smartphone's NFC reader which gives you contact info for the owner.

Beyond that, all tracker devices are proprietary. The Tile/Chipolo network only works with Tile/Chipolo; Samsung's SmartTag only works with the SmartThings app which is exclusive to Samsung phones; etc.

I just don't think the comparison is apt


New chipolos work in both Apple Find My and Android Find Hub. Most third-party tags seem to support multiple platforms now.


Warning: many don't support both at the same time (setting up for one disables the other until reset)


Loads of 3rd party tags that work on the other networks though.

That was the point they were making, not that they had some NFC extra-functionality


What are some examples of 3rd party tags? I know the Find My network itself has been hacked around with for a 3rd party tag. I've always dreamed of an open-source tag that integrates with all the major networks so I can put a little tag on my cat's collar and track his travels.


Ugreen tags work with both Google's Find network and Apple's FindMy network.


Do they work with the network or with the app?

If it's the first, that seems like they would be a vastly superior offering to AirTags, especially outside America, where people aren't so Apple-crazy.


I don't know, I have an android phone. I just know it says it's compatible with the airtag network on the box, so I suspect it's the first. Ugreen doesn't even have any specific app for these, so I assume it goes through the primary setup method.


Well there you go. I guess its not so proprietary after all.


But its still what he thought...


This angle is touched on in the article, as the words on a page script vs Robin's performance of it which is drawing from his unique human experience which would have been different from another actor (the lived experience throughline the article is making, not necessarily the experiences being described in the script, mind).

I do think though that article is a bit nebulous in parts, in the sense that articles and books we read are also just words on a page and its those mediums LLMs are similarly using, which is why I think they attempted to morph into a point later in the article about the acted performance.

I still get the gist the author is trying to convey though, in that through lived experiences we crystalize and hone in on the things that matter which allows us to have actual first-hand opinions rather than just second-hand ones from others. It's those first-hand experiences that are often most valuable to others and drowning them out in an avalanche of either stylistic or wholly generated slop makes them more difficult to find.


Are you referring to this part?

> Where did his performance come from? How did those moments take shape? Only the actor could tell you, and actually, he probably couldn't. It was sensed more than it was consciously considered, but the alchemy required his lived experiences

That just seemed to me like a cop-out - what exactly is it about his lived experiences that made him well suited to effectively convey the experiences of a made up man written up by two others? Because it's clearly not "write what you know".


Was about this part:

> Step out of the story and examine the acting. Robin Williams was given a script. Any other actor could have been handed that script, but ZERO other actors would have performed it like that. The script has all the words, but he brought the words to life. What's more, he did so by drawing on his own life.

The part I agree with in the article is that a Robin Williams performance is something unique that is an amalgamation of his lived experiences (whether or not related to specific scenarios he's portraying). All actors are drawing from things differently (even their own meta acting experiences). The part I was agreeing with you about is the article's premise being based on an analogy from a film is harder to sustain.

They're not wrong though that reading about some experience is different than experiencing it first-hand and the value that can bring, it's just how that ties in with LLMs while making an analogy about a script of fiction is obviously stretching it a bit in making a sharper takeaway.


Pretty sure Robin Williams did live the human experience.


But the whole argument in the monologue is that Will Hunting hasn't - that it's possible to have had some of "the human experience" but still be unqualified to talk about the parts you haven't.


> Decide whether it's accurate, insightful, worth thinking about and researching further, etc. based on its substance

This is the part the original human poster is assumed to have screened as a first step, not the audience, particularly if the audience is unfamiliar with the subject (such as a guide, etc).

I literally came across a guide online from a user who wasn't a spammer, who disclaimed they haven't even read the very guide they posted as an article on their website, as it was LLM generated. At least that user put up a disclaimer but why would I trust such a guide, given my and others' extremely inconsistent experience with the veracity of LLM output and as someone coming to the guide to learn (ie: not a domain expert)? Overwhelmingly other users don't put up such disclaimers so we don't even get to know whether they've vetted anything.

Trust is the key thing. To continually erode reader trust means you're putting the burden at every step on the reader. Sure, one should always apply critical thinking to even human output but there is an implicit, baseline assumption that with human output they're at least familiar with what they've output (whether they're lying or telling the truth or ignorant but honest). LLMs meanwhile handle ground truths in a flaky way, such as when they'll hallucinate quotes from even articles they claim to have read and cited. And the most common models users are using are the cheapest/free ones anyway, only compounding the accuracy issues.

Imagine you went to a library assuming authors, publishers and library staff have done some minimum due diligence only to find the library is being replaced rapidly with books that no one in the chain has read.

No one can be a domain expert in every single thing they encounter, which is why we place trust in others to varying degrees to fill in the gaps based on their experience and knowledge, even if you're a dyed in the wool skeptic. When increasingly what we encounter isn't being vetted as a basic first step then it's a waste of time and rude to the audience, which only decreases peoples' tolerance for bullshit and increases cynicism (which we could use less of).


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: