Hacker Newsnew | past | comments | ask | show | jobs | submit | richardfey's commentslogin

Looking forward to giving this a try with llama.cpp. I’m watching the open-weights competition with high expectations.

I think that would spoil a lot of the fun. I would favour visual and audio feedback loops to interactively learn how things work under the hood.

Sounds like a great way to inflate your success metrics for AI queries

> Part of this can be good (you talk about what they care about, where 90% of broadcast messaging might not apply) and part of it can be bad (manipulation.)

Side note: it's manipulation either ways because you chose what to talk about, with a goal in mind.


Actually, no. "Manipulation" is a negatively loaded word, and you wouldn't use that word if f.ex. someone helpfully & truthfully helps others see they've misunderstood sth.

I disagree. Whether something is manipulation depends on whether you are trying to change someone’s opinion or behavior, not on whether the manipulator has “good” or “bad” intentions, since those judgments are not objectively universal.

You can look up the word in a dictionary, instead of arguing. Bye & have a nice day.

Serious question: how do we verify claims like these on the effectiveness of a harness?


Has anyone tried Kimi K3 against gpt-5.6-sol on real projects?


Exactly; this is a no-go for me, I will wait for an independent provider to sell the service, which is possible thanks to the open weights.


They could use an agent to summarise the source material, and then train models on those summaries, and claim that some sort of clean-room training has happened?


You've been to too many meetings with PMs and directors saying "An agent could very easily do this"


I understand you're trying to be funny, but my point is that with novel technology there are novel ways to claim innocence in courts because of the legislative void.


I can't find on their website some indication of what kind of usage I can get out it, otherwise I'd be interested.


$6 a month I plan to use deepseek v4 flash mainly which should provide closer to 5x the usage on the cheaper ones but no set number


This is a great statistical analysis and it was a pleasure to read, but I wasn't expecting the claims to be so poorly supported. There's also a reply from one of the Meta authors there, worth checking out.


Which claims do you think are poorly supported? I'm the author; I tried to include everything other people need to repeat the same experiments. I've even had two people write in directly to me, stating that they have been able to replicate my findings.

Or are you referring to the claims from Meta, Google, and Adobe -- which failed to hold up under independent evaluation.

You are also correct that one of the meta authors wrote in a comment. However, he demonstrated a clear lack of understanding regarding what makes the bits "independent" or how to resolve the independence problem.


> Or are you referring to the claims from Meta, Google, and Adobe -- which failed to hold up under independent evaluation.

This.

> However, he demonstrated a clear lack of understanding regarding what makes the bits "independent" or how to resolve the independence problem.

Yes, I didn’t want to call it out explicitly, but this is exactly the kind of thing that would have made my undergraduate statistics professor lose patience.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: