Hacker Newsnew | past | comments | ask | show | jobs | submit | crthpl's commentslogin

what specifically has changed about Claude Code?


https://cchistory.mariozechner.at

and as is normal for hosted models, almost everything... based on load flucation they may even send your prompt to a quantised model


it does not use haiku. you can just ask Claude (possibly not fable BC of cyber) to reverse engineer the obfuscated JS.


I've had Opus refuse to do work due to cybersecurity concerns as well. Quite frustrating when you're trying to fix an exploit--as soon as it reads the bad code, it just shuts down! I ended up cutting out more and more bits of the code until eventually Opus relented.


For the US models, you can look at the rewritten CoTs or just ask them what they think is happening.


Asking them what they think is happening is actually not reliable, though? They don't always know how they came to a conclusion after the fact. It is probable that it's roughly similar to the path they took to get there, since it's the same weights, but it's not certain. And, if you make it standard practice to always collect that data (e.g. if you have an automated tool to ask the model to explain itself after every action to log it), it seems like you might find yourself being blocked for violating terms of service. It looks like "distilling".

In short, there are workarounds, but they're not guaranteed to work forever and they're likely to bump into terms of service.


where the heck are you seeing OpenAI racing to IPO before anthropic?


on the API, their margins are very high


no, it is discounted expected future cash flows


Yes it can! That's the whole point of RL! it generates slightly out of distribution rollouts, and rewards good rollouts to change the distribution of the output


That's not out of distributíon, that's inside the distribution of the rollout. If you don't create rollouts for the game of Chess then it doesn't know how to play Chess no matter how smart it is at tasks you've created rollouts for. It's structurally stuck in its distribution.


where the heck did you get those parameter numbers from?


Sonnet and Opus are from Elon Musk (given the people he's hired it seems likely it is approximately true). Mythos is quite widely spoken about.


are you sure it's 60 fps and not uncapped?


have you tried pangram? it's basically the only good AI detector, and they have nearly 0 false positives


> and they have nearly 0 false positives

I really don't see how this can be possible unless they're accepting abysmal recall? Perhaps I'm missing something fundamental here, but the idea that AI and non-AI assisted text can be separated with "nearly 0 false positives" just says to me that it's really just a filter for the weakest, most obvious AI generated text. Is that valuable?


Pangram's explicit pitch is extremely low false positives, accepting that a higher rate of false negatives is acceptable.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: