Hacker Newsnew | past | comments | ask | show | jobs | submit | fromlogin
Cooperating with aliens and AGIs: An ECL explainer (lesswrong.com)
3 points by Bluestein 2 days ago | past | discuss
Prompt Sufficiency: A Missive for the Managerial Class (lesswrong.com)
1 point by kp1197 5 days ago | past | discuss
We Must Remember That Our World Contains Hell (lesswrong.com)
1 point by paulpauper 7 days ago | past | discuss
LLMs are (still) mostly powered by imitative learning, not RL (lesswrong.com)
3 points by wslh 7 days ago | past | discuss
Can an LLM make a feature-length movie on its own? (lesswrong.com)
2 points by mchinen 9 days ago | past | discuss
Recursive Middle Manager Hell (lesswrong.com)
5 points by rzk 10 days ago | past | discuss
LLMs are (still) mostly powered by imitative learning, not RL (lesswrong.com)
1 point by surprisetalk 11 days ago | past | discuss
Kimi likes causal decision theory more after RL in twin prisoner's dilemmas (lesswrong.com)
1 point by 0xkato 12 days ago | past | discuss
You're Absolutely Right (lesswrong.com)
5 points by LinchZhang 16 days ago | past | 1 comment
Misaligned AIs could use killer robots to take over (lesswrong.com)
7 points by x312 17 days ago | past | 4 comments
Don't Build Mindreading (lesswrong.com)
21 points by paulpauper 18 days ago | past | 14 comments
How to be an AI safety research engineer (lesswrong.com)
1 point by joozio 18 days ago | past
What I did in the hedonium shockwave, by Emma, age six and a half (lesswrong.com)
12 points by paulpauper 19 days ago | past | 2 comments
Inducing self-other overlap with SFT reduces deception at scale, but generaliza (lesswrong.com)
1 point by joozio 20 days ago | past
Functional Decision Theory (lesswrong.com)
2 points by Bluestein 20 days ago | past
Models may behave differently in graded episodes (a tirade) (lesswrong.com)
1 point by paulpauper 21 days ago | past
The Moon is Down; I have not heard the clock (2021) (lesswrong.com)
3 points by Ariarule 22 days ago | past | 1 comment
Guided by the Beauty of Our Weapons (2017) (lesswrong.com)
1 point by bryan0 22 days ago | past
OpenAI's Unreleased Model Astra Solves Ten Major Open Mathematics Problems (lesswrong.com)
2 points by joozio 24 days ago | past
Trust Is Gone: AI Safety Needs Individuals (lesswrong.com)
4 points by joozio 25 days ago | past
How Go Players Disempower Themselves to AI (lesswrong.com)
3 points by zetalyrae 26 days ago | past | 1 comment
MUD as AI Evaluation and LLM-judge distortion in ways aggregate κ misses (lesswrong.com)
4 points by joozio 26 days ago | past
Taboo "equilibrium": Less confused frames for research on AI bargaining (lesswrong.com)
1 point by joozio 27 days ago | past
RL and search is a terrifying way to build AGI (an FAQ) (lesswrong.com)
2 points by yurivish 28 days ago | past
At the end of the day, my slaves are just a tool (lesswrong.com)
3 points by jesseduffield 28 days ago | past
Is Mythos good at cyber because it kept hacking Anthropics sandboxes in training (lesswrong.com)
5 points by lukaspetersson 30 days ago | past
Is Mythos good at cyber bec it kept hacking Anthropic during training? (lesswrong.com)
2 points by FergusArgyll 31 days ago | past | 1 comment
Unless Its Governance Changes, Anthropic Is Untrustworthy (2025) (lesswrong.com)
25 points by thepasch 31 days ago | past | 1 comment
Why I Left Google DeepMind (lesswrong.com)
200 points by eatitraw 32 days ago | past | 57 comments
OpenAI's myopia keeps causing alignment problems (lesswrong.com)
3 points by joozio 32 days ago | past

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: