It reads to me like "We did all the work you'd do to figure out how to fix the benchmark, then we decided to throw out the benchmark". Is there some reason the underlying data is so golden that it can't be patched? At the end they argue for a slightly more curated approach to benchmark generation, but my gut is that using messy ill-specified tests taken from real world data and patching them into fairness would be a pretty solid path to take.
And it reads to me like they have some other reason to move on from SWE Bench Pro, but they don't want to say what it is. They say right up top, "~30% of the tasks are broken." But that leaves ~70% un-broken, which seems pretty good to me. It would be nice if they would also say: "Here's the list of instances that are broken: <CSV>". Or, "Here's the subset of SWE Bench Pro we will use going forward." They're letting the perfect be the enemy of the good.
Pointing out problems (e.g., hidden tests that assume narrow implementation details) is much easier than fixing them (e.g., creating tests that work for any possible choice of implementation).
If they fixed it, then it wouldn't be SWE-Bench Pro anymore, right? It'd be "SWE-Bench-Pro-Fixed-OpenAI." I think it's better optics for the independence of the benchmark if the OpenAI team lets some third party do the fixing and release the improved benchmark.
...Although OpenAI did exactly that when they released SWE-Bench Verified, so maybe I'm talking out of my butt here.
There's a book covering this and more from 1993 called "Strange Attractors:
Creating Patterns in Chaos" by Julian C. Sprott that's freely available here: https://sprott.physics.wisc.edu/SA.HTM
It's fun (errr... for me at least) to translate the ancient basic code into a modern implementation and play around.
The article mentions that it's interesting how the 2d functions can look 3d. That's definitely true. But, there's also no reason why you can't just add on however many dimensions you want and get real many-dimensioned structures with which you can noodle around with visualizations and animations.
As an undergraduate I worked with some other Physics students to construct an analog circuit using op amps that modeled one of Sprott’s equations and we confirmed experimentally that the system exhibited chaotic behavior. We also used a transconductance amplifier as a control parameter and swept through the different states (chaotic, period windows) of the circuit. We did not go as far as comparing the experimental and predicted period windows while I was there but it was an interesting project for us. At one point I turned up an article in Physica D describing how to calculate the first Lyapunov exponent using small data sets which we used to compute whether we were in a period window or not.
Sibling comments are giving good answers about Silksong in particular. But, a more general point to keep in mind is that gaming by revenue is an order of magnitude larger than movies. In that light it's a little strange that gaming news/events don't hit the general attention sphere more than they do.
Check out the game Bombe [1]. It's a minesweeper variant where instead of directly flagging or uncovering cells, you define rules for when cells can be flagged. As it gets more advanced you end up building lemmas that implicitly chain off each other. Then as _you_ get more advanced (and the game removes some arbitrary restrictions around your toolset) you can generalize your rules and golf down what you've already constructed.
Also, China's battery production is described as a "battery complex" while US battery production is described as a "battery industry" or "battery industrial base".
The video meanders for a bit at the start, but about a third of the way through turns into a pretty interesting breakdown of how deferred rendering works in general, and then specifically Breath of the Wild's deferred rendering passes.
Waypoint was lost. Good gaming media is like a big water droplet that keeps on getting smushed by a the big dumb thumb of capitalism. It still shifts to a new spot, but ever diminishing. One of the rare ones that even made money. Still, gone.
Should thinkers read? As a thinker, I thought I'd venture here to the reading room to see if I could glean any data-driven insights. What I found was a new form of Readable Thinking Words (aka "writing") called a blog. After a moments perusal I realized this was a kind of web log, something I'd come across in my last reading room excursion 20 years ago. Despite my ignorance, I knew I'd be able to leverage my thinkability into something positive for the readers, so I decided to leave this enriching comment. What a huge improvement for you all! Thinkers really should read every now and then.
This is something I call External Thought Driven Decision Infrastructure. If you sign up for my newsletter, you'll receive a free 14 page pdf outlining the groundbreaking method.
> Less than 1% is radioactive for 10,000 years. This portion can be easily isolated and shielded
How would this work? My assumption was that the pellets are fairly homogeneous. Does the decay happen faster in exterior of the pellet? Or is there some process to concentrate the radiation?
The pellets are homogeneous. The reprocessing of fuel involves melting or dissolving them and chemically separating out the waste products, with the remaining unused fuel going back into new pellets.
The waste products are spread throughout the fuel pellets evenly, so the pellets have to be deconstructed to remove them.
I am wondering the same question. I suppose if the decay is totally random throughout any given volume of uranium, then separating it out would have to be chemical or electromagnetic or something?