German law forces you to provide address/contact info on any commercial website (just running ads qualifies).
But this is often "abused" by lawyers to basically send cease-and-desist letters to small websites that don't have one (typically because they're unaware of the law, not because of malicious intent).
On any website, except one that can be proven to have absolutely no commercial utility whatsoever.
Your blog qualifies as commercial if you write about software at all, because someone might see your software posts and use it to influence a hiring decision.
I think you point out something really important: there real value isn't the data, it's in surfacing the useful data.
For amateur astronomers, an interesting question is whether there is a satellite above me that I can see tonight (and when / where)? For professionals, the literal million dollar questions are more like: are there any satellites on a collision course? Which satellites have moved recently and why? Is there a dime-sized piece of metal somewhere out there that could hit my (employer's) satellite?
Seems to have taken several open-source pieces of data (live satellite feeds, NASA UAP files) and combined them into a vaguely conspiracy-ish website. It's really odd, but has the feel of old-timey conspiracy websites spruced up with vibe coding.
I expect that building the satellite viewer is now just a few dollars worth of tokens. I built something like this a few years ago, and it took a week of work. It's cool that the marginal cost of satisfying your curiousity about where things are in space has dropped so close to zero, but don't be fooled into thinking this website lets you in on some vast conspiracy.
Are you saying that "entropy" = "worse" and that somehow these human stories of earlier golden ages reflect a law of physics? Seems much more likely that they, like this article, reflect the ingrained human psychological tendency to remember the past as better than it actually was...
> remember the past as better than it actually was
Yes, I very clearly remember the not so faraway past when shoes last decades, when non-filled soles (as opposed to e.g. grids of blades) would have been considered a crime and an obvious matter of public health, and similarly for the use of toxic materials.
I also remember the times when machinery used quality material and lasted accordingly; when blinding lights were obviously considered a basic matter of health and when beeps and bebibeep melodies were considered a nuisance. And much more.
I grant you: the past offered people and systems with common sense - now largely gone.
Depends on your definition of what "real" means. Does it mean that people really have the experience of seeing little people? If so, then in that sense it's real. Or does it mean that these little people exist in some other way, outside of the mind and potentially accessible to other physical means of intervention? Probably not.
Accurate weather forecasting has been one of the major achievements of the 20th and 21st century. Computing power is a central piece of this story, but it's also important to remember that the government infrastructure in place to collect ground-truth current weather data is utterly critical to these model's successes. From launching weather balloons to running global weather-monitoring satellites, the scientists and systems at NOAA/NWS (and in this case, the UK counterparts) provide critical expertise and data.
I say this because it seems that earlier announcements where industrial deep neural nets "outperformed NOAA" likely encouraged the slash-and-burn Trump administration in its gutting of critical activities and centers of expertise at NOAA. The impression that industry can predict weather better than the government agencies totally misses that the industrial models utterly rely on government data for inputs. In fact, almost all weather reports you see---weather.com, TV, etc.---are just lightly repackaged products that NOAA provides for free on weather.gov (which you can access for free without ads).
Clench the other fist while writing. This loosens the writing hand. Something in the subconscious releases tension in the other hand. It's a trick I've used with reasonable success and it's immediately effective. It does require a bit of conscious effort, but overall it's a great hack.
What I think is really going on is an attempt to segment the market in favor of Google's strengths. It's a bet that models are "good enough" for many use cases even before they reach human-level intelligence, and Google is trying to capture workflows where quantity beats quality.
They are likely deliberately avoiding the SoTA race for a few reasons:
1. Their best models are marginally better than current SoTA releases.
2. They'd like to let Ant/OAI make mistakes with safeguards / let them get the regulatory heat. The unknown unknowns are huge with SoTA models (eg OAI accidentally hacking huggingface) and they are protecting their reputation.
3. They want to encourage companies to become cost conscious because they can likely win on price in the long run. Getting market share in "quantity beats quality" workflows forces companies to establish processes to choose the "cheapest acceptable model", which is a good environment for Google.
I've heard of capture technologies that literally pump organic matter down far enough that it sinks to the seabed (high enough pressure collapses air bubbles and results negative buoyancy). The problem is that your pumps need to operate using less carbon than you capture (easy with solar?) and that they need to be durable enough to pump billions of gallons of seawater with low maintenance (much harder, salt is terrible for machines).
Even if you get it to the seabed, the majority will be remineralised there. And if you pump enough down there to bury significant amounts of organic carbon, the bottom waters will likely become hypoxic as a result.
I know a bit about this field. This conjecture reads as somewhat more niche than the cyclic double cover conjecture recently proved by OpenAI, but nevertheless represents a real contribution.
You want to know how long it takes to solve an optimization problem, in this case over convex, lipschitz functions. (The restriction to a spherical domain is not really a restriction, you can just change variables for any bounded domain.) Anyway, showing upper bounds on time complexity is "easy" because it's just the runtime of your algorithm. Showing (nontrivial) lower bounds is usually much harder because it requires constraining all algorithms.
This proof apparently shows that the lower bound time complexity is equal to the time complexity of an existing 30-year old algorithm: it requires Omega(d^2) function evaluations to solve over this class of functions.
My gut says likely implies that d is the minimal number of evaluations if you have a gradient oracle because you can approximate a gradient with d function evaluations, but I'm not sure how hard it is to make that rigorous.
> so advanced that it's just as readable to me as Greek
I used to feel this way about statistics.
The language and terms are hard to understand and many of the formulas are taught as "just memorize this" instead of building up from first principles.
But then I started using statistics to analyze something I cared a lot about (paintball) and I quickly realized it's like learning anything new:
- there is jargon
- and core concepts
- when you learn the above, it suddenly makes a lot more sense.
I gotta know what you use stats for regarding paintball. I haven't played in years but I loved playing back in the tipman 98 custom era (not sure if that's still a popular marker).
- I started with just pen, paper and a stopwatch (as a college coach)
- I assumed paintball would be more like football where it's hard to track individual effects
- Turns out it's a surprisingly simple and stable "state machine". e.g. the odds of winning with +1 body (e.g. 5v4, 4v3 etc) is, in college, about ~75%
- Paintball is one of those sports where "the weakest player determines the outcome". Why? b/c if 1 player gets out early, you are fighting out of a hole.
It also made me appreciate that as good a book as Moneyball is, reading it after you try to create analytics for your own sport makes it 3x as enjoyable/insightful.
One downside though:
I would watch games and I got so good at internalizing the stats per state of the game that it was like watching the world series of poker where I could see both player odds of getting eliminated and probability of winning over time charts as I watched the games. Made it harder to be the "come on guys! we can win this" coach when we were down on points + bodies.
> Made it harder to be the "come on guys! we can win this" coach when we were down on points + bodies.
Would you say that paintball on that level is almost deterministic?
Compared to, for example, (American) football where momentum and perseverance can turn games around completely.
This is my exact experience right now trying to wade through the research on psychometrics and skill/knowledge assessment design. It’s mostly just applied statistics but like all such fields, over decades of specialization it acquired its own jargon for abstractions that are quickly recognizable to anyone with a sufficiently developed nose for modern mathematics. But you still have to wade through all the definitions to make those connections before you actually understand everything.
Not to diminish the comment, but most things are not as complex as they sound when phrased in everyday language or sound much more complex than they are when phrased in technical language.
Technical language is a tool that allows insiders to say less and refer to more, and to be specific, but it's just a tool. Most things can be described in accessible ways.
I think you'd be surprised at what you could understand and at just how few domains are truly complex enough that a layman couldn't understand with a little bit of patience and an accessible summary.
It took me a while to understand a lot of these math concepts.
Turns out people doing Engineering research are using a very small but powerful bag of tricks from a handful of few famous Mathematicians. The concepts are named after them!
I don’t think OP made much effort to make the comment accessible to non-experts, and so it should be taken as a gauge of the fundamental difficulty of the topic.
Very confused by this comment. The older (poorer) parts of the ML literature focus on models with convex and (gradient-)Lipschitz objectives, but that's not representative of reality, not even close. Modern objectives for AI models are famously nonconvex (catastrophically, from the point of view of classical optimisation theory), and that's where the interesting research is.
I'd push back on this. Most of the core optimization techniques (eg, ADAM, stochastic gradient descent) are straight out of the convex optimization literature. Generally you need to use optimizers that work well on convex objectives because near minimizers, functions tend to be convex. (Proof by contradiction: a non-convex point has a strict descent direction.)
The fact that neural networks are highly nonconvex has encouraged a lot of research, but it's more of the kind aimed at resolving tension: these methods are probably good for convex functions, why do they continue to work for nonconvex problems, and are there tweaks we can make to improve them in that setting? It's not a lot of de novo theory; more standing on the shoulders of giants, etc etc.
No, I have to push back as well, sorry. It takes a very long time to get to the "near-minimizer" stage when training a neural network, and in practice, you never get there (see neural scaling law regimes). What you are saying is the viewpoint from 6-7 years ago. Things have changed.
The reasons why optimizers work well for neural networks in their highly nonconvex landscapes has absolutely nothing to do with their performance in convex landscapes. If that were true, everyone would be using Newton-CG. These optimizers were born in the convex optimization literature as a consequence of the genetic optimization nature of incremental publication (and because that was all we had), but their modern study is through the lens of implicit regularization (their preferences for certain minima) and their stepwise vs. continuous rates for feature learning in multilayer models.
This is completely new theory by the way, and requires painful reinvention of the field. It does not stand on the shoulders of convex optimization. The nonconvex setting is assuredly not a perturbation of the convex setting, and those that do continue to work on deep learning optimization from the convex optimization perspective are well behind the times.
It seems that we have two different stories here: in one, the new optimization theory represents a stark departure from the prior art, a sort of revolutionary new view of the understanding of optimization as applied to neural networks.
In the other story, the current understanding of optimization is a natural evolution of past work, where a new generation of researchers respond to social and technological changes, adapting and building on the work of the past, taking what's useful, downplaying the importance of some ideas, and inventing new language to describe concepts that seem most relevant to the current situation.
Both stories tell some of the truth. A revolution or evolution? Looking at the literature (eg the sibling comment here) shows that even today, convexity is used as an intuition pump for modern optimization techniques. But there are also new ideas that apply to the specific exigencies of neural nets, and downplayed ideas (eg convergence rates) that seem less relevant.
The optimizers are lifted from convex optimization, but the point above was that they are applied to highly non-convex problems. They work for finding local minima, but a lot of the deeper literature does not translate (e.g. the conjecture being discussed in this post).
I'll point out that "does not work" is not the same as "not as efficient" :) But it does seem the Adam paper had an error.
I think that Nesterov's first order method is the most efficient general first order algorithm on convex problems, so anything else is in some sense worse. (Edit: removed incorrect ADAM comment.)
Yours' "not as efficient" in [2] means that, sometimes, ADAM "does not work." Look at figure 2, ADAM literally does not work in the case of "true model."
Yes, apologies, I didn't read the articles you linked before posting this. I did update the comment.
I don't think this changes the point, which is that most optimization methods used in AI owe a substantial intellectual debt to convex optimization theory.
I love convex optimization and there are a few SciML projects I am on where I really need results from there. But in AI research with deep neural networks, it's become a liability, because people will just not let go. I'm getting tired of reviewing convex optimization theory papers in ML conferences that are still trying to wave away the obvious issues with their application to deep learning. It's harsh, but I do feel we can only start talking about an intellectual debt once that stops being the case.
Another intuition is that near a minimum you can Taylor expand the function and show that the higher order coefficients (past the square) are negligible.
It's not a matter of whether the theory "works"; it's a matter of whether one is asking the right questions. Convex optimization studies how quickly an optimizer can reach the optimum. In the non-convex case, there are many basins containing their own local minima. The more sensible questions there are "which basin is it likely to go into?" and "how do I steer it to go where I want?". Global convergence rates are largely irrelevant by comparison.
Yes, order d is the minimal number of evaluations of gradients needed for the same problem! That has actually been known since 1979 (Nemirovsky and Yudin showed that), and there are methods with the same complexity so this question in the gradient model has been solved for a long time. "because you can approximate a gradient with d function evaluations" was exactly why d^2 made sense as a lower bound for this case! Basically, the lower bound question can also be thought about as "can you do better than approxing a gradient?", so this result says no.
I'm sorry this comment didn't sit well with you. My goal was to induce discussion by describing the claimed result (which was buried in the post), not to discourage it.
If you have more specific feedback on what you found distasteful, I'd be happy to hear it.
I did not see _alternator_'s comment as asinine. I like a venue where people who have some expertise feel comfortable enough to share it, and are not criticized for doing so
reply