Which model best allows me to transcribe speech that uses a lot of domain-specific terms? For example, when I say "Claude Code", it often gets transcribed as "Cloud Code", and I have to go back and edit or do a second pass with a traditional LLM (which can introduce additional errors).
agree with omneity here. Whisper's initial-prompt trick is exactly that, and several hosted vendors have equivalents (custom vocabulary / keyword prompting). Domain vocabulary is where STT models separate the most in our runs. for example, on medical terms the field spreads from about 8% to 19% WER across models: https://benchmarks.speko.ai/blog/what-a-voice-agent-hears.
We often find that models that wins on clean speech are often not the one that wins on your terms, so test with your own vocabulary rather than a headline number.
I’ve had a lot of success in the past with fine tuning STT using synthetic data.
I was doing it for Veterinary (ambient recording -> SOAP notes) which has tons of complex domain-specific language AND it is critically important to get right.
“CPR” transcribing as “see pee are” just doesn’t cut it in that industry.
Synthetic-data fine-tuning is the other credible answer to domain vocabulary. Curious whether you re-benchmark the fine-tune when new base models ship?
You can't fight this battle in the current judicial system. If before anyone can successfully bring a lawsuit in front of a judge the company has unlimited money, then at least one of these three things will be true:
a. the damages won't make a dent
b. they can throw lawyers at the problem for the win
c. they will have grafted their way into the economy so deeply that no one will want to kill their main source of income
This battle is already lost, if they find economical success then they are untouchable.
Can confirm in the Netherlands also. I'm not an influencer so I don't know exactly, but a few years ago there was either a crackdown, an expansion of the rules, or at least a big media campaign by the government to make the influences that focus on a Dutch audience aware of the rules surrounding promoted content
As long as the fine is cheaper than the earnings, they'll keep on doing it. And fines are retroactive, at that point they've already broadcasted to and influenced a bunch of people.
I am also worried about this. Typical ads are already devastatingly effective against most people, causing incalculable distortions in market dynamics and incalculable losses to consumers. This is an order of magnitude worse.
But there is something even more insidious: elections. Ads will be purchased to shape public perceptions. Those with the deepest pockets will have ads which are charismatic, convincing, and unbelievably effective. This is the death of democracy if regulation isn't adopted right now.
Unfortunately this isn't straightforward. The distinction between a product, service, person, and party are impossible to clearly delineate. Is a SpaceX ad an endorsement of Elon Musk? Is an Elon Musk ad an endorsement of the U.S. Republicans? The famous Citizens United ruling determined that there is no practical legal distinction in this domain, meaning that without a Constitutional amendment (which will never happen in this political climate), the U.S. is going to barrel headlong into some kind of dystopia.
My only ray of sunshine here is competition. If OpenAI gets a reputation for untrustworthy AI, and there are viable competitors, people will hopefully switch. It's messy and imperfect and will require constant vigilance (because they're not going to tell us when they're misleading us). I think open weights models have become very good, and will improve much more over time. When we can reasonably run these models locally, ubiquitously, we have cost effective options.
"Once men turned their thinking over to machines in the hope that this would set them free. But that only permitted other men with machines to enslave them." -Dune
I wish we would stop speedrunning this Dune/Blade Runner future we keep trying to attain.
Look into the Adbusters organization. I'm not sure I agree with them on everything, but they have created some excellent satire over the years and been involved in prominent protests.
ChatGPT is way above typical CPM, maybe $60-$100. That puts it in the same category as premium YT verticals. If they can sell enough at this price point it would make a big impact, but I'm super-doubtful they can do that; campaigns are basically blind at this stage, and I would not be surprised if this was intentional.
From the article, it looks like ads are separate from answers in a clearly delineated way. It's not quite as nefarious as what you're suggesting. Frankly, I'm surprised.
This is only round one however, and I could see a future where they offer more "intimate" experiences. At least that's not today.
But it what's the solution to free services? There is many people that can't afford to have to pay for search, emails, maps. On top of paying for accessing content, micro transactions. For Internet related companies what if the country the person is from, doesn't have access to the financial system of the country that company is operating from.
I think It might be better to regulate them than remove them. Or to find a solution for people that can't access it.
(Another is higher tiers, free for personal users etc)
I hate ads but I also think about what's the alternative
If it's like the commercials in the movie the Truman show I think this could be a nightmare.
I think it all boils down to how (if?) this will be transparent to the user.
If the response suggests a sponsored product/service to the user it could either be the specific product/service that ChatGPT is biased to suggest or the whole idea of using a product/service.
As a user I would like to know both of these biases to understand that using any product/service for my need might not even be the best thing.
Ah, so you are against ads! At least the way ads are today. With that I mean that the way ads today mostly focus on manipulating over informing is effectively hindering your ability to choose.
> Ah, so you are against ads! At least the way ads are today. With that I mean that the way ads today mostly focus on manipulating over informing is effectively hindering your ability to choose.
Speaking as someone who worked in advertising for a looooonnnngggg time, there are three ways that ads seem to work (above a randomised holdout).
1. The ad is something that the person actually wants, and they then convert.
2. The ad is something that the person was about to buy, and they convert.
3. The ad is part of a long running campaign that shifts (some) consumer's behaviour over time (ads for Coke etc fall in this category).
Ads aren't magic, and most people ignore them most of the time. It turns out that enough people don't that they can be profitable sometimes, but mostly that's just targeting people with more money than sense.
This HN notion that ads shape everyone's behaviour (except me, obviously) is so disconnected from my experience that I am sometimes confused by it.
When I walk around my increasingly billboard-infested hometown. I observe that almost every ad actually nudges me in some way: whether I should switch internet providers, whether my financial strategy is fine, what kind of scam this tabloid newspaper is running, etc...
I believe most people are influenced by these constant micro-impressions far more than they realize.
My best defense is to build a resentment towards the companies being advertised.
> I believe most people are influenced by these constant micro-impressions far more than they realize.
Your feelings are fair, but if this were the case then we'd see far more consumers changing their preferred phone/electricity/TV provider on a regular basis, and we could correlate this with (lagged) advertising spend.
But there is very little evidence for this. Like, I have a doctorate in psychology, and behaviour change is really, really hard and only works in a small number of cases. That model does seem to have predictive power, in that very few people convert as a result of seeing an ad (less than 1/1000 in the very best cases), and most of those that do don't keep buying products for long.
Now, if you already do something (like order takeaway pizza) it's much easier to get someone to do more of that, but its very very very hard to establish a new behaviour.
Anecdotally, lots of intelligent people I know like to pretend to be completely uninfluenced by advertising, which I think is delusional.
In my view advertising is in big parts (your category 3) basically a constant side-channel attack against peoples minds, eroding decisionmaking and purchasing decisions (mostly!) to peoples detriment.
The net-beneficial aspects (product information, and, arguably, boosted general demand) could be better achieved otherwise.
Your explanation should be common sense for anyone who has been around long enough and I agree that it seems to be fashionable to claim that ads have a nefarious, hypnotic effect on The Masses. The thing that does seem to have an outsized effect on how people think is their desire to be seen as part of a group and they tend to signal this by repeating arguments they've heard members of that group make. That's where advertising/propaganda has a real foothold in everyone's psychology. If you form a parasocial bond with an influencer or a group it does have a strong effect on your behaviour since you want to be liked by them so you will say and do things you think they'll approve of and agree with what they say and do. No one is really exempt from this since being exempt involves complete isolation and/or extreme disagreeableness. The only real remedy is having strong bonds with other people who aren't hopelessly beholden to outside influences. No small task since it only takes one member of the group to act as a channel.
> The thing that does seem to have an outsized effect on how people think is their desire to be seen as part of a group and they tend to signal this by repeating arguments they've heard members of that group make. That's where advertising/propaganda has a real foothold in everyone's psychology.
Monkey see, monkey do. This is basically how human society works, and is impossible to stop. The best you could do would be control who people hang around with, but even that is almost impossibly hard to make work.
> This HN notion that ads shape everyone's behaviour (except me, obviously) is so disconnected from my experience that I am sometimes confused by it.
Same. They've also never really bothered me, especially billboards and street advertisements. I was actually sad when they got rid of the neon signs in Hong Kong.
Which category is it when the user asks which is the best brand for the thing they want to buy and the platform lies to them and recommends the brand that paid the most money?
Concealed advertising is highly problematic and deserves discussion. It is not, however, what OpenAI introduced in the article above. Their current ad formats are, apparently, transparently labeled and clearly separated.
It is very likely that this will adjust in the future. Just look at how indistinguishable from organic content some prominent ad placements have become over time. But: There is merit in discussing the actual product, not some hypothetical version of it. If we do discuss hypotheticals, we should make that explicit.
This also helps the whole debate (and therefore, I believe, ad critics): arguments grounded in reality are usually far more effective than those that aren't. ... And there might come a point where OpenAI introduces some kind of native ad format. We would be well advised to preserve our credibility for that discussion.
edit: Paid content in the organic LLM output is a current problem already. It ist just not what this announcement is about.
> Unpopular opinion: I think users should be free to choose.
There is nothing unpopular about your opinion - silent majority definitely wants ad-subsidized free tiers (that is why they are so popular and ads are not banned anywhere in the western world), it is just that arguing with radicals on HN or reddit is exhausting and a waste of time in general.
If you are implying that this is the case because this is also how leaders are selected in a democracy, then of course the sarcasm is noted - but here's my real unpopular opinion: Democracy only has value when the majority of voters are not morons. Unfortunately, I can't come up with a better solution - maybe create a model for which there is a wide agreement on and let it rule?
I simply meant that users ought to be free to choose between a product that offers a free ad-supported tier and one that is paid and ad-free. And I do believe that letting morons choose for themselves is the lesser evil.
It doesn't really matter what the system of government is, power-hungry psychopaths will always attempt to game the system in their favour. The history of governance is just a long arms race of inventing new systems to thwart them while they do everything they can to corrupt those systems and concentrate power in themselves. It's like fending off griefers in an online game, there's an endless supply of them, they don't care about the rules or the game, they just want to dominate other players for their own perverse enjoyment. The only lasting solution would either be some form of genetic modification to remove that trait from humanity entirely or some perfect way to identify them but you'd have to be extremely careful that they didn't gain access to and corrupt those things before they were completely eliminated.
I prefer "Democracy only has value when the majority of voters are informed". This covers your case, but it also redresses the balance to shift blame away from everyday people and onto society as a whole.
Isn't society made up of everyday people? While I agree that the public being informed is a mutual responsibility, placing that responsibility on the government (which in many cases has plenty of incentive to keeping people ignorant) will not work either.
> placing that responsibility on the government will not work either.
I agree, but a) the responsibility is on us all: citizens, government, and the media b) we can definitely introduce regulations that force government to be more open, etc.
That's not a unique attribute of democracy nor is it guaranteed. Any system of government that hasn't succumbed to corruption and factionalism will have clear rules of succession. Even an absolute monarchy can go on for many generations without bloodshed.
The author only compared output token costs -- but for typical agentic workloads, input tokens dominate the costs by a large margin. Running inference locally, input tokens are, to first order, free. (They only generate implicit costs through higher time-to-first-token, higher power use, and lower token output speed).
Even ignoring superior caching on a local setup, Mac hardware can often process input token around 10x as quickly as they produce output tokens. Openrouter seems to have only a 2x difference on the same
models.
For larger contexts (eg. 20,000+ token agent workflows), being 10x faster still isn't enough. You have to be close to ~100x faster at crunching contexts for it to feel like realtime.
+1 for introducing them as real-valued functions over cartesian coordinates!
Typically, spherical harmonics are introduced as a complex function over spherical coordinates, which makes them much easier to derive, but imo hides their beauty.
The real-valued, cartesian form of regular spherical harmonics is also called "solid harmonics" or "harmonic polynomials", in case you want to dig deeper.
> At a pragmatic level, can't you say, hey here's something thats probably nothing, let's scan it again in 6 months
If a doctor even _hints_ there might be cancer, the patient will have a terrible 6 months (with actual, measurable negative health impacts of the added stress). Also, at some uncertainy-level (say, 10% chance of cancer) the doctor _has_ to say something and has to schedule expensive followups to not risk liability, even though in 90% of the cases it is not only unnecessary, but actively harmful to the patients.
When, on average, the cost of the screening + the harm done by a false positive outweighs the benefits of an early detection, you shouldn't do the screening in the first place.
At $enterprise, we were just looking for a proper term that sets "responsible vibing" apart from "YOLO vibe coding". We landed on "agent assisted coding".
It's a bit more technical. And it has a three-letter acronym. Gotta have a three letter acronym.
Yes, please don't push "vibe engineering" to mean how you defined it in your blog post. To me, it means exactly the opposite.
I see "vibe" as pejorative. Adding "engineering" does not elevate it from "vibe coding", as I think is your intention in the post, it just shifts "vibe" term to a different domain.
To me, "vibe engineering" means using LLM to develop "design" with no care as to its validity just like "vibe coding" means for "code".
"Agentic xyz" or "Agent assisted xyz" is more fitting.
FWIW, I do not see "vibe" as always pejorative, rather it depends on goals. When quick results and not long term quality matter, "vibing" is a legit tactic.
Anyways, just my interpretations. Please, keep up the good work. Remember, the two hardest things in software are naming, cache invalidation and off-by-one errors. It's good you continue to tackle the zeroth one.
I really like "agent assisted coding". I think the word "vibe" is gonna always swing in a yolo direction, so having different words is helpful for differentiating fundamentally different applications of the same agentic coding tools.
I've used this in the past for collaborative diagramming sessions and love its ease and simplicity, but the point of Mermaid is its portability - ie. can be embedded in Markdown docs and viewed in various editors/platforms.
Thanks @maho! We're hoping to keep the improvements flowing. I'm non-technical but from my perspective I thought Mermaid sequence diagram functionality really shines! Would love to fill the gap in my knowledge. What is better about https://sequencediagram.org/ than Mermaid sequence diagrams?
It's mostly broadband noise that can be simulated by simpler methods, but visualizing possible resonance patterns for the low-frequency emissions from the compressor (which typically runs at 20Hz, 40Hz, ..., 120 Hz) would be good to know.
Although I am not sure how the 2d simulation result carries over to the 3d world...
reply