Hacker Newsnew | past | comments | ask | show | jobs | submit | PeterSmit's commentslogin

Cool idea. Crashes on iOS though, after downloading the 1.7B model

Which makes total sense given the current political situation.


With the huge hoods these things have the driver has a hard time seeing what is right in front of them, and when they hit a pedestrian (kid or adult) they are much more likely to die.

https://www.carscoops.com/2024/12/suvs-and-pickup-trucks-2-3...


I’ve had this same idea, and it doesn’t work. Or at least: it works quote well, but the problem is that you get hallucinations. And it can be incredibly discouraging to find out the flashcards you’ve been cramming are completlh wrong.


I've had this same problem using ChatGPT and German. Even for basic German hallucinations can be unexpected and problematic. (I don't recall the model, but it was a recent one.)

In one instance, I was having it correct akkusativ/dativ/nominativ sentences and it would say the sentence is in one case when I knew it was in another case. I'd ask ChatGPT if it was sure, and then it would change its answer. If pressed further, it would again change its answer.

I was originally quite excited about using an LLM for my language practice, but now I'm pretty cautious with it.

It is also why I'm very skeptical of AI-based language learning apps, especially if the creator is not a native speaker.


Would agentic workflows come in handy in these cases? I mean having a controller agent after the sentence is created, where this agent would be able to search the web or have access to a database? or personal notes and ensure everything is correct.


Maybe? I suppose it depends on the quality of the controller agent, which then comes back to the quality of the original LLM.


What models have you been using for that? While I haven’t tried automating the production of vocabulary lists through an API, within the last few weeks I have had the chat versions of ChatGPT 4o, Claude Sonnet 3.5, and one of the latest Gemini models produce annotated vocabulary lists based on literary texts in English, Russian, and Latin. I didn’t spot any hallucinations.

I was asking only for the meanings of the words and phrases, though. I didn’t ask for things like pronunciations, grammatical categories, etc. In the past, when I’ve tried to get that kind of granular information from LLMs, there were indeed errors, presumably because of tokenization issues.

A few days ago, I ran some similar tests with Japanese, asking for readings of kanji and jukugo in an extended text. All of the models I had tried before for such tasks had screwed up. This time, however, ChatGPT o1 scored 100%. It also was able to analyze sentence grammar accurately, unlike the other models I tried. I was impressed.

At current API prices, though, o1 might be a bit too expensive for such a task.


I wonder if there are any benchmarks specifically designed to evaluate LLMs' performance in language learning tasks


I haven’t heard of any. It would great if there were....


I had this problem initially but found that if you use these then hallucinations mostly go away.

1. Role based "agents" with a router and logs (for auditing reasoning and decision making).

2. Cross validation and redundancy with the translation "agent" using a 2nd language (that is not English) that you are also native in to check if the translation carries the same "meaning" (sentiment) and cultural significance (Turkish is especially rich in symbolism and cultural memes).

YMMV: I am a car salesman irl and have no formal training.


I'm not sure if you're from the Netherlands, but I can assure you it's more nuanced that this. Mixing only works when cars are not dominant, so you need low car volumes and low speed in these areas. Residential areas in cities are an example of this: no through traffic, max 30kmh limit.

Most of (new) Dutch road design is designed to give pedestrians and cyclists multiple safe options, while cars have to take the long way round. You can in theory still get basically anywhere with a car if you need, but often (especially in cities) it easier to walk/cycle/take the train/tram/metro. The result is that things can be closer to each other (no parking moat everywhere) so in the end the trip is shorter and safer for everyone, including people choosing to take the car.


As an example: More and more "cars are guests" roads are being added. These are usually cycling dominant routes and while completely removing cars might be preferable it's not always possible. Due to the roads being designed as widened cycling paths (and look like it) which barely fit a car you can have cars there but you'd think twice driving there, which makes the drivers more cautious and lowers the car traffic volume a lot. Note: the throughput of a cycling path far exceeds that of a normal road per surface area used (about ab order of magnitude vs cars).


I would argue the point of the article isn’t “we need more bollards everywhere “, it’s “our regard for pedestrian safety is absurdly low, even cheap tools to increase pedestrian safety (like bollards) are uncommon / controversial"


If the harm to society is 1M$ per car, should we be driving them at all?


No, we shouldn't. But the cost of installing bollards and the "harm to society" are two distinct costs. There are about 2.37 pedestrian deaths per billion vehicle miles traveled in America. Even if we assign a generous cost of 5M$ per life lost, that only amounts to a 0.01$ per mile driven, which is probably not enough to cover the cost of installing bollards all over the place.


How about 10 bollards per sold car, you can probably get away with $ 1k per bollard (including installation), most cars cost a multiple of 10k. Let‘s see how far that gets you. You can of course modify the bollard tax by car weight or by price.


So $10K extra per car? Assuming we're talking the US here where the average price of a new car is under $50K, that's more than a 20% bollard tax.


The average price is really that low? Do you have any pointers/data?


This Fortune article[0] lists an average as of January of $47,338, I believe based on Kelley Blue Book.

[0] https://fortune.com/2024/02/28/how-expensive-new-used-cars-o...


Not in the cloud.


It's all fun and games until the websites you need (bank, local government, etc) only support the pipes build by a tech corp in a faraway country. Diversity is good and healthy.


>until the websites you need (bank, local government, etc) only support the pipes build by a tech corp in a faraway country

they don't get to decide this if push comes to shove. Banks and governments in European jurisdictions obviously can be forced to comply with European laws and if there was some geopolitical question about security you can just force them to switch to a local fork of Chromium which given that it's open source is technically relatively trivial.

It's the same as Linux essentially. The overwhelming majority of commits comes from RedHat, Huawei and Samsung or other international corps which is fine because there's always the implicit option to fork it. We don't need fifty different kernels given that we're talking about open source software. In the olden days of Internet explorer and dependence on proprietary software this argument made sense because you could theoretically be squeezed without an ad-hoc alternative, but that's not the case any more.


We are on the same page here.


And also basic levels of German and French. Outside of tech I actually hear from multiple sources German is an important language for trade between medium-sided companies, to Germany of course but also a large portion of the other eastern EU members.


Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: