It's not clear how to decide which threads to bundle together. Checking identical URLs is too narrow (e.g. the three threads mentioned have different URLs) and how to match by content isn't so easy.
With modern NLP models being extremely good, can't you bundle together all sites whose content embeddings are very close? Presumably there's an acceptable threshold somewhere that achieves a good tradeoff of false positives vs negatives?
https://news.ycombinator.com/item?id=34681636
https://news.ycombinator.com/item?id=34850571