Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

For web scraping I used htmltidy (https://en.wikipedia.org/wiki/HTML_Tidy) which cleaned it sufficiently that I could run XSLT over it (gags at the memory)


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: