Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Again, this is not true of lxml. XPath was made for structured data, but in lxml you can give it the HTMLParser factory and it will handle rubbish HTML just fine. I use lxml/XPath professionally to scrape economic data for an ibank, and you'd be surprised at the Microsoft Frontpage-type spaghetti HTML I parse with it -- and it all works great.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: