screenshots.debian.net

Description

How is html_text different from .xpath('//text()') from LXML or .get_text() from Beautiful Soup ?

* Text extracted with html_text does not contain inline styles,

javascript, comments and other text that is not normally visible to

users;

* html_text normalizes whitespace, but in a way smarter than

.xpath('normalize-space()), adding spaces around inline elements (which

are often used as block elements in html markup), and trying to avoid

adding extra spaces for punctuation;

* html-text can add newlines (e.g. after headers or paragraphs), so that

the output text looks more like how it is rendered in browsers.


Upload more screenshots

Please help extend the collection of screenshots. Just make a screenshot and upload it here. You don't need to register or anything.

Upload a screenshot

Hint: upload an image here from your clipboard with Ctrl-V


Homepage

https://github.com/zytedata/html-text


Install this software package

If the package is available for the distribution you are currently using on your computer then install the software by clicking on…

Install python3-html-text

Read the original on screenshots.debian.net ↗