bastimeyer · GitHub

This fixes the character encoding in parsed HTML documents. This should've been caught way sooner.

lxml's XML parser requires bytes as input, hence why strings must be encoded to bytes first (default utf8 encoding). The content is then decoded and parsed correctly as utf8 regardless whether the XML declaration with the document's encoding is missing or not.

The HTML parser on the other hand treats byte inputs differently, but only when the <meta charset="utf8"> tag is missing in the first X bytes of the document. So if we encode input strings here to bytes as well (default utf8 encoding) and if the tag is missing too, then this will lead to decoding errors, as the parser won't treat the input as utf8 encoded data.

See the added tests.

Read the original on github.com ↗