fix: backslashes in uploaded HTML files turn into line breaks or break the upload (#31450)

With the default content extraction engine, backslashes in an uploaded .html or .htm file were read as escape sequences. A path like C:\new\table was saved with a line break and a tab in it, and a page containing C:\Users failed to upload with a 'unicodeescape' codec error. HTML files are now read the same way as .txt and .md uploads, so the saved text matches the page.

Fixes #31440
This commit is contained in:
Classic298 2026-09-27 21:08:14 +02:00 • committed by GitHub
parent b39abad5c2
commit 59ea3b7c2c
No known key found for this signature in database
GPG key ID: B5690EEEBB952194

View file

@ -745,7 +745,7 @@ class Loader:
)
loader = TextLoader(file_path, encoding=self._detect_text_encoding(file_path))
elif file_ext in ['htm', 'html']:
loader = HTMLLoader(file_path, encoding='unicode_escape')
loader = HTMLLoader(file_path, encoding=self._detect_text_encoding(file_path))
elif file_ext == 'md':
loader = TextLoader(file_path, encoding=self._detect_text_encoding(file_path))
elif file_content_type == 'application/epub+zip':