langchain

mirror of https://github.com/hwchase17/langchain synced 2024-11-06 03:20:49 +00:00

History

Daniel Chalef b157e0c1c3 Add HTML document_loader that includes page title metadata (#1720 ) This `BSHTMLLoader` document_loader loads an HTML document, extracts text and adds the page title to the returned Document's metadata. The loader uses the already installed bs4 package to extract both text content and the page title. Included in this PR is an example HTML file and an integration test that tests against this file. --------- Co-authored-by: Daniel Chalef <daniel.chalef@private.org>		2023-03-16 21:47:17 -07:00
..
integration_tests	Add HTML document_loader that includes page title metadata (#1720 )	2023-03-16 21:47:17 -07:00
unit_tests	sql: do not hard code the LIMIT clause in the table_info section (#1563 )	2023-03-13 23:08:27 -07:00
__init__.py	initial commit	2022-10-24 14:51:15 -07:00