You cannot select more than 25 topics Topics must start with a letter or number, can include dashes ('-') and can be up to 35 characters long.
langchain/libs/community/tests/examples
Eugene Yurtsev cd52433ba0
community[minor]: Add `SQLDatabaseLoader` document loader (#18281)
- **Description:** A generic document loader adapter for SQLAlchemy on
top of LangChain's `SQLDatabaseLoader`.
  - **Needed by:** https://github.com/crate-workbench/langchain/pull/1
  - **Depends on:** GH-16655
  - **Addressed to:** @baskaryan, @cbornet, @eyurtsev

Hi from CrateDB again,

in the same spirit like GH-16243 and GH-16244, this patch breaks out
another commit from https://github.com/crate-workbench/langchain/pull/1,
in order to reduce the size of this patch before submitting it, and to
separate concerns.

To accompany the SQLAlchemy adapter implementation, the patch includes
integration tests for both SQLite and PostgreSQL. Let me know if
corresponding utility resources should be added at different spots.

With kind regards,
Andreas.


### Software Tests

```console
docker compose --file libs/community/tests/integration_tests/document_loaders/docker-compose/postgresql.yml up
```

```console
cd libs/community
pip install psycopg2-binary
pytest -vvv tests/integration_tests -k sqldatabase
```

```
14 passed
```



![image](https://github.com/langchain-ai/langchain/assets/453543/42be233c-eb37-4c76-a830-474276e01436)

---------

Co-authored-by: Andreas Motl <andreas.motl@crate.io>
3 months ago
..
README.org community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
README.rst community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
brandfetch-brandfetch-2.0.0-resolved.json community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
default-encoding.py community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
docusaurus-sitemap.xml community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
duplicate-chars.pdf community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
example-utf8.html community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
example.html community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
example.json community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
example.mht community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
facebook_chat.json community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
factbook.xml community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
fake-email-attachment.eml community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
fake.odt community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
fake.vsdx community[minor]: New documents loader for visio files (with extension .vsdx) (#16171) 4 months ago
hello.msg community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
hello.pdf community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
hello_world.js community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
hello_world.py infra: add print rule to ruff (#16221) 4 months ago
layout-parser-paper.pdf community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
mlb_teams_2012.csv community[minor]: Add `SQLDatabaseLoader` document loader (#18281) 3 months ago
mlb_teams_2012.sql community[minor]: Add `SQLDatabaseLoader` document loader (#18281) 3 months ago
multi-page-forms-sample-2-page.pdf community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
non-utf8-encoding.py community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
sample_rss_feeds.opml community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
sitemap.xml community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
slack_export.zip community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
stanley-cups.csv community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
stanley-cups.tsv community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
stanley-cups.xlsx community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago
test_empty.csv community[minor]: Add pebblo safe document loader (#16862) 4 months ago
test_nominal.csv community[minor]: Add pebblo safe document loader (#16862) 4 months ago
whatsapp_chat.txt community[major], core[patch], langchain[patch], experimental[patch]: Create langchain-community (#14463) 6 months ago

README.rst

Example Docs
------------

The sample docs directory contains the following files:

-  ``example-10k.html`` - A 10-K SEC filing in HTML format
-  ``layout-parser-paper.pdf`` - A PDF copy of the layout parser paper
-  ``factbook.xml``/``factbook.xsl`` - Example XML/XLS files that you
   can use to test stylesheets

These documents can be used to test out the parsers in the library. In
addition, here are instructions for pulling in some sample docs that are
too big to store in the repo.

XBRL 10-K
^^^^^^^^^

You can get an example 10-K in inline XBRL format using the following
``curl``. Note, you need to have the user agent set in the header or the
SEC site will reject your request.

.. code:: bash

   curl -O \
     -A '${organization} ${email}'
     https://www.sec.gov/Archives/edgar/data/311094/000117184321001344/0001171843-21-001344.txt

You can parse this document using the HTML parser.